For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).
Mojo struct
Inner_matmul_default
struct Inner_matmul_default
Generic CPU matmul microkernel using scalar FMA accumulation.
Implements InnerMatmulKernel for the fallback path used when no
architecture-specific kernel (VNNI, NEON, I8MM) applies. Accumulates
partial products from a packed B tile into a local SIMD register buffer
and writes the result back to the C matrix with optional boundary checks.
Implemented traitsβ
AnyType,
Copyable,
ImplicitlyCopyable,
ImplicitlyDeletable,
InnerMatmulKernel,
Movable
Methodsβ
__inner_matmul__β
def __inner_matmul__[kernel_rows: Int, kernel_cols: Int, simd_size: Int](self, c: TileTensor[Storage=c.Storage, address_space=c.address_space, linear_idx_type=c.linear_idx_type], a: TileTensor[Storage=a.Storage, address_space=a.address_space, linear_idx_type=a.linear_idx_type], b_packed: TileTensor[Storage=b_packed.Storage, address_space=b_packed.address_space, linear_idx_type=b_packed.linear_idx_type], global_offset: GemmShape, global_bound: GemmShape, tile_n_k: IndexList[Int(2)], skip_boundary_check: Bool)
Utility function on the inner loop. Run the inner kernel on the whole (kernel_rows, TileN, TileK) tile.
Parameters:
- βkernel_rows (
Int): Number of rows in the microkernel tile along the M dimension. - βkernel_cols (
Int): Number of columns in the microkernel tile along the N dimension. - βsimd_size (
Int): SIMD vector width used for packed B loads and accumulation.
Args:
- βc (
TileTensor[Storage=c.Storage, address_space=c.address_space, linear_idx_type=c.linear_idx_type]): Output C matrix tile receiving the accumulated products. - βa (
TileTensor[Storage=a.Storage, address_space=a.address_space, linear_idx_type=a.linear_idx_type]): Input A matrix tile in row-major (non-transposed) layout. - βb_packed (
TileTensor[Storage=b_packed.Storage, address_space=b_packed.address_space, linear_idx_type=b_packed.linear_idx_type]): Packed B matrix tile in cache-friendly rank-3 layout. - βglobal_offset (
GemmShape): Global (M, N, K) coordinate offset of this tile within the full matrices. - βglobal_bound (
GemmShape): Global (M, N, K) extent of the full matrices used for boundary checks. - βtile_n_k (
IndexList[Int(2)]): Tile extents in the (N, K) dimensions to iterate over. - βskip_boundary_check (
Bool): Whether to skip out-of-bounds checks on C loads and stores.
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!