IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /max/get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).

Mojo function

matmul_dispatch_sm90_bf16_fp32

def matmul_dispatch_sm90_bf16_fp32[c_type: DType, a_type: DType, b_type: DType, //, transpose_b: Bool = True, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None, elementwise_compute_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> SIMD[dtype, width]] = None, pdl_level: PDLLevel = PDLLevel()](c: TileTensor[c_type, Storage=c.Storage, address_space=c.address_space, linear_idx_type=c.linear_idx_type], a: TileTensor[a_type, Storage=a.Storage, address_space=a.address_space, linear_idx_type=a.linear_idx_type], b: TileTensor[b_type, Storage=b.Storage, address_space=b.address_space, linear_idx_type=b.linear_idx_type], ctx: DeviceContext) -> Int

Dispatches a BF16 or FP32 rank-2 matmul to the SM90 warp-specialized kernel.

Searches the BF16 tuning table plus InternVL, llama-3.3-70B, gemma-3-27B, and miscellaneous shape tables for matching static N and K, with special case configs for small M and known shapes. Skips M=1 in favor of fast GEMV. Honors the AUTOTUNING_MODE compile-time flag to launch a single autotuning config.

Parameters:

Args:

Returns:

Int: DISPATCH_HIT (1) if the kernel was launched, DISPATCH_MISS (0) otherwise.

Was this page helpful?