For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).
Mojo function
heuristic_and_outliers_dispatch
def heuristic_and_outliers_dispatch[c_type: DType, a_type: DType, b_type: DType, scales_dtype: DType, //, SF_VECTOR_SIZE: Int, transpose_b: Bool = True, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None, elementwise_compute_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> SIMD[dtype, width]] = None, pdl_level: PDLLevel = PDLLevel()](c: TileTensor[c_type, Storage=c.Storage, address_space=c.address_space, linear_idx_type=c.linear_idx_type], a: TileTensor[a_type, Storage=a.Storage, address_space=a.address_space, linear_idx_type=a.linear_idx_type], b: TileTensor[b_type, Storage=b.Storage, address_space=b.address_space, linear_idx_type=b.linear_idx_type], a_scales: TileTensor[scales_dtype, Storage=a_scales.Storage, address_space=a_scales.address_space, linear_idx_type=a_scales.linear_idx_type], b_scales: TileTensor[scales_dtype, Storage=b_scales.Storage, address_space=b_scales.address_space, linear_idx_type=b_scales.linear_idx_type], tensor_sf: Float32, ctx: DeviceContext) -> Int
Dispatches an SM100 block-scaled matmul by selecting a tuning config from per-format outlier tables for specific M ranges, falling back to a small-BN config for GEMVs (m == 1) and a heuristic config table for the remaining cases. Returns DISPATCH_HIT when a matching config is found and launched, or DISPATCH_MISS when no config matches.
Parameters:
- βc_type (
DType): Element type of the output tensorc(inferred). - βa_type (
DType): Element type of the LHS input tensora(inferred). - βb_type (
DType): Element type of the RHS input tensorb(inferred). - βscales_dtype (
DType): Element type of the per-block scale tensorsa_scalesandb_scales(inferred). - βSF_VECTOR_SIZE (
Int): Number of elements each scale factor covers. Must match the format: 16 for NVFP4, 32 for MXFP4, or 32 for MXFP8. - βtranspose_b (
Bool): Whetherbis stored transposed. Must beTrue(defaults toTrue). - βelementwise_lambda_fn (
Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None]): Optional epilogue applied to the matmul resultcin a separate kernel after the matmul completes (defaults toNone). - βelementwise_compute_lambda_fn (
Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> SIMD[dtype, width]]): Optional compute function fused into the matmul kernel epilogue (defaults toNone). - βpdl_level (
PDLLevel): Programmatic Dependent Launch scheduling level for overlapping this kernel with prior GPU work (defaults toPDLLevel()).
Args:
- βc (
TileTensor[c_type, Storage=c.Storage, address_space=c.address_space, linear_idx_type=c.linear_idx_type]): Output TileTensor accumulating the matmul result. - βa (
TileTensor[a_type, Storage=a.Storage, address_space=a.address_space, linear_idx_type=a.linear_idx_type]): LHS input TileTensor. - βb (
TileTensor[b_type, Storage=b.Storage, address_space=b.address_space, linear_idx_type=b.linear_idx_type]): RHS input TileTensor (must be transposed). - βa_scales (
TileTensor[scales_dtype, Storage=a_scales.Storage, address_space=a_scales.address_space, linear_idx_type=a_scales.linear_idx_type]): Per-block scales fora. - βb_scales (
TileTensor[scales_dtype, Storage=b_scales.Storage, address_space=b_scales.address_space, linear_idx_type=b_scales.linear_idx_type]): Per-block scales forb. - βtensor_sf (
Float32): Global tensor scaling factor applied asalpha. - βctx (
DeviceContext): Device context used to launch the kernel.
Returns:
Int: DISPATCH_HIT if a config was selected and the kernel launched, otherwise DISPATCH_MISS.
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!