For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).
Mojo function
matmul_Q6_K
def matmul_Q6_K[elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None](a_tt: TileTensor[DType.float32, Storage=a_tt.Storage, linear_idx_type=a_tt.linear_idx_type], b_tt: TileTensor[DType.uint8, Storage=b_tt.Storage, linear_idx_type=b_tt.linear_idx_type], c_tt: TileTensor[DType.float32, Storage=c_tt.Storage, linear_idx_type=c_tt.linear_idx_type], ctx: Optional[DeviceContext] = None)
Computes a matrix multiplication with Q6_K block-quantized weights.
Dispatches to an x86 or ARM NEON implementation at compile time; other targets fail to compile.
Parameters:
- βelementwise_lambda_fn (
Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None]): Optional epilogue applied to each output element.
Args:
- βa_tt (
TileTensor[DType.float32, Storage=a_tt.Storage, linear_idx_type=a_tt.linear_idx_type]): Left-hand operand tensor in float32. - βb_tt (
TileTensor[DType.uint8, Storage=b_tt.Storage, linear_idx_type=b_tt.linear_idx_type]): Right-hand operand tensor holding Q6_K quantized uint8 weights. - βc_tt (
TileTensor[DType.float32, Storage=c_tt.Storage, linear_idx_type=c_tt.linear_idx_type]): Output tensor in float32. - βctx (
Optional[DeviceContext]): Optional device context for parallel execution.
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!