IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

Struct_grouped_matmul_block_scaled_mxfp4

struct Struct_grouped_matmul_block_scaled_mxfp4[preshuffled_b: Bool = False, lane_bytes: Int = Int(16)]

MOGG wrapper for grouped block-scaled matrix multiplication.

Provides graph compiler integration for block-scaled grouped matmul operations used in Mixture of Experts (MoE) layers on AMD GPUs.

Parameters​

  • ​preshuffled_b (Bool): When True, dispatches to mxfp4_grouped_matmul_amd_preb which expects B in the 5D preshuffled layout from Shuffler.preshuffle_b_5d (typically produced by the model's weight adapter at load time, e.g. Kimi K2.5). When False (default), dispatches to the dense mxfp4_grouped_matmul_amd kernel that reads B row-major. The caller is responsible for preparing B in the matching layout.
  • ​lane_bytes (Int): Element packing of A and B β€” 16 for MXFP4 (default) or 32 for MXFP8. The kernel reads a/b as raw bytes, so this rather than the operand dtype selects the format, and with it the K extent (K at MXFP8, K // 2 at MXFP4). Preshuffled-B path only.

Implemented traits​

AnyType, Deinitable, Movable

Methods​

execute​

static def execute[c_type: DType, a_type: DType, b_type: DType, //, target: StringSpan[ImmStaticOrigin]](c: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=c.static_spec], a: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=a.static_spec], b: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=b.static_spec], a_scales: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=a_scales.static_spec], b_scales: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=b_scales.static_spec], expert_start_indices: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=expert_start_indices.static_spec], expert_ids: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=expert_ids.static_spec], max_num_tokens_per_expert: UInt32, num_active_experts: UInt32, estimated_total_m: UInt32, decode_grid_m_cap: UInt32, context: DeviceContext)

Executes grouped block-scaled matrix multiplication.

Computes C = A @ B^T for multiple expert groups where A and B are block-scaled (e.g. MXFP4: 4-bit floating point packed as uint8).

Parameters:

Args: