For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).
Python class
MoEQuantized
MoEQuantizedβ
class max.nn.MoEQuantized(devices, hidden_dim, num_experts, num_experts_per_token, moe_dim, num_logical_experts=None, gate_cls=<class 'max.nn.moe.moe.MoEGate'>, mlp_cls=<class 'max.nn.linear.MLP'>, shared_mlp_cls=None, has_shared_experts=False, shared_experts_dim=0, ep_size=1, dtype=bfloat16, apply_router_weight_first=False, use_swigluoai=False, swiglu_alpha=0.0, swiglu_limit=0.0, gated_activation_fn=None, pre_expert_norm_cls=None, ep_batch_manager=None, quant_config=None, shared_experts_dtype=None, shared_experts_quant_config=None, is_sharding=False)
Bases: MoE
Mixture of Experts with FP8 or NVFP4 quantization.
-
Parameters:
-
- devices (list[DeviceRef])
- hidden_dim (int)
- num_experts (int)
- num_experts_per_token (int)
- moe_dim (int)
- num_logical_experts (int | None)
- gate_cls (Callable[..., MoEGate])
- mlp_cls (Callable[..., MLP])
- shared_mlp_cls (Callable[..., MLP] | None)
- has_shared_experts (bool)
- shared_experts_dim (int)
- ep_size (int)
- dtype (DType)
- apply_router_weight_first (bool)
- use_swigluoai (bool)
- swiglu_alpha (float)
- swiglu_limit (float)
- gated_activation_fn (Callable[[TensorValue, int], TensorValue] | None)
- pre_expert_norm_cls (Callable[[], Module] | None)
- ep_batch_manager (EPBatchManager | None)
- quant_config (QuantConfig | None)
- shared_experts_dtype (DType | None)
- shared_experts_quant_config (QuantConfig | None)
- is_sharding (bool)
configure_ep_scale_fusion()β
configure_ep_scale_fusion(dispatch_supports_fold)
Enable the MXFP4 EP A-scale preshuffle fold on the shared EP config so the dispatch ops emit slot-sized scales. Must run BEFORE the dispatch op (the dispatch output shape depends on this flag); the EP forward driver calls it once per layer before dispatch.
The fold writes the up-proj A-scale during dispatch and the down-proj
A-scale during fused_silu directly into the grouped-matmul slot
layout, dropping the standalone preshuffle kernels from the decode
critical path.
-
Parameters:
-
dispatch_supports_fold (bool) β Whether the selected dispatch path threads the fold parameters.
-
Return type:
-
None
down_proj_scalesβ
property down_proj_scales: TensorValue
Returns stacked down-projection weight scales.
gate_up_proj_scalesβ
property gate_up_proj_scales: TensorValue
Returns stacked gate/up weight scales for grouped matmul.
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!