For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).
Mojo module
mla_decode_sparse_kv_bf16
SM100 (B200) sparse MLA decode kernel with BF16 KV cache.
K is loaded by a BF16 + SWIZZLE_128B gather4 TMA covering the full
576-element row (tile_width=576, box_w=64). OffsetPosition[sparse=True]
overrides num_keys with the sparse topk. A dedicated 4-warp gather WG
(warps 8-11) decodes each tile's row indices cooperatively into the
double-buffered idx_smem, then round-robins the tile's 144 gather4
issues across its 4 elected lanes under a single expect_bytes.
Supports NullMask and CausalMask. Sliding-window attention is FP8-only.
Structs
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!