For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).
Mojo module
mla_decode_kv_fp8
Implements the SM100 MLA decode kernel variant that loads KV cache in FP8 and converts to BF16 in shared memory before MMA.
Structs
-
MLA_SM100_Decode_KV_FP8: FP8 KV decode kernel for MLA attention on SM100 GPUs.
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!