Mojo module
mla
Functionsβ
- β
copy_fn_unified: - β
flare_mla_decoding: MLA decoding kernel that would only be called in the optimized compute graph. - β
flare_mla_decoding_dispatch: - β
flare_mla_prefill: MLA prefill kernel that would only be called in the optimized compute graph. Only supports ragged Q/K/V inputs. - β
flare_mla_prefill_dispatch: - β
mla_decoding: - β
mla_decoding_single_batch: Flash attention v2 algorithm. - β
mla_prefill: - β
mla_prefill_plan: This calls a GPU kernel that plans how to process a batch of sequences with varying lengths using a fixed-size buffer. - β
mla_prefill_plan_kernel: - β
mla_prefill_single_batch: MLA for encoding where seqlen > 1. - β
q_block_idx: - β
set_buffer_lengths_to_zero:
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!