For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).
Mojo function
naive_fa_decode_apple
def naive_fa_decode_apple[output_type: DType, k_t: MHAOperand, v_t: MHAOperand, mask_t: MHAMask, //, ragged: Bool = False, sink: Bool = False, _use_valid_length: Bool = False, _is_cache_length_accurate: Bool = False](q: LayoutTensor[element_layout=q.element_layout, layout_int_type=q.layout_int_type, linear_idx_type=q.linear_idx_type, masked=q.masked, alignment=q.alignment], k: k_t, v: v_t, mask_functor: mask_t, output: LayoutTensor[output_type, element_layout=output.element_layout, layout_int_type=output.layout_int_type, linear_idx_type=output.linear_idx_type, masked=output.masked, alignment=output.alignment], valid_length: LayoutTensor[DType.uint32, element_layout=valid_length.element_layout, layout_int_type=valid_length.layout_int_type, linear_idx_type=valid_length.linear_idx_type, masked=valid_length.masked, alignment=valid_length.alignment], scale: Float32, batch_size: Int, max_prompt_len: Int, max_cache_size: Int, num_heads: Int, depth: Int, group: Int, ctx: DeviceContext, sink_weights: OptionalReg[LayoutTensor[q.dtype, Layout.row_major(Int(-1)), ImmutAnyOrigin]] = None)
Host launcher for the Apple split-K decode attention pair (decode-only).
Parameters:
- βoutput_type (
DType): The element type of theoutputtensor (inferred). Unused by the kernels; mirrorsmha_gpu_naivefor dispatch uniformity. - βk_t (
MHAOperand): TheMHAOperandtype of the key cache operand (inferred). - βv_t (
MHAOperand): TheMHAOperandtype of the value cache operand (inferred). - βmask_t (
MHAMask): TheMHAMaskfunctor type applied to attention scores (inferred). - βragged (
Bool): Whether sequences are ragged with variable lengths and row offsets invalid_length(defaults toFalse). - βsink (
Bool): Whether attention sink is enabled, pre-seeding split 0 with per-head sink weights (defaults toFalse). - β_use_valid_length (
Bool): Whether to usevalid_lengthfor KVCache decode as per-sequence query lengths (defaults toFalse). - β_is_cache_length_accurate (
Bool): Whether the cache length equals the query length, so no new-token KV is added (defaults toFalse).
Args:
- βq (
LayoutTensor[element_layout=q.element_layout, layout_int_type=q.layout_int_type, linear_idx_type=q.linear_idx_type, masked=q.masked, alignment=q.alignment]): The query tensor; one token per sequence (decode). - βk (
k_t): The key cache operand implementing theMHAOperandcontract. - βv (
v_t): The value cache operand implementing theMHAOperandcontract. - βmask_functor (
mask_t): The mask functor applied to each attention score. - βoutput (
LayoutTensor[output_type, element_layout=output.element_layout, layout_int_type=output.layout_int_type, linear_idx_type=output.linear_idx_type, masked=output.masked, alignment=output.alignment]): The output tensor; written by the stitch kernel with the normalized attention output. - βvalid_length (
LayoutTensor[DType.uint32, element_layout=valid_length.element_layout, layout_int_type=valid_length.layout_int_type, linear_idx_type=valid_length.linear_idx_type, masked=valid_length.masked, alignment=valid_length.alignment]): Per-sequence row offsets or query lengths (uint32); meaning depends onraggedand_use_valid_length. - βscale (
Float32): The softmax scale factor applied toQ.K^Tscores. - βbatch_size (
Int): Number of sequences in the batch. - βmax_prompt_len (
Int): Maximum prompt length; the dense decode path's query length. - βmax_cache_size (
Int): Full key count for the dense decode path (the K tensor's seq dim). - βnum_heads (
Int): Number of query attention heads. - βdepth (
Int): Head dimension; must be a multiple ofWARP_SIZEand at mostNAIVE_FA_DECODE_APPLE_MAX_HEAD_DIM. - βgroup (
Int): Number of query heads per KV head (GQA group size). - βctx (
DeviceContext): The device context used to enqueue kernels and allocate partial buffers. - βsink_weights (
OptionalReg[LayoutTensor[q.dtype, Layout.row_major(Int(-1)), ImmutAnyOrigin]]): Per-head sink weights (shape[num_heads]); read only whensinkisTrue(defaults toNone).
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!