IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

DecodeKVConsumer

struct DecodeKVConsumer[dtype: DType, config: MLA_SM100_Decode_Config, num_producer: Int = Int(1), num_consumer: Int = Int(2)]

Consumer side of the decode KV pipeline that waits for and releases KV stages.

Parameters​

  • ​dtype (DType): Element type of the KV tiles stored in SMEM.
  • ​config (MLA_SM100_Decode_Config): Decode config supplying num_kv_stages, BN_QK, and q_depth for pipeline stage sizing.
  • ​num_producer (Int): Number of producer threads arriving on each producer mbarrier (defaults to 1).
  • ​num_consumer (Int): Number of consumer threads arriving on each consumer mbarrier (defaults to 2, matching the standard mmaQK+mmaPV dual-consumer KV pipeline).

Fields​

  • ​pipe (DecodeKVConsumer[dtype, config, num_producer, num_consumer].KVPipeType):
  • ​smem (Pointer[Scalar[dtype], MutAnyOrigin, address_space=AddressSpace.SHARED]):

Implemented traits​

AnyType, Copyable, ImplicitlyCopyable, ImplicitlyDeletable, Movable, RegisterPassable, TrivialRegisterPassable

comptime members​

kv_stage_elems​

comptime kv_stage_elems = (config * config)

KVPipeType​

comptime KVPipeType = KVPipelineGeneric[config.num_kv_stages, Int(1), num_producer, num_consumer]

Methods​

__init__​

def __init__(pipe: KVPipelineGeneric[config.num_kv_stages, Int(1), num_producer, num_consumer], smem: Pointer[Scalar[dtype], MutAnyOrigin, address_space=AddressSpace.SHARED]) -> Self

stage_base_ptr​

def stage_base_ptr[*, qk_stage: Int = Int(0)](self) -> Pointer[Scalar[dtype], MutAnyOrigin, address_space=AddressSpace.SHARED]

Returns:

Pointer[Scalar[dtype], MutAnyOrigin, address_space=AddressSpace.SHARED]

stage_index​

def stage_index[*, qk_stage: Int = Int(0)](self) -> UInt32

Returns:

UInt32

wait​

def wait[*, qk_stage: Int = Int(0)](self)

release​

def release[*, qk_stage: Int = Int(0)](mut self, e: Int32)

release_all​

def release_all(mut self)

Explicit-arrive release for a non-MMA (independent-thread) consumer.

Every thread of the num_consumer-wide consumer role (e.g. a full warpgroup doing a manual SMEM re-swizzle read) calls this once; the num_consumer independent arrive() calls satisfy the mbar's expected count. Use this instead of release[qk_stage](e) (which uses elect_mma_arrive for a single elected thread per warp) when every thread independently participates, not just one MMA-eligible lane per warp.