For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).

Python module

max.pipelines.architectures.eagle3_deepseekV3

DeepseekV3 + Eagle3 speculator pipeline.

`Eagle3DeepseekV3`

class max.pipelines.architectures.eagle3_deepseekV3.Eagle3DeepseekV3(config)

source

Bases: Module

Eagle3 draft model paired with a DeepseekV3 target.

Parameters:: config (DeepseekV3Config)

`input_types()`

input_types(kv_params)

source

Parameters:: kv_params (KVCacheParamInterface)
Return type:: tuple[TensorType | BufferType, …]

`Eagle3DeepseekV3Inputs`

class max.pipelines.architectures.eagle3_deepseekV3.Eagle3DeepseekV3Inputs(tokens, input_row_offsets, signal_buffers, host_input_row_offsets, batch_context_lengths, draft_tokens=None, draft_kv_blocks=None, seed=None, temperature=None, top_k=None, max_k=None, top_p=None, min_top_p=None, in_thinking_phase=None, pinned_bitmask=None, wait_payload=None, device_bitmask_scratch=None, *, kv_cache_inputs=None, lora=None, hidden_states=None, return_n_logits, data_parallel_splits, ep_inputs=())

source

Bases: DeepseekV3Inputs

Inputs for the Eagle3 + DeepseekV3 unified model.

Parameters:

tokens (Buffer)
input_row_offsets (Buffer)
signal_buffers (list[Buffer])
host_input_row_offsets (Buffer)
batch_context_lengths (list[Buffer])
draft_tokens (Buffer | None)
draft_kv_blocks (list[Buffer] | None)
seed (Buffer | None)
temperature (Buffer | None)
top_k (Buffer | None)
max_k (Buffer | None)
top_p (Buffer | None)
min_top_p (Buffer | None)
in_thinking_phase (Buffer | None)
pinned_bitmask (Buffer | None)
wait_payload (Buffer | None)
device_bitmask_scratch (Buffer | None)
kv_cache_inputs (KVCacheInputs[Buffer, Buffer] | None)
lora (LoRAInputs | None)
hidden_states (Buffer | list[Buffer] | None)
return_n_logits (Buffer)
data_parallel_splits (Buffer)
ep_inputs (tuple[Buffer, ...])

`buffers`

property buffers: tuple[Buffer, ...]

source

Returns positional Buffer inputs for model ABI calls.

`device_bitmask_scratch`

device_bitmask_scratch: Buffer | None = None

source

Device scratch buffer that receives the in-graph H2D from pinned_bitmask; the acceptance sampler reads from it. Only set when structured output is enabled.

`draft_kv_blocks`

draft_kv_blocks: list[Buffer] | None = None

source

`draft_tokens`

draft_tokens: Buffer | None = None

source

`in_thinking_phase`

in_thinking_phase: Buffer | None = None

source

Per-batch bool flag set by the pipeline for relaxed acceptance during thinking. Not consumed by the eagle3_deepseekV3 graph today, but the field is required to satisfy the _UnifiedSpecDecodeInputs protocol used by OverlapTextGenerationPipeline.

`max_k`

max_k: Buffer | None = None

source

`min_top_p`

min_top_p: Buffer | None = None

source

Per-batch sampling parameters consumed by the stochastic acceptance sampler. max_k and min_top_p are 0-d CPU scalars; the rest are [batch_size] tensors on the primary device.

`pinned_bitmask`

pinned_bitmask: Buffer | None = None

source

Pinned host bitmask for constrained decoding.

Shape [batch_size, num_speculative_tokens + 1, vocab_size]. Position i contains the valid-token mask given the FSM state after consuming draft[0:i-1]; position num_speculative_tokens is for the bonus token. None when structured output is disabled.

`seed`

seed: Buffer | None = None

source

Per-execute uint64 [1] seed consumed by the stochastic acceptance sampler (and, when enabled, the synthetic benchmarking sampler).

`temperature`

temperature: Buffer | None = None

source

`top_k`

top_k: Buffer | None = None

source

`top_p`

top_p: Buffer | None = None

source

`wait_payload`

wait_payload: Buffer | None = None

source

CPU int64[2] payload = [flag._unsafe_ptr, 1] consumed by the in-graph mo.wait_host_value_with_dep op. Only set when structured output is enabled.

`Eagle3DeepseekV3Model`

class max.pipelines.architectures.eagle3_deepseekV3.Eagle3DeepseekV3Model(pipeline_config, session, devices, kv_cache_config, weights, adapter=None, return_logits=ReturnLogits.VARIABLE, return_hidden_states=ReturnHiddenStates.SELECTED_LAYERS)

source

Bases: DeepseekV3Model

Eagle3 + DeepseekV3: target + draft in one compiled graph.

Loads target weights from a DeepseekV3-shaped main checkpoint and draft weights from a separate Eagle3 checkpoint (pipeline_config.draft_model).

Parameters:

pipeline_config (PipelineConfig)
session (InferenceSession)
devices (list[Device])
kv_cache_config (KVCacheConfig)
weights (Weights)
adapter (WeightsAdapter | None)
return_logits (ReturnLogits)
return_hidden_states (ReturnHiddenStates)

`execute()`

execute(model_inputs)

source

Execute and return all graph outputs for speculative decoding.

Parameters:: model_inputs (ModelInputs)
Return type:: UnifiedEagleOutputs

`load_model()`

load_model(session)

source

Load the model with the given weights.

Parameters:: session (InferenceSession)
Return type:: Model

`prepare_initial_token_inputs()`

prepare_initial_token_inputs(replica_batches, kv_cache_inputs=None, return_n_logits=1, draft_tokens=None, draft_kv_cache_buffers=None, **kwargs)

source

Prepares the initial inputs to be passed to execute().

The inputs and functionality can vary per model. For example, model inputs could include encoded tensors, unique IDs per tensor when using a KV cache manager, and kv_cache_inputs (or None if the model does not use KV cache). This method typically batches encoded tensors, claims a KV cache slot if needed, and returns the inputs and caches.

Parameters:

replica_batches (Sequence[Sequence[TextContext]])
kv_cache_inputs (KVCacheInputs[Buffer, Buffer] | None)
return_n_logits (int)
draft_tokens (Buffer | None)
draft_kv_cache_buffers (list[Buffer] | None)

Return type:

Eagle3DeepseekV3Inputs

`Eagle3DeepseekV3Unified`

class max.pipelines.architectures.eagle3_deepseekV3.Eagle3DeepseekV3Unified(config, draft_config=None, speculative_config=None, enable_structured_output=False)

source

Bases: Module

Fused nn.Module: merge + target forward + greedy rejection + shift.

The target model returns concatenated hidden states from 3 intermediate layers (first, middle, last). The draft model fuses these via fc and generates the next speculative token.

Parameters:

config (DeepseekV3Config)
draft_config (DeepseekV3Config | None)
speculative_config (SpeculativeConfig | None)
enable_structured_output (bool)

`input_types()`

input_types(kv_params, draft_kv_params=None)

source

Input types for the Eagle3 unified graph.

Order: tokens, device_offsets, host_offsets, return_n_logits,: data_parallel_splits, signal_buffers, target_kv_cache, batch_context_lengths, target_ep_inputs, draft_tokens, draft_kv_blocks_per_device, seed, temperature, top_k, max_k, top_p, min_top_p[, pinned_bitmask, wait_payload, device_bitmask_scratch].

The bitmask triple is appended last when enable_structured_output is True.

Parameters:

kv_params (KVCacheParamInterface)
draft_kv_params (KVCacheParams | None)

Return type:

tuple[TensorType | BufferType, …]

`Eagle3MHADeepseekV3Inputs`

class max.pipelines.architectures.eagle3_deepseekV3.Eagle3MHADeepseekV3Inputs(tokens, input_row_offsets, signal_buffers, host_input_row_offsets, batch_context_lengths, draft_tokens=None, draft_kv_blocks=None, seed=None, temperature=None, top_k=None, max_k=None, top_p=None, min_top_p=None, in_thinking_phase=None, pinned_bitmask=None, wait_payload=None, device_bitmask_scratch=None, *, kv_cache_inputs=None, lora=None, hidden_states=None, return_n_logits, data_parallel_splits, ep_inputs=())

source

Bases: DeepseekV3Inputs

Inputs for the Eagle3 MHA-draft + DeepseekV3 unified model.

The draft owns its own KVCacheInputs (separate from the target’s MLA cache) so MHA dispatch metadata can be plumbed at both prefill (q_max_seq_len = 1 + num_speculative_tokens) and decode (q_max_seq_len = 1) widths.

Parameters:

tokens (Buffer)
input_row_offsets (Buffer)
signal_buffers (list[Buffer])
host_input_row_offsets (Buffer)
batch_context_lengths (list[Buffer])
draft_tokens (Buffer | None)
draft_kv_blocks (list[Buffer] | None)
seed (Buffer | None)
temperature (Buffer | None)
top_k (Buffer | None)
max_k (Buffer | None)
top_p (Buffer | None)
min_top_p (Buffer | None)
in_thinking_phase (Buffer | None)
pinned_bitmask (Buffer | None)
wait_payload (Buffer | None)
device_bitmask_scratch (Buffer | None)
kv_cache_inputs (KVCacheInputs[Buffer, Buffer] | None)
lora (LoRAInputs | None)
hidden_states (Buffer | list[Buffer] | None)
return_n_logits (Buffer)
data_parallel_splits (Buffer)
ep_inputs (tuple[Buffer, ...])

`buffers`

property buffers: tuple[Buffer, ...]

source

Returns positional Buffer inputs for model ABI calls.

`device_bitmask_scratch`

device_bitmask_scratch: Buffer | None = None

source

Device scratch buffer that receives the in-graph H2D from pinned_bitmask (None when off).

`draft_kv_blocks`

draft_kv_blocks: list[Buffer] | None = None

source

One persistent kv_blocks Buffer per device. Other fields (cache_lengths, lookup_table, max_lengths, attention_dispatch_metadata) are borrowed from the target’s kv_cache_inputs at graph build time.

`draft_tokens`

draft_tokens: Buffer | None = None

source

`in_thinking_phase`

in_thinking_phase: Buffer | None = None

source

Per-batch bool flag for relaxed-acceptance gating. Not consumed by the graph today but required by the _UnifiedEagleInputs protocol used by OverlapTextGenerationPipeline.

`max_k`

max_k: Buffer | None = None

source

`min_top_p`

min_top_p: Buffer | None = None

source

`pinned_bitmask`

pinned_bitmask: Buffer | None = None

source

Pinned host bitmask for constrained decoding (None when off).

`seed`

seed: Buffer | None = None

source

`temperature`

temperature: Buffer | None = None

source

`top_k`

top_k: Buffer | None = None

source

`top_p`

top_p: Buffer | None = None

source

`wait_payload`

wait_payload: Buffer | None = None

source

CPU int64[2] payload consumed by mo.wait_host_value_with_dep (None when off).

`Eagle3MHADeepseekV3Model`

class max.pipelines.architectures.eagle3_deepseekV3.Eagle3MHADeepseekV3Model(pipeline_config, session, devices, kv_cache_config, weights, adapter=None, return_logits=ReturnLogits.VARIABLE, return_hidden_states=ReturnHiddenStates.SELECTED_LAYERS)

source

Bases: DeepseekV3Model

Eagle3 MHA-draft + DeepseekV3: target + draft in one compiled graph.

Parameters:

pipeline_config (PipelineConfig)
session (InferenceSession)
devices (list[Device])
kv_cache_config (KVCacheConfig)
weights (Weights)
adapter (WeightsAdapter | None)
return_logits (ReturnLogits)
return_hidden_states (ReturnHiddenStates)

`execute()`

execute(model_inputs)

source

Executes the graph with the given inputs.

Parameters:: model_inputs (ModelInputs) – The model inputs to execute, containing tensors and any other required data for model execution.
Returns:: ModelOutputs containing the pipeline’s output tensors.
Return type:: UnifiedEagleOutputs

This is an abstract method that must be implemented by concrete PipelineModels to define their specific execution logic.

`load_model()`

load_model(session)

source

Load the model with the given weights.

Parameters:: session (InferenceSession)
Return type:: Model

`prepare_initial_token_inputs()`

prepare_initial_token_inputs(replica_batches, kv_cache_inputs=None, return_n_logits=1, draft_tokens=None, draft_kv_cache_buffers=None, **kwargs)

source

Prepares the initial inputs to be passed to execute().

Parameters:

replica_batches (Sequence[Sequence[TextContext]])
kv_cache_inputs (KVCacheInputs[Buffer, Buffer] | None)
return_n_logits (int)
draft_tokens (Buffer | None)
draft_kv_cache_buffers (list[Buffer] | None)

Return type:

Eagle3MHADeepseekV3Inputs

`convert_eagle3_draft_state_dict()`

max.pipelines.architectures.eagle3_deepseekV3.convert_eagle3_draft_state_dict(state_dict, huggingface_config=None, pipeline_config=None)

source

Convert an Eagle3 draft checkpoint to Eagle3DeepseekV3 module keys.

The Eagle3 checkpoint (nvidia/Kimi-K2.5-Thinking-Eagle3) has keys: fc.*, layers.0.*, norm.*, lm_head.*.

All weights are loaded; norm and lm_head are kept independent from the target model. Only embed_tokens is shared from the target.

Parameters:

state_dict (dict[str, Weights]) – Raw Eagle3 checkpoint.
huggingface_config (PreTrainedConfig | None)
pipeline_config (PipelineConfig | None)

Returns:

State dict with keys matching Eagle3DeepseekV3 module hierarchy.

Return type:

dict[str, WeightData]

Eagle3DeepseekV3​

input_types()​

Eagle3DeepseekV3Inputs​

buffers​

device_bitmask_scratch​

draft_kv_blocks​

draft_tokens​

in_thinking_phase​

max_k​

min_top_p​

pinned_bitmask​

seed​

temperature​

top_k​

top_p​

wait_payload​

Eagle3DeepseekV3Model​

execute()​

load_model()​

prepare_initial_token_inputs()​

Eagle3DeepseekV3Unified​

input_types()​

Eagle3MHADeepseekV3Inputs​

buffers​

device_bitmask_scratch​

draft_kv_blocks​

draft_tokens​

in_thinking_phase​

max_k​

min_top_p​

pinned_bitmask​

seed​

temperature​

top_k​

top_p​

wait_payload​

Eagle3MHADeepseekV3Model​

execute()​

load_model()​

prepare_initial_token_inputs()​

convert_eagle3_draft_state_dict()​

`Eagle3DeepseekV3`

`input_types()`

`Eagle3DeepseekV3Inputs`

`buffers`

`device_bitmask_scratch`

`draft_kv_blocks`

`draft_tokens`

`in_thinking_phase`

`max_k`

`min_top_p`

`pinned_bitmask`

`seed`

`temperature`

`top_k`

`top_p`

`wait_payload`

`Eagle3DeepseekV3Model`

`execute()`

`load_model()`

`prepare_initial_token_inputs()`

`Eagle3DeepseekV3Unified`

`input_types()`

`Eagle3MHADeepseekV3Inputs`

`buffers`

`device_bitmask_scratch`

`draft_kv_blocks`

`draft_tokens`

`in_thinking_phase`

`max_k`

`min_top_p`

`pinned_bitmask`

`seed`

`temperature`

`top_k`

`top_p`

`wait_payload`

`Eagle3MHADeepseekV3Model`

`execute()`

`load_model()`

`prepare_initial_token_inputs()`

`convert_eagle3_draft_state_dict()`