For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).
Python module
max.pipelines.architectures.bert
BERT sentence transformer architecture for embeddings generation.
BertInputsβ
class max.pipelines.architectures.bert.BertInputs(next_tokens_batch: 'Buffer', attention_mask: 'Buffer', *, kv_cache_inputs: 'KVCacheInputsInterface[Buffer, Buffer] | None' = None, lora: 'LoRAInputs | None' = None, vision_embeddings: 'list[Buffer]' = <factory>, vision_scatter_indices: 'list[Buffer]' = <factory>, hidden_states: 'Buffer | list[Buffer] | None' = None)
Bases: ModelInputs
-
Parameters:
attention_maskβ
attention_mask: Buffer
next_tokens_batchβ
next_tokens_batch: Buffer
BertModelConfigβ
class max.pipelines.architectures.bert.BertModelConfig(*, dtype, device, pool_embeddings, huggingface_config, max_seq_len)
Bases: ArchConfigWithBoundedMaxSeqLen, ArchConfig
Configuration for Bert models.
-
Parameters:
deviceβ
device: DeviceRef
dtypeβ
dtype: DType
huggingface_configβ
huggingface_config: AutoConfig
initialize()β
classmethod initialize(pipeline_config, model_config=None)
Initializes a BertModelConfig instance from pipeline configuration.
-
Parameters:
-
- pipeline_config (PipelineConfig) β The MAX Engine pipeline configuration.
- model_config (MAXModelConfig | None)
-
Returns:
-
An initialized BertModelConfig instance.
-
Return type:
max_seq_lenβ
max_seq_len: int
pool_embeddingsβ
pool_embeddings: bool
BertPipelineModelβ
class max.pipelines.architectures.bert.BertPipelineModel(pipeline_config, session, devices, kv_cache_config, weights, adapter=None, return_logits=ReturnLogits.ALL, max_batch_size=1)
Bases: GraphPipelineModel[TextContext]
-
Parameters:
-
- pipeline_config (PipelineConfig)
- session (InferenceSession)
- devices (list[Device])
- kv_cache_config (KVCacheConfig)
- weights (Weights)
- adapter (WeightsAdapter | None)
- return_logits (ReturnLogits)
- max_batch_size (int)
batch_processor_clsβ
batch_processor_cls
alias of BertBatchProcessor
execute()β
execute(model_inputs)
Executes the graph with the given inputs.
-
Parameters:
-
model_inputs (ModelInputs) β The model inputs to execute, containing tensors and any other required data for model execution.
-
Returns:
-
ModelOutputs containing the pipelineβs output tensors.
-
Return type:
This is an abstract method that must be implemented by concrete PipelineModels to define their specific execution logic.
modelβ
model: Model
model_config_clsβ
model_config_cls
alias of BertModelConfig
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!