For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
compute_log_probabilities_1tok
def compute_log_probabilities_1tok[target: StringSlice[ImmStaticOrigin], levels: Int](output_token_index: Int, lp_logits: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=lp_logits.static_spec], lp_tokens: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=lp_tokens.static_spec], logits: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=logits.static_spec], tokens: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=tokens.static_spec], sampled_tokens: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=sampled_tokens.static_spec], logit_row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=logit_row_offsets.static_spec], token_row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=token_row_offsets.static_spec], lp_output_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=lp_output_offsets.static_spec])
Computes log probabilities for a single token position from its logits row.
Parameters:
- βtarget (
StringSlice[ImmStaticOrigin]): Compilation target string, such as"cpu"or"gpu". - βlevels (
Int): Number of heap levels controlling the top-k width; the output stores2**levelsentries per token. Whenlevels <= 0, only the sampled token's log probability is emitted.
Args:
- βoutput_token_index (
Int): Global row index intolp_logitsandlp_tokensfor this output position. - βlp_logits (
ManagedTensorSlice[IOSpec[_, _].Output, static_spec=lp_logits.static_spec]): Output tensor of shape[num_output_tokens, 2**levels]receiving the log probabilities. - βlp_tokens (
ManagedTensorSlice[IOSpec[_, _].Output, static_spec=lp_tokens.static_spec]): Output tensor of shape[num_output_tokens, 2**levels]receiving the token ids paired withlp_logits. - βlogits (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=logits.static_spec]): Input logits ragged by batch, shape[total_rows, vocab_size]. - βtokens (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=tokens.static_spec]): Previously generated tokens across all batches, ragged. - βsampled_tokens (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=sampled_tokens.static_spec]): Most recently sampled token for each batch, used when this position is the last token in its sequence. - βlogit_row_offsets (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=logit_row_offsets.static_spec]): Per-batch start offsets into the first axis oflogits. - βtoken_row_offsets (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=token_row_offsets.static_spec]): Per-batch start offsets intotokens. - βlp_output_offsets (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=lp_output_offsets.static_spec]): Per-batch start offsets into the output row axis, mapping each output index to its batch.