pub struct Tokenizer {
pub bos_token_id: Option<u32>,
pub eos_token_id: Option<u32>,
pub pad_token_id: Option<u32>,
pub im_start_id: Option<u32>,
pub im_end_id: Option<u32>,
pub chat_template: Option<String>,
pub extra_eos: HashSet<u32>,
pub add_bos: bool,
/* private fields */
}Expand description
A loaded BPE tokenizer.
Fields§
§bos_token_id: Option<u32>Special tokens
eos_token_id: Option<u32>§pad_token_id: Option<u32>§im_start_id: Option<u32>Chat template special tokens
im_end_id: Option<u32>§chat_template: Option<String>Jinja chat template carried by the container (spec §6.1); None → hardcoded ChatML fallback.
extra_eos: HashSet<u32>Extra stop ids from the container’s generation config.
add_bos: boolGeneration prepends BOS (llama post_processor semantics).
Implementations§
Source§impl Tokenizer
impl Tokenizer
Sourcepub fn from_file(path: impl AsRef<Path>) -> Result<Self, TokenizerError>
pub fn from_file(path: impl AsRef<Path>) -> Result<Self, TokenizerError>
Load tokenizer from HuggingFace tokenizer.json file.
Sourcepub fn from_bytes(bytes: &[u8]) -> Result<Self, TokenizerError>
pub fn from_bytes(bytes: &[u8]) -> Result<Self, TokenizerError>
Load tokenizer from raw tokenizer.json bytes (CMF VOCAB section).
Sourcepub fn from_json(json: &str) -> Result<Self, TokenizerError>
pub fn from_json(json: &str) -> Result<Self, TokenizerError>
Load tokenizer from JSON string.
Sourcepub fn byte_level() -> Self
pub fn byte_level() -> Self
Create a minimal tokenizer for testing (byte tokens, no merges).
Sourcepub fn encode_plain(&self, text: &str) -> Vec<u32>
pub fn encode_plain(&self, text: &str) -> Vec<u32>
Encode text as PLAIN text: the whole input is one added-token-free
segment (NFC → split → byte-map → BPE), so a literal <|im_end|>
in it stays bytes instead of becoming the special id. This is the
trainer’s Bpe::encode (user text, never the template frame) —
router v2 tokenizes the user message with it (spec §9.4).
Sourcepub fn decode(&self, ids: &[u32]) -> String
pub fn decode(&self, ids: &[u32]) -> String
Decode token IDs back to text. Special tokens are skipped; added tokens are raw text; everything else reverses the byte-level map.
Sourcepub fn decode_token(&self, id: u32) -> String
pub fn decode_token(&self, id: u32) -> String
Streaming decode of ONE token: no sequence-level Strip — a per-token strip would eat the ▁-spaces of every SP word.
Sourcepub fn decode_token_for_hash(&self, id: u32) -> String
pub fn decode_token_for_hash(&self, id: u32) -> String
Decode one vocabulary entry for Engram’s compressed token map while retaining special tokens.
Sourcepub fn decode_for_protocol(&self, ids: &[u32]) -> String
pub fn decode_for_protocol(&self, ids: &[u32]) -> String
Decode generated protocol text while retaining special markers. The
V4.1 harmony parser needs <think>, EOS, and spaced DSML tags.
Sourcepub fn raw_token_for_hash(&self, id: u32) -> String
pub fn raw_token_for_hash(&self, id: u32) -> String
Return the backend vocabulary spelling for an Engram map entry.
Sourcepub fn apply_chat_template(&self, messages: &[(String, String)]) -> Vec<u32>
pub fn apply_chat_template(&self, messages: &[(String, String)]) -> Vec<u32>
Render the container’s Jinja chat template (HF semantics: trim_blocks + lstrip_blocks + loop controls) and encode it. Falls back to hardcoded ChatML when the file carries none.
Sourcepub fn apply_chat_template_json(
&self,
messages: &[Value],
tools: Option<&[Value]>,
enable_thinking: Option<bool>,
) -> Vec<u32>
pub fn apply_chat_template_json( &self, messages: &[Value], tools: Option<&[Value]>, enable_thinking: Option<bool>, ) -> Vec<u32>
Like apply_chat_template, with an explicit enable_thinking value for
reasoning-model templates (Qwen3/3.5 emit an empty None leaves the variable
undefined — the template’s own default applies.
Chat template with the FULL message shape and a tool list.
The pair-based API below flattens every message to (role, text),
which silently drops exactly what agentic use needs: the tools
array, role: "tool" results, and tool_calls on assistant
turns. The templates this format embeds — Qwen-family, Nanbeige —
have carried a {%- if tools %} branch all along; this is the
call that finally feeds it. Messages arrive as JSON objects in
the OpenAI shape and pass through to minijinja unflattened, so a
template sees the same fields a Python apply_chat_template
would.
Sourcepub fn try_apply_chat_template_json(
&self,
messages: &[Value],
tools: Option<&[Value]>,
enable_thinking: Option<bool>,
) -> Result<Vec<u32>, String>
pub fn try_apply_chat_template_json( &self, messages: &[Value], tools: Option<&[Value]>, enable_thinking: Option<bool>, ) -> Result<Vec<u32>, String>
Like Self::apply_chat_template_json, but a template that fails
to render is an ERROR instead of a quiet ChatML approximation.
The fallback flattens every message to (role, text): it has no
place for tools, tool_calls or role: "tool". For a plain chat
that is a tolerable degradation; for a request with tools it means
the model never sees the functions and answers as if none were
offered — a failure no client can detect. The server calls this
variant when tools are present and reports the error instead.
Files without a template still take the ChatML path (Ok).
Sourcepub fn render_chat_json(
&self,
messages: &[Value],
tools: Option<&[Value]>,
enable_thinking: Option<bool>,
) -> Option<String>
pub fn render_chat_json( &self, messages: &[Value], tools: Option<&[Value]>, enable_thinking: Option<bool>, ) -> Option<String>
Render the template against JSON-shaped messages (parity surface).
pub fn apply_chat_template_opts( &self, messages: &[(String, String)], enable_thinking: Option<bool>, ) -> Vec<u32>
Sourcepub fn with_bos(&self, ids: Vec<u32>) -> Vec<u32>
pub fn with_bos(&self, ids: Vec<u32>) -> Vec<u32>
Prepend BOS when the tokenizer declares it (llama family).
Sourcepub fn render_chat(&self, messages: &[(String, String)]) -> Option<String>
pub fn render_chat(&self, messages: &[(String, String)]) -> Option<String>
Render the carried template to text (parity-testable surface).
Sourcepub fn render_chat_opts(
&self,
messages: &[(String, String)],
enable_thinking: Option<bool>,
) -> Option<String>
pub fn render_chat_opts( &self, messages: &[(String, String)], enable_thinking: Option<bool>, ) -> Option<String>
Render the carried template to text with explicit thinking mode.
Sourcepub fn vocab_size(&self) -> usize
pub fn vocab_size(&self) -> usize
Vocabulary size.
Sourcepub fn token_to_id(&self, token: &str) -> Option<u32>
pub fn token_to_id(&self, token: &str) -> Option<u32>
Return the ID for an exact token spelling, including added/special tokens. Multimodal prompt preparation uses this to validate the image placeholder against the model configuration.
Sourcepub fn convert_tokens_to_ids(&self, token: &str) -> Option<u32>
pub fn convert_tokens_to_ids(&self, token: &str) -> Option<u32>
Alias matching the HuggingFace tokenizer API used by the official DeepSeek image processor.