Skip to main content

Tokenizer

Struct Tokenizer 

Source
pub struct Tokenizer {
    pub bos_token_id: Option<u32>,
    pub eos_token_id: Option<u32>,
    pub pad_token_id: Option<u32>,
    pub im_start_id: Option<u32>,
    pub im_end_id: Option<u32>,
    pub chat_template: Option<String>,
    pub extra_eos: HashSet<u32>,
    pub add_bos: bool,
    /* private fields */
}
Expand description

A loaded BPE tokenizer.

Fields§

§bos_token_id: Option<u32>

Special tokens

§eos_token_id: Option<u32>§pad_token_id: Option<u32>§im_start_id: Option<u32>

Chat template special tokens

§im_end_id: Option<u32>§chat_template: Option<String>

Jinja chat template carried by the container (spec §6.1); None → hardcoded ChatML fallback.

§extra_eos: HashSet<u32>

Extra stop ids from the container’s generation config.

§add_bos: bool

Generation prepends BOS (llama post_processor semantics).

Implementations§

Source§

impl Tokenizer

Source

pub fn from_file(path: impl AsRef<Path>) -> Result<Self, TokenizerError>

Load tokenizer from HuggingFace tokenizer.json file.

Source

pub fn from_bytes(bytes: &[u8]) -> Result<Self, TokenizerError>

Load tokenizer from raw tokenizer.json bytes (CMF VOCAB section).

Source

pub fn from_json(json: &str) -> Result<Self, TokenizerError>

Load tokenizer from JSON string.

Source

pub fn byte_level() -> Self

Create a minimal tokenizer for testing (byte tokens, no merges).

Source

pub fn encode(&self, text: &str) -> Vec<u32>

Encode text to token IDs.

Source

pub fn encode_plain(&self, text: &str) -> Vec<u32>

Encode text as PLAIN text: the whole input is one added-token-free segment (NFC → split → byte-map → BPE), so a literal <|im_end|> in it stays bytes instead of becoming the special id. This is the trainer’s Bpe::encode (user text, never the template frame) — router v2 tokenizes the user message with it (spec §9.4).

Source

pub fn decode(&self, ids: &[u32]) -> String

Decode token IDs back to text. Special tokens are skipped; added tokens are raw text; everything else reverses the byte-level map.

Source

pub fn decode_token(&self, id: u32) -> String

Streaming decode of ONE token: no sequence-level Strip — a per-token strip would eat the ▁-spaces of every SP word.

Source

pub fn decode_token_for_hash(&self, id: u32) -> String

Decode one vocabulary entry for Engram’s compressed token map while retaining special tokens.

Source

pub fn decode_for_protocol(&self, ids: &[u32]) -> String

Decode generated protocol text while retaining special markers. The V4.1 harmony parser needs <think>, EOS, and spaced DSML tags.

Source

pub fn raw_token_for_hash(&self, id: u32) -> String

Return the backend vocabulary spelling for an Engram map entry.

Source

pub fn apply_chat_template(&self, messages: &[(String, String)]) -> Vec<u32>

Render the container’s Jinja chat template (HF semantics: trim_blocks + lstrip_blocks + loop controls) and encode it. Falls back to hardcoded ChatML when the file carries none.

Source

pub fn apply_chat_template_json( &self, messages: &[Value], tools: Option<&[Value]>, enable_thinking: Option<bool>, ) -> Vec<u32>

Like apply_chat_template, with an explicit enable_thinking value for reasoning-model templates (Qwen3/3.5 emit an empty block when it is false, so the model answers directly). None leaves the variable undefined — the template’s own default applies. Chat template with the FULL message shape and a tool list.

The pair-based API below flattens every message to (role, text), which silently drops exactly what agentic use needs: the tools array, role: "tool" results, and tool_calls on assistant turns. The templates this format embeds — Qwen-family, Nanbeige — have carried a {%- if tools %} branch all along; this is the call that finally feeds it. Messages arrive as JSON objects in the OpenAI shape and pass through to minijinja unflattened, so a template sees the same fields a Python apply_chat_template would.

Source

pub fn try_apply_chat_template_json( &self, messages: &[Value], tools: Option<&[Value]>, enable_thinking: Option<bool>, ) -> Result<Vec<u32>, String>

Like Self::apply_chat_template_json, but a template that fails to render is an ERROR instead of a quiet ChatML approximation.

The fallback flattens every message to (role, text): it has no place for tools, tool_calls or role: "tool". For a plain chat that is a tolerable degradation; for a request with tools it means the model never sees the functions and answers as if none were offered — a failure no client can detect. The server calls this variant when tools are present and reports the error instead. Files without a template still take the ChatML path (Ok).

Source

pub fn render_chat_json( &self, messages: &[Value], tools: Option<&[Value]>, enable_thinking: Option<bool>, ) -> Option<String>

Render the template against JSON-shaped messages (parity surface).

Source

pub fn apply_chat_template_opts( &self, messages: &[(String, String)], enable_thinking: Option<bool>, ) -> Vec<u32>

Source

pub fn with_bos(&self, ids: Vec<u32>) -> Vec<u32>

Prepend BOS when the tokenizer declares it (llama family).

Source

pub fn render_chat(&self, messages: &[(String, String)]) -> Option<String>

Render the carried template to text (parity-testable surface).

Source

pub fn render_chat_opts( &self, messages: &[(String, String)], enable_thinking: Option<bool>, ) -> Option<String>

Render the carried template to text with explicit thinking mode.

Source

pub fn vocab_size(&self) -> usize

Vocabulary size.

Source

pub fn token_to_id(&self, token: &str) -> Option<u32>

Return the ID for an exact token spelling, including added/special tokens. Multimodal prompt preparation uses this to validate the image placeholder against the model configuration.

Source

pub fn convert_tokens_to_ids(&self, token: &str) -> Option<u32>

Alias matching the HuggingFace tokenizer API used by the official DeepSeek image processor.

Source

pub fn is_eos(&self, id: u32) -> bool

Check if token ID is EOS.

Trait Implementations§

Source§

impl Debug for Tokenizer

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self> ⓘ

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self> ⓘ

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, !>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self> ⓘ
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self> ⓘ

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more