Expand description
Tokenizer caching layer (L1: prefix matching at special-token boundaries).
Wraps a cache-compatible Tokenizer in a cache that records prefix
tokenizations at every special-token boundary. On a hit, the cached prefix
tokens are merged with a fresh encode of the trailing suffix only — turning
O(N) tokenization work into O(suffix_len) when prompts share a system prefix.
§Correctness
Boundaries are taken only at positions immediately following a registered
special token (e.g. <|im_start|>, <|im_end|>, <s>, </s>). Special tokens
are atomic in BPE (special: true, normalized: false), so splitting there
preserves the invariant tokenize(prefix) + tokenize(suffix) == tokenize(prefix + suffix).
No fallback to whitespace or punctuation — better to miss than to corrupt.
Callers must supply actual atomic special tokens recognized by the inner tokenizer;
adding arbitrary strings to the boundary list is unsafe.
Atomicity alone is insufficient when registered special-token strings can overlap.
CachedTokenizer::new disables L1 for such sets because the boundary scanner could
otherwise split inside the token selected by the underlying tokenizer.
§Storage normalization
When L1 is enabled, every encode returns Encoding::Sp (token-ids only) —
hits merge cached prefix ids with a fresh suffix encode, and misses assemble the ids
from the per-boundary segment encodes (see L1Cache::populate_and_encode) — even
when the inner tokenizer would have produced Encoding::Hf (rich offsets/attention/
etc). All current downstream consumers in Dynamo only call Encoding::token_ids, so
this lossy normalization is safe; revisit if a caller starts reading offsets or
attention masks from encodings produced through the cache.
§Configuration
special_tokens: Vec<String>— must be supplied at construction (theTokenizertrait is intentionally minimal and does not expose them). An empty list disables L1:encode/encode_batchshort-circuit straight to the inner tokenizer with no lookup, no miss-counter bump, and no insert attempt. A list whose members can overlap disables L1 identically.encode_segmentsalways passes through to the inner tokenizer without caching. Flattening segments for L1 would discard their special-token trust boundaries.max_memory_bytes— token-ID payload byte budget, excluding keys and metadata. Moka shares admission and eviction via W-TinyLFU. Deferred maintenance makes the capacity approximate; it is not a limit on total process memory.CachedTokenizer::newowns a private cache.CachedTokenizer::new_with_cacheshares storage across wrappers; entries survive a wrapper being dropped while shared storage remains alive. Equal namespaces must identify identical tokenizer behavior, including tokenizer files, backend, and options that affect token IDs.
§Provenance
Adapted from llm-tokenizer v1.3.2 (cache/l1.rs, cache/mod.rs). L0 and
upstream fingerprinting were dropped; L1 covers the multi-turn-chat workload.
Shared caches use caller-supplied namespaces to separate tokenizer identities.
Structs§
- Cache
Token Usage - Token-level cache usage for one successful encode.
- Cached
Tokenizer - Caching wrapper around an inner tokenizer.
- L1Cache
- L1 cache: prefix matching at special-token boundaries, backed by a weighted W-TinyLFU
mokacache that owns storage, recency/frequency tracking, and eviction. Hit/miss counts (our notion of a prefix hit) are tracked separately for metrics. - L1Cache
Stats - Shared
Tokenizer Cache - Shared storage and eviction budget for any number of cached tokenizers.
- Shared
Tokenizer Cache Stats - Storage statistics after pending maintenance. Concurrent writes can change them.
Type Aliases§
- Cache
Event Fn - Optional per-event observer.
on_hitruns after each cache hit,on_missafter each miss — wired byCachedTokenizer::with_observerto push events straight into Prometheus counters without a periodic sampling step. - Cache
Token Usage Fn - Optional observer for token-level cache usage.