Expand description
Tokenizer caching layer (L1: prefix matching at special-token boundaries).
Wraps a cache-compatible Tokenizer in a cache that records prefix
tokenizations at every special-token boundary. On a hit, the cached prefix
tokens are merged with a fresh encode of the trailing suffix only — turning
O(N) tokenization work into O(suffix_len) when prompts share a system prefix.
§Correctness
Boundaries are taken only at positions immediately following a registered
special token (e.g. <|im_start|>, <|im_end|>, <s>, </s>). Special tokens
are atomic in BPE (special: true, normalized: false), so splitting there
preserves the invariant tokenize(prefix) + tokenize(suffix) == tokenize(prefix + suffix).
No fallback to whitespace or punctuation — better to miss than to corrupt.
Atomicity alone is insufficient when registered special-token strings can overlap.
CachedTokenizer::new disables L1 for such sets because the boundary scanner could
otherwise split inside the token selected by the underlying tokenizer.
§Storage normalization
When L1 is enabled, every encode returns Encoding::Sp (token-ids only) —
hits merge cached prefix ids with a fresh suffix encode, and misses assemble the ids
from the per-boundary segment encodes (see L1Cache::populate_and_encode) — even
when the inner tokenizer would have produced Encoding::Hf (rich offsets/attention/
etc). All current downstream consumers in Dynamo only call Encoding::token_ids, so
this lossy normalization is safe; revisit if a caller starts reading offsets or
attention masks from encodings produced through the cache.
§Configuration
special_tokens: Vec<String>— must be supplied at construction (theTokenizertrait is intentionally minimal and does not expose them). An empty list disables L1:encode/encode_batchshort-circuit straight to the inner tokenizer with no lookup, no miss-counter bump, and no insert attempt. A list whose members can overlap disables L1 identically.encode_segmentsalways passes through to the inner tokenizer without caching. Flattening segments for L1 would discard their special-token trust boundaries.max_memory_bytes— L1 byte budget; entries evicted via approximate LRU.
§Provenance
Adapted from llm-tokenizer v1.3.2 (cache/l1.rs, cache/mod.rs). L0 and
fingerprinting were dropped; L1 alone covers the headline multi-turn-chat
workload, and the in-memory cache lifetime is bound to a single tokenizer
instance so fingerprint-based invalidation is unnecessary.
Structs§
- Cache
Token Usage - Token-level cache usage for one successful encode.
- Cached
Tokenizer - Caching wrapper around an inner tokenizer.
- L1Cache
- L1 cache: prefix matching at special-token boundaries, backed by a weighted W-TinyLFU
mokacache that owns storage, recency/frequency tracking, and eviction. Hit/miss counts (our notion of a prefix hit) are tracked separately for metrics. - L1Cache
Stats
Type Aliases§
- Cache
Event Fn - Optional per-event observer.
on_hitruns after each cache hit,on_missafter each miss — wired byCachedTokenizer::with_observerto push events straight into Prometheus counters without a periodic sampling step. - Cache
Token Usage Fn - Optional observer for token-level cache usage.