Skip to main content

Module cache

Module cache 

Source
Expand description

Tokenizer caching layer (L1: prefix matching at special-token boundaries).

Wraps a cache-compatible Tokenizer in a cache that records prefix tokenizations at every special-token boundary. On a hit, the cached prefix tokens are merged with a fresh encode of the trailing suffix only — turning O(N) tokenization work into O(suffix_len) when prompts share a system prefix.

§Correctness

Boundaries are taken only at positions immediately following a registered special token (e.g. <|im_start|>, <|im_end|>, <s>, </s>). Special tokens are atomic in BPE (special: true, normalized: false), so splitting there preserves the invariant tokenize(prefix) + tokenize(suffix) == tokenize(prefix + suffix). No fallback to whitespace or punctuation — better to miss than to corrupt. Callers must supply actual atomic special tokens recognized by the inner tokenizer; adding arbitrary strings to the boundary list is unsafe.

Atomicity alone is insufficient when registered special-token strings can overlap. CachedTokenizer::new disables L1 for such sets because the boundary scanner could otherwise split inside the token selected by the underlying tokenizer.

§Storage normalization

When L1 is enabled, every encode returns Encoding::Sp (token-ids only) — hits merge cached prefix ids with a fresh suffix encode, and misses assemble the ids from the per-boundary segment encodes (see L1Cache::populate_and_encode) — even when the inner tokenizer would have produced Encoding::Hf (rich offsets/attention/ etc). All current downstream consumers in Dynamo only call Encoding::token_ids, so this lossy normalization is safe; revisit if a caller starts reading offsets or attention masks from encodings produced through the cache.

§Configuration

  • special_tokens: Vec<String> — must be supplied at construction (the Tokenizer trait is intentionally minimal and does not expose them). An empty list disables L1: encode/encode_batch short-circuit straight to the inner tokenizer with no lookup, no miss-counter bump, and no insert attempt. A list whose members can overlap disables L1 identically.
  • encode_segments always passes through to the inner tokenizer without caching. Flattening segments for L1 would discard their special-token trust boundaries.
  • max_memory_bytes — token-ID payload byte budget, excluding keys and metadata. Moka shares admission and eviction via W-TinyLFU. Deferred maintenance makes the capacity approximate; it is not a limit on total process memory.
  • CachedTokenizer::new owns a private cache. CachedTokenizer::new_with_cache shares storage across wrappers; entries survive a wrapper being dropped while shared storage remains alive. Equal namespaces must identify identical tokenizer behavior, including tokenizer files, backend, and options that affect token IDs.

§Provenance

Adapted from llm-tokenizer v1.3.2 (cache/l1.rs, cache/mod.rs). L0 and upstream fingerprinting were dropped; L1 covers the multi-turn-chat workload. Shared caches use caller-supplied namespaces to separate tokenizer identities.

Structs§

CacheTokenUsage
Token-level cache usage for one successful encode.
CachedTokenizer
Caching wrapper around an inner tokenizer.
L1Cache
L1 cache: prefix matching at special-token boundaries, backed by a weighted W-TinyLFU moka cache that owns storage, recency/frequency tracking, and eviction. Hit/miss counts (our notion of a prefix hit) are tracked separately for metrics.
L1CacheStats
SharedTokenizerCache
Shared storage and eviction budget for any number of cached tokenizers.
SharedTokenizerCacheStats
Storage statistics after pending maintenance. Concurrent writes can change them.

Type Aliases§

CacheEventFn
Optional per-event observer. on_hit runs after each cache hit, on_miss after each miss — wired by CachedTokenizer::with_observer to push events straight into Prometheus counters without a periodic sampling step.
CacheTokenUsageFn
Optional observer for token-level cache usage.