What is splintr?
Splintr loads a tokenizer from three sources and dispatches it to four backends, all behind a single AnyTokenizer type — the calling code never changes with the vocabulary:
| Source | Loads | Backends |
|---|---|---|
Bundled (from_pretrained) |
8 vocabularies: OpenAI, Llama 3, DeepSeek, Mistral, Whisper | byte-level BPE, SPM-BPE |
tokenizer.json (from_json) |
Any HuggingFace file — normalizers, pre-tokenizers, decoders | byte-level BPE, Unigram, WordPiece |
GGUF vocab (from_gguf_vocab) |
The tokenizer.ggml.* keys, parsed by your GGUF loader |
byte-level BPE, SPM-BPE, Unigram, WordPiece |
Correctness is differential: every family is fuzzed id-for-id against its reference implementation using strings built from each vocabulary's own added and special tokens. See CONTRIBUTING.md for how that is established.
Why it exists
Tokenization sits on the hot path of every LLM application — prompts, training corpora, RAG chunks, token counting for billing. Python-based tokenizers cannot use all your cores, so batch preprocessing turns into wall-clock latency. The usual escape is one library per format — tiktoken, sentencepiece, tokenizers — three dependencies, three APIs, no common handle, and no answer at all for a GGUF vocabulary. Splintr's answer is one handle over every format, at Rust speed, with reference implementations as the correctness oracle.
Performance
Batch encoding parallelizes across texts, which is where the gap is widest — and it widens with batch size, as the fixed cost of spinning up the pool is amortized over more work:

Single texts stay on the sequential path, and still lead across every content type:

Call it ~10x tiktoken on batches, ~2-2.6x on single texts. The ballpark holds across machines; the exact figure does not, since absolute throughput moves with hardware, CPU architecture and the versions compared against.
The measured table below is a separate, more recent run — on an AMD Ryzen 9 5900X (24 cores, Linux), CPython 3.12, against tiktoken 0.8.0, HuggingFace tokenizers 0.22.1 and TokenDagger 0.1.1. Splintr, tiktoken and TokenDagger run cl100k_base; the HuggingFace column is gpt2, so read it as a scale rather than a like-for-like. Where it disagrees with the charts above, which were plotted on different hardware, the table is the measured one:
| Configuration | Splintr | Tiktoken | HuggingFace | TokenDagger | vs tiktoken |
|---|---|---|---|---|---|
| 1,000 texts | 56.4 MB/s | 5.8 MB/s | 14.7 MB/s | 4.7 MB/s | 9.8x |
| 500 texts | 65.1 MB/s | 5.5 MB/s | 14.8 MB/s | 5.4 MB/s | 11.9x |
| 100 texts | 50.3 MB/s | 4.0 MB/s | 11.3 MB/s | 3.8 MB/s | 12.7x |
Reproduce it with benchmarks/benchmark_batch.py. See docs/benchmarks.md for per-content-type latency, methodology and the PCRE2 backend.
Quick Start
Python
# Load a pretrained vocabulary
= # OpenAI GPT-4/3.5
# tokenizer = Tokenizer.from_pretrained("llama3") # Meta Llama 3 family
# tokenizer = Tokenizer.from_pretrained("deepseek_v3") # DeepSeek V3/R1
# Encode and decode
=
=
# Batch encode (parallel across texts)
=
See the API Guide for complete documentation and examples.
Rust
use from_pretrained;
let tokenizer = from_pretrained?;
let tokens = tokenizer.encode;
let batch_tokens = tokenizer.encode_batch;
let text = tokenizer.decode?;
See the API Guide and docs.rs for complete documentation.
Key Features
- Four backends, one handle — Byte-level/raw BPE, SentencePiece BPE, Unigram, and WordPiece all load as
AnyTokenizer, so calling code stays the same whichever vocabulary you use - Parallel batch encoding — Rayon across texts; sequential for single texts based on empirical benchmarking
- Three loading sources — Bundled vocabularies (8 supported), any HuggingFace
tokenizer.json, or a GGUF vocabulary - Streaming decoder — Real-time LLM output with proper UTF-8 boundary handling; one decoder per tokenizer (guide)
- 54 agent tokens — ChatML, thinking, ReAct, tool-calling, RAG citation tokens, built-in across all vocabularies (docs)
- Special-token policy —
encode_ordinary/encode_allowed_specialso untrusted text cannot forge a control token - Cross-platform — Python bindings via PyO3 (Linux, macOS, Windows), CPython 3.8+; native Rust library
Supported Vocabularies
| Vocabulary | Used By | base_vocab_size |
|---|---|---|
| cl100k_base | GPT-4, GPT-3.5-turbo | 100,277 |
| o200k_base | GPT-4o | 200,019 |
| llama3 | Llama 3, 3.1, 3.2, 3.3 | 128,256 |
| deepseek_v3 | DeepSeek V3, DeepSeek R1 | 128,815 |
| mistral_v1 | Mistral 7B v0.1/v0.2 | 32,000 |
| mistral_v2 | Mistral 7B v0.3, Codestral | 32,768 |
| mistral_v3 | Mistral NeMo, Large 2, Pixtral | 131,072 |
| whisper | OpenAI Whisper multilingual | 51,865–51,866 |
All bundled vocabularies include 54 agent tokens (except Whisper, which includes 1608 standard Whisper tokens). Load any other model with from_json("tokenizer.json") or from_gguf_vocab(). See docs/vocabularies.md for complete details and standard token lists.
Streaming Decoder
For real-time LLM output where tokens arrive one at a time:
=
BPE tokens don't align with UTF-8 boundaries. A multi-byte character might split across tokens. The streaming decoder buffers incomplete sequences and only outputs complete characters. One decoder per tokenizer, built by that tokenizer, so "".join(chunks) + flush() equals decode(ids) for any vocabulary. See API Guide for details and best practices.
Special Tokens in Untrusted Text
A tokenizer that matches special tokens will promote text that spells a control token to that token's real id. <|im_start|> typed by a user becomes the same id the server emits — downstream, nothing can tell them apart. Encoding takes an explicit mode:
| Mode | Behaviour |
|---|---|
encode_with_special(text) / All |
Match every configured special token found in text |
encode_ordinary(text) / Ordinary |
Match none — special spellings stay ordinary content |
encode_allowed_special(text, allowed) / Allow |
Match only the named tokens; raise on any other |
All three are on every Python tokenizer type — Tokenizer, AnyTokenizer, SpmTokenizer, SentencePieceTokenizer, WordPieceTokenizer — alongside encode (model-ready with boundary template), encode_raw (content tokens only), and encode_batch.
=
=
# Default: literal control token becomes real control-token id
# Ordinary: never match special tokens
# Allow-list: reject anything outside it
See docs/special_tokens.md for detailed guidance and a guide to the token list per vocabulary.
How It Works
Pre-tokenization runs on regexr, a pure-Rust regex engine with JIT and SIMD, and special tokens are matched with Aho-Corasick in a single pass. Merging uses a linked list rather than a vector, so pathological inputs stay linear, with an LRU cache over repeated chunks and FxHashMap for rank lookups. Batches are encoded in parallel with Rayon; single texts stay sequential, which measures faster below roughly 1 MB.
The other three backends are real implementations of their algorithms, not approximations: SentencePiece Unigram uses Viterbi maximum-score segmentation, SentencePiece BPE merges by score, and WordPiece does greedy longest-match with the ## continuation prefix.
Contributing
Bug reports, feature suggestions and pull requests are welcome — see CONTRIBUTING.md for development setup, the checks CI runs, and how correctness is established against the reference tokenizers.
Acknowledgments
Splintr builds on concepts from tiktoken, SentencePiece and tokenizers — which also serve as the reference implementations its output is checked against.
Citation
If you use Splintr in your research, please cite:
License
MIT — see LICENSE.