token-estimate
Estimate LLM token counts from text without loading a tokenizer.
Real tokenizers need a multi-megabyte vocabulary and a non-trivial dependency tree. That is right when you need an exact count, and wrong when you only need to answer "will this roughly fit in the context window" — the common case in CLI tooling, log processing and pre-flight budget checks.
use ;
let n = estimate;
assert!;
// Profile inferred from punctuation density
let _ = estimate_auto;
Accuracy, stated honestly
For ordinary English prose this lands within roughly 10-15% of a BPE tokenizer. It degrades predictably:
| Input | Behaviour |
|---|---|
| Code and markup | Tokenizes denser than prose; estimate runs low. Use Profile::Code. |
| Non-Latin scripts | Often one token per character; a chars/4 heuristic underestimates badly. Density is detected and adjusted, but treat CJK as a rough floor. |
| Long whitespace or punctuation runs | Compress well; estimate runs high. Whitespace runs are collapsed to compensate. |
If you need an exact count, use a real tokenizer. This is for budgeting.
Use estimate_with_margin when overflowing a context window costs more than
sending less.
License
MIT