Expand description
Tokenizer implementations: whitespace, CJK-character, keyword, identifier, unicode-word.
Structs§
- CjkChar
Tokenizer - Emits each CJK character as its own token; non-CJK runs split on whitespace with ASCII punctuation stripped.
- Identifier
Tokenizer - Identifier-aware tokenizer: emits lowercased original + split parts for identifiers,
falls back to
WhitespaceTokenizerfor plain words. - Keyword
Tokenizer - Returns the entire (whitespace-trimmed) input as a single token. Empty input → empty vec.
- Whitespace
Tokenizer - Splits on ASCII whitespace, trims leading/trailing ASCII punctuation from each token, and drops empty results.