Skip to main content

Module tokenizer

Module tokenizer 

Source
Expand description

Tokenizer implementations: whitespace, CJK-character, keyword, identifier, unicode-word.

Structs§

CjkCharTokenizer
Emits each CJK character as its own token; non-CJK runs split on whitespace with ASCII punctuation stripped.
IdentifierTokenizer
Identifier-aware tokenizer: emits lowercased original + split parts for identifiers, falls back to WhitespaceTokenizer for plain words.
KeywordTokenizer
Returns the entire (whitespace-trimmed) input as a single token. Empty input → empty vec.
WhitespaceTokenizer
Splits on ASCII whitespace, trims leading/trailing ASCII punctuation from each token, and drops empty results.