khive-text
Text analysis primitives for khive: tokenization, normalization, and token
filtering. A pluggable Tokenizer -> TokenFilter chain composed into an
Analyzer, plus named presets for common cases (standard English, CJK,
identifiers, KG entity names).
Features
unicode— addsUnicodeWordTokenizer(Unicode word-boundary splitting viaunicode-segmentation)stem— addsSnowballStemmer, aTokenFilterbacked byrust-stemmers(English by default viaSnowballStemmer::english(), or anyrust_stemmers::Algorithmviafor_algorithm)full— both of the above
Usage
use ;
let tokens = standard.analyze;
assert_eq!;
preset::standard() chains WhitespaceTokenizer with LowercaseFilter,
a BM25-compatible stop-word filter, and MinLengthFilter(1) — producing the
same token stream as khive-bm25's SimpleTokenizer::default() (same stop
list, no max-length cap). This is intentionally different from the public
StopWordFilter in khive_text::filter, which has its own stop list for
callers building custom pipelines. Other named presets: preset::simple()
(whitespace + lowercase only, keeps stop words),
preset::keyword() (whole input as one token), preset::cjk()
(character-level unigrams for CJK scripts, whitespace for Latin), and
preset::kg_name() (identifier-aware splitting tuned for entity names like
"bert-base-uncased"). Build a custom pipeline with the same builder:
use StandardAnalyzer;
use LowercaseFilter;
use IdentifierTokenizer;
let analyzer = with_tokenizer
.filter;
Script detection
is_cjk_char, contains_cjk, and is_meaningful_query (a ScriptProfile-based
heuristic for rejecting queries that are too short or punctuation-only to search
on) live in the lang module for callers that need to branch on script before
choosing an analyzer.
Where this sits
khive-text has no required khive-* dependencies and no current in-workspace
consumers — it is a standalone tokenization library extracted for reuse by any
future retrieval crate that needs a pluggable analyzer independent of
khive-bm25's built-in SimpleTokenizer.
License
Apache-2.0.