Skip to main content

Crate khive_text

Crate khive_text 

Source
Expand description

Text analysis primitives: tokenization, normalization, filtering.

Re-exports§

pub use analyzer::StandardAnalyzer;
pub use lang::contains_cjk;
pub use lang::contains_cjk_f32;
pub use lang::is_cjk_char;
pub use lang::is_meaningful_query;
pub use lang::ScriptProfile;

Modules§

analyzer
Composable text analysis pipeline.
filter
Token filters: lowercase, stop words, length constraints, stemming.
identifier
Script/identifier detection and splitting for code-aware tokenization.
lang
Script/alphabet identification for per-language routing decisions.
preset
Named analyzer constructors for common use cases.
tokenizer
Tokenizer implementations: whitespace, CJK-character, keyword, identifier, unicode-word.

Traits§

Analyzer
Full analysis pipeline: tokenize + filter chain.
TokenFilter
Transforms or drops a single token. Returns None to drop.
Tokenizer
Splits a string into raw tokens. Must be deterministic and stateless.

Type Aliases§

BoxedAnalyzer
BoxedTokenizer