pub struct TextIndex { /* private fields */ }Expand description
A parsed text index: the token table (always resident) plus a posting source.
Implementations§
Source§impl TextIndex
impl TextIndex
Sourcepub fn from_section(section: &[u8], codec: u8) -> Result<TextIndex, FileError>
pub fn from_section(section: &[u8], codec: u8) -> Result<TextIndex, FileError>
Parse a whole section (local): token table + the full postings blob resident.
Sourcepub fn from_token_table(
prefix: &[u8],
codec: u8,
loader: Box<dyn Fn(u64, u64) -> Option<Vec<u8>> + Send + Sync>,
) -> Result<TextIndex, FileError>
pub fn from_token_table( prefix: &[u8], codec: u8, loader: Box<dyn Fn(u64, u64) -> Option<Vec<u8>> + Send + Sync>, ) -> Result<TextIndex, FileError>
Build from a section prefix holding the token table, plus a loader that
fetches a posting (offset_within_postings_blob, len) on demand (remote).
Sourcepub fn postings_base(section_prefix: &[u8]) -> Option<usize>
pub fn postings_base(section_prefix: &[u8]) -> Option<usize>
Byte offset of the postings blob within the section (token-table end). The remote opener uses this to base its posting-range loader.
pub fn token_count(&self) -> usize
Sourcepub fn lookup(&self, token: &str) -> Vec<u32>
pub fn lookup(&self, token: &str) -> Vec<u32>
Subjects whose literals contain the exact token (case-insensitive — the
caller passes a tokenized word), sorted ascending.
Sourcepub fn substring(&self, piece: &str) -> Option<Vec<u32>>
pub fn substring(&self, piece: &str) -> Option<Vec<u32>>
Subjects whose literals contain a word that CONTAINS piece as a
substring — the union over every matching table token. The token table
is always resident so the scan is in-memory; each matching token costs
one posting read (a range fetch on the remote path), so an unselective
piece matching more than [SUBSTRING_TOKEN_CAP] tokens returns None
and the caller falls back to its non-indexed path.