pub struct SparseVectorQuery {
pub field: Field,
pub vector: Vec<(u32, f32)>,
pub combiner: MultiValueCombiner,
pub heap_factor: f32,
pub weight_threshold: f32,
pub max_query_dims: Option<usize>,
pub pruning: Option<f32>,
pub min_query_dims: usize,
pub over_fetch_factor: f32,
pub lsp_gamma: Option<usize>,
/* private fields */
}Expand description
Sparse vector query for similarity search
Fields§
§field: FieldField containing the sparse vectors
vector: Vec<(u32, f32)>Query vector as (dimension_id, weight) pairs
combiner: MultiValueCombinerHow to combine scores for multi-valued documents
heap_factor: f32Approximate search factor (1.0 = exact, lower values = faster but approximate) Controls MaxScore pruning aggressiveness in block-max scoring
weight_threshold: f32Minimum abs(weight) for query dimensions (0.0 = no filtering) Dimensions below this threshold are dropped from candidate generation. BMP still uses the bounded full query when scoring visited candidates.
max_query_dims: Option<usize>Maximum candidate-generation dimensions (None = implementation cap).
Keeps only the top-k dimensions by abs(weight); BMP final scoring uses
up to MAX_QUERY_TERMS dimensions from the full query.
pruning: Option<f32>Fraction of query dimensions to keep (0.0-1.0), same semantics as
indexing-time pruning: sort by abs(weight) descending,
keep top fraction. BMP applies it to candidate generation and scores
visited candidates with the bounded full query. None or 1.0 = no pruning.
min_query_dims: usizeMinimum number of query dimensions before pruning and weight_threshold filtering are applied. Protects short queries from losing signal. Default: 4. Set to 0 to always apply.
over_fetch_factor: f32Multiplier on executor limit for ordinal deduplication (1.0 = no over-fetch)
lsp_gamma: Option<usize>LSP/0 γ. None is depth-derived; Some(0) is exhaustive.
Implementations§
Source§impl SparseVectorQuery
impl SparseVectorQuery
Sourcepub fn new(field: Field, vector: Vec<(u32, f32)>) -> Self
pub fn new(field: Field, vector: Vec<(u32, f32)>) -> Self
Create a new sparse vector query
Default combiner is LogSumExp { temperature: 0.7 } — a
softmax-weighted smooth maximum. A document’s score follows its
strongest ordinals; ordinal count contributes nothing on its own,
so many-chunk documents cannot outrank a focused strong match.
Sourcepub fn with_combiner(self, combiner: MultiValueCombiner) -> Self
pub fn with_combiner(self, combiner: MultiValueCombiner) -> Self
Set the multi-value score combiner
Sourcepub fn with_over_fetch_factor(self, factor: f32) -> Self
pub fn with_over_fetch_factor(self, factor: f32) -> Self
Set executor over-fetch factor for multi-valued fields. After MaxScore execution, ordinal combining may reduce result count; this multiplier compensates by fetching more from the executor. (1.0 = no over-fetch, 2.0 = fetch 2x then combine down)
Sourcepub fn with_heap_factor(self, heap_factor: f32) -> Self
pub fn with_heap_factor(self, heap_factor: f32) -> Self
Set the heap factor for approximate search
Controls the trade-off between speed and recall:
- 1.0 = exact search (default)
- 0.8-0.9 = ~20-40% faster with minimal recall loss
- Lower values = more aggressive pruning, faster but lower recall
Sourcepub fn with_weight_threshold(self, threshold: f32) -> Self
pub fn with_weight_threshold(self, threshold: f32) -> Self
Set minimum weight threshold for query dimensions Dimensions with abs(weight) below this are dropped before search.
Sourcepub fn with_max_query_dims(self, max_dims: usize) -> Self
pub fn with_max_query_dims(self, max_dims: usize) -> Self
Set maximum number of query dimensions (top-k by weight)
Sourcepub fn with_pruning(self, fraction: f32) -> Self
pub fn with_pruning(self, fraction: f32) -> Self
Set pruning fraction (0.0-1.0): keep top fraction of query dims by weight.
Same semantics as indexing-time pruning.
Sourcepub fn with_min_query_dims(self, min_dims: usize) -> Self
pub fn with_min_query_dims(self, min_dims: usize) -> Self
Set minimum query dimensions before pruning/filtering are applied. Queries with fewer dimensions than this skip weight_threshold and pruning.
Sourcepub fn with_lsp_gamma(self, gamma: usize) -> Self
pub fn with_lsp_gamma(self, gamma: usize) -> Self
Select at most the top-γ superblocks by SBMax using LSP/0. Zero retains exhaustive SBMax-ordered traversal.
Sourcepub fn from_indices_weights(
field: Field,
indices: Vec<u32>,
weights: Vec<f32>,
) -> Self
pub fn from_indices_weights( field: Field, indices: Vec<u32>, weights: Vec<f32>, ) -> Self
Create from separate indices and weights vectors
Sourcepub fn from_text(
field: Field,
text: &str,
tokenizer_name: &str,
weighting: QueryWeighting,
sparse_index: Option<&SparseIndex>,
) -> Result<Self>
pub fn from_text( field: Field, text: &str, tokenizer_name: &str, weighting: QueryWeighting, sparse_index: Option<&SparseIndex>, ) -> Result<Self>
Create from raw text using a HuggingFace tokenizer (single segment)
This method tokenizes the text and creates a sparse vector query.
For multi-segment indexes, use from_text_with_stats instead.
§Arguments
field- The sparse vector field to searchtext- Raw text to tokenizetokenizer_name- HuggingFace tokenizer path (e.g., “bert-base-uncased”)weighting- Weighting strategy for tokenssparse_index- Optional sparse index for IDF lookup (required for IDF weighting)
Sourcepub fn from_text_with_stats(
field: Field,
text: &str,
tokenizer: &HfTokenizer,
weighting: QueryWeighting,
global_stats: Option<&GlobalStats>,
) -> Result<Self>
pub fn from_text_with_stats( field: Field, text: &str, tokenizer: &HfTokenizer, weighting: QueryWeighting, global_stats: Option<&GlobalStats>, ) -> Result<Self>
Create from raw text using global statistics (multi-segment)
This is the recommended method for multi-segment indexes as it uses aggregated IDF values across all segments for consistent ranking.
§Arguments
field- The sparse vector field to searchtext- Raw text to tokenizetokenizer- Pre-loaded HuggingFace tokenizerweighting- Weighting strategy for tokensglobal_stats- Global statistics for IDF computation
Sourcepub fn from_text_with_tokenizer_bytes(
field: Field,
text: &str,
tokenizer_bytes: &[u8],
weighting: QueryWeighting,
global_stats: Option<&GlobalStats>,
) -> Result<Self>
pub fn from_text_with_tokenizer_bytes( field: Field, text: &str, tokenizer_bytes: &[u8], weighting: QueryWeighting, global_stats: Option<&GlobalStats>, ) -> Result<Self>
Create from raw text, loading tokenizer from index directory
This method supports the index:// prefix for tokenizer paths,
loading tokenizer.json from the index directory.
§Arguments
field- The sparse vector field to searchtext- Raw text to tokenizetokenizer_bytes- Tokenizer JSON bytes (pre-loaded from directory)weighting- Weighting strategy for tokensglobal_stats- Global statistics for IDF computation
Trait Implementations§
Source§impl Clone for SparseVectorQuery
impl Clone for SparseVectorQuery
Source§fn clone(&self) -> SparseVectorQuery
fn clone(&self) -> SparseVectorQuery
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read moreSource§impl Debug for SparseVectorQuery
impl Debug for SparseVectorQuery
Source§impl Display for SparseVectorQuery
impl Display for SparseVectorQuery
Source§impl Query for SparseVectorQuery
impl Query for SparseVectorQuery
Source§fn scorer<'a>(
&self,
reader: &'a SegmentReader,
limit: usize,
) -> ScorerFuture<'a>
fn scorer<'a>( &self, reader: &'a SegmentReader, limit: usize, ) -> ScorerFuture<'a>
Source§fn scorer_with_options<'a>(
&self,
reader: &'a SegmentReader,
limit: usize,
options: ScorerOptions,
) -> ScorerFuture<'a>
fn scorer_with_options<'a>( &self, reader: &'a SegmentReader, limit: usize, options: ScorerOptions, ) -> ScorerFuture<'a>
Source§fn scorer_sync<'a>(
&self,
reader: &'a SegmentReader,
limit: usize,
) -> Result<Box<dyn Scorer + 'a>>
fn scorer_sync<'a>( &self, reader: &'a SegmentReader, limit: usize, ) -> Result<Box<dyn Scorer + 'a>>
Source§fn scorer_sync_with_options<'a>(
&self,
reader: &'a SegmentReader,
limit: usize,
options: ScorerOptions,
) -> Result<Box<dyn Scorer + 'a>>
fn scorer_sync_with_options<'a>( &self, reader: &'a SegmentReader, limit: usize, options: ScorerOptions, ) -> Result<Box<dyn Scorer + 'a>>
Query::scorer_with_options.Source§fn count_estimate<'a>(&self, _reader: &'a SegmentReader) -> CountFuture<'a>
fn count_estimate<'a>(&self, _reader: &'a SegmentReader) -> CountFuture<'a>
Source§fn decompose(&self) -> QueryDecomposition
fn decompose(&self) -> QueryDecomposition
Source§fn is_filter(&self) -> bool
fn is_filter(&self) -> bool
Source§fn as_doc_predicate<'a>(
&self,
_reader: &'a SegmentReader,
) -> Option<DocPredicate<'a>>
fn as_doc_predicate<'a>( &self, _reader: &'a SegmentReader, ) -> Option<DocPredicate<'a>>
Source§fn as_doc_bitset(&self, _reader: &SegmentReader) -> Option<DocBitset>
fn as_doc_bitset(&self, _reader: &SegmentReader) -> Option<DocBitset>
Source§fn bitset_cardinality_estimate(&self, _reader: &SegmentReader) -> Option<u64>
fn bitset_cardinality_estimate(&self, _reader: &SegmentReader) -> Option<u64>
None = unknown (treated as matching everything).Auto Trait Implementations§
impl Freeze for SparseVectorQuery
impl RefUnwindSafe for SparseVectorQuery
impl Send for SparseVectorQuery
impl Sync for SparseVectorQuery
impl Unpin for SparseVectorQuery
impl UnsafeUnpin for SparseVectorQuery
impl UnwindSafe for SparseVectorQuery
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
Source§impl<T> DropFlavorWrapper<T> for T
impl<T> DropFlavorWrapper<T> for T
Source§impl<T, W> HasTypeWitness<W> for Twhere
W: MakeTypeWitness<Arg = T>,
T: ?Sized,
impl<T, W> HasTypeWitness<W> for Twhere
W: MakeTypeWitness<Arg = T>,
T: ?Sized,
Source§impl<T> Identity for Twhere
T: ?Sized,
impl<T> Identity for Twhere
T: ?Sized,
Source§impl<T> Instrument for T
impl<T> Instrument for T
Source§fn instrument(self, span: Span) -> Instrumented<Self> ⓘ
fn instrument(self, span: Span) -> Instrumented<Self> ⓘ
Source§fn in_current_span(self) -> Instrumented<Self> ⓘ
fn in_current_span(self) -> Instrumented<Self> ⓘ
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§impl<T> Pointable for T
impl<T> Pointable for T
Source§impl<T> PolicyExt for Twhere
T: ?Sized,
impl<T> PolicyExt for Twhere
T: ?Sized,
impl<T> Read<Exclusive, BecauseExclusive> for Twhere
T: ?Sized,
impl<E> ResultError for E
impl<T> ResultType for T
Source§impl<T> ToCompactString for Twhere
T: Display,
impl<T> ToCompactString for Twhere
T: Display,
Source§fn try_to_compact_string(&self) -> Result<CompactString, ToCompactStringError>
fn try_to_compact_string(&self) -> Result<CompactString, ToCompactStringError>
ToCompactString::to_compact_string() Read moreSource§fn to_compact_string(&self) -> CompactString
fn to_compact_string(&self) -> CompactString
CompactString. Read more