Skip to main content

Searcher

Struct Searcher 

Source
pub struct Searcher { /* private fields */ }
Expand description

Independent search worker over a shared immutable Corpus.

Owns an ordered stream, compiled kernels, and reusable query/result workspace. A searcher is Send but not Sync; operations require &mut self. Create multiple workers from the same corpus to submit queries independently.

Implementations§

Source§

impl Searcher

Source

pub fn batch_workspace_bytes(&self) -> usize

Bytes reserved in batch query, score, selection, and readback buffers.

This is additional to the corpus and single-query workspace. Returns zero before batch workspace is reserved; excludes compiler/runtime bookkeeping and returned host neighbor vectors.

Source

pub fn reserve_batch(&mut self, query_count: usize, k: usize) -> Result<()>

Reserve and compile workspace for a batch of query_count queries.

search_batch calls this automatically. Call it during setup to exclude allocation and compilation from the first batch. Query widths are rounded to 8, 16, 32, or 64; compiled widths are cached and workspace is retained. Score storage covers at most 262,144 corpus rows, independent of index size. At width 64 this uses about 66 MiB for k=5 or 322 MiB for k=1,024, plus query storage (96 KiB at 384 dimensions). A one-query batch uses search.

§Errors

query_count must be 1–64 and k must be 1–1,024. Allocation, synchronization, and compiler failures propagate as errors.

Source

pub fn search_batch( &mut self, queries: &[f32], k: usize, ) -> Result<Vec<Vec<Neighbor>>>

Search up to 64 row-major FP32 queries, sharing corpus reads across them.

queries contains batch_size * self.dimensions() components. Results retain query order and use the same score/ID ordering as Self::search. A tiled FP16 × FP32 matrix kernel reuses each corpus tile across the batch; device selection maintains one running top-k per query. No full corpus × batch score matrix or host-side selection is used. Queries are normalized to FP32 without additional quantization. Accumulation order differs from search, so rounding can change rankings of near-ties.

An empty batch returns an empty vector. First use of a query width may allocate and compile; Self::reserve_batch can prepare it in advance.

let device = Device::open(0)?;
let mut db = Searcher::build(&device, 3, [[1.0, 0.0, 0.0], [0.0, 1.0, 0.0]])?;
db.reserve_batch(2, 5)?;
let queries = [1.0, 0.0, 0.0, 0.0, 1.0, 0.0];
let results = db.search_batch(&queries, 5)?;
assert_eq!(results[0][0].id, 0);
assert_eq!(results[1][0].id, 1);
§Errors

Rejects incomplete rows, more than 64 queries, k outside 1–1,024, and nonfinite or zero-norm queries. All queries are validated before execution. Runtime and compilation failures propagate as errors.

Source

pub fn search_batch_excluding( &mut self, queries: &[f32], k: usize, excluded: &[u32], ) -> Result<Vec<Vec<Neighbor>>>

Search a batch with one shared set of excluded insertion IDs.

Exclusions may be unsorted and repeated, apply only to this call, and compose with sharding and all supported k values. Each query returns min(k, remaining_rows) matches. Only final neighbors are read back. Input requirements and errors match Self::search_batch; out-of-range excluded IDs are also rejected.

Source

pub fn search_batch_into( &mut self, queries: &[f32], k: usize, output: &mut Vec<Vec<Neighbor>>, ) -> Result<()>

Replace a reusable batch output, retaining the capacity of surviving rows. Invalid input leaves output unchanged; runtime failures may partially write it.

Source

pub fn search_batch_excluding_into( &mut self, queries: &[f32], k: usize, excluded: &[u32], result: &mut Vec<Vec<Neighbor>>, ) -> Result<()>

Batch search with shared exclusions into reusable host output.

Source§

impl Searcher

Source

pub fn corpus(&self) -> &Corpus

Shared immutable storage. Clone this handle to retain a snapshot or build another searcher without copying vectors.

Source

pub fn stream(&mut self) -> &mut Stream

This searcher’s ordered execution stream. Use it to submit producers or consumers around device search. Do not replace or synchronize other queues implicitly; use HRX events for cross-stream dependencies.

Examples found in repository?
examples/device_search.rs (line 13)
4fn main() -> hrxdb::Result<()> {
5    let device = Device::open(0)?;
6    let corpus = Corpus::build(
7        &device,
8        3,
9        [[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [-1.0, 0.0, 0.0]],
10    )?;
11    let mut worker = corpus.searcher()?;
12    worker.reserve_device(1, 1)?;
13    let query = worker.stream().allocate(12)?;
14    worker.stream().upload(
15        query.binding(),
16        &[1.0f32, 0.0, 0.0]
17            .into_iter()
18            .flat_map(f32::to_le_bytes)
19            .collect::<Vec<_>>(),
20    )?;
21    let mut visited = DeviceExclusions::new(worker.stream(), corpus.len())?;
22    let mut result = DeviceNeighbors::new(worker.stream(), 1, 1)?;
23    for _ in 0..2 {
24        let _done = worker.search_device(
25            DeviceQueries::new(query.binding(), 1, 3, 3)?,
26            Some(visited.binding()),
27            &mut result,
28        )?;
29        visited.insert_device(worker.stream(), result.ids(), 1)?;
30    }
31    let done = worker.search_device(
32        DeviceQueries::new(query.binding(), 1, 3, 3)?,
33        Some(visited.binding()),
34        &mut result,
35    )?;
36    let mut consumer = corpus.stream()?;
37    consumer.wait_event(&done)?;
38    let matches = result.read(&mut consumer)?;
39    assert_eq!(matches[0][0].id, 2);
40    println!("third unvisited match: {:?}", matches[0][0]);
41    Ok(())
42}
Source§

impl Searcher

Source

pub fn memory_usage(&self) -> SearcherMemory

Separate shared corpus storage from this searcher’s reserved workspace. Excludes native allocation granularity, compiler code, runtime staging, and caller-owned inputs/outputs. No synchronization or GPU work occurs.

Source§

impl Searcher

Source

pub fn reserve_device(&mut self, batch: usize, k: usize) -> Result<()>

Prepare device normalization, scan, and selection for repeated submissions. First use can allocate/compile and wait while replacing host-imported batch workspace. Fully reserved device submissions perform no host wait.

Examples found in repository?
examples/device_search.rs (line 12)
4fn main() -> hrxdb::Result<()> {
5    let device = Device::open(0)?;
6    let corpus = Corpus::build(
7        &device,
8        3,
9        [[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [-1.0, 0.0, 0.0]],
10    )?;
11    let mut worker = corpus.searcher()?;
12    worker.reserve_device(1, 1)?;
13    let query = worker.stream().allocate(12)?;
14    worker.stream().upload(
15        query.binding(),
16        &[1.0f32, 0.0, 0.0]
17            .into_iter()
18            .flat_map(f32::to_le_bytes)
19            .collect::<Vec<_>>(),
20    )?;
21    let mut visited = DeviceExclusions::new(worker.stream(), corpus.len())?;
22    let mut result = DeviceNeighbors::new(worker.stream(), 1, 1)?;
23    for _ in 0..2 {
24        let _done = worker.search_device(
25            DeviceQueries::new(query.binding(), 1, 3, 3)?,
26            Some(visited.binding()),
27            &mut result,
28        )?;
29        visited.insert_device(worker.stream(), result.ids(), 1)?;
30    }
31    let done = worker.search_device(
32        DeviceQueries::new(query.binding(), 1, 3, 3)?,
33        Some(visited.binding()),
34        &mut result,
35    )?;
36    let mut consumer = corpus.stream()?;
37    consumer.wait_event(&done)?;
38    let matches = result.read(&mut consumer)?;
39    assert_eq!(matches[0][0].id, 2);
40    println!("third unvisited match: {:?}", matches[0][0]);
41    Ok(())
42}
Source

pub fn search_device( &mut self, queries: DeviceQueries<'_>, excluded: Option<View<'_>>, output: &mut DeviceNeighbors, ) -> Result<Event>

Queue GPU normalization, exhaustive cosine search, and top-k into owned device outputs. Return a completion event without reading results back.

excluded is an optional shared bitmap (u32 words, least-significant bit first) over insertion IDs, at least ceil(corpus.len()/32)*4 bytes. It is read in place, allowing device producers to update it incrementally. Output batch count must match input. output.k() selects k. Invalid queries set output status=1 and count=0; DeviceNeighbors::read reports these as errors. Inputs/bitmap must be ready on this searcher’s stream. Keep buffers read-only until completion. Slots beyond counts are invalid.

Scaled FP32 norm accumulation handles extreme finite magnitudes without overflow. Its rounding can differ from host FP64 normalization; very tiny relative components may underflow. Scores use FP32 accumulation.

Examples found in repository?
examples/device_search.rs (lines 24-28)
4fn main() -> hrxdb::Result<()> {
5    let device = Device::open(0)?;
6    let corpus = Corpus::build(
7        &device,
8        3,
9        [[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [-1.0, 0.0, 0.0]],
10    )?;
11    let mut worker = corpus.searcher()?;
12    worker.reserve_device(1, 1)?;
13    let query = worker.stream().allocate(12)?;
14    worker.stream().upload(
15        query.binding(),
16        &[1.0f32, 0.0, 0.0]
17            .into_iter()
18            .flat_map(f32::to_le_bytes)
19            .collect::<Vec<_>>(),
20    )?;
21    let mut visited = DeviceExclusions::new(worker.stream(), corpus.len())?;
22    let mut result = DeviceNeighbors::new(worker.stream(), 1, 1)?;
23    for _ in 0..2 {
24        let _done = worker.search_device(
25            DeviceQueries::new(query.binding(), 1, 3, 3)?,
26            Some(visited.binding()),
27            &mut result,
28        )?;
29        visited.insert_device(worker.stream(), result.ids(), 1)?;
30    }
31    let done = worker.search_device(
32        DeviceQueries::new(query.binding(), 1, 3, 3)?,
33        Some(visited.binding()),
34        &mut result,
35    )?;
36    let mut consumer = corpus.stream()?;
37    consumer.wait_event(&done)?;
38    let matches = result.read(&mut consumer)?;
39    assert_eq!(matches[0][0].id, 2);
40    println!("third unvisited match: {:?}", matches[0][0]);
41    Ok(())
42}
Source§

impl Searcher

Source

pub fn build<I, R>(device: &Device, dimensions: usize, rows: I) -> Result<Self>
where I: IntoIterator<Item = R>, I::IntoIter: ExactSizeIterator, R: AsRef<[f32]>,

Build from a sized stream of rows without retaining a host corpus copy. A row is accepted as any AsRef<[f32]>, including borrowed slices.

Rows are normalized and quantized to FP16. Dimensions must be 1–16,384; the iterator must accurately report at most 2^30 rows. Larger corpora are split internally into allocations of at most 2^32 FP16 elements, keeping scan addresses within the compiler’s element-index limit. IDs and query results span the entire corpus. Vector and inverse-norm allocations reserve whole 256-row tiles; extra rows are readable slack, not indexed data. See CorpusShardView::capacity_rows. The default scan schedule is tuned for 10M × 384 vectors. Construction allocates device storage, uploads the corpus, and compiles the kernels synchronously.

§Errors

Returns an error for an unsupported device, invalid dimensions/row count, nonfinite or zero-norm rows, mismatched row lengths, or runtime/compiler failures. Allocation and compiler limits depend on the system and shape.

Source

pub fn build_fp16<I, R>( device: &Device, dimensions: usize, rows: I, ) -> Result<Self>
where I: IntoIterator<Item = R>, I::IntoIter: ExactSizeIterator, R: AsRef<[u8]>,

Build from rows of little-endian IEEE 754 binary16 bytes.

Each row must contain exactly dimensions * 2 bytes, without padding. Values are copied unchanged, zero padding is added, and inverse norms are computed from the supplied FP16 values. Rows need not be normalized. This avoids an FP32 row allocation and re-quantization during ingestion. Scores describe the supplied representation, which may differ from normalizing FP32 values before rounding in Self::build.

let device = Device::open(0)?;
// Two 2D rows: [1, 0] and [0, 1], stored as little-endian FP16.
let bytes = [0x00, 0x3c, 0, 0, 0, 0, 0x00, 0x3c];
let mut db = Searcher::build_fp16(&device, 2, bytes.as_chunks::<4>().0)?;
let mut scores = vec![0.0; db.len()];
db.scores_into(&[1.0, 0.0], &mut scores)?;
assert_eq!(scores, [1.0, 0.0]);
§Errors

The shape, iterator, device, and runtime requirements are the same as Self::build. Incorrect byte lengths, nonfinite values, and zero-norm rows are rejected.

Source

pub fn build_fp16_with_config<I, R>( device: &Device, dimensions: usize, rows: I, config: ScanConfig, ) -> Result<Self>
where I: IntoIterator<Item = R>, I::IntoIter: ExactSizeIterator, R: AsRef<[u8]>,

Build from FP16 byte rows with an explicit scan schedule.

Has the same input requirements and errors as Self::build_fp16, and also rejects unsupported ScanConfig field values.

Source

pub fn build_with_config<I, R>( device: &Device, dimensions: usize, rows: I, config: ScanConfig, ) -> Result<Self>
where I: IntoIterator<Item = R>, I::IntoIter: ExactSizeIterator, R: AsRef<[f32]>,

Build an index with an explicit scan schedule.

Has the same input requirements and errors as Self::build, and also rejects unsupported ScanConfig field values.

Source

pub fn new(corpus: Corpus) -> Result<Self>

Create an independent searcher over shared corpus storage.

Source

pub fn with_config(corpus: Corpus, config: ScanConfig) -> Result<Self>

Create a searcher with a chosen single-query scan schedule.

Source

pub fn on_stream( corpus: Corpus, stream: Stream, config: ScanConfig, ) -> Result<Self>

Take ownership of a caller’s stream to compose custom kernels and search. The stream must belong to the corpus’s device. Pending work stays ordered.

Source

pub fn len(&self) -> usize

Number of indexed vectors.

Source

pub fn is_empty(&self) -> bool

Whether the index has no vectors.

Source

pub fn dimensions(&self) -> usize

Logical vector length supplied at construction.

Source

pub fn padded_dimensions(&self) -> usize

Stored vector length, rounded up to a multiple of 128 with zeros.

Source

pub fn shard_count(&self) -> usize

Number of independently addressed corpus allocations (zero when empty).

Source

pub fn config(&self) -> ScanConfig

Currently selected scan schedule.

Source

pub fn compilation_reports(&self) -> &[Compilation]

Scoring/read-control report pairs in shard order, followed by small-k selection/merge, large-k sorting/merge, and exclusion masking reports. Batch kernel reports are appended as query widths are prepared. Artifact paths refer to the local HRX cache.

Source

pub fn configure(&mut self, config: ScanConfig) -> Result<()>

Compile a new scan configuration outside query timing; retains the corpus. Returns an error for an invalid schedule or synchronization/compilation failure. The previous schedule is retained if compilation fails. This configures single-query scans; the batch matrix kernel uses its own schedule, independent of ScanConfig.

Reserve selection scratch for queries returning up to k neighbors.

Construction reserves k=32. Searches grow this capacity automatically; call this during setup to keep allocation out of the first larger query. Scratch is retained for subsequent queries. No kernels are compiled.

§Errors

k must be 1–1,024. Runtime allocation or synchronization errors propagate.

Source

pub fn search(&mut self, query: &[f32], k: usize) -> Result<Vec<Neighbor>>

Search with device-side top-k, without compiling kernels.

Returns min(k, self.len()) matches ordered by descending computed cosine score, then ascending insertion ID. The query is normalized to FP32. FP16 storage and FP32 arithmetic can change rankings from the original vectors. This method blocks until results are readable on the host. Selection scratch grows on the first larger-k query and is reused; Self::reserve_search can reserve it before serving queries.

§Errors

The query must have Self::dimensions finite components and nonzero norm; k must be 1–1,024, including for an empty index. Runtime failures propagate as errors.

Examples found in repository?
examples/search.rs (line 17)
4fn main() -> hrxdb::Result<()> {
5    let device = Device::open(0)?;
6    let corpus = Corpus::build(
7        &device,
8        3,
9        [[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [0.8, 0.2, 0.0]],
10    )?;
11    let external_ids = ["document-a", "document-b", "document-c"];
12    let mut index = corpus.searcher()?;
13    drop(device);
14    // The worker and this snapshot share vector allocations.
15    let snapshot = corpus.clone();
16    let worker = std::thread::spawn(move || -> hrxdb::Result<()> {
17        for neighbor in index.search(&[1.0, 0.0, 0.0], 2)? {
18            println!(
19                "{}: {:.6}",
20                external_ids[neighbor.id as usize], neighbor.similarity
21            );
22        }
23        Ok(())
24    });
25    assert_eq!(snapshot.len(), 3);
26    worker.join().expect("query worker panicked")
27}
Source

pub fn search_excluding( &mut self, query: &[f32], k: usize, excluded: &[u32], ) -> Result<Vec<Neighbor>>

Search while excluding insertion IDs, with device-side selection.

Returns up to k remaining rows, ordered as in Self::search. Excluded IDs may be unsorted and repeated. Exclusions apply only to this query; pass the growing visited set on each step of a similarity walk. Only the selected neighbors are read back, including when k exceeds 32.

if let Some(next) = db.search_excluding(query, 1, visited)?.first() {
    visited.push(next.id);
}
§Errors

In addition to Self::search’s requirements, all excluded IDs must be less than Self::len. Invalid input leaves the index usable.

Source

pub fn search_into( &mut self, query: &[f32], k: usize, output: &mut Vec<Neighbor>, ) -> Result<()>

Reuse a host output vector; successful calls replace its contents while retaining capacity. Invalid inputs leave it unchanged; runtime errors may leave partial output. This waits for GPU completion.

Source

pub fn search_excluding_into( &mut self, query: &[f32], k: usize, excluded: &[u32], output: &mut Vec<Neighbor>, ) -> Result<()>

Search with exclusions into reusable host output. See search_into.

Source

pub fn measure(&mut self, query: &[f32], k: usize) -> Result<Measurement>

Measure separate, completed control/scan/full-search invocations.

Uses host wall time rather than GPU timestamps. Control and scan exclude query preparation; full search includes it. Inputs and errors are the same as Self::search. No warmups are performed by this method.

Source

pub fn measure_excluding( &mut self, query: &[f32], k: usize, excluded: &[u32], ) -> Result<Measurement>

Measure a query with exclusions, using separate control/scan/search runs.

Timing has the same semantics as Self::measure; the complete search includes bitmap preparation and device masking. Call Self::reserve_search first to exclude scratch growth from the measurement. Inputs and errors are the same as Self::search_excluding.

Source

pub fn scores(&mut self, query: &[f32]) -> Result<Vec<f32>>

Return every cosine score in insertion-ID order for host-side selection.

Returns one FP32 score per insertion ID and allocates a full host score array. Use Self::scores_into to reuse a caller-owned buffer across queries. Both methods support larger result sets, exclusions, and grouped results by letting the caller select from the scores on the host. The query has the same validation requirements as Self::search.

Source

pub fn scores_into(&mut self, query: &[f32], output: &mut [f32]) -> Result<()>

Write every cosine score in insertion-ID order into a reusable buffer.

Transfers self.len() * 4 bytes from the device and blocks until the output is readable. No host score array or GPU buffer is allocated by this method. Selection, exclusions, and grouping are left to the caller; Self::search uses device selection for its supported k range.

§Errors

output.len() must equal Self::len, including for an empty index. The query has the same validation requirements as Self::search. Invalid inputs leave the output unchanged. Runtime failures propagate as errors and may leave the output partially written.

Trait Implementations§

Source§

impl Drop for Searcher

Source§

fn drop(&mut self)

Executes the destructor for this type. Read more
Source§

fn pin_drop(self: Pin<&mut Self>)

🔬This is a nightly-only experimental API. (pin_ergonomics)
Execute the destructor for this type, but different to Drop::drop, it requires self to be pinned. Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<T> DropFlavorWrapper<T> for T

Source§

type Flavor = MayDrop

The DropFlavor that wraps T into Self
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T, W> HasTypeWitness<W> for T
where W: MakeTypeWitness<Arg = T>, T: ?Sized,

Source§

const WITNESS: W = W::MAKE

A constant of the type witness
Source§

impl<T> Identity for T
where T: ?Sized,

Source§

const TYPE_EQ: TypeEq<T, <T as Identity>::Type> = TypeEq::NEW

Proof that Self is the same type as Self::Type, provides methods for casting between Self and Self::Type.
Source§

type Type = T

The same type as Self, used to emulate type equality bounds (T == U) with associated type equality constraints (T: Identity<Type = U>).
Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self> ⓘ

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self> ⓘ

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> IntoEither for T

Source§

fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ

Converts self into a Left variant of Either<Self, Self> if into_left is true. Converts self into a Right variant of Either<Self, Self> otherwise. Read more
Source§

fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
where F: FnOnce(&Self) -> bool,

Converts self into a Left variant of Either<Self, Self> if into_left(&self) returns true. Converts self into a Right variant of Either<Self, Self> otherwise. Read more
Source§

impl<T> PolicyExt for T
where T: ?Sized,

Source§

fn and<P, B, E>(self, other: P) -> And<T, P>
where T: Sized + Policy<B, E>, P: Policy<B, E>,

Create a new Policy that returns Action::Follow only if self and other return Action::Follow. Read more
Source§

fn or<P, B, E>(self, other: P) -> Or<T, P>
where T: Sized + Policy<B, E>, P: Policy<B, E>,

Create a new Policy that returns Action::Follow if either self or other returns Action::Follow. Read more
Source§

impl<T> Read<Exclusive, BecauseExclusive> for T
where T: ?Sized,

Source§

impl<T> Same for T

Source§

type Output = T

Should always be Self
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, !>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self> ⓘ
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self> ⓘ

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more