Skip to main content

Crate hrxdb

Crate hrxdb 

Source
Expand description

A GPU-powered vector database for unified-memory systems, built on HRX and Loom.

HRX manages GPU buffers, streams, and execution; Loom kernels provide exhaustive cosine search and composable GPU scoring on AMD Strix Halo. The Rust API exposes FP16 corpus storage and FP32 scores.

A shared Corpus owns immutable storage. Each Searcher owns its stream and reusable workspace. TopK selects from application-defined GPU scores. Scores describe the quantized corpus, not the original FP32 rows.

Execution requires Linux x86_64, a gfx1151 AMD GPU, and the native runtime prerequisites in the hrx-rs documentation. Building and generating documentation do not initialize GPU hardware.

use hrxdb::{Corpus, Device};

let device = Device::open(0)?;
let corpus = Corpus::build(&device, 3, [
    [1.0, 0.0, 0.0],
    [0.0, 1.0, 0.0],
])?;
let mut index = corpus.searcher()?;
let external_ids = ["document-a", "document-b"];
for neighbor in index.search(&[1.0, 0.0, 0.0], 2)? {
    println!("{}: {}", external_ids[neighbor.id as usize], neighbor.similarity);
}

IDs are insertion positions, not application IDs. The index does not persist vectors, map external IDs, or support updates, metadata filters, or approximate search. Query-time exclusions use insertion IDs. Searcher::corpus provides borrowed storage bindings for application-owned GPU kernels, side arrays, and reductions; see Corpus for the contract.

Structs§

Compilation
Compiler artifacts and resource reports for reproducible tuning.
Corpus
Cheaply cloned, immutable GPU-resident FP16 corpus. Clones share allocations.
CorpusBuildMemory
Logical allocation extents used to reserve a corpus before loading it. Excludes compiler/driver memory, allocation granularity and caller-owned input mappings. Peak includes host conversion and native upload staging.
CorpusMemory
Bytes in one shared corpus allocation set; cloned handles do not duplicate it.
CorpusShardView
Borrowed bindings for a contiguous range of insertion IDs. Bindings include readable slack; only rows in row_range() are corpus data.
Device
A borrowed device registry entry. Opening a stream never acquires ownership of the native device, and dropping a model never shuts down HRX globally.
DeviceExclusions
Reusable device bitmap over insertion IDs. Updates are incremental and may consume top-k IDs directly, keeping a similarity walk entirely on the GPU. Use one ordered stream, or establish event dependencies between streams.
DeviceNeighbors
Caller-owned fixed-stride device results, reusable across submissions.
DeviceQueries
Borrowed FP32 queries in row-major device storage. Values are normalized on the GPU; nonfinite or zero-norm rows produce status=1 and zero matches. Stride is in FP32 elements. All accessed rows must be ready on the searcher’s stream (use an event for a producer on another stream).
Measurement
Host-completion measurements; these are not GPU timestamps.
Neighbor
A ranked match, ordered by descending score then ascending ID.
PreparedSearch
A fixed query shape with shared corpus storage and a bounded worker count.
ResidentShard
One owned resident FP16 allocation to adopt without copying its vectors. Rows use the corpus’s padded stride; capacity must be a multiple of 256. Logical rows must be finite, nonzero, and have zero dimension padding.
ScanConfig
Scan parameters. The benchmark can sweep all 16 supported configurations.
ScoreBatch
A borrowed, contiguous row-major FP32 score matrix. Rows are independent queries; columns are candidate IDs. NaN and negative infinity are excluded.
SearchInference
Resident search results; only each row’s counts entries are valid.
Searcher
Independent search worker over a shared immutable Corpus.
SearcherMemory
Shared corpus bytes and private workspace bytes for one searcher.
TopK
Reusable GPU top-k over application-defined score matrices, independent of a corpus. Queues work on caller-owned streams of one device without host readback. Ties prefer lower column IDs. NaN and -infinity are absent; +infinity is valid. Large matrices are selected in bounded 262,144-column tiles with running GPU top-k. Scores are read in place, including when a matrix exceeds 4 GiB.
WorkspaceMemory
Private buffer bytes owned by one searcher, excluding shared storage and caller-owned device outputs. Host-imported buffers include page rounding.

Enums§

Error
A runtime, validation, provisioning or compiler failure.

Constants§

MAX_BATCH
Largest supported query count for batched search and standalone selection. This does not limit the number of rows requested by Corpus::gather_into.
MAX_K
Largest supported top-k result count for search and standalone selection.

Type Aliases§

Result
The result of an HRX operation.