hrxdb
A GPU-powered vector database for unified-memory systems, built on HRX and Loom.
hrxdb is an embedded vector database designed for unified-memory systems like AMD Strix Halo. HRX manages GPU buffers, streams, and execution; Loom kernels implement exact cosine search and top-k. Your own Loom kernels can work directly with the same resident corpus. Gather named rows, compute application-specific scores, and select top-k without copying intermediate vectors or score matrices back to the CPU.
The library exposes a Rust API through
hrx-rs. Execution currently targets
AMD Strix Halo (gfx1151) on Linux x86_64. Rust 1.91+ is required; other GPUs
and operating systems are not supported for execution in this release.
Unified memory makes large resident collections possible without a separate discrete GPU memory pool. hrxdb builds around that model: ingest once, reuse the corpus across searches and custom scoring, and read back only what the application needs. Host ingestion still copies into owned runtime buffers; unified memory does not make every operation zero-copy or remove bandwidth limits.
- Exact cosine search: exhaustive FP16 corpus scans, FP32 scores, deterministic ties, exclusions, and top-k up to 1,024.
- Shared-read batches: up to 64 queries per batch with bounded score tiles and running GPU top-k.
- Composable storage: cheap-clone
Corpushandles, independentSearcherworkspaces, read-only shard bindings, and device-resident row gathering. - Custom GPU pipelines: standalone
TopK, device query/result buffers, completion events, and persistent exclusion bitmaps. - Large resident corpora: automatic sharding, direct FP16 ingestion, validated zero-copy adoption of device buffers, and memory accounting.
hrxdb is an embedded vector index and GPU computation library. Persistence, embedding generation, application metadata, updates, approximate search, and server APIs belong to the application. Search is exact over the stored quantized representation; FP16 rounding and FP32 arithmetic can change rankings relative to the original vectors.
Quick start
Add the dependency:
[]
= "0.3.1"
use ;
IDs are zero-based insertion positions. Keep external IDs and metadata in
parallel application-owned storage. Corpus::build accepts a sized iterator
of FP32 rows, normalizes and quantizes them, and uploads in bounded chunks.
Corpus::build_fp16 preserves supplied little-endian FP16 values and computes
their inverse norms. Finite, nonzero rows are required.
When an encoder already has an HRX ModelContext, use
Corpus::build_fp16_in(&context, dimensions, rows) (or build_in for FP32).
The corpus retains that context and uses its GPU and memory budget for native
storage, upload staging and subsequent workspaces. corpus.context() returns
the retained context. Prepare search with that context to submit encoder
tensors directly. Existing device-based constructors remain available.
From a checkout, run the search example:
Building does not initialize GPU hardware or download native code. First GPU
or compiler use provisions the native bundle pinned by the resolved HRX release.
Execution needs the AMD kernel driver, /dev/kfd and render-device permissions,
compatible C/C++ runtime libraries and libatomic, and glibc 2.43+
(the bundle's baseline is Ubuntu 26.04). See the
HRX setup guide.
After provisioning, HRX_OFFLINE=1 prevents runtime downloads.
Choose your operation
| Task | API |
|---|---|
| Search one query or a batch | Searcher::search, search_batch |
| Reuse host result capacity | search_into, search_batch_into, scores_into |
| Share one corpus across workers | Corpus::clone, Corpus::searcher |
| Gather selected rows for a custom kernel | Corpus::gather_into |
| Bind resident vectors without a copy | Corpus::shards |
| Adopt owned FP16 device allocations | Corpus::from_device |
| Search with GPU-produced queries | Searcher::search_device, DeviceQueries |
| Rank application-defined GPU scores | TopK, ScoreBatch |
| Keep results and exclusions on the GPU | DeviceNeighbors, DeviceExclusions |
| Account for shared storage and workspace | Corpus::memory_usage, Searcher::memory_usage |
| Reserve loader peaks and pin a shared corpus | Corpus::build_memory, Corpus::load_resident_fp16 |
| Load and cache within an encoder's context | Corpus::build_fp16_in, Corpus::load_resident_fp16_in |
The optional residency API uses an explicit HRX ResidencyManager. It reserves
corpus storage and bounded conversion/upload staging before allocating, rolls
back failed loads, and shares storage for the same immutable artifact/device/
shape key. A lease, lease.pin(), exported corpus clone or live searcher prevents
eviction. Attach manager.budget() with Corpus::with_workspace_budget to charge
subsequent searcher, batch and custom scoring allocations on corpus streams to
the same ceiling. Coordinated prepare_search uses its context's budget for
private workers. Caller mappings and compiler/driver memory remain separate;
this is not a process-wide memory ceiling.
Corpus::load_resident_fp16_in(&context, artifact, dimensions, rows) uses the
live residency manager attached to the context's memory budget. Its cache keys
include runtime identity, and individual allocations carry their charges, so
corpus bytes are not reserved twice. Native shard bindings retain their existing
stream-ordering contract; they are not coordinated BufferViews. See the
shared-context example.
Corpus is Clone + Send + Sync; clones share immutable allocations. Each
Searcher owns its stream and workspace. Independent workers allow independent
submission, but do not guarantee overlapping execution, higher throughput,
priority, or latency isolation on a saturated GPU. Use batching to share corpus
reads across queries. Use events to order work between streams.
Public MAX_BATCH and MAX_K expose the supported query and result ceilings.
Logical dimensions are 1–16,384; corpora support up to 230 rows, subject to
available memory. Internal shards contain at most 232 FP16 elements and include
readable row slack for custom tiled kernels. Only logical rows participate in
search; custom kernels must mask slack.
See the API and execution guide for buffer layouts,
normalization, exclusions, memory limits, and synchronization contracts.
See the API reference, or generate it locally
with cargo doc --locked --no-deps --open.
Build your own scoring pipeline
resident corpus → gather named subsets → custom scoring kernel → GPU top-k
↓
selected IDs and scores
gather_into writes compact normalized FP32 rows, preserving requested order
and duplicates while resolving shards internally. TopK consumes your own
row-major score matrix and returns score-column IDs. Album grouping, custom
reductions, and metadata stay in application code.
Runnable examples:
- Subset comparisons: gather two blocks, score them on the GPU, and read back only the winners.
- Album ranking: bind corpus vectors and an album side array, reduce scores by album, then select top-k.
- Device search: queue queries and update visited IDs on the GPU, with event-ordered result consumption.
The custom scoring kernels illustrate interoperability; they are not tuned GEMMs.
Applications using HRX types directly should also depend on
hrx = { package = "hrx-rs", version = "0.7.0" }.
Measured performance
Recorded on local gfx1151 with generated data, 2026-09-11:
| Workload | Median host-completion latency |
|---|---|
| One top-10 query over 10M × 384 vectors | 33.7 ms |
| Sixty top-5 queries over 6,909,092 × 384 vectors | 80.7 ms |
These are separate runs, not a batching speedup comparison. The sixty-query run was 1.42× faster than the preceding batch implementation in an interleaved comparison with identical returned IDs and score bits. Smaller shapes can be slower. Timings exclude ingestion and compilation; clocks and other system activity were not controlled. They are not embedding-quality measurements or guarantees for a particular collection.
The results and raw samples include methodology, memory use, compiler reports, and reproduction commands. To benchmark locally:
Development and release
CPU tests and documentation work without GPU hardware:
On a prepared gfx1151 host, run hardware correctness tests separately from performance benchmarks:
HRX_OFFLINE=1
The largest test allocates about 9.4 GB. See CONTRIBUTING.md for setup, validation, and reporting issues; CHANGELOG.md for the release contents; and RELEASE.md for publication steps.
MIT licensed. The separately distributed HRX runtime and compiler have their own third-party notices.