hrxdb 0.3.1

A GPU-powered vector database for unified-memory systems, built on HRX and Loom
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
# hrxdb API and execution guide

For a project overview and quick start, see the [README](../README.md).

hrxdb is an embedded GPU-powered vector database for unified-memory systems,
built on HRX and Loom, with a Rust API through
[hrx-rs](https://github.com/zacharydenton/hrx-rs). It exhaustively
searches FP16 vectors using FP32 cosine scores and returns the best 1–1,024 matches,
with optional excluded IDs. Large corpora are sharded internally.
Execution currently targets Linux x86_64 with a gfx1151 GPU (AMD Strix Halo).
The library exposes resident storage, search, and selection for application-owned
pipelines. Persistence, updates, metadata filters, and server APIs are outside
the current scope.

## Quick start

Add the dependency to your application:

```toml
[dependencies]
hrxdb = "0.3.1"
```

```rust
use hrxdb::{Corpus, Device};

fn main() -> hrxdb::Result<()> {
    let device = Device::open(0)?;
    let corpus = Corpus::build(&device, 3, [
        [1.0, 0.0, 0.0],
        [0.0, 1.0, 0.0],
        [0.8, 0.2, 0.0],
    ])?;
    let mut db = corpus.searcher()?;
    let neighbors = db.search(&[1.0, 0.0, 0.0], 2)?;
    assert_eq!(neighbors[0].id, 0);
    Ok(())
}
```

`build` accepts an `ExactSizeIterator` of anything implementing `AsRef<[f32]>`,
including borrowed slices or generated rows. Ingestion converts and uploads in
approximately 4 MiB chunks; it does not retain a second host corpus. IDs are
zero-based insertion positions, not external IDs. Keep application IDs in a
parallel array and resolve only the returned matches:

```rust
let external_ids = ["document-a", "document-b", "document-c"];
for neighbor in neighbors {
    println!("{}: {}", external_ids[neighbor.id as usize], neighbor.similarity);
}
```

For a corpus already stored as FP16, use `Corpus::build_fp16`. Each row is an `AsRef<[u8]>` containing exactly
`dimensions * 2` little-endian IEEE 754 binary16 bytes, without padding:

```rust
let corpus = Corpus::build_fp16(&device, dimensions, corpus_bytes.chunks_exact(dimensions * 2))?;
```

The bytes are preserved, zero padding is added, and inverse norms are computed
in one pass. Rows need not be normalized. This avoids allocating FP32 rows and
rounding them back to FP16. Reject incomplete trailing rows when reading a file:
`chunks_exact` omits any remainder. Both constructors require an accurately
sized iterator and upload in bounded chunks.

For an application with an existing HRX `ModelContext` (for example, one shared
by its encoders), use the context constructors:

```rust
let corpus = Corpus::build_fp16_in(&context, dimensions, rows)?;
let search = corpus.prepare_search(&context, 1, 10, 1)?;
let result = search.submit(&encoder_output)?;
```

`build_in` provides the equivalent FP32 path. The corpus retains the context;
`corpus.context()` exposes it, and its clones and prepared searches keep the
runtime alive. Prepared searches reject another runtime even on the same GPU.
The context's budget, if configured, charges native corpus storage, bounded host
conversion buffers, upload staging and subsequent corpus-stream workspaces.
This preserves the native `shards()`, `stream()` and custom-kernel API: shard
buffers are not HRX coordinated `BufferView`s. Coordinated encoder tensors enter
search through `prepare_search`'s scoped GPU handoff. Standalone stream work still
uses the documented explicit synchronization rules.

A kernel can read native corpus buffers and write a context tensor in the same
dispatch. Use `context.runtime().graph().gpu_scoped`: declare the tensor binding
as `Access::Write`, capture a corpus clone and private stream, and bind the
corpus's read-only shard views alongside the tensor's scoped native view. Drain
the stream before returning and never retain a scoped view outside the callback.
Wrap the output with `context.tensor(desc, binding, completion)` to carry that
graph's producer dependency into downstream inference. Ordinary tracked graph
nodes cannot directly accept native corpus bindings. The
[context search example](../examples/context_search.rs) demonstrates this with
the gather kernel writing a context tensor, then using it as a resident query.

For cached loading, call
`Corpus::load_resident_fp16_in(&context, artifact, dimensions, rows)`. Keep the
`ResidencyManager` used by `RuntimeOptions::memory_budget` alive. The loader
recovers this manager, charges allocations individually, and isolates cache keys
by runtime, shape and immutable artifact identity. Cloned contexts reuse the
entry without consuming its row iterator. A different runtime gets a separate
entry. Leases, pins, exported corpus clones and searchers prevent eviction.
`lease.bytes()` is zero for this allocation-budgeted cache; use
`lease.memory_usage().total()` for corpus storage and manager statistics for the
shared ceiling. Unlike the device-based loader, this path acquires charges as
allocations are made rather than reserving the entire loader peak up front.
Failures roll back charges. See the runnable
[context search example](../examples/context_search.rs).

`Corpus` owns immutable storage and is `Clone + Send + Sync`. Cloning retains
those allocations without copying vectors. Each `corpus.searcher()` creates a
`Send` worker with its own ordered stream, kernels, and workspace; queries take
`&mut self`. Workers can run independently on different threads. Keep a corpus
clone as a snapshot, or replace the application's current corpus while existing
workers finish using the old one. The original `Device` may be dropped.

```rust
let mut first = corpus.searcher()?;
let mut second = corpus.searcher()?; // Shared vectors, independent scratch.
let snapshot = corpus.clone();      // No new GPU allocation.
```

`Searcher::new(corpus)` takes ownership of a handle. `Searcher::on_stream`
accepts an existing stream and `ScanConfig` for applications with an established
queue. `Searcher::build` and `build_fp16` are convenience constructors combining
storage and one worker; `Corpus::build` alone allocates no search workspace.
See the runnable [search example](../examples/search.rs).

`searcher.corpus()` returns `&Corpus`, which callers can clone to retain storage
independently of the worker. Independent streams do not guarantee concurrent
execution, more memory bandwidth, priority, or latency isolation.

The FP32 constructor normalizes each row, rounds it to FP16, and stores an FP32
inverse norm for the rounded row. Queries are normalized to FP32. Scores are
FP32 dot products corrected by that stored inverse norm. Search is exhaustive
over this **quantized representation**; FP16 rounding and FP32 arithmetic can
change rankings relative to the original vectors. Scores are not clamped.
For `build_fp16`, the quantized representation is the supplied values; preserving
them may produce different scores from normalizing FP32 rows before rounding.
Equal computed scores prefer the lower ID. Small indexes return `min(k, len)`
neighbors, and an empty index returns an empty vector for a valid query.

Zero-norm and nonfinite vectors, dimension mismatches, and k outside 1–1,024 are
errors. Logical dimensions are 1–16,384 and padded to a multiple of 128; the
primary 384-dimensional case needs no padding. Row count is limited to 2^30,
and each internal allocation contains at most **2^32 FP16 elements** (8 GiB
of vector storage), respecting the compiler's 32-bit element-index limit.
Shard capacity is rounded to 256 rows, including readable tail slack. This
permits 8,388,608 logical rows per shard at 512 dimensions or 11,184,640 at 384.
Both constructors automatically split larger corpora. For example, 8,942,135 ×
512 faces use two shards while retaining one index and global insertion IDs.
Scans write each shard's scores into its range of a shared score array; GPU
selection finds the best k across the entire corpus. `shard_count()` reports
the number of corpus allocations.
Available GPU memory and compiler resources may impose tighter limits for large
shapes. Hardware coverage includes dimensions 1, 3, 127, 129, 384, 512, and 769.

## Larger result sets and exclusions

Request up to 1,024 candidates and collapse them using application metadata.
For example, keep the highest-scoring photo from each album until a page is full:

```rust
db.reserve_search(1024)?; // Optional setup: reserve scratch before serving queries.
let candidates = db.search(query, 1024)?;
let mut seen_albums = std::collections::HashSet::new();
let page: Vec<_> = candidates.into_iter()
    .filter(|n| seen_albums.insert(album_ids[n.id as usize]))
    .take(60)
    .collect();
```

The page can contain fewer than 60 albums if the candidates do not cover that
many distinct albums. A greedy similarity walk can exclude its visited IDs
directly, without reading back the full score array:

```rust
let next = db.search_excluding(query, 1, &visited_ids)?;
if let Some(neighbor) = next.first() {
    visited_ids.push(neighbor.id);
}
```

`search_excluding(query, k, excluded)` accepts unsorted, repeated insertion IDs.
Out-of-range IDs are errors. It returns `min(k, remaining_rows)` matches, or an
empty vector when every row is excluded. Exclusions apply to one query and
compose with the full k range and internal sharding. Both search methods run
selection on the GPU and read back at most k score/ID pairs.

`search_into(query, k, &mut results)` and
`search_excluding_into(query, k, excluded, &mut results)` reuse the capacity of
a caller-owned `Vec<Neighbor>`. Their allocating counterparts are convenience
wrappers.

Searcher construction reserves single-query selection scratch for k=32. The first larger query grows
it to the next power of two, then reuses that capacity. `reserve_search(k)` lets
applications make that allocation during setup. Single-query searches compile no kernels.

## Batched queries

Use `search_batch` for up to **64 queries over the same corpus**, such as sixty
sampled album photos each requesting five neighbors:

```rust
db.reserve_batch(60, 5)?; // Optional: allocate and compile before serving requests.
let matches = db.search_batch(&queries, 5)?;
// queries is 60 * db.dimensions() FP32 components in row-major order.
// matches[i] contains the neighbors for query i.
```

`search_batch_excluding(&queries, k, &excluded_ids)` applies one shared exclusion
set to the batch. Both methods support k=1–1,024 and internal corpus sharding.
Exclusions may be unsorted and repeated, and affect only that call. Every query
is validated before GPU execution; incomplete rows, nonfinite or zero-norm
queries, more than 64 queries, and out-of-range excluded IDs are errors. Empty
batches return an empty vector; an empty index returns one empty result per
valid query. A one-query batch uses the single-query scan.

`search_batch_into` and `search_batch_excluding_into` accept a reusable
`Vec<Vec<Neighbor>>`, retaining capacities of rows that survive a batch resize.

The [matrix kernel](../kernels/batch_scan.loom) loads a tile of 64 corpus rows into workgroup memory and
reuses it across the queries. Corpus values stay FP16 in workgroup memory and
expand to FP32 in registers; normalized queries and accumulation stay FP32. A
compiler scheduling fence bounds operand lifetimes to one column at a time.
The increasing-component accumulation order differs from the single-query scan,
so very close scores can change rank.
Equal computed scores still prefer the lower insertion ID.

Scores are materialized for at most 262,144 corpus rows at a time. GPU selection
keeps each query's tile top-k, then merges it with that query's running top-k.
The full corpus × batch score matrix is never allocated, and only final results
are read back (2,400 bytes for sixty top-5 queries).

Batch workspace grows to accommodate the largest reserved query width and k,
and is reused. Widths round up to 8, 16, 32, or 64, with compiled kernels cached
per width. At width 64 the workspace uses approximately **66 MiB for k=5** or
**322 MiB for k=1,024**, plus query storage (96 KiB at dimension 384). Smaller
corpora need less. This storage is additional to the index and single-query
scratch; `batch_workspace_bytes()` reports the reserved buffer bytes. First use
can allocate and compile, so use `reserve_batch(query_count, k)` during setup
when first-request latency matters. `ScanConfig` tunes the single-query scan;
the matrix kernel has its own schedule.

On 6,909,092 × 384 generated rows, sixty top-5 queries took **80.7 ms** with
the optimized batch scan versus **114.7 ms** with its predecessor (1.42×),
with identical returned IDs and scores. Workspace remains 66.11 MiB. See the
[interleaved comparison and samples](https://github.com/zacharydenton/hrxdb/blob/main/results/README.md#batch-scan-optimization);
the earlier batch-versus-individual measurement is retained separately.

## Full-score queries

`scores` and `scores_into` are supported query entry points for larger result
sets beyond 1,024 or custom host-side ranking.
They return the same cosine scores used by `search`, in insertion-ID order.
`scores` allocates a `Vec<f32>`; `scores_into` writes into an existing slice of
exactly `db.len()` entries so it can be reused across queries:

```rust
let mut scores = vec![0.0; db.len()]; // Allocate once per worker, outside its query loop.
db.scores_into(query, &mut scores)?;
```

Readback transfers four bytes per row and blocks until the output is ready.
`scores_into` allocates no host score array or GPU buffer, but full score transfer
and host selection still cost more than device top-k. Score readback always
returns all rows, regardless of exclusions supplied to previous searches.

## Custom GPU computations

`db.corpus()` returns `&Corpus` over the existing GPU allocations.
It performs no copy, allocation, compilation, or GPU work. Applications can bind
their own Loom kernels to these vectors, together with their own side arrays
and output buffers. Add `hrx = { package = "hrx-rs", version = "0.7.0" }` to
the application's dependencies to use the matching runtime and compiler API.

```rust
let corpus = db.corpus();
let stream = corpus.stream()?; // A new stream on the index's device.
let compiler_options = hrx::loom::CompilerOptions {
    target: stream.target().clone(),
    ..Default::default()
};
for shard in corpus.shards() {
    let rows = shard.row_range(); // Global insertion IDs; exclusive end.
    let capacity = shard.capacity_rows(); // Readable rows, including tail slack.
    let vectors = shard.vectors(); // Borrowed hrx::View, local rows 0..capacity.
    let inverse_norms = shard.inverse_norms(); // Same local row numbering.
    // This caller-owned side array must also have enough readable tail space.
    let albums = album_buffer.try_slice(rows.start * 4, capacity * 4)?;
    // Bind vectors, inverse_norms, albums, and application-owned output.
    // Only local rows < rows.len() may contribute to the result.
}
```

`Corpus` exposes `stream`, `len`, `is_empty`, `dimensions`, `padded_dimensions`,
`capacity_range`, and `shards`. `CorpusShardView` exposes `row_range`,
`capacity_range`, `capacity_rows`, `vectors`, and `inverse_norms`.
Shard ranges cover all insertion IDs in order without gaps; an empty corpus
has no shards. Both bindings cover `capacity_rows()`, the logical shard length
rounded up to a multiple of 256. Each logical vector row contains little-endian
FP16 with zero dimension padding and a byte stride of `padded_dimensions() * 2`.
Each logical norm entry is one little-endian FP32 inverse norm. Multiply the dot
with a normalized query by this inverse norm for cosine similarity. These are
the stored values: `build` normalizes before FP16 conversion, whereas
`build_fp16` preserves its input values.

Rows between `row_range().len()` and `capacity_rows()` are readable slack with
**unspecified contents**, not indexed vectors. They have no insertion IDs.
A fixed 256-row kernel can read its final tile in full, but must mask slack by
local row index before reduction or selection. Do not rely on zero contents:
a zero score could beat every real negative cosine. Side-array loads must also
be in bounds; padding the ordinal array with a sentinel is one option. For other
tile sizes, check `rows.len().div_ceil(tile_rows) * tile_rows <= capacity_rows()`.
Each shard adds at most 255 rows of vector and norm storage; that capacity still
respects the 2^32-element limit. Built-in searches process only logical rows.

`shard.capacity_range()` gives the global row-coordinate range covered by its
readable bindings, including slack. `corpus.capacity_range()` gives `0..end`,
where `end` is the largest shard capacity end (zero for an empty corpus). Use
this extent to allocate global side arrays that cover full-tile loads. Interior
slack may overlap later shards' logical rows; summing capacities is not the
same calculation. Imported shards can reserve more slack than host ingestion.

`corpus.stream()` creates an independent, owned stream on the original device,
including after the original `Device` is dropped or a corpus clone moves to
another thread. Configure the compiler with `stream.target()` and reuse the
stream for repeated work. The stream can outlive the corpus, but borrowed
bindings require a live corpus handle.

`db.stream()` borrows the searcher's **existing ordered stream**. Submit custom
producers, device searches, and consumers there to preserve execution order.
For separate streams, record a producer event and call `wait_event` on the
consumer stream. No accessor implicitly waits for work on another queue.

Corpus uploads have completed when construction returns. Callers own compilation,
submission, side data, scratch, and completion; borrowing a view does not wait
for asynchronous GPU work. **Treat both corpus bindings as read-only.** HRX's
binding type does not enforce this: writes are unsupported and invalidate the
index's data invariants. Normal HRX dispatch safety requirements still apply.

The runnable [album example](../examples/album_scores.rs) binds an
[application-owned kernel](../examples/album_scores.loom) and an album ordinal per
corpus row. It computes the best cosine for each `(query row, album)` directly
on the device, merges maxima across shards, and passes the score matrix to
`TopK`. Only the selected albums are read back. Its custom kernel reads full
256-row tiles and masks slack, including when an interior shard's side-array
slack overlaps the next shard's logical rows. Missing albums produce negative
infinity. The example favors clarity over throughput; album grouping and
reduction remain application code. It needs a corpus and an ordered stream,
and drops the original device before calling the operation.

```sh
HRX_OFFLINE=1 cargo run --locked --release --example album_scores
```

## Gathering named rows on the device

Use `Corpus::gather_into` when a custom kernel needs two subsets of the stored
vectors, rather than a query against the whole corpus:

```rust
let left_ids = [7, 2, 7]; // Order and duplicates are preserved.
let right_ids = [9, 3];
let mut stream = corpus.stream()?;
let left = stream.allocate(left_ids.len() * corpus.dimensions() * 4)?;
let right = stream.allocate(right_ids.len() * corpus.dimensions() * 4)?;
let left_done = corpus.gather_into(&mut stream, &left_ids, left.binding())?;
let right_done = corpus.gather_into(&mut stream, &right_ids, right.binding())?;
// Bind left and right to an application GEMM on this stream.
// On another stream, wait for right_done, which also follows the first gather.
```

The output is a compact row-major FP32 matrix with exactly `dimensions()`
components per row. Each stored FP16 component is expanded and multiplied by
its stored inverse norm, so a downstream dot product computes cosine similarity
subject to FP32 rounding. There is no dimension or row padding. Pairwise dots
can round differently from built-in search. The destination must be four-byte
aligned and have enough bytes; an oversized binding's suffix is left untouched.

hrxdb resolves shards internally and reads only the selected rows. IDs and
output extents are checked before submission; invalid arguments do not modify
the output. Empty requests record an event without writing. Inputs are insertion
IDs, and the output must not alias corpus storage. First use compiles a kernel
shared by corpus clones. Later calls can reuse routing scratch through the
caller's stream pool. The operation uploads ID routing data, performs no vector
readback or host completion wait, and returns an HRX event. Order previous output
uses on the same stream, or establish event dependencies before overwriting it.

Gather is independent of search limits: it supports up to 2^30 requested rows
and 2^32 output components, subject to available memory. Public `MAX_BATCH`
(64 queries) and `MAX_K` (1,024 results) describe search and standalone `TopK`.
Applications can use these constants without duplicating the supported limits.

The runnable [subset comparison example](../examples/subset_scores.rs) gathers two
blocks, computes their dense score matrix with an application-owned kernel,
and selects winners with `TopK`. Only final matches reach the host. The example
kernel illustrates composition; it is not a tuned GEMM. Its selection IDs are
columns in the right-hand subset and must be mapped through `right_ids`.

```sh
HRX_OFFLINE=1 cargo run --locked --release --example subset_scores
```

## Device inputs, outputs, and iterative retrieval

See the runnable [device search example](../examples/device_search.rs).

`DeviceQueries` describes borrowed row-major FP32 input, including a row stride
in elements. `search_device` normalizes on the GPU and writes into reusable
`DeviceNeighbors`. The returned HRX event marks completion; submission performs
no result readback. Reserve anticipated shapes before latency-sensitive work:

```rust
use hrxdb::{DeviceExclusions, DeviceNeighbors, DeviceQueries};

let mut db = corpus.searcher()?;
db.reserve_device(1, 5)?;
let mut results = DeviceNeighbors::new(db.stream(), 1, 5)?;
let mut visited = DeviceExclusions::new(db.stream(), corpus.len())?;
// query_buffer contains dimensions FP32 values produced on db.stream().
let done = db.search_device(
    DeviceQueries::new(query_buffer.binding(), 1, corpus.dimensions(), corpus.dimensions())?,
    Some(visited.binding()),
    &mut results,
)?;
// Same-stream consumers can run immediately, without a host wait.
visited.insert_device(db.stream(), results.ids(), results.batch_size() * results.k())?;
// A separate stream must explicitly wait for the result producer.
let mut consumer = corpus.stream()?;
consumer.wait_event(&done)?;
let matches = results.read(&mut consumer)?;
```

`DeviceNeighbors` owns FP32 scores and u32 IDs in `[batch, k]` order, plus one
u32 count and status per query. These regions share one allocation and a single
host readback transfer. Slots after the valid count contain negative
infinity and `u32::MAX`. Status is 0 on success; a nonfinite or zero-norm device
query produces status 1 and zero matches. Host `read`/`read_into` reports invalid
queries as errors; kernels can inspect status directly. `read_into` retains its
host staging and surviving result-row capacities. GPU-only consumers allocate
no host readback staging. Do not overwrite inputs or outputs until prior uses
complete; views and mutable Rust borrows do not establish GPU completion.

Device normalization uses scaled FP32 reductions to handle extreme finite
magnitudes, including entirely subnormal rows. Its rounding can differ from
host FP64 normalization, and tiny relative components can underflow. Both paths
search exhaustively with FP32 scores; very close rankings can differ. Up to
64 queries and k=1–1,024 are supported. First use may allocate, compile, or wait
when replacing imported host workspace; fully reserved submissions do not wait
on the host.

`DeviceExclusions` retains one bit per insertion ID. `insert`/`remove` validate
host IDs and upload only those updates; `insert_device`/`remove_device` consume
u32 IDs without host transfers, ignoring out-of-range values including the
result sentinel. Atomic updates preserve duplicate IDs and shared bitmap words.
`clear` resets the set. Use the same stream, or events between streams. Device
search also accepts an application-owned bitmap with this layout.

## Standalone selection

`TopK` consumes a `ScoreBatch`: a contiguous row-major FP32 matrix with up to
64 queries and 2^30 candidate columns. It requires no corpus or cosine metric:

```rust
use hrxdb::{DeviceNeighbors, ScoreBatch, TopK};

let mut topk = TopK::new(&stream)?;
topk.reserve(&stream, query_count, album_count, 60)?;
let mut winners = DeviceNeighbors::new(&stream, query_count, 60)?;
// Run application scoring on this stream, writing score_buffer.
let done = topk.select(
    &mut stream,
    ScoreBatch::new(score_buffer.binding(), query_count, album_count)?,
    &mut winners,
)?;
```

IDs are score-column ordinals: album IDs in this example. Higher scores win;
equal scores prefer lower columns. NaN and negative infinity are absent;
positive infinity is valid. A row with fewer than k eligible values has a
smaller count. The scores remain in their original allocation. Large matrices
use bounded 262,144-column tiles and running GPU top-k; scratch does not grow
with the total column count beyond that tile size. Plans and buffers are reused.
Order a `TopK` worker's submissions on one stream, or establish completion
before reusing its scratch on another stream.

## Adopting resident vectors

`Corpus::from_device(&device, &mut stream, dimensions, shards)` consumes owned
FP16 `ResidentShard { vectors, rows }` allocations. This avoids copying a
resident corpus through host memory or into replacement vector buffers.

Each allocation uses `ceil(dimensions/128)*128` FP16 elements per row, with
capacity divisible by 256, enough room for its logical rows, and at most 2^32
FP16 elements. Logical rows must be finite and nonzero, with zero dimension
padding; tail row slack is unspecified. Shard order determines insertion IDs.
An empty shard list creates an empty corpus; individual shards cannot be empty.

The import kernel validates logical rows and computes inverse norms using FP32
accumulation. It reads back only a four-byte validation result and completes
the supplied stream before publishing the corpus. FP32 norm rounding can differ
from `build_fp16`'s host calculation. Producer writes on another stream require
an event dependency first. Ownership transfers even if validation fails.

## Memory accounting

`corpus.memory_usage()` separates logical FP16 vectors, dimension padding,
vector slack, logical FP32 norms, and norm slack. `db.memory_usage()` reports
that shared corpus separately from the worker's query, score, selection,
readback, exclusion, batch, and device-query allocations:

```rust
let shared_bytes = corpus.memory_usage().total(); // Count once across workers.
let worker_bytes = db.memory_usage().workspace.total();
```

Reports describe reserved buffer bytes, including page rounding of imported
host buffers. They exclude native allocation granularity, compiler/runtime
bookkeeping and staging, and caller-owned inputs/outputs. `DeviceNeighbors`,
`DeviceExclusions`, and `TopK` expose their own device-buffer byte totals.
Accounting does not synchronize or query the GPU.

## Kernel families

[`scan_family.loom`](../kernels/scan_family.loom) uses Loom configuration, dependent
vector/view types, template providers, and explicit loop unrolling. Rust embeds
the authored source and supplies specialization values; it generates no kernel
source strings. Cosine and checksum providers share the same memory schedule.

The tuned default uses 128 threads, wave32, two rows per wave, and four adjacent
FP16 components per lane per load. At dimension 384 each lane reads 12 components
of each row, reuses query values, accumulates in FP32, and reduces within its
wave. Corpus values go directly into registers, with no LDS staging. Loom's
multidimensional views use 64-bit byte pointers with a 32-bit element index;
tests place unique matches on both sides of 4 GiB and at the end of 7.68 GB.
The 8,942,135 × 512 test covers the 8 GiB shard boundary and corpus tail.

[`select_family.loom`](../kernels/select_family.loom) partitions the score array
into groups of 1,024. Each 256-thread group holds four candidates per thread,
repeatedly reduces to the maximum score and minimum matching ID, and invalidates
the winner. Loom workgroup reductions handle subgroup exchange, LDS publication,
and barriers. The same family selects from intermediate candidate pairs until
only the final top-k remains. Keeping each partition's top-k preserves the
global top-k. This path serves k=1–32. Invalid lanes use negative infinity and
a sentinel ID.

For k=33–1,024, [`sort_family.loom`](../kernels/sort_family.loom) bitonic-sorts
1,024-row tiles in 8 KiB of workgroup memory and keeps each tile's top k.
Pairwise merges rank each candidate by binary-searching the other sorted list,
then retain the best k. Every round halves the number of lists, including at
k=1,024. Both selection paths order equal scores by ascending insertion ID.

Excluded IDs populate a reusable, imported host bitmap (one bit per row).
[`mask.loom`](../kernels/mask.loom) marks excluded scores as negative infinity on
the GPU before selection. Search omits these entries from readback results.
The following scan overwrites the scores, so exclusions cannot leak into later
queries. Unmasked searches skip this dispatch entirely.

At 10M × 384, storage is 7.68 GB of vectors, 40 MB of inverse norms, 40 MB of
scores, a 1.25 MB exclusion bitmap, and about 5 MB of initial selection scratch
(about 160 MB at k=1,024). Scoring writes and initial selection reads
40 MB each, approximately 1% additional traffic relative to vectors. Corpus,
query, result, and scratch allocations are reused. Query/result host memory is
page-aligned and imported once into HRX, with explicit completion before host
access. Once selection capacity is reserved, searches allocate no new GPU buffers. Allocating convenience methods return fresh host vectors; `_into` methods
reuse caller-owned capacity.

## Benchmark

Rust 1.91+, Linux x86_64, an AMD kernel driver, KFD/render permissions, and
a gfx1151 GPU are required for execution. The pinned native HRX bundle needs
glibc 2.43 or newer (Ubuntu 26.04 is its distribution baseline); see the
[HRX installation guide](https://github.com/zacharydenton/hrx-rs/blob/v0.7.0/docs/GPU-NPU.md#native-setup).
Building and generating documentation do not initialize the GPU. The crate
uses the published `hrx-rs 0.6` dependency and needs no sibling checkout. HRX provisions its pinned runtime automatically; use
`HRX_OFFLINE=1` after provisioning to prohibit downloads.

```sh
cargo run --release --bin hrxdb-bench -- --output default.json
cargo run --release --bin hrxdb-bench -- --sweep --output sweep.json
# A smaller run:
cargo run --release --bin hrxdb-bench -- --rows 100000 --samples 10
# Larger result sets, with the first 100,000 IDs excluded:
cargo run --release --bin hrxdb-bench -- --k 1024 --exclude-first 100000 --output large-k.json
# Compare sixty top-5 queries with one shared-read batch:
cargo run --release --bin hrxdb-bench -- --rows 6900000 --batch 60 --k 5 --samples 10 --output batch.json
```

Defaults: 10M rows, 384 dimensions, k=10, three warmups and 30 measured queries.
`--k` accepts 1–1,024. `--exclude-first N` excludes IDs 0 through N−1 and must
leave at least one row. Full-search timings include bitmap preparation and
device masking. Scratch reservation happens before validation and timing.
`--batch B` compares individual queries with a tiled batch, alternating timing
order and checking results against individual GPU queries and CPU scores.
It records workspace bytes and counts rank differences within numerical
tolerance. The batch benchmark supports exclusions but cannot use `--sweep`.
`--sweep` compares all 16 combinations of 128/256 threads, 1/2/4/8 rows per wave,
and 2/4 components per load on the same corpus. It selects the minimum median
full-search latency, preferring fewer VGPRs for results within 1%. Optional
`llvm-readobj` diagnostics supply register/LDS/scratch usage; set
`HRXDB_LLVM_READOBJ` to override its executable. Without it, selection uses
latency alone and resource fields are null. Saved benchmark reports use cache-relative
artifact identifiers so they contain no user-specific filesystem paths.

Each sample measures separately completed read-control, scoring, and full-search
invocations. Full search includes query normalization/publication, scan, GPU
selection, and result readback. Compilation and ingestion are excluded. HRX has
no GPU timestamps, so all timings use host wall time with explicit completion.
Queries vary between samples; the default single-query mode uses no batching. Useful GB/s divides
**unpadded FP16 vector bytes** by elapsed time, using decimal GB. It excludes
padding, norms, query, scores, and candidate traffic from the numerator. The
read control consumes each FP16 value into a checksum; it is a matching kernel
baseline, not a hardware memory-controller counter.

Before timing each configuration, the benchmark compares GPU selection with a
CPU top-k over the entire score array. It also compares sampled dots, including
the tail and 4 GiB boundary, and all returned similarities with CPU calculations
over identically quantized rows. It separately reports original-vs-quantized
top-k overlap on a generated subset of up to 8,192 rows; that statistic is not a
quality estimate for real embedding datasets.

See [measured results](https://github.com/zacharydenton/hrxdb/blob/main/results/README.md). The 256 GB/s value is a stretch target;
the practical acceptance criterion is at least 90% of the matching read control.

## Validation

```sh
cargo test --release
cargo test --release -- --ignored --test-threads=1 --skip batch::bench
cargo clippy --all-targets -- -D warnings
```

Hardware tests cover all 16 schedules, CPU score/ranking comparisons, padded
dimensions, tails, empty/small indexes, invalid vectors, repeated queries, ties,
negative scores, all k values, and multi-level selection with exact synthetic
scores. FP16 ingestion tests cover byte preservation, norm correction, invalid
rows, chunk boundaries, and reusable score readback. CPU tests check shard
layouts without allocating large corpora. Hardware tests compare sharded and
single-allocation results and exercise exclusions, growing walks, all-excluded
queries, large k, and odd merge tails. The 10M-row test allocates approximately
7.8 GB and checks unique matches around the 4 GiB boundary and at the allocation
tail. The faces-shape test allocates about 9.4 GB and crosses the compiler's
2^32-element boundary. Run hardware tests
separately from performance measurements to avoid GPU contention.

Batch tests cover query widths, row-major input validation, shared exclusions,
FP16 extremes and padded dimensions, exact ties, running selection across tiles
and shards, workspace reuse, and ownership across threads. The faces-shape test
also runs sixty queries across the 8 GiB shard boundary with bounded workspace.

Corpus-binding tests read and check every exported value, padding component, and
inverse norm for both ingestion paths. They run the album example on a separate
stream across unaligned shard boundaries, compare against built-in cosine
scores, and verify subsequent searches still return the same results. Empty
corpora expose no shards; borrowed bindings remain tied to a live corpus handle.
Tiled tests cover 201- and 55-row tails, exact tile boundaries, and positive
slack values that would otherwise outrank all-negative corpus scores. CPU tests
check rounded shard capacities against the address limit at every padded dimension.
Stream tests drop the original device, move the index to another thread, verify
device identity and independent streams, run a custom kernel, and use a returned
stream after the index itself has been dropped.

Device pipeline tests cover independent shared-storage workers, snapshot
ownership, zero-copy FP16 import and rejection of invalid rows, strided inputs,
widths through 64, extreme magnitudes, device status, cross-stream event
consumption, incremental GPU exclusions, and standalone selection across large
matrix tiles. They verify bounded workspace and repeated output reuse.

Gather tests cover requested order and duplicates, shard boundaries, normalized
FP32 output, cached-kernel reuse, stream-owned scratch, output validation, and
the subset-comparison example.

The repository's [release guide](../RELEASE.md) covers packaging and publication.
See [CHANGELOG.md](../CHANGELOG.md) for version history. Contributions should include
appropriate CPU or opt-in GPU coverage; report bugs with the GPU target, driver,
crate version, and a minimal reproducer. Hardware tests do not run in ordinary CI.