Recern Vector
A single-file vector database. Embedded, inspectable, boring in the best way.
Status: Phase 1 prototype (0.0.x). The file format and API will change between releases.
Project page: recern.net/vector · Benchmark report: recern.net/vector/benchmarks
Recern Vector stores vectors, JSON metadata and an HNSW index in one file — no server, no configuration. Its internals are part of the API: every query can explain how it was executed, every collection reports its graph structure and memory, and recall can be measured against exact search at any time.
Quick start (CLI)
Python
# saves on clean exit
=
=
Build and test from crates/recern-vector-py (see its README). One abi3 wheel covers CPython 3.11+. Calls release the GIL during search, batch inserts, recall estimation and saving.
Rust API
use ;
let mut db = open_or_create?;
let chunks = db.create_collection?;
chunks.upsert?;
chunks.upsert_many?; // atomic batch, index built on all cores
let options = default.filter;
let report = chunks.explain?; // hits + strategy, visited nodes, timing
let recall = chunks.estimate_recall?;
db.save?;
What is inside
| Part | Prototype implementation |
|---|---|
| Metrics | cosine (vectors normalized on insert), L2, dot product |
| Index | HNSW with the neighbor-selection heuristic; exact scan on request |
| Updates | upsert and delete by string id; deleted nodes stay as graph waypoints until compact |
| Batch build | upsert_many links records in parallel with one lock per node (as in hnswlib); nodes still being inserted are never used as descent points, and a final pass re-links any record search could not reach. Batches are atomic; with one thread the graph is identical to sequential upserts |
| Filters | MongoDB-style: equality, $in, $gt/$gte/$lt/$lte, dotted paths, AND |
| Filter planning | estimates selectivity from a sample; very selective filters scan matching records exactly, others filter during graph traversal |
| Introspection | stats() (layers, degree, unreachable records, memory), explain(), estimate_recall() |
| Storage | one file: header, collections, CRC32 footer; saved atomically via temp file + fsync + rename |
| Python | PyO3 bindings; float32 NumPy arrays are read through the buffer protocol |
The file layout is documented in storage.rs.
Benchmark
Standard ANN-Benchmarks datasets, top-10, one query at a time from Python (after a warm-up pass), Apple M3 Max (14 cores). Full method, sweeps and reproduction: bench/RESULTS.md.
SIFT1M (1M × 128, L2) — QPS at a recall@10 target:
| Engine | ≥ 0.90 | ≥ 0.95 | ≥ 0.99 | Build (threads) |
|---|---|---|---|---|
| Recern Vector HNSW | 15,973 | 9,041 | 4,409 | 35 s (14) |
| faiss HNSWFlat | 13,012 | 6,961 | 3,571 | 32 s (14) |
| LanceDB IVF_HNSW_SQ | 765 | 731 | — | 28 s (all) |
LanceDB IVF_PQ (num_sub_vectors=32) |
245 | 245 | 245 | 11 s (all) |
| sqlite-vec, exact scan | 12 | 12 | 12 | 5 s (1) |
GloVe-100 (1.18M × 100, cosine):
| Engine | ≥ 0.90 | ≥ 0.95 | Build (threads) |
|---|---|---|---|
| Recern Vector | 2,373 | 635 | 48 s (14) |
| faiss HNSWFlat | 1,687 | 356 | 43 s (14) |
| LanceDB IVF_HNSW_SQ | — (max 0.879) | — | 44 s (all) |
LanceDB IVF_PQ (num_sub_vectors=25) |
271 | 167 | 8 s (all) |
| sqlite-vec, exact scan | 7 | 7 | 4 s (1) |
Searches run on one thread for Recern Vector and faiss.
What this says about the prototype:
- Faster than the reference HNSW at equal recall: 1.2–1.3× on SIFT and 1.4–1.8× on GloVe. Recall per
efis slightly higher than faiss (denser layer-0 links), and search prefetches every neighbor's vector before computing distances, so memory loads overlap. - Build is within 1.1–1.2× of faiss on the same cores and scales ~10× on 14 cores.
- Single-query latency is 4–21× lower than LanceDB from Python. Against sqlite-vec's exact scan it is ~760× faster on SIFT and ~90× on GloVe at recall ≥ 0.95.
- The Python layer adds about 1 µs per query (
bench/profile_search.py).
cargo run --release --example bench -- [records] [dim] [noise] runs a quick synthetic sanity check without downloading datasets.
Prototype limitations
- The whole database is loaded into memory;
save()rewrites the entire file (no WAL yet). - Single writer; single upserts link sequentially (use
upsert_manyfor bulk loads). f32only; no quantization yet.- No Node.js bindings yet.
Next steps
- Tune filter planning: the exact-scan threshold is currently conservative
- WAL for incremental writes,
int8quantization - Flat layer-0 adjacency to cut pointer chasing further
Contributing
Issues and pull requests are welcome. Before sending a change, run:
License
Licensed under either of Apache License, Version 2.0 or MIT license at your option.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in this project by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.