Recern Vector
A single-file vector database. Embedded, inspectable, boring in the best way.
Version 0.3.0: portable Python API, embedded/Qdrant backends and checked, resumable migration. File format 2 and format-1 compatibility are unchanged.
Release: v0.3.0 (PyPI, crates.io) · Project page: recern.net/vector · Benchmark report: recern.net/vector/benchmarks
Documentation: recern.net/vector/docs (sources in docs/) · Examples: examples/ — quickstart, semantic search, RAG retrieval, tuning recall, CLI and Rust
Recern Vector stores vectors, JSON metadata and an HNSW index in a local snapshot with a WAL for incremental commits. A checkpoint produces a standalone database file. Its internals are part of the API: every query can explain how it was executed, every collection reports its graph structure and memory, and recall can be measured against exact search at any time.
Quick start (CLI)
Python
# saves on clean exit
=
=
Build and test from crates/recern-vector-py (see its README). One abi3 wheel covers CPython 3.11+. Calls release the GIL during search, batch inserts, recall estimation and saving.
Portable Python API (0.3.0)
Keep application data operations the same while changing backend configuration:
=
# client = Client.qdrant("https://vectors.example.com", api_key=key, namespace="app")
=
=
The portable contract defines IDs, metadata, filters, distances, errors and backend capabilities. Export/import transfers a consistent embedded snapshot to new embedded or Qdrant collections, with checksums, dry-run, resumable batches and read-back verification. Qdrant 1.19.2 is the tested external backend; it does not promise atomic remote batches or consistent portable exports.
Install or upgrade with pip install --upgrade recern-vector to use the portable API. Runnable examples: portable.py, migrate.py.
Rust API
use ;
let mut db = open_or_create?;
let chunks = db.create_collection?;
chunks.upsert?;
chunks.upsert_many?; // atomic batch, index built on all cores
let options = default.filter;
let report = chunks.explain?; // hits + strategy, visited nodes, timing
let recall = chunks.estimate_recall?;
db.save?;
What is inside
| Part | Implementation |
|---|---|
| Metrics | cosine (vectors normalized on insert), L2, dot product |
| Index | HNSW with the neighbor-selection heuristic; exact scan on request |
| Updates | upsert and delete by string id; deleted nodes stay as graph waypoints until compact |
| Batch build | upsert_many links records in parallel with one lock per node (as in hnswlib); nodes still being inserted are never used as descent points, and a final pass re-links any record search could not reach. Batches are atomic; with one thread the graph is identical to sequential upserts |
| Filters | equality, $in, $gt/$gte/$lt/$lte, dotted paths, $and, $or, $not |
| Filter planning | sampled selectivity and operation-cost estimates choose exact scan or HNSW with an adaptive candidate budget |
| Introspection | stats() (layers, degree, unreachable records, memory), explain(), estimate_recall() |
| Storage | stable format 2, format-1 compatibility, checksummed page WAL, atomic checkpoints, stale-writer protection |
| Precision | f32 or per-vector int8 quantization; distance calculations operate directly on stored codes |
| Python | PyO3 bindings; float32 NumPy arrays are read through the buffer protocol |
| Node.js | native Node-API bindings, Float32Array input and asynchronous search |
The file layout is documented in storage.rs.
Benchmark
Standard ANN-Benchmarks datasets, top-10, one query at a time from Python (after a warm-up pass), Apple M3 Max (14 cores). Full method, sweeps and reproduction: bench/RESULTS.md.
SIFT1M (1M × 128, L2) — QPS at a recall@10 target:
| Engine | ≥ 0.90 | ≥ 0.95 | ≥ 0.99 | Build (threads) |
|---|---|---|---|---|
| Recern Vector HNSW | 15,973 | 9,041 | 4,409 | 35 s (14) |
| faiss HNSWFlat | 13,012 | 6,961 | 3,571 | 32 s (14) |
| LanceDB IVF_HNSW_SQ | 765 | 731 | — | 28 s (all) |
LanceDB IVF_PQ (num_sub_vectors=32) |
245 | 245 | 245 | 11 s (all) |
| sqlite-vec, exact scan | 12 | 12 | 12 | 5 s (1) |
GloVe-100 (1.18M × 100, cosine):
| Engine | ≥ 0.90 | ≥ 0.95 | Build (threads) |
|---|---|---|---|
| Recern Vector | 2,373 | 635 | 48 s (14) |
| faiss HNSWFlat | 1,687 | 356 | 43 s (14) |
| LanceDB IVF_HNSW_SQ | — (max 0.879) | — | 44 s (all) |
LanceDB IVF_PQ (num_sub_vectors=25) |
271 | 167 | 8 s (all) |
| sqlite-vec, exact scan | 7 | 7 | 4 s (1) |
Searches run on one thread for Recern Vector and faiss.
What this says about the prototype:
- Faster than the reference HNSW at equal recall: 1.2–1.3× on SIFT and 1.4–1.8× on GloVe. Recall per
efis slightly higher than faiss (denser layer-0 links), and search prefetches every neighbor's vector before computing distances, so memory loads overlap. - Build is within 1.1–1.2× of faiss on the same cores and scales ~10× on 14 cores.
- Single-query latency is 4–21× lower than LanceDB from Python. Against sqlite-vec's exact scan it is ~760× faster on SIFT and ~90× on GloVe at recall ≥ 0.95.
- The Python layer adds about 1 µs per query (
bench/profile_search.py).
cargo run --release --example bench -- [records] [dim] [noise] runs a quick synthetic sanity check without downloading datasets.
Version 0.2.0
- Versioned format 2, frozen compatibility fixtures and format-1 migration.
- Checksummed WAL with changed-page writes, crash recovery, checkpoint and stale-writer protection.
- Optional
Quantization::Int8/ Pythonquantization="int8"per collection. $and,$or,$notfilters with cost-based exact/HNSW selection and adaptive candidate budgets.- Native Node.js bindings, including asynchronous search.
A database now uses a snapshot plus WAL. Call checkpoint() before copying a standalone .rvec. save() makes pending changes durable; it does not checkpoint. See the format and durability contract and current limitations. Historical benchmark results above describe the pre-0.2 f32 engine and have not been rerun for int8.
0.3.0 includes the first Python portable client and embedded-to-Qdrant migration path. 0.4.0 remains planned for a separate Recern server and portable Rust/Node clients; 0.5.0 for incremental writes, checkpoints, benchmarks and package distribution. See the API portability and migration roadmap, including the gates for 1.0.0.
Contributing
Issues and pull requests are welcome. Before sending a change, run:
License
Licensed under either of Apache License, Version 2.0 or MIT license at your option.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in this project by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.