recern-vector 0.2.0

A single-file, embedded, inspectable vector database
Documentation

Recern Vector

PyPI crates.io crates.io CLI CI License: MIT OR Apache-2.0

A single-file vector database. Embedded, inspectable, boring in the best way.

Version 0.2.0. Stable file format 2; reads existing format-1 files. WAL, int8, logical filters and native Node.js bindings are included.

pip install recern-vector           # Python
cargo add recern-vector             # Rust library
cargo install recern-vector-cli     # the recern-vector command

Release: v0.2.0 (PyPI, crates.io) · Project page: recern.net/vector · Benchmark report: recern.net/vector/benchmarks

Documentation: recern.net/vector/docs (sources in docs/) · Examples: examples/ — quickstart, semantic search, RAG retrieval, tuning recall, CLI and Rust

Recern Vector stores vectors, JSON metadata and an HNSW index in a local snapshot with a WAL for incremental commits. A checkpoint produces a standalone database file. Its internals are part of the API: every query can explain how it was executed, every collection reports its graph structure and memory, and recall can be measured against exact search at any time.

Quick start (CLI)

cargo build --release
alias rv=target/release/recern-vector

rv init docs.rvec
rv create-collection docs.rvec chunks --dim 384 --metric cosine
rv insert docs.rvec chunks chunks.jsonl     # {"id": "...", "vector": [...], "metadata": {...}} per line; --threads N

rv query docs.rvec chunks --like chunk-42 -k 5 --explain
rv query docs.rvec chunks --vector '[0.1, ...]' --filter '{"lang": "en", "year": {"$gte": 2020}}'

rv inspect docs.rvec                        # layers, degree, reachability, memory, hints
rv recall docs.rvec chunks --ef 16,32,64,128
rv delete docs.rvec chunks chunk-1 chunk-2
rv compact docs.rvec chunks

Python

import numpy as np
import recern_vector as rv

with rv.Database.open_or_create("docs.rvec") as db:          # saves on clean exit
    chunks = db.create_collection("chunks", dim=384, metric="cosine")
    chunks.upsert_many(ids, np.asarray(embeddings, dtype=np.float32), metadatas)
    hits = chunks.search(query, k=5, filter={"lang": "en"})
    print(chunks.explain(query, k=5), chunks.stats())

Build and test from crates/recern-vector-py (see its README). One abi3 wheel covers CPython 3.11+. Calls release the GIL during search, batch inserts, recall estimation and saving.

Rust API

use recern_vector::{CollectionConfig, Database, Filter, Metric, RecallOptions, SearchOptions};

let mut db = Database::open_or_create("docs.rvec")?;
let chunks = db.create_collection("chunks", CollectionConfig::new(384, Metric::Cosine))?;
chunks.upsert("chunk-1", &embedding, Some(serde_json::json!({"lang": "en"})))?;
chunks.upsert_many(records)?;                          // atomic batch, index built on all cores

let options = SearchOptions::default().filter(Filter::eq("lang", "en"));
let report = chunks.explain(&query, 10, &options)?;   // hits + strategy, visited nodes, timing
let recall = chunks.estimate_recall(&RecallOptions::default())?;
db.save()?;

What is inside

Part Implementation
Metrics cosine (vectors normalized on insert), L2, dot product
Index HNSW with the neighbor-selection heuristic; exact scan on request
Updates upsert and delete by string id; deleted nodes stay as graph waypoints until compact
Batch build upsert_many links records in parallel with one lock per node (as in hnswlib); nodes still being inserted are never used as descent points, and a final pass re-links any record search could not reach. Batches are atomic; with one thread the graph is identical to sequential upserts
Filters equality, $in, $gt/$gte/$lt/$lte, dotted paths, $and, $or, $not
Filter planning sampled selectivity and operation-cost estimates choose exact scan or HNSW with an adaptive candidate budget
Introspection stats() (layers, degree, unreachable records, memory), explain(), estimate_recall()
Storage stable format 2, format-1 compatibility, checksummed page WAL, atomic checkpoints, stale-writer protection
Precision f32 or per-vector int8 quantization; distance calculations operate directly on stored codes
Python PyO3 bindings; float32 NumPy arrays are read through the buffer protocol
Node.js native Node-API bindings, Float32Array input and asynchronous search

The file layout is documented in storage.rs.

Benchmark

Standard ANN-Benchmarks datasets, top-10, one query at a time from Python (after a warm-up pass), Apple M3 Max (14 cores). Full method, sweeps and reproduction: bench/RESULTS.md.

SIFT1M (1M × 128, L2) — QPS at a recall@10 target:

Engine ≥ 0.90 ≥ 0.95 ≥ 0.99 Build (threads)
Recern Vector HNSW 15,973 9,041 4,409 35 s (14)
faiss HNSWFlat 13,012 6,961 3,571 32 s (14)
LanceDB IVF_HNSW_SQ 765 731 — 28 s (all)
LanceDB IVF_PQ (num_sub_vectors=32) 245 245 245 11 s (all)
sqlite-vec, exact scan 12 12 12 5 s (1)

GloVe-100 (1.18M × 100, cosine):

Engine ≥ 0.90 ≥ 0.95 Build (threads)
Recern Vector 2,373 635 48 s (14)
faiss HNSWFlat 1,687 356 43 s (14)
LanceDB IVF_HNSW_SQ — (max 0.879) — 44 s (all)
LanceDB IVF_PQ (num_sub_vectors=25) 271 167 8 s (all)
sqlite-vec, exact scan 7 7 4 s (1)

Searches run on one thread for Recern Vector and faiss.

What this says about the prototype:

  • Faster than the reference HNSW at equal recall: 1.2–1.3× on SIFT and 1.4–1.8× on GloVe. Recall per ef is slightly higher than faiss (denser layer-0 links), and search prefetches every neighbor's vector before computing distances, so memory loads overlap.
  • Build is within 1.1–1.2× of faiss on the same cores and scales ~10× on 14 cores.
  • Single-query latency is 4–21× lower than LanceDB from Python. Against sqlite-vec's exact scan it is ~760× faster on SIFT and ~90× on GloVe at recall ≥ 0.95.
  • The Python layer adds about 1 µs per query (bench/profile_search.py).

cargo run --release --example bench -- [records] [dim] [noise] runs a quick synthetic sanity check without downloading datasets.

Version 0.2.0

  • Versioned format 2, frozen compatibility fixtures and format-1 migration.
  • Checksummed WAL with changed-page writes, crash recovery, checkpoint and stale-writer protection.
  • Optional Quantization::Int8 / Python quantization="int8" per collection.
  • $and, $or, $not filters with cost-based exact/HNSW selection and adaptive candidate budgets.
  • Native Node.js bindings, including asynchronous search.

A database now uses a snapshot plus WAL. Call checkpoint() before copying a standalone .rvec. save() makes pending changes durable; it does not checkpoint. See the format and durability contract and current limitations. Historical benchmark results above describe the pre-0.2 f32 engine and have not been rerun for int8.

Contributing

Issues and pull requests are welcome. Before sending a change, run:

cargo fmt --all --check
cargo clippy --all-targets -- -D warnings
cargo test

License

Licensed under either of Apache License, Version 2.0 or MIT license at your option.

Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in this project by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.