velesdb-core
The embedded tri-engine of VelesDB: vector, graph and columnar metadata in one Rust database.
Objective
Semantic retrieval usually means running a vector store, a graph database and a relational store side by side, then stitching their results together in application code — three deployments, three consistency stories, three query languages.
velesdb-core is the embedded engine that collapses those three into one
process and one language. Vectors (HNSW + SIMD), typed graph edges and typed
columnar metadata live in the same collection and are queried together with
VelesQL. No server, no network hop, no external dependency: it is a Rust
library that reads and writes a directory on your disk.
If you do not need retrieval over your own data, you do not need this crate.
Use cases
- A desktop or CLI application that must search its own documents offline, with no service to install and no data leaving the machine.
- A RAG pipeline that filters candidates on structured metadata (
tenant,date,status) in the same query as the vector search, instead of post-filtering results and losing recall. - A recommendation feature where "similar to this item" must be combined with "and connected to the user by at most 2 hops" — vector plus graph traversal in one statement.
- An AI agent that needs durable memory (facts, events, learned procedures) with TTL and snapshots, embedded in the agent process itself.
- An embedded/edge deployment where a 32x-compressed index must fit in RAM on constrained hardware.
Prerequisites
| Requirement | Minimum version | Note |
|---|---|---|
| Rust | 1.90 | Workspace MSRV, pinned in rust-toolchain.toml |
| Cargo | shipped with Rust | No other build tool required |
| Disk | writable directory | The persistence feature (on by default) memory-maps files there |
| Embeddings | any source | This crate does not compute embeddings — you supply the vectors |
| GPU | optional | Only for the gpu feature; falls back to SIMD when absent |
Installation
For WASM or any target without a filesystem, disable the default feature:
First success in 60 seconds
Create a project, add the two dependencies, paste this into src/main.rs, run
it.
&&
use json;
use ;
$ cargo run
id=1 title=rust score=1.0000
id=3 title=cargo score=0.9939
Success looks exactly like that: two lines, id=1 first with a cosine score of
1.0000 (the query vector is identical to point 1), id=3 second. Point 2 is
orthogonal to the query and is correctly excluded from the top 2.
Anything else is a failure — in particular, a second cargo run prints
Error: CollectionExists("documents") because the collection is already
persisted in ./veles-quickstart. Delete that directory to start over, or skip
create_collection when the collection already exists.
Configuration
Compile-time features (Cargo.toml):
| Feature | Default | Effect |
|---|---|---|
persistence |
on | mmap storage, WAL, rayon parallelism, tokio. Turn off for WASM. |
gpu |
off | wgpu compute pipeline for batch distance kernels; falls back to SIMD |
openapi |
off | utoipa::ToSchema derives on the api_types DTOs |
update-check |
off | HTTP client for automatic version checking |
internal-bench |
off | Exposes internal hooks used by some benches |
bench-sift1m |
off | SIFT1M benchmark. Links ureq/TLS as a regular dependency — never enable in a shipping build |
loom |
off | Loom concurrency testing (nightly only) |
test-fault-injection |
off | RAII guards forcing internal failures in tests. Never enable in production |
Runtime settings (HNSW parameters, limits, storage, logging) are read from
velesdb.toml — see the configuration guide.
Examples
examples/—crash_driver(crash-recovery test driver),profile_batch_insert(flamegraph target for HNSW batch insert),simd_precision_check(SIMD vs scalar validation). These are engine tooling, not tutorials.examples/rust/—multimodel_search.rs, a runnable vector + graph + metadata query.examples/mini_recommender/andexamples/ecommerce_recommendation/— complete standalone applications.
API / commands
Generated reference: docs.rs/velesdb-core. Import map (where each type lives): Core API map.
Task guides, all moved out of this README so it stays readable:
| Guide | What it covers |
|---|---|
| Collections, metrics, storage | Collection model, the 5 distance metrics, embedding dimensions, payload format, quantized storage modes, bulk ingestion, durability |
VelesQL reference |
Vector/text/hybrid queries, metadata filters, WITH options, operator table, JOIN limit, EXPLAIN |
| Sparse vectors and fusion | Named sparse indexes, DAAT MaxScore, RRF and Relative Score fusion |
| Streaming inserts | StreamIngester, backpressure, delta buffer (insert-and-search) |
| Query plan cache | Two-tier LRU cache, write-generation invalidation, EXPLAIN cache fields, metrics |
| Agent Memory SDK (Rust) | Semantic, episodic and procedural memory, TTL, eviction, snapshots |
| Core performance | Every published number, its measurement context, and how to reproduce it |
| Graph patterns · Multi-model queries | Graph modelling and cross-engine statements |
| Search modes · Tuning guide · Quantization | Recall/latency trade-offs |
| Write concurrency · Concurrency and locking | The write model and file locking |
Performance
Two headline numbers, both measured rather than estimated. Every figure, its hardware and its reproduction command live in Core performance.
| Claim | Measured | Context |
|---|---|---|
| Native HNSW search with AVX-512/AVX2/NEON SIMD | 450µs p50 end-to-end | 10K points, 384D, WAL on, recall ≥ 96% |
ColumnStore filtering vs. scanning JSON payloads |
up to 130x faster | integer equality, 100K rows |
Reproduce with cargo bench -p velesdb-core --bench hnsw_benchmark and
cargo bench -p velesdb-core --bench column_filter_benchmark.
Numbers move with hardware and dataset. Treat them as the shape of the engine's cost, not a guarantee for your workload — measure on yours.
Known limits
- No embedding generation. You bring the vectors; the crate never calls a model or the network to produce them.
- No clustering, sharding or replication.
velesdb-coreis a single-process embedded engine. One process at a time may open a database directory: a second one fails withDatabaseLocked. - One writer per collection. Concurrent readers are fine; concurrent writers to the same collection serialize — see Write concurrency.
- Metric and dimension are immutable. Changing either means creating a new collection and reindexing.
JOIN ... USING (...)supports one column only. Multi-columnUSINGparses but does not execute; useJOIN ... ON left = rightinstead.- No agent-memory service layer here. The explainable
MemoryService,why()and the deterministic context compiler (compile_context) live one level up invelesdb-memory, which depends on this crate — never the reverse. - WASM builds are read/compute only.
--no-default-featuresremoves mmap storage, WAL, rayon and tokio along with thepersistencefeature.
Compatibility
velesdb-core is a library, not an agent or MCP surface, so this table lists
the platforms and toolchains the project builds and tests on.
| Environment | Status | Note |
|---|---|---|
| Rust 1.90 (pinned) | Supported | rust-toolchain.toml; CI uses the same version |
Linux x86_64 |
Supported | CI: cargo check --workspace --all-targets --all-features |
| Linux aarch64 | Supported | CI: dedicated ARM64 benchmark runner (ubuntu-24.04-arm) |
Windows x86_64 (MSVC) |
Supported | CI: --all-features check on windows-latest |
macOS aarch64 / x86_64 |
Supported | Release pipeline builds both Darwin targets |
wasm32-unknown-unknown |
Supported, restricted | CI checks --no-default-features only; no filesystem persistence |
| Rust nightly | Build-checked | Only for the loom concurrency feature |
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
Error: CollectionExists("documents") |
create_collection re-run against an existing on-disk collection |
Delete the database directory, or call get_vector_collection first and only create when it returns None |
[VELES-004] Vector dimension mismatch: expected 4, got 3 |
The vector (on insert or query) does not match the dimension fixed at creation | Use your embedding model's exact output size; the dimension cannot be changed after creation |
[VELES-031] Database is already opened by another process: <path> |
A second process tried to open the same directory | Close the first process, or point the second one at another directory — one writer process per database |
get_vector_collection returns None for a name you created |
The name belongs to a graph or metadata-only collection | Use get_graph_collection / get_metadata_collection, or get_any_collection for the type-erased handle |
Data missing after a crash or kill -9 |
upsert updates in-memory/WAL state; destructors are best-effort |
Call flush() as your explicit commit boundary |
License
VelesDB Core License 1.0 — see LICENSE.
Last updated: 2026-07-25 · Applies to: velesdb-core 5.0.0 · Report a docs error