velesdb-memory
The explainable, local-first memory engine for AI agents, as a single MCP server.
Portability: ✅ the server and all of its tools work in any MCP client over stdio · ⚙️ sharing one store between several clients at once requires the HTTP transport (
--features http) · ⚠️ the.mcpbbundles are a one-click packaging format (Claude Desktop and registry-aware clients), not a transport, and are built stdio-only · ⚠️ the automatic agent hooks are harness-specific, and tool-result replacement (updatedToolOutput) is Claude Code only.
Objective
A coding agent forgets everything between sessions, and a vector store only gives it back text that looks like the question. Neither can answer "why did we do this?", because the answer is usually a fact that shares no words with the question — the ticket behind the decision, the incident behind the constant.
velesdb-memory gives an agent durable memory that never leaves the machine: it remembers facts, recalls them semantically, connects them with typed links, and walks those links to return the evidence trail behind an answer. It also ships a deterministic context compiler that shrinks an agent's prompt under a hard token budget with no model call at all.
What you actually gain
Two problems cost you real money and real quality every day, and this fixes both.
Your agent forgets. Close the session, and everything it learned about your codebase is gone. Tomorrow you explain it again. velesdb-memory keeps those facts on your disk and hands them back at the start of the next session.
Every turn re-sends the whole conversation. That is what you are billed for, and a context stuffed with repeated logs is also a context where the model pays less attention to what matters. The compiler shrinks that payload before it is sent — deterministically, with no AI call of its own.
| What improves | Measured | How it was measured |
|---|---|---|
| Context sent to the model | 82.5 % smaller on a 12-turn coding session (80–87 % per turn as it grows) | committed corpus, real cl100k tokenizer — recompiled twice per run, byte-identical |
| Your actual bill | 10.9 % to 21.9 % saved on the same session, A/B, real billing | billed campaign |
| Cost of storing a memory | zero AI calls — nothing is sent anywhere | the write path never calls a model |
| A 55 KB build log entering context | 767 characters, error and file:line kept | the PostToolUse hook |
The honest part: those percentages come from our corpus on our sessions. The spread is published as prominently as the best figure, and every number above is pinned to its source by a contract the CI enforces — if a figure drifts from what the code produces, the build goes red.
How it works, in four steps
Nothing here needs an AI provider. Everything runs on your machine.
1. It stores facts, not transcripts. You (or your agent) call remember
with one fact: "the API port is 6333 because 3000 collided with the web UI".
It is written to a local file store. No model call, no network.
2. It finds them by meaning, not keywords. recall matches on sense, so
asking about "which port did we settle on" finds that fact even though the
words differ. This uses a local embedding model of your choosing — and the
honest caveat is that until you pick one, the out-of-the-box default is a
deterministic offline embedder that matches surface form, not meaning
(great for zero-dependency demos, wrong for real recall). Picking a real
model is one env var, no rebuild: two commands, below.
3. It connects them, which is the part that matters. Facts are linked to
the topics they mention. why starts from the best match and then walks those
links, so it returns the answer plus the facts that explain it — including
ones sharing no words with your question. A plain search cannot do that.
Those links have to exist. If you only ever call
remember, the graph stays flat andwhybehaves like a search. Pointremember_extractedat a paragraph and it splits it into facts and wires the links for you.
4. It compresses what is too big, at the right moment. Five agent
hooks fire automatically in Claude
Code: two remind it to save and reload its state around a session, one does the
same before a compaction, PreToolUse requires successful recall before an
opted-in repository edit, and PostToolUse is the one that replaces an
oversized Bash result with a schema-compatible compiled view — so the bulky
text never enters the conversation when Claude accepts the replacement.
Nothing is deleted: the complete original Bash output object is serialized as
JSON and its path is quoted in the replacement, so the agent can read the full
thing whenever the compiled view is not enough.
Use cases
- A coding agent that must still know, three weeks later, why a timeout is set to 8 s — and can show the PR and the incident it came from.
- Regulated or air-gapped work where context cannot transit a third-party LLM API, and "show why it recalled that" has to be answerable.
- Long agent sessions that hit the context window: compile the prompt instead of summarizing and restarting.
Prerequisites
| Requirement | Minimum version | Note |
|---|---|---|
| Rust | 1.90 | Only to install or build. The binary itself has no runtime dependency. |
| An MCP client | — | Claude Code, Claude Desktop, Codex CLI, Cursor, Cline, Zed, opencode, Windsurf, Devin CLI. |
| Ollama, or any OpenAI-compatible server | any | Optional — only for real semantic recall and for model-based extraction. Both backends are compiled into the default binary: enabling one is a runtime env-var switch (VELESDB_MEMORY_EMBEDDER), never a rebuild. The default embedder is offline and dependency-free. openai names a protocol, not a vendor: oMLX, llama.cpp, LM Studio, vLLM and hosted providers all speak it, and each is reached by URL rather than by a backend name of its own. See MCP_SERVER_SETUP.md. |
| Node.js | any LTS | Optional — only for the Claude Desktop stdio→HTTPS bridge. |
Installation
No Rust toolchain? velesdb-memory is on the
MCP registry as
io.github.cyberlife-coder/velesdb-memory, with prebuilt .mcpb bundles on
each release. Full
options: MCP server setup.
First success in 60 seconds
Wire it into Claude Code:
Then ask your agent to store a decision and recall it. These are the MCP calls it makes, and exactly what comes back:
remember { "fact": "we chose parking_lot to avoid lock poisoning",
"metadata": { "project": "checkout" } }
→ { "id": 9876543210, "id_str": "9876543210" }
recall { "query": "locking strategy", "limit": 5 }
→ { "memories": [ { "id": 9876543210, "id_str": "9876543210",
"score": 0.59,
"content": "we chose parking_lot to avoid lock poisoning",
"metadata": { "project": "checkout", "_veles_date": 20260725 } } ] }
A non-empty memories array means the server is wired and the store is
writable. (_veles_date is stamped automatically — see
automatic dating.)
Every other client — Cursor, Zed, Codex CLI, Claude Desktop, Windsurf, Devin CLI — is one config block away in MCP server setup.
Real semantic recall in 5 minutes
The 60-second setup above runs the offline hash embedder: deterministic,
zero-dependency — and lexical, not semantic. Both semantic backends are
compiled into the binary you just installed, so upgrading is configuration,
not a rebuild. With Ollama installed, the recommended
model is bge-m3 (multilingual, 1024-dim):
No Ollama? Any OpenAI-compatible server (oMLX, llama.cpp, LM Studio, vLLM)
works with VELESDB_MEMORY_EMBEDDER=openai and a URL — the model still runs
on your machine, and memory still never leaves it. All options:
embedding backend.
Two things to know when switching:
- Your agent can check. The
memory_statustool reports which embedder actually runs and whether recall is semantic — ask "call memory_status" and readembedder.semantic. The server also flags a degraded (hash) run in its instructions to every connecting client. - Your memories survive the switch. The store records which model filled
it; on a mismatch the server refuses to serve nonsense and names the
migration command (
velesdb-memory migrate-embeddings, dry-run first), which re-embeds every fact under its original id. Switching embedders never costs you your memories.
See the wedge (offline, one command)

recall("why we chose parking_lot") [vector similarity only]
0.47 we chose parking_lot to avoid lock poisoning after a panic
0.18 PR #42 swaps the std Mutex for parking_lot
└─ EPIC-317 is nowhere here — it shares no words with the question.
why("why we chose parking_lot") [vector seed + graph traversal]
hop 0 we chose parking_lot ...
hop 1 PR #42 ...
hop 2 EPIC-317: intermittent CI hang under load
└─ the graph reached the very ticket the decision fixed.
A vector search ranks by resemblance, so it is blind to the ticket. why()
follows the typed links and reaches it. That gap is the product.
| Runnable demo | What it shows |
|---|---|
why_across_sessions.py |
the reason survives a process restart — recall of the top 5 of 16 memories stays blind, why() reaches it |
why_magic_constant.py |
why a magic constant has its value — a business reason sharing no words with the code |
memory_builds_its_own_graph.py |
paste raw prose → a local model auto-wires the graph (no relate()), why() walks it to the root cause |
why_magic_constant.mjs |
the same engine and wedge in the Node binding |
Not a weak-embedder trick. In each retrieval demo, recall stays blind to the reason even under a real semantic embedder (
ollama/all-minilm), not just the offlinehashdefault.
The graph's contribution, isolated
cargo run --release -p velesdb-memory --example bench_multihop runs 24
decision → PR → problem chains with the same embedder throughout and only the
graph toggled. Each question ("why did we adopt <tech>") has a 1-hop answer
(the decision, which shares words) and a 2-hop answer (the original problem,
which shares none):
| embedder | direct recall | multi-hop, vector-only | multi-hop, vector + graph |
|---|---|---|---|
hash (deterministic) |
100% | 0% | 100% |
real model (Ollama all-minilm) |
100% | 33% | 100% |
The direct control confirms the vector engine is healthy — it aces look-alike retrieval. On multi-hop, a real semantic embedder still recovers only a third of the answers; the graph recovers all of them, +67 pp with a real model. Run that arm yourself:
&&
VELESDB_MEMORY_EMBEDDER=ollama \
bench_multihop measures the engine's contribution on controlled data with
the graph pre-wired, so the numbers reflect retrieval, not an LLM. The
end-to-end extraction comparison on the real
LoCoMo dataset lives in
examples/locomo/.
What the server exposes
22 MCP tools in the default build, in three families:
| Family | Tools |
|---|---|
| Durable memory | remember, recall, recall_where, recall_fused, relate, unrelate, forget, entity, why, feedback, remember_extracted, memory_status, list_memories |
| Context compiler | compile_context, compile_transcript, explain_compilation, retrieve_context_source, context_savings, suggest_budget |
| Session resumption | save_working_context, load_working_context, list_working_contexts |
Parameters, returns, and error codes for every one: MCP tool reference.
By design the server exposes memory semantics only — never raw database
capabilities (query, create_collection, upsert, traverse).
Where to go next
| Guide | What it covers |
|---|---|
| MCP server setup | every client config, the shared HTTPS daemon, the local CA, Windows, embedding and extraction backends |
| MCP tool reference | one section per tool: parameters, returns, limits, error model |
| Context compiler | budgets, preservation rules, risk, retrieval handles, media, path ingestion, transcripts, the compile-stdin CLI and the PostToolUse hook |
| Agent Memory SDK | the other path: the embedded, language-native AgentMemory API |
| Migrating embedding models | migrate-embeddings end to end: regimes, the journal, crash recovery, the switch, and what it costs |
BENCHMARK.md |
every published retrieval number, its method, and how to reproduce it |
POSITIONING.md |
honest comparison against Mem0 and Zep/Graphiti, and where local-first is a hard requirement |
CHANGELOG.md |
what changed in each release |
Measured, generation-free retrieval lift against a pure-vector baseline on
public datasets — HotpotQA +7.2 pp, 2WikiMultiHopQA +2.1 pp overall,
TimeQA +9.7 pp, tri-engine +29 pp — is tabulated with its full method
in BENCHMARK.md.
Compatibility
| Environment | Status | Note |
|---|---|---|
| Any MCP client | Supported | stdio by default; streamable-HTTP with --features http. |
| Claude Code | Supported, with a caveat | claude mcp add, stdio or --transport http. Also the only harness with the PostToolUse replacing hook. Over HTTP, its native client does not re-initialize on an expired-session 404 within a call: that one call surfaces as a timeout and writes nothing, and the next call re-initializes and succeeds (proven from the daemon's own request log — see the idle-sessions guide). After any timeout, verify the write before trusting it. |
| Claude Desktop | Supported, with a caveat | Its config file accepts stdio only; for the shared daemon the installers wire a pinned mcp-remote stdio→HTTPS bridge, whose current dependency tree needs Node.js 20.18.1 or newer. That bridge does not yet recover transparently from an idle-expired session; restart Desktop after such a timeout. |
| Codex CLI | Supported | codex mcp add, or a [mcp_servers.*] table. The shared-daemon installers require Codex 0.113+ and use its native Streamable HTTP transport so an expired-session 404 is re-initialized instead of hanging behind a bridge. Four lifecycle hooks ship: SessionStart resumes rolling context, PreToolUse/PostToolUse require a successful recall before an opted-in apply_patch, and Stop saves context with the learning-loop checklist. PreCompact/PostCompact are not wired — they have no documented context channel; SessionStart handles the post-compaction continuation. |
| Windsurf | Supported | stdio (mcp_config.json) or serverUrl against the daemon. One advisory pre_user_prompt hook is wired; it is shown to the user, not injected into the model context. |
Other verified clients: Cursor, Cline, Zed, opencode, Devin CLI.
Known limits
- Memory semantics only. No
query,create_collection,upsert, ortraverse— deliberate, and enforced by the tool surface. - One process per store. The store takes a single-writer
flock, so two stdio clients against the sameVELESDB_MEMORY_PATHfail withStorage(DatabaseLocked). Run the HTTP daemon to share one memory across clients. - The HTTP transport has no authentication. It binds loopback-only by default; a non-loopback bind is refused unless you explicitly allow it, and is only safe behind an authenticating reverse proxy.
- A store is fixed to one embedder. The embedding dimension is probed from the model, so do not switch embedders on an existing store.
- Bring-your-own-links by default. The graph is built by
relateandlinks; automatic extraction needs--features extractplus a local model. pathingestion is off unless allowlisted viaVELESDB_MEMORY_INGEST_ROOTS.- Binding parity is incomplete.
compile_transcriptis MCP-only; the Node binding has nocontext_savings/explain_compilation; the published PyPI wheel predates the compiler entirely — see the surface matrix. - No selective source purge. Compiled sources are kept permanently by default; you can set a TTL going forward, or delete the whole store.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
Storage(DatabaseLocked) |
Two processes opened the same store — usually a second client, or a stray stdio process next to the daemon. | Run one --http daemon and point every client at it, or give the second client its own VELESDB_MEMORY_PATH. |
| The server never starts from a JSON/TOML config | ~ is not expanded: those configs spawn the binary without a shell. |
Use an absolute path, e.g. /home/you/.cargo/bin/velesdb-memory. |
extraction backend not configured from remember_extracted |
The call omitted extractor and VELESDB_MEMORY_EXTRACTOR is unset. |
Pass extractor: "outline", or configure the daemon default; see auto-extraction. |
IngestDisabled on a path fragment |
VELESDB_MEMORY_INGEST_ROOTS is unset or empty — path ingestion is off by default. |
Start the server with an allowlist of absolute directories. |
relate / forget reports a missing id from a JS client |
Ids exceed 2^53 and lose precision as JSON numbers. | Relay the id_str field, or set "policy": {"ids_as_strings": true} on compiler calls. |
License
The distributed binary embeds velesdb-core and is governed by the VelesDB
Core License 1.0 (source-available): a derivative of the Elastic License 2.0,
not an OSI-approved license. The wrapper source in this crate is intentionally
readable and forkable.
- Can you use it at work, or in a commercial product? Yes. Running the server locally, or embedding the library inside your own application where your users only ever receive results, is the license's expressly-permitted embedded, local-first use.
- What is forbidden? Re-hosting VelesDB as a multi-tenant service where
third parties drive the database. This server makes that impossible by
design: memory semantics only, never raw
query/create_collection/upsert/traverse. - What do you owe when redistributing? Keep the LICENSE file and copyright notices, and add a velesdb.com attribution in any public app that ships the binary. Internal, dev, and test use need no attribution.
Full terms and the canonical FAQ: LICENSE. Questions: contact@wiscale.fr.
velesdb-memory v0.12.0 · Last updated: 2026-08-08 · Report a docs error