Kimetsu
Give your coding agent a memory that gets sharper every run.
Evidence-first memory for coding agents and Kimetsu's own terminal chat. Kimetsu sits beside your AI agent, watches what actually solves problems, remembers it, and feeds the high-signal context back — so the next run starts where the last one left off.
At a glance
Measured, not claimed (every number is reproducible with kimetsu brain bench / the ROI ledger):
| Metric | Result | How it's measured |
|---|---|---|
| Cost per win | $0.19 vs $2.47 (~13× cheaper) | 16-task Terminal-Bench slice, Kimetsu-on vs no-brain baseline |
| Retrieval quality | recall@4 0.949 · MRR 0.914 @ ~138 ms (default), up to 0.975 · 0.933 | kimetsu brain bench, 100-memory / 210-case dataset, jina-v2-base-code × cross-encoder rerank |
| Scale | ~1M memories in ~3 GB RAM, sub-2s retrieval | usearch HNSW ANN, O(log N) |
| Footprint | one SQLite file per project, no cloud, no telemetry | back it up with cp |
Configurable surface (all in .kimetsu/project.toml, all optional with safe defaults):
| Section | Controls | Default |
|---|---|---|
embedder |
semantic model + cross-encoder reranker (or FTS-only lean build) | jina-v2-base-code × ms-marco-tinybert |
[cheap_model] |
local Ollama or cloud model for digest / resume / ask / skill-draft — optional, degrades gracefully |
off (rule-based) |
[storage] backend |
flat (FTS + ANN) · graph-lite (+ typed-edge hops) · graph (remote) |
flat |
[broker] |
warm_start, compress_capsules, session_dedupe, answer_grade_min_score, proactive_prefetch |
sensible on/off |
[sync] |
server-less multi-machine sync via a shared dir |
off |
[learning] |
store_queries for the self-tuning eval set |
on (local only) |
Full reference: How Kimetsu Works · local models: docs/LOCAL-MODELS.md.
Why Kimetsu
LLM coding agents are brilliant and forgetful. Every session starts from zero — the same wrong turns, the same re-explaining of your conventions, the same expensive exploration you already paid for last week.
Kimetsu fixes the forgetting. It's a sidecar brain: a single Rust binary that runs next to any supported host agent through MCP (Claude Code, Codex, Pi, OpenClaw, Cursor, Gemini CLI) or as its own terminal chat — or, in beta, server-hosted over HTTP MCP and shared across a team. It learns which memories the model actually used to win, and lets that knowledge compound across runs.
- It remembers. Project conventions, failure patterns, the exact command that regenerates your schema — captured once, retrieved automatically.
- It learns what helps. Memories that the model cites before solving a problem get promoted. Silent passengers and stale advice decay and get pruned.
- It never explores twice (v2.0). A session-start digest plus an episodic resume mean the agent's first turn already knows the repo and what you were doing last time — no re-deriving the basics, no "where was I."
- It answers, not just injects (v2.0).
kimetsu ask "…"composes a grounded, cited answer from memory using a local model — zero frontier tokens, works offline. Lessons cited often graduate into runnable skills. - It's cheap to be right. On a recorded 16-task Terminal-Bench slice,
Kimetsu-enabled runs cost ~13x less per win than the no-brain host-agent
baseline: $0.19/win vs $2.47/win — and the
kimetsu brain roiledger shows the token savings per memory on your own work. - It gets smarter, not just bigger. Semantic retrieval finds the right memory even when you used different words; it self-tunes retrieval against your own query history; and brain insights show hit-rate, citation rate, and token economy so the value is measurable.
- It's yours, on your machine. The whole brain is one SQLite file per
project. No external vector DB, no cloud, no telemetry. Back it up with
cp.
Kimetsu (鬼滅) — "demon slayer." It slays the demon every agent fights: amnesia.
How it works
Host agent (Claude / Codex / Pi / OpenClaw / kimetsu chat)
│ asks for context ▲ cites what helped
▼ │
MCP tools ──► Broker ──► top memories ──► agent run
│ scores candidates by relevance ×
│ usefulness × freshness × scope
▼
brain.db — one SQLite file: FTS5 + semantic ANN (usearch HNSW)
- Before a task, the broker walks your project brain and your cross-project user brain, scores every candidate, and injects the top few inside an adaptive token budget. The semantic build matches by meaning (O(log N) ANN — scales to ~1M memories in ~3 GB RAM, sub-2s retrieval).
- While it works, Kimetsu surfaces known pitfalls before the first attempt, and the model cites the memories that actually help.
- After the task, cited memories get promoted, unused advice decays on a half-life curve, and non-trivial sessions auto-harvest their lessons.
Full mechanics — scoring, citations, decay, conflict detection, the daemon — in HOW-KIMETSU-WORKS.
Quickstart
# 1. Install — no Rust toolchain needed (cargo + prebuilt archives in docs/INSTALL.md)
# 2. Wire it into your host agent — init + install + selftest in one shot
# 3. Prove the brain works
# ✓ recorded a memory and retrieved it — the brain works
From here your agent banks memories automatically — record one yourself and watch it come back:
New in v2.0 — never explore twice:
Prefer a standalone REPL? kimetsu chat --workspace . --project . is a full
terminal assistant with the same brain. Every install path (npm, prebuilt
archives), host-wiring details, the auto-harvest/distiller setup, and
maintenance commands live in docs/INSTALL.md.
Retrieval quality — benchmarked, not vibes
The semantic build retrieves with jina-v2-base-code + a
cross-encoder reranker, chosen with kimetsu brain bench on a
100-memory / 210-case dataset seeded from real exported memories. The
latency-optimized default (ms-marco-tinybert-l-2-v2) lands
recall@4 0.949, MRR 0.914 at ~138 ms per retrieval+rerank; the
quality-best rerankers reach recall@4 0.975, MRR 0.933. Swap embedder and
reranker with one config key each and re-judge on your own corpus — full grid
in Retrieval models & benchmarking.
Kimetsu Remote (beta)
Share one brain per repository from a server, over HTTP MCP — for a team, or for yourself across machines:
# server
# each client
Bearer auth, per-repo brains, optional shared org-brain, server-side repo ingest, TLS, Prometheus metrics, and a server-side reranker — full setup in docs/REMOTE.md.
What's in the box
| Surface | What it is |
|---|---|
kimetsu chat |
A full terminal coding assistant — slash commands, skills, hooks, background tasks, MCP, agents. Runs against your workspace, no Harbor required. |
kimetsu brain |
Durable, auto-migrating project + user memory in a single SQLite file. Citations, decay, conflict detection, FTS + optional semantic (usearch HNSW ANN, scales to ~1M memories) retrieval, pluggable retrieval backends (flat / graph-lite), and kimetsu brain insights effectiveness analytics. |
kimetsu ask / warm-start (v2.0) |
Grounded, cited Q&A from memory via a local model (zero frontier tokens); a session-start digest + episodic resume so the first turn already knows the repo and your last task. |
kimetsu brain skills (v2.0) |
Often-cited lessons graduate into runnable, provenance-linked skills installed into your host's native skill dir — propose-only. |
kimetsu brain sync (v2.0) |
Server-less multi-machine sync via event-log replication over a shared folder (Dropbox / Syncthing / NAS). |
kimetsu bridge |
Cross-harness skill portability — import/export skills between supported hosts such as Claude Code, Codex, Agents, and Kimetsu. |
| MCP sidecar | kimetsu mcp serve exposes the brain to any MCP host as kimetsu_* tools, including kimetsu_brain_answer for mid-task grounded synthesis. |
| Kimetsu Remote (beta) | kimetsu-remote — the brain over HTTP MCP, one per repository, shared from a server (separate package). |
Built as a small Rust workspace (kimetsu-cli, -chat, -agent, -brain,
-core, and -remote). Lint + tests run clean on every change.
Docs
- Install & host wiring — every install path, plugin install/uninstall/status, auto-harvest + distiller, maintenance commands.
- How Kimetsu Works — the conceptual reference: the brain, the broker, citations, decay, conflict detection, the MCP surface, retrieval models & benchmarking, the bridge, doctor, and config.
- Kimetsu Remote — server setup, org brain, server-side ingest, TLS, client wiring.
- CHANGELOG — what shipped in each release.
- Per-crate
src/lib.rsdoc comments for module-level detail.
License
Dual-licensed under MIT or Apache-2.0 — your choice.