plugmem-mcp
⚠️ Experimental. plugmem is mostly an AI-built experiment — written with the help of a small local model (Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf) and various Claude models, in roughly equal measure. Expect non-professional design choices, rough edges, broken behavior, or mistakes. Use it at your own risk.
plugmem-mcp is the Model Context Protocol
server over the plugmem temporal-memory engine
— a thin, long-lived shell around
plugmem-host that exposes a memory to
AI agents and any non-Rust program as MCP tools over stdio JSON-RPC. The
engine stays resident for the process's lifetime, so every call is host speed.
The installed binary is plugmem-mcp.
Install
Prebuilt for Linux, Windows and macOS (x64 & arm64) on every tagged release.
Pick one method — you don't need more than one; they install the same
plugmem-mcp binary.
Homebrew (macOS / Linux)
From the m62624/homebrew-plugmem
tap; brew upgrade / brew uninstall then manage it like any formula:
$ brew install m62624/plugmem/plugmem-mcp
Installer script (no Rust toolchain)
latest always points at the newest tag on the
Releases page:
# Linux / macOS (POSIX sh)
$ curl --proto '=https' --tlsv1.2 -LsSf https://github.com/m62624/plugmem/releases/latest/download/plugmem-mcp-installer.sh | sh
# Windows (PowerShell) — alternative to the .msi
> powershell -ExecutionPolicy Bypass -c "irm https://github.com/m62624/plugmem/releases/latest/download/plugmem-mcp-installer.ps1 | iex"
Windows .msi
Download plugmem-mcp-*.msi from the
Releases page. Double-click to
install; it registers in "Add or remove programs" for normal upgrades and
uninstalls.
cargo binstall
cargo-binstall downloads the prebuilt binary instead of compiling — it just works on every OS/arch above:
$ cargo binstall plugmem-mcp
From source
Needs a Rust toolchain. From crates.io:
$ cargo install plugmem-mcp
…or from a local checkout of this repo:
$ cargo install --path crates/plugmem-mcp
# or, to build without installing:
$ cargo build --release -p plugmem-mcp # binary at target/release/plugmem-mcp
Uninstall
cargo uninstall plugmem-mcp (for cargo install/binstall);
brew uninstall plugmem-mcp (Homebrew); "Add or remove programs" (.msi). The
shell/PowerShell installers ship no uninstaller — remove ~/.cargo/bin/plugmem-mcp
and ~/.config/plugmem-mcp (Windows: %USERPROFILE%\.cargo\bin\plugmem-mcp.exe
and %LOCALAPPDATA%\plugmem-mcp) by hand. See the
workspace README for the full matrix.
Which door is this? (read before reaching for MCP)
plugmem is embedded-first, like SQLite — the fastest, simplest path is to link the engine into your process, not to talk to a server.
| You are… | Use | Why |
|---|---|---|
| an agent, or a program in another language (Python, Node, Go…) that wants a memory | plugmem-mcp (this binary) |
Spawn the process, speak JSON-RPC on its stdin/stdout. Language-independent; the memory stays resident. |
| writing Rust | plugmem-host — embed it as a dependency |
The engine in your process, like linking SQLite. Maximum speed, no pipe, no second process. Don't front your own Rust with MCP. |
| a person at a terminal or shell script | plugmem-cli |
The human/scripting door. Not the door for programmatic or cross-language access — that's MCP. |
| JavaScript / TypeScript (Node) | plugmem-napi |
The engine as a native Node addon (napi-rs), in-process; on npm as plugmem. |
So: another language → MCP; Rust → embed the host lib; a human → the CLI. The MCP server's main consumer is the agent itself. And whichever door you use, you are tending your own memory file — plugmem keeps no server of its own.
What is (and isn't) MCP here
- A sidecar process, not a daemon. The host (Claude Desktop, an IDE, an
agent runner) spawns
plugmem-mcpand talks to it over stdin/stdout. It listens on no port and serves one memory file for its lifetime. When the host goes away, so does the sidecar. - Many readers / many languages = many processes, coordinated by plugmem's file-level MVCC (immutable snapshot generations + an advisory writer lock), not one network server. There is deliberately no network mode: that would add ports, auth and a connection pool for no embedded-use benefit.
- Concurrent, on plain threads. Independent requests run in parallel on a small worker pool; the only thing a request ever waits on is the embedder's HTTP call, kept off the shared lock. No async runtime — the engine's work is in-memory and fast, so threads are the whole story (see Concurrency).
What recall does
Recall fuses four sources by reciprocal-rank fusion with a recency boost (tags filter; they are not a source):
| Source | Algorithm | What it finds |
|---|---|---|
| Lexical | BM25 over a Unicode (UAX #29) tokenizer | exact terms / keyword overlap |
| Semantic | int8-quantized cosine — flat below a threshold, an HNSW graph above | meaning / nearest neighbours |
| Graph | entity graph with typed edges, breadth-first from query anchors | relational knowledge |
| Temporal | range scans over a recorded_at-ordered index; bitemporal validity |
"what was true then", time windows |
Tools
Every tool is named plugmem_* (so it never collides with another server's
tools) and takes an optional format argument: "json" (default) returns
compact machine JSON; "human" pretty-prints it (and, for plugmem_recall,
returns the engine's prompt-ready block instead of the structured result).
Result payloads ride in the MCP content[].text field; a tool-level failure
sets isError: true so the model can read and react to it.
Writer mode (default — a read-write memory of its own):
| tool | what it does |
|---|---|
plugmem_remember |
store a fact (text, optional entity, tags[], links[] of {rel, entity}, valid_from); returns the id + similar/conflicting facts |
plugmem_recall |
ranked, token-budgeted recall (query, tags[], entities[], as_of, range [from,to], k, closed) |
plugmem_revise |
close fact id, record the successor (same args as remember + id) |
plugmem_forget |
tombstone fact id (purged at the next maintain) |
plugmem_link |
upsert a typed edge src -rel-> dst |
plugmem_show |
one fact's full card by id |
plugmem_stats |
engine size counters |
plugmem_export |
every open fact as a JSON array |
plugmem_maintain |
purge tombstones, compact, build the vector index |
plugmem_checkpoint |
flush the journal into a fresh snapshot |
plugmem_verify |
content-integrity check |
plugmem_version / plugmem_about |
the running version; a pointer to the plugmem skill |
Read-only mode (--read-only — observe another process's writer over a
shared snapshot): plugmem_recall, plugmem_show, plugmem_stats,
plugmem_export, plugmem_verify, plus plugmem_generation (the pinned
snapshot generation) and plugmem_refresh (advance to the writer's latest
published checkpoint). Write tools are refused with a tool-level error.
The fact id on each plugmem_recall line (the [fN] in the human block, or
the "id" field in JSON) is how you address a fact in plugmem_revise,
plugmem_forget and plugmem_show — the usual "recall, then act" flow.
No import tool — that's the CLI's job. plugmem_export returns the facts
inline (no file needed), but bulk-loading from a backup.jsonl means reading
a file on the server's disk — which a sandboxed or remote server can't see.
So restoring/migrating a memory from a file is done with
plugmem-cli import (it has the disk, and
streams the file in batches), not over MCP. An agent doesn't bulk-load anyway —
it remembers facts one at a time with plugmem_remember as the conversation goes.
Usage
The host spawns the binary and wires its arguments once, in its MCP config:
plugmem-mcp [--db PATH] [--config PATH] [--read-only] [--workers N]
--db— the memory file (else$PLUGMEM_DB, else./plugmem.db).--config— aconfig.toml(else$PLUGMEM_CONFIG, else the XDG default).--read-only— observe another process's writer (requires a checkpointed database).--workers N— worker threads (else[server].workers, else half the cores).
A Claude Desktop / MCP-client config entry looks like:
A failure to start (bad config, or the file already locked by another writer) is reported to stderr with a non-zero exit, so the spawning host sees the server did not come up.
Concurrency
One reader thread pulls stdin lines into a channel; a pool of worker threads
drains it, dispatches, and writes replies under a single stdout lock (so lines
never interleave). Each worker holds a cheap handle — a writer clones the
Database (an Arc around the engine's RwLock), a reader shares one snapshot
(its own RwLock for concurrent reads, a Mutex for the embedder) — so
independent requests overlap, with the embedder's HTTP call outside the engine
lock. Replies carry their JSON-RPC id, so a client correlates them regardless
of completion order.
The pool defaults to max(1, available_parallelism() / 2) — half the machine's
cores, leaving room for the agent, the OS and a local embedder rather than
monopolizing the box. Override it with [server].workers or --workers.
Configuration
Optional config.toml, found by --config PATH, then $PLUGMEM_CONFIG, then
$XDG_CONFIG_HOME/plugmem/config.toml (all optional). The engine, embedder and
maintenance sections are the same shared loader the CLI uses; MCP adds one
[server] section.
[]
= 4 # worker threads (default: half the cores)
[]
= 768 # embedding size (0 = vectors off)
[] # default: none — lexical/tags/graph/time still work
= "ollama" # ollama | openai | lmstudio | vllm | llamacpp | none
= "http://localhost:11434/v1"
= "nomic-embed-text"
= "OPENAI_API_KEY"
[]
= 1024
= 100
The embedder unlocks the vector recall source; with kind = "none" (the
default) recall still answers from lexical, tag, graph and temporal evidence.
One OpenAI-compatible client covers Ollama, OpenAI, LM Studio, vLLM and
llama.cpp-server. $PLUGMEM_EMBEDDER overrides [embedder].kind.
License
MIT.