plugmem-mcp 0.1.3

plugmem MCP server (stdio JSON-RPC).
plugmem-mcp-0.1.3 is not a library.

plugmem-mcp

⚠️ Experimental. plugmem is mostly an AI-built experiment — written with the help of a small local model (Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf) and various Claude models, in roughly equal measure. Expect non-professional design choices, rough edges, broken behavior, or mistakes. Use it at your own risk.

plugmem-mcp is the Model Context Protocol server over the plugmem temporal-memory engine — a thin, long-lived shell around plugmem-host that exposes a memory to AI agents and any non-Rust program as MCP tools over stdio JSON-RPC. The engine stays resident for the process's lifetime, so every call is host speed.

The installed binary is plugmem-mcp.

Install

Prebuilt for Linux, Windows and macOS (x64 & arm64) on every tagged release. Pick one method — you don't need more than one; they install the same plugmem-mcp binary.

Homebrew (macOS / Linux)

From the m62624/homebrew-plugmem tap; brew upgrade / brew uninstall then manage it like any formula:

$ brew install m62624/plugmem/plugmem-mcp

Installer script (no Rust toolchain)

latest always points at the newest tag on the Releases page:

# Linux / macOS  (POSIX sh)
$ curl --proto '=https' --tlsv1.2 -LsSf https://github.com/m62624/plugmem/releases/latest/download/plugmem-mcp-installer.sh | sh
# Windows (PowerShell) — alternative to the .msi
> powershell -ExecutionPolicy Bypass -c "irm https://github.com/m62624/plugmem/releases/latest/download/plugmem-mcp-installer.ps1 | iex"

Windows .msi

Download plugmem-mcp-*.msi from the Releases page. Double-click to install; it registers in "Add or remove programs" for normal upgrades and uninstalls.

cargo binstall

cargo-binstall downloads the prebuilt binary instead of compiling — it just works on every OS/arch above:

$ cargo binstall plugmem-mcp

From source

Needs a Rust toolchain. From crates.io:

$ cargo install plugmem-mcp

…or from a local checkout of this repo:

$ cargo install --path crates/plugmem-mcp
# or, to build without installing:
$ cargo build --release -p plugmem-mcp    # binary at target/release/plugmem-mcp

Uninstall

cargo uninstall plugmem-mcp (for cargo install/binstall); brew uninstall plugmem-mcp (Homebrew); "Add or remove programs" (.msi). The shell/PowerShell installers ship no uninstaller — remove ~/.cargo/bin/plugmem-mcp and ~/.config/plugmem-mcp (Windows: %USERPROFILE%\.cargo\bin\plugmem-mcp.exe and %LOCALAPPDATA%\plugmem-mcp) by hand. See the workspace README for the full matrix.

Which door is this? (read before reaching for MCP)

plugmem is embedded-first, like SQLite — the fastest, simplest path is to link the engine into your process, not to talk to a server.

You are… Use Why
an agent, or a program in another language (Python, Node, Go…) that wants a memory plugmem-mcp (this binary) Spawn the process, speak JSON-RPC on its stdin/stdout. Language-independent; the memory stays resident.
writing Rust plugmem-host — embed it as a dependency The engine in your process, like linking SQLite. Maximum speed, no pipe, no second process. Don't front your own Rust with MCP.
a person at a terminal or shell script plugmem-cli The human/scripting door. Not the door for programmatic or cross-language access — that's MCP.
JavaScript / TypeScript (Node) plugmem-napi The engine as a native Node addon (napi-rs), in-process; on npm as plugmem.

So: another language → MCP; Rust → embed the host lib; a human → the CLI. The MCP server's main consumer is the agent itself. And whichever door you use, you are tending your own memory file — plugmem keeps no server of its own.

What is (and isn't) MCP here

  • A sidecar process, not a daemon. The host (Claude Desktop, an IDE, an agent runner) spawns plugmem-mcp and talks to it over stdin/stdout. It listens on no port and serves one memory file for its lifetime. When the host goes away, so does the sidecar.
  • Many readers / many languages = many processes, coordinated by plugmem's file-level MVCC (immutable snapshot generations + an advisory writer lock), not one network server. There is deliberately no network mode: that would add ports, auth and a connection pool for no embedded-use benefit.
  • Concurrent, on plain threads. Independent requests run in parallel on a small worker pool; the only thing a request ever waits on is the embedder's HTTP call, kept off the shared lock. No async runtime — the engine's work is in-memory and fast, so threads are the whole story (see Concurrency).

What recall does

Recall fuses four sources by reciprocal-rank fusion with a recency boost (tags filter; they are not a source):

Source Algorithm What it finds
Lexical BM25 over a Unicode (UAX #29) tokenizer exact terms / keyword overlap
Semantic int8-quantized cosine — flat below a threshold, an HNSW graph above meaning / nearest neighbours
Graph entity graph with typed edges, breadth-first from query anchors relational knowledge
Temporal range scans over a recorded_at-ordered index; bitemporal validity "what was true then", time windows

Tools

Every tool is named plugmem_* (so it never collides with another server's tools) and takes an optional format argument: "json" (default) returns compact machine JSON; "human" pretty-prints it (and, for plugmem_recall, returns the engine's prompt-ready block instead of the structured result). Result payloads ride in the MCP content[].text field; a tool-level failure sets isError: true so the model can read and react to it.

Writer mode (default — a read-write memory of its own):

tool what it does
plugmem_remember store a fact (text, optional entity, tags[], links[] of {rel, entity}, valid_from); returns the id + similar/conflicting facts
plugmem_recall ranked, token-budgeted recall (query, tags[], entities[], as_of, range [from,to], k, closed)
plugmem_revise close fact id, record the successor (same args as remember + id)
plugmem_forget tombstone fact id (purged at the next maintain)
plugmem_link upsert a typed edge src -rel-> dst
plugmem_show one fact's full card by id
plugmem_stats engine size counters
plugmem_export every open fact as a JSON array
plugmem_maintain purge tombstones, compact, build the vector index
plugmem_checkpoint flush the journal into a fresh snapshot
plugmem_verify content-integrity check
plugmem_version / plugmem_about the running version; a pointer to the plugmem skill

Read-only mode (--read-only — observe another process's writer over a shared snapshot): plugmem_recall, plugmem_show, plugmem_stats, plugmem_export, plugmem_verify, plus plugmem_generation (the pinned snapshot generation) and plugmem_refresh (advance to the writer's latest published checkpoint). Write tools are refused with a tool-level error.

The fact id on each plugmem_recall line (the [fN] in the human block, or the "id" field in JSON) is how you address a fact in plugmem_revise, plugmem_forget and plugmem_show — the usual "recall, then act" flow.

No import tool — that's the CLI's job. plugmem_export returns the facts inline (no file needed), but bulk-loading from a backup.jsonl means reading a file on the server's disk — which a sandboxed or remote server can't see. So restoring/migrating a memory from a file is done with plugmem-cli import (it has the disk, and streams the file in batches), not over MCP. An agent doesn't bulk-load anyway — it remembers facts one at a time with plugmem_remember as the conversation goes.

Usage

The host spawns the binary and wires its arguments once, in its MCP config:

plugmem-mcp [--db PATH] [--config PATH] [--read-only] [--workers N]
  • --db — the memory file (else $PLUGMEM_DB, else ./plugmem.db).
  • --config — a config.toml (else $PLUGMEM_CONFIG, else the XDG default).
  • --read-only — observe another process's writer (requires a checkpointed database).
  • --workers N — worker threads (else [server].workers, else half the cores).

A Claude Desktop / MCP-client config entry looks like:

{
  "mcpServers": {
    "plugmem": {
      "command": "plugmem-mcp",
      "args": ["--db", "/home/me/agent.plugmem"]
    }
  }
}

A failure to start (bad config, or the file already locked by another writer) is reported to stderr with a non-zero exit, so the spawning host sees the server did not come up.

Concurrency

One reader thread pulls stdin lines into a channel; a pool of worker threads drains it, dispatches, and writes replies under a single stdout lock (so lines never interleave). Each worker holds a cheap handle — a writer clones the Database (an Arc around the engine's RwLock), a reader shares one snapshot (its own RwLock for concurrent reads, a Mutex for the embedder) — so independent requests overlap, with the embedder's HTTP call outside the engine lock. Replies carry their JSON-RPC id, so a client correlates them regardless of completion order.

The pool defaults to max(1, available_parallelism() / 2) — half the machine's cores, leaving room for the agent, the OS and a local embedder rather than monopolizing the box. Override it with [server].workers or --workers.

Configuration

Optional config.toml, found by --config PATH, then $PLUGMEM_CONFIG, then $XDG_CONFIG_HOME/plugmem/config.toml (all optional). The engine, embedder and maintenance sections are the same shared loader the CLI uses; MCP adds one [server] section.

[server]
workers = 4            # worker threads (default: half the cores)

[engine]
dim = 768              # embedding size (0 = vectors off)

[embedder]             # default: none — lexical/tags/graph/time still work
kind = "ollama"        # ollama | openai | lmstudio | vllm | llamacpp | none
url = "http://localhost:11434/v1"
model = "nomic-embed-text"
api_key_env = "OPENAI_API_KEY"

[maintenance]
snapshot_every_ops = 1024
maintain_every_forgets = 100

The embedder unlocks the vector recall source; with kind = "none" (the default) recall still answers from lexical, tag, graph and temporal evidence. One OpenAI-compatible client covers Ollama, OpenAI, LM Studio, vLLM and llama.cpp-server. $PLUGMEM_EMBEDDER overrides [embedder].kind.

License

MIT.