# agentplane
**A durable, replayable, policy-governed runtime for AI agents โ in Rust.** ๐ฆ
[](#-license)
[](#-status)
[](#-status)
Not a prompt framework. Not an agent library. The layer *beneath* those โ the
thing that makes an agent's actions survivable, auditable, and governable when it
is calling real systems that move real money.
```rust
// Performs its effects once, and journals everything.
let outcome = runtime.run("reconcile", input).await?;
// Replay re-executes the logic and reads every effect back from the journal.
// No tool is called again. No clock is read again. No invoice is issued twice.
runtime.replay(outcome.run_id, Mode::Strict).await?;
```
---
## ๐ฅ The problem
Production agents fail in ways a better model does not fix:
- A 40-minute run dies at minute 38, and the retry re-issues every invoice.
- *"Why did the agent refund โฌ4,200?"* has no answer, because the reasoning was
prose in a log line.
- Untrusted tool output steers the next tool call.
- A prompt change ships with no way to know what it broke.
These are **runtime** problems. agentplane is a runtime.
## ๐ก The idea
> **The journal is the plan of record.** Orchestration is deterministic and
> replayable. Everything non-deterministic โ model inference, tool calls, the
> clock, randomness โ is an *effect*: performed at most once, written to an
> append-only hash-chained log, and read back on replay.
Get that right and six things fall out of **one** mechanism: crash recovery,
audit, cost accounting, regression testing, tamper evidence, and regulatory
record-keeping. They stop being six subsystems that can each rot independently.
And critically: **the audit trail is also the recovery mechanism**, so it cannot
quietly stop working โ the system would stop working with it. Logging that exists
only to satisfy an auditor always rots.
## ๐ Try it
```sh
cargo run --example durable_pipeline # crash, resume, divergence
cargo run --example clearing_case # correlation, obligations, human tasks
cargo run --example plan_graph # multi-step plans, contract, provenance
# Calls a model and replays without calling it again โ no API key, no network.
cargo run --example model_run --features redb,testkit
```
`durable_pipeline` prints the whole claim in four steps: a live run, a strict
replay that touches nothing, a crash that resumes without repeating work, and a
changed build that is **quarantined instead of quietly rewriting history**.
New here? โ **[docs/getting-started.md](https://hupe1980.github.io/agentplane/docs/getting-started/)**
## ๐ฆ What you get
| ๐งพ | **A journal you can audit** โ append-only, hash-chained, per-record signatures naming the workload that wrote them, and a per-plane Merkle log so deleting a whole run is detectable |
| โฑ๏ธ | **Durable execution** โ crash mid-run and resume from the last completed effect; a suspended run costs a row on disk, not a task |
| ๐๏ธ | **Cases, not long-lived workflows** โ runs stay minutes, business processes span months, so a deploy never has to migrate an in-flight workflow |
| ๐ก๏ธ | **Policy before every effect** โ a total, I/O-free gate; a run denied at step 7 never starts at step 1 |
| ๐ท๏ธ | **Information-flow labels** โ *may this principal act* and *may this value go there* are different questions, and both are answerable |
| ๐ธ | **Budgets that bind** โ a failed model call is billed for what it burned, because the provider bills for it too |
| ๐ค | **Human oversight** โ durable worklists with four-eyes, declared expiry behaviour, and an operator who can *stop* a run and have it unwind |
| ๐ | **Real wires** โ MCP tools, A2A peers, Anthropic and OpenAI drivers, each with a failure mapping that says whether the call landed |
Full inventory, including what is **not** built โ
**[docs/status.md](https://hupe1980.github.io/agentplane/docs/status/)**
## ๐ Documentation
| ๐ | [Getting started](https://hupe1980.github.io/agentplane/docs/getting-started/) โ first run, first skill, first replay |
| ๐ง | [Concepts](https://hupe1980.github.io/agentplane/docs/concepts/) โ the ideas the rest is built from |
| ๐๏ธ | [Architecture](https://hupe1980.github.io/agentplane/docs/architecture/) โ how it actually works, mechanism by mechanism |
| ๐ณ | [Cookbook](https://hupe1980.github.io/agentplane/docs/cookbook/) โ task-shaped recipes |
| ๐ | [Security model](https://hupe1980.github.io/agentplane/docs/security/) โ the trust boundary, and what it does not cover |
| โ๏ธ | [Operations](https://hupe1980.github.io/agentplane/docs/operations/) โ deploying, HA, retention, observability |
| โ๏ธ | [Regulation](https://hupe1980.github.io/agentplane/docs/regulation/) โ EU AI Act obligation by obligation, and what is missing |
| ๐ | [Status](https://hupe1980.github.io/agentplane/docs/status/) โ built vs designed-not-built |
| ๐ค | [Contributing](CONTRIBUTING.md) โ the assurance ladder, and how to run it |
## ๐งช Assurance
Each layer answers a question the others structurally cannot.
```sh
just # list every check
just ci # lint ยท 3 feature configs ยท examples ยท docs ยท packaging
just ci-full # the above, plus TLA+ specs and the full mutation sweep
```
Two are unusual enough to name:
**๐ฌ Formal specs.** Six TLA+ specifications are model-checked on every push โ
the effect protocol, retry safety, sagas, fencing, authorization, delegation. And
because a spec whose invariants cannot be violated proves nothing, each is
re-checked against 18 deliberately broken copies of itself; every mutant must be
caught by the *specific* invariant written for it.
**๐งฌ Mutation testing over the code.** 106 guarantees are broken on purpose, and
the test *named for each one* must fail. A mutation caught by some other test is
reported **weak**, not passing โ that usually means the guarantee has no test of
its own and is being held up by one that could be rewritten without anyone
noticing what it protected.
This is not decoration. The project shipped an unfalsifiable guarantee once: the
refusal to replan on untrusted data was implemented, tested, and green โ and
deleting it would have failed no test, because the fixtures laundered the taint
before it reached the check. It was found by accident. The sweep is so the next
one is not.
## ๐ซ Non-goals
| Ship a prompt library or IDE | Your manifests; agentplane hashes and versions them |
| Route or proxy model traffic | LiteLLM, Bifrost, your own `ModelProvider` |
| Implement a vector database | LanceDB / pgvector behind a seam |
| Replace a deterministic protocol engine | Keep it; agentplane sits *beside* it, never inside it |
| Require Kubernetes | One static binary |
| Train, fine-tune, or serve models | Permanently out of scope |
| Grade output quality | It emits replayable traces; grade them elsewhere |
| Interpret payload contents | Payloads are opaque, and labeled |
| Claim regulatory compliance | It provides technical means; compliance is the deployer's |
**Who should not use this:** a team running three agents against low-stakes data.
The complexity is justified when agents touch money, meters, or regulated
records.
## ๐ Status
**Pre-alpha, pre-release, no API stability.** Breaking changes land without
deprecation. The journal record format and the storage schema will change.
Rust **1.94+**. `#![forbid(unsafe_code)]`. One crate, feature-gated: an embedded
[redb](https://github.com/cberner/redb) store by default โ pure Rust, two crates
deep, no C toolchain โ with everything else opt-in.
Honest framing on regulation: agentplane is not "compliant" and cannot be.
Compliance attaches to a system in a context, assessed by its provider or
deployer. What this gives you is the **technical means** to discharge EU AI Act
Articles 12 and 14 โ means that are already load-bearing for recovery and
testing, and therefore cannot quietly rot. [Regulation](https://hupe1980.github.io/agentplane/docs/regulation/) maps
obligation to mechanism, names what is *not* built, and notes that the Digital
Omnibus moved the high-risk dates to December 2027 without amending the
articles.
## ๐ License
MIT OR Apache-2.0, at your option.