agentplane
A durable, replayable, policy-governed runtime for AI agents โ in Rust. ๐ฆ
Not a prompt framework. Not an agent library. The layer beneath those โ the thing that makes an agent's actions survivable, auditable, and governable when it is calling real systems that move real money.
// Performs its effects once, and journals everything.
let outcome = runtime.run.await?;
// Replay re-executes the logic and reads every effect back from the journal.
// No tool is called again. No clock is read again. No invoice is issued twice.
runtime.replay.await?;
๐ฅ The problem
Production agents fail in ways a better model does not fix:
- A 40-minute run dies at minute 38, and the retry re-issues every invoice.
- "Why did the agent refund โฌ4,200?" has no answer, because the reasoning was prose in a log line.
- Untrusted tool output steers the next tool call.
- A prompt change ships with no way to know what it broke.
These are runtime problems. agentplane is a runtime.
๐ก The idea
The journal is the plan of record. Orchestration is deterministic and replayable. Everything non-deterministic โ model inference, tool calls, the clock, randomness โ is an effect: performed at most once, written to an append-only hash-chained log, and read back on replay.
Get that right and six things fall out of one mechanism: crash recovery, audit, cost accounting, regression testing, tamper evidence, and regulatory record-keeping. They stop being six subsystems that can each rot independently.
And critically: the audit trail is also the recovery mechanism, so it cannot quietly stop working โ the system would stop working with it. Logging that exists only to satisfy an auditor always rots.
๐ Try it
# field provenance, protected arguments
# env-gated Amazon Bedrock Converse call
# env-gated OpenAI Responses call
# a model choosing tools, and four refusals
# plan once, execute without the model โ
# a prompt injection with no reader
# erase a case: every copy unreadable,
# and the chain still verifies
# Calls a model and replays without calling it again โ no API key, no network.
# Digest-only multimodal dispatch and zero-I/O replay โ also fully offline.
# An agent whose prompt, model, result shape and ceilings come from a file.
# A real MCP server in this process beside a typed Rust tool โ one agent
# reaching both, and a strict replay that calls neither.
# Four agents, one plane: a coded editor that dictates the sequence, and a
# YAML desk that consults the same specialists as tool://agent/... grants.
# This plane served as an A2A 1.0 agent, called the way a peer would call it:
# a public card, authenticated methods, and a message that arrives untrusted.
# Live tokens for a human, one journaled completion for the machine โ and a
# replay that performs neither.
# One customer's approved โฌ500, spent across two separate runs, then revoked โ
# with the terms still readable afterwards.
Or skip Rust entirely โ a file and a key are the whole agent, and a file may
hold a whole room: several manifests separated by ---, the Kubernetes
packaging convention. Each document keeps its own digest โ the file is
packaging, not identity โ and a run starts at the room's declared orchestrator:
Or without a Rust toolchain at all โ needing one to run a YAML file rather defeats the point of the file:
Distroless, nonroot, no shell. It runs --read-only --network none because the
default journal is genuinely in memory and the example's provider is the
deterministic fake โ so the first run needs neither a disk nor the internet.
:slim (the default, and :latest) carries every model provider โ Anthropic,
OpenAI, Gemini, Bedrock, any OpenAI-compatible server; :full adds MCP, the A2A
peer server, the operator HTTP API, Cedar, key rings, governed media and
Postgres. Both are multi-arch, cosign-signed keylessly, and carry SLSA build
provenance and an SBOM attached to the digest:
durable_pipeline prints the whole claim in four steps: a live run, a strict
replay that touches nothing, a crash that resumes without repeating work, and a
changed build that is quarantined instead of quietly rewriting history.
A first program is one import and about forty lines:
use *;
And when it goes wrong, the plane answers with what it does have rather than
with a variant name โ fn main() reports through Debug, so on the errors you
hold, Debug is the message:
Error: no skill provides capability 'demo.greeet' โ this plane provides:
demo.greet. `run` takes a capability, not a skill name; a skill declares its own
with `SkillDescriptor::new(..).provides(..)`
A declarative tool loop runs from a file too โ the manifest grants
tool://tickets/read, and which transport reaches tickets is deployment
wiring rather than part of the reviewed declaration:
A grant naming a server nobody wired is refused at build, not on every run.
Needs --features cli,mcp-stdio, or the :full image.
And an agent can be hosted from the same file โ the A2A 1.0 server that passes the protocol project's own conformance kit, started without writing Rust:
Add --operator-addr 127.0.0.1:9090 and the worklist, task decisions and
GET /runs?outcome=quarantined are served too โ on their own listener, off
unless asked for, and separated from the peer surface by policy (peer reaches
a2a:*, operator reaches api:*) rather than by the port. A served plane also
sweeps deadlines, task expiry, dead letters, due timers and abandoned runs
โ a lease that expired while still naming an owner is an instance that died
holding the run, and the sweep takes it over and resumes it โ so a run that
sleeps, waits or loses its instance actually finishes.
Both --policy and --tokens are required and have no defaults. That is the
design rather than an inconvenience: a permissive engine and no engine are the
same behaviour, and a server that authenticates nobody has no actor to record a
decision against. Needs --features cli,a2a-server,cedar, or the :full image.
New here? โ docs/getting-started.md
๐ฆ What you get
| ๐งพ | A journal you can audit โ append-only, hash-chained, per-record signatures naming the workload that wrote them, and a per-plane Merkle log so deleting a whole run is detectable |
| โฑ๏ธ | Durable execution โ crash mid-run and resume from the last completed effect; a suspended run costs a row on disk, not a task. Recovery is initiated, not merely possible: the sweep finds every run whose owner died holding it โ an expired, unreleased lease โ and resumes it, journaling the takeover in its own sealed run |
| ๐๏ธ | Cases, not long-lived workflows โ runs stay minutes, business processes span months, so a deploy never has to migrate an in-flight workflow |
| ๐ก๏ธ | Policy before live dispatch โ a total, I/O-free gate; denials are journaled, strict replay never re-judges history, and plan authority is checked before step 1 |
| ๐ท๏ธ | Field-level information flow โ exact outbound arguments are bound to hierarchical provenance; recipient, amount, path, URL and other authority-bearing fields can require trusted or named sources while ordinary content remains untrusted |
| ๐ธ | Budgets that bind โ a failed model call is billed for what it burned, because the provider bills for it too |
| ๐งฌ | Effects that take together, or not at all โ a group declares the resources it touches and refuses any member outside them. Each reversible member records the concrete call that undoes it, built from what that call actually returned rather than reconstructed later from state that has moved โ the gap a per-step saga leaves, since compensate is handed the output of a step that failed and therefore has none. commit is the frontier: invariants are checked there because it is the last instant at which failing them is free, and only then are deferred members released. That is what makes an irreversible send safe โ an aborted group never sends it, which beats sending and apologising. Doubt reverses nothing |
| ๐ค | Human oversight on the call, not a summary of it โ requires_approval: true on a tool grant opens a task carrying the exact tool and arguments about to be dispatched, and nothing happens until somebody approves. Gating the agent's answer instead is a review that arrives after the money moved. Durable worklists, four-eyes, declared expiry behaviour, and an operator who can stop a run and have it unwind |
| ๐ | Erasure that reaches the backups โ deleting clears the live store; the backup taken an hour earlier still has everything, and backups are offsite and often immutable by design. So payload bytes are sealed under a per-case data key wrapped by a key the crate never holds: erasing a case destroys the key, and every copy becomes unreadable at once โ including the ones nobody can reach. Rotation re-wraps without rewriting bulk data; an erased case never comes back. VaultTransit speaks Vault's transit engine, so the wrapping key never leaves Vault. And the journal too: SealedJournal::wrap(store, keys, tenant) seals the journal's payload fields โ run input, prompts and tool-call arguments, effect and reconciliation outputs, failure messages, notes and frozen plans โ under the same per-case scope, so one erasure reaches blobs and journal alike. Only the payload is sealed, so exactly-once and every index keep working with no key; and the chain commits to the ciphertext, so an auditor holding no keys still verifies the history of a run whose data is gone โ erasure and keys |
| ๐ | An agent that is only a file โ agentplane run agent.yaml. No Rust, no main, no skill. The digest covers the agent in its entirety rather than only its boundary, and the run is journaled and deterministically replayable. Declarative formats give you the first; durable platforms give you the second; the pairing is what makes the evidence about something you can actually read |
Ten rows, not the inventory. The full surface โ the export/audit/restore
toolchain, typed release, standing authorities, effect groups that commit with
the journal, the emergency stop, the audited sweeper, model drivers and
streaming, MCP and A2A on both sides, signed Agent Cards, governed media and
memory, multi-tenancy, quotas, witnessing, break-glass, and why there is no
AllowAll anywhere โ is documented mechanism by mechanism on the site:
what you get, in full.
What is deliberately not built, and what will move โ docs/status.md
๐ Documentation
| ๐ | Getting started โ first run, first skill, first replay |
| ๐ง | Concepts โ the ideas the rest is built from |
| ๐๏ธ | Architecture โ how it actually works, mechanism by mechanism |
| ๐ณ | Cookbook โ task-shaped recipes, including wiring an MCP server beside typed tools |
| ๐ | Manifest reference โ every field, what enforces it, and what an absent value means |
| ๐งช | Testing agents โ the fake provider, fault injection, and proving a replay actually replayed |
| ๐ | Security model โ the trust boundary, and what it does not cover |
| ๐๏ธ | Erasure and keys โ erasure that reaches backups, key rotation and revocation, and how tenants are kept apart |
| โ๏ธ | Operations โ deploying, HA, retention, observability |
| โ๏ธ | Regulation โ EU AI Act obligation by obligation, and what is missing |
| ๐ | Status โ what is pre-alpha, what to pin, what is deliberately absent |
| โฌ๏ธ | Upgrading โ what breaks between pre-alpha releases, and the shortest correct fix |
| ๐ | Changelog โ what changed and when, including every mechanism's reasoning as it landed |
| ๐ค | Contributing โ the assurance ladder, and how to run it |
๐งช Assurance
Each layer answers a question the others structurally cannot.
Two are unusual enough to name:
๐ฌ Formal specs. TLA+ specifications are model-checked on every push โ the effect protocol, retry safety, sagas, fencing, authorization, delegation. And because a spec whose invariants cannot be violated proves nothing, each is re-checked against deliberately broken copies of itself; every mutant must be caught by the specific invariant written for it.
๐งพ Conformance by the protocol's own kit. just test-a2a-tck runs the
official a2a-tck against this crate's
A2A server on a live socket. Every other A2A test drives this server with this
crate's own client, which proves symmetry, not conformance โ a client and
server written from the same misreading agree everywhere. The kit's first run
found five defects no in-repo test could reach.
๐ Tests against a real provider. just test-live runs the OpenAI and
Gemini drivers against the actual APIs. They are gated twice โ an explicit AGENTPLANE_LIVE=1
and a key โ because a credential being available is not a decision to spend
money with it, and they are never part of ci. They exist because a stubbed
provider is structurally unable to have the defects a real one finds: it never
rejects a malformed request and never returns a shape the driver mis-reads.
Writing them found two, both of which every offline test had passed. The Gemini
battery is the sharpest case: a thought signature is minted and validated by
Google, so a canned server accepts whatever a fixture tells it to and says
nothing about whether Gemini takes the signature back โ the one check that
distinguishes a driver carrying the model's turn verbatim from one rebuilding
it, which is where the rest of the ecosystem has been losing this.
๐งฌ Mutation testing over the code. Every load-bearing guarantee is broken on purpose, and the test named for each one must fail. A mutation caught by some other test is reported weak, not passing โ that usually means the guarantee has no test of its own and is being held up by one that could be rewritten without anyone noticing what it protected.
This is not decoration. The project shipped an unfalsifiable guarantee once: the refusal to replan on untrusted data was implemented, tested, and green โ and deleting it would have failed no test, because the fixtures laundered the taint before it reached the check. It was found by accident. The sweep is so the next one is not.
It runs on every push, sharded six ways. It was gated to pull requests, on
the reasoning that a push to main had already passed it โ true of a repository
that merges, and this one has never opened a pull request, so the gate switched
the sweep off rather than making it cheaper. Three mutations then rotted into
code that no longer compiled, leaving three guarantees unfalsifiable with
every check green. just anchors reported all three present, correctly: it
checks that a mutation still matches the code it names, which is text and
not types.
MUTANTS_SHARD=k/n takes a round-robin slice โ round-robin because the table
groups mutations by subject, so a contiguous split would hand one shard every
expensive target. Each line carries a [current/total] progress counter, and
each shard needs its own checkout: the sweep rewrites source in place.
๐ซ Non-goals
| agentplane does not | Use instead |
|---|---|
| Ship a prompt library or IDE | Your prompts; agentplane pins the manifest that governs them by digest |
| Route or proxy model traffic | LiteLLM, Bifrost โ the drivers themselves ship: OpenAI Responses, Anthropic Messages and Bedrock Converse are here, with streaming, structured output, reasoning continuation and per-provider failure mappings. What is out of scope is choosing between them at runtime |
| Implement a vector database | LanceDB / pgvector behind the SemanticRetriever seam; embedding is a journaled effect so the query vector is history rather than a recomputation |
| Ship a built-in tool catalogue | Write a typed Tool, or wire an MCP server. The tools other frameworks ship โ web search, code interpreter โ are mostly provider-hosted: they run during generation, so the call is not announced, authorized, metered or replayable. That is a world-visible action outside the journal, which is the one thing this runtime is for. Governed URL fetching is the media feature; untrusted code belongs behind a process boundary |
| Replace a deterministic protocol engine | Keep it; agentplane sits beside it, never inside it |
| Require Kubernetes | One static binary |
| Train, fine-tune, or serve models | Permanently out of scope |
| Grade output quality | It emits replayable traces; grade them elsewhere |
| Interpret payload contents | Payloads are opaque, and labeled |
| Claim regulatory compliance | It provides technical means; compliance is the deployer's |
Who should not use this: a team running three agents against low-stakes data. The complexity is justified when agents touch money, meters, or regulated records.
๐ Status
Pre-alpha, pre-release, no API stability. Breaking changes land without deprecation. The journal record format and the storage schema will change.
Rust 1.94.1+. #![forbid(unsafe_code)]. One crate, feature-gated: an embedded
redb store by default โ pure Rust, two crates
deep, no C toolchain โ with everything else opt-in.
Honest framing on regulation: agentplane is not "compliant" and cannot be. Compliance attaches to a system in a context, assessed by its provider or deployer. What this gives you is the technical means to discharge EU AI Act Articles 12 and 14 โ means that are already load-bearing for recovery and testing, and therefore cannot quietly rot. Regulation maps obligation to mechanism, names what is not built, and notes that the Digital Omnibus moved the high-risk dates to December 2027 without amending the articles.
๐ License
MIT OR Apache-2.0, at your option.