agentplane
A durable, replayable, policy-governed runtime for AI agents β in Rust. π¦
Documentation Β· Getting started Β· API reference Β· What is built
Not a prompt framework. Not an agent library. The layer beneath those β the thing that makes an agent's actions survivable, auditable, and governable when it is calling real systems that move real money.
// Performs its effects once, and journals everything.
let outcome = runtime.run.await?;
// Replay re-executes the logic and reads every effect back from the journal.
// No tool is called again. No clock is read again. No invoice is issued twice.
runtime.replay.await?;
π₯ The problem
Production agents fail in ways a better model does not fix:
- A 40-minute run dies at minute 38, and the retry re-issues every invoice.
- "Why did the agent refund β¬4,200?" has no answer, because the reasoning was prose in a log line.
- Untrusted tool output steers the next tool call.
- A prompt change ships with no way to know what it broke.
These are runtime problems. agentplane is a runtime.
π‘ The idea
The journal is the plan of record. Orchestration is deterministic and replayable. Everything non-deterministic β model inference, tool calls, the clock, randomness β is an effect: performed at most once, written to an append-only hash-chained log, and read back on replay.
Get that right and six things fall out of one mechanism: crash recovery, audit, cost accounting, regression testing, tamper evidence, and regulatory record-keeping. They stop being six subsystems that can each rot independently.
And critically: the audit trail is also the recovery mechanism, so it cannot quietly stop working β the system would stop working with it. Logging that exists only to satisfy an auditor always rots.
π Try it
# field provenance, protected arguments
# re-settles nothing, and partial failure
# as a terminal state
# resumes it, on the record, nothing repeated
# supplies the fact, the runtime keeps the
# verdict, and giving up leaves a finding
# tenant, an agent or one revision, and
# nothing new starts
# gauges from the census, one alert
# predicate β and the OTLP wiring
# sweep finds it and finishes it
# env-gated Amazon Bedrock Converse call
# env-gated OpenAI Responses call
# a model choosing tools, and four refusals
# a person approves the exact call β
# suspend, worklist, approve or refuse
# plan once, execute without the model β
# a prompt injection with no reader, and
# an invented recipient refused
# erase a case: every copy unreadable,
# and the chain still verifies
# Calls a model and replays without calling it again β no API key, no network.
# Digest-only multimodal dispatch and zero-I/O replay β also fully offline.
# An agent whose prompt, model, result shape and ceilings come from a file.
# A real MCP server in this process beside a typed Rust tool β one agent
# reaching both, and a strict replay that calls neither.
# Four agents, one plane: a coded editor that dictates the sequence, and a
# YAML desk that consults the same specialists as tool://agent/... grants.
# This plane served as an A2A 1.0 agent, called the way a peer would call it:
# a public card, authenticated methods, and a message that arrives untrusted.
# Two planes in one process: a served reviewer and a desk that consults it
# through `cx.call_peer` β the peer sees the run's chain plus one link, and a
# strict replay of the desk's run never reaches the reviewer.
# Live tokens for a human, one journaled completion for the machine β and a
# replay that performs neither.
# One customer's approved β¬500, spent across two separate runs, then revoked β
# with the terms still readable afterwards.
Or skip Rust entirely β a file and a key are the whole agent, and a file may
hold a whole room: several manifests separated by ---, the Kubernetes
packaging convention. Each document keeps its own digest β the file is
packaging, not identity β and a run starts at the room's declared orchestrator:
Or without a Rust toolchain at all β needing one to run a YAML file rather defeats the point of the file:
Distroless, nonroot, no shell. It runs --read-only --network none because the
default journal is genuinely in memory and the example's provider is the
deterministic fake β so the first run needs neither a disk nor the internet.
:slim (the default, and :latest) carries every model provider β Anthropic,
OpenAI, Gemini, Bedrock, any OpenAI-compatible server; :full adds MCP, the A2A
peer server, the operator HTTP API, Cedar, key rings, governed media and
Postgres. Both are multi-arch, cosign-signed keylessly, and carry SLSA build
provenance and an SBOM attached to the digest:
durable_pipeline prints the whole claim in four steps: a live run, a strict
replay that touches nothing, a crash that resumes without repeating work, and a
changed build that is quarantined instead of quietly rewriting history.
A first program is one import and about forty lines:
use *;
And when it goes wrong, the plane answers with what it does have rather than
with a variant name β fn main() reports through Debug, so on the errors you
hold, Debug is the message:
Error: no skill provides capability 'demo.greeet' β this plane provides:
demo.greet. `run` takes a capability, not a skill name; a skill declares its own
with `SkillDescriptor::new(..).provides(..)`
A declarative tool loop runs from a file too β the manifest grants
tool://tickets/read, and which transport reaches tickets is deployment
wiring rather than part of the reviewed declaration:
A grant naming a server nobody wired is refused at build, not on every run.
Needs --features cli,mcp-stdio, or the :full image.
And an agent can be hosted from the same file β the A2A 1.0 server that passes the protocol project's own conformance kit, started without writing Rust:
Add --operator-addr 127.0.0.1:9090 and the operator surface is served too β
the worklist and task decisions, plus the backlogs an on-call person asks for by
question rather than by id: what is quarantined, what is escalated, which
obligations were missed, which messages reached nobody, and which webhook
receivers stopped accepting (the full table).
Every backlog that is work has a verb that empties it, including the hard one:
a run stopped on an effect nobody can account for names the call, takes a
person's answer about what actually happened, and is then judged again by the
runtime β or written off, in which case what it left standing becomes an audit
finding rather than leaving with the status
(answering a quarantine).
A missed obligation is answered the same way, by an account that names who
looked. The one listing with no verb is dead letters, and it is a diagnosis
rather than a queue β the fix is a correlation key in somebody's emitter β so it
is ordered newest-first, which is what a page onto a list nothing removes from
has to be.
On their own listener, off
unless asked for, and separated from the peer surface by policy (peer reaches
a2a:*, operator reaches api:*) rather than by the port. A served plane also
sweeps deadlines, task expiry, dead letters, due timers and abandoned runs
β a lease that expired while still naming an owner is an instance that died
holding the run, and the sweep takes it over and resumes it β so a run that
sleeps, waits or loses its instance actually finishes.
Both --policy and --tokens are required and have no defaults. That is the
design rather than an inconvenience: a permissive engine and no engine are the
same behaviour, and a server that authenticates nobody has no actor to record a
decision against. A token may carry its caller's own scope and not_after;
every run that caller starts is then admitted under a chain rooted at the
caller β checked against the plan, refused once expired β and the journal
names the caller, never the plane, as who the run acted for. Needs
--features cli,a2a-server,cedar, or the :full image.
New here? β docs/getting-started.md
π¦ What you get
| π§Ύ | A journal you can audit β append-only, hash-chained, per-record signatures naming the workload that wrote them, and a per-plane Merkle log so deleting a whole run is detectable |
| β±οΈ | Durable execution β crash mid-run and resume from the last completed effect. Recovery is initiated, not merely possible: a sweep finds every run whose owner died holding it and takes it over |
| ποΈ | Cases, not long-lived workflows β runs stay minutes, business processes span months, so a deploy never migrates an in-flight workflow. Admission claims an idempotency key in the transaction that writes the first record, so a redelivery is answered with the original run |
| π‘οΈ | Policy before live dispatch β a total, I/O-free gate; denials are journaled, strict replay never re-judges history, and plan authority is checked before step 1 |
| π·οΈ | Field-level information flow β outbound arguments carry hierarchical provenance, so an authority-bearing field can require a trusted or named source while ordinary content stays untrusted |
| πΈ | Budgets and tenant quotas that bind β a failed model call is billed for what it burned, because the provider bills for it too, and a replayed run reaches the same tally at the same point β budgets |
| 𧬠| Effects that take together, or not at all β each reversible member records the concrete call that undoes it, built from what that call actually returned; an irreversible send is deferred to commit, so an aborted group never sends it β effects |
| π€ | Human oversight on the call, not a summary of it β a task carries the exact tool and arguments about to be dispatched, and a read-only preview puts four thousand records on the reviewer's screen instead of older_than: "2024-01-01" β worklists |
| π | Erasure that reaches the backups β payload bytes are sealed under a per-case key the crate never holds, so erasing a case destroys the key and the backup taken an hour ago becomes unreadable too. The chain commits to the ciphertext, so an auditor holding no keys still verifies the run β erasure and keys |
| π | An agent that is only a file β agentplane run agent.yaml. No Rust, no main, no skill. The digest covers the agent in its entirety, and the run is journaled and deterministically replayable |
Ten rows, not the inventory. The full surface β the export/audit/restore
toolchain, a durable manifest registry with an enumerable inventory, typed
release, standing authorities, effect groups that commit with the journal,
batch runs over 10β΅ items with per-item journals and an item-granular resume,
the scoped emergency stop, the audited sweeper, a scheduled recovery drill and a
retention pass that says what it could not reach, model drivers and streaming,
MCP and A2A on both sides, signed Agent Cards, governed media and memory,
multi-tenancy, quotas, witnessing, break-glass, and why there is no AllowAll
anywhere β is documented mechanism by mechanism on the site:
what you get, in full.
What is deliberately not built, and what will move β docs/status.md
π Documentation
| π | Getting started β first run, first skill, first replay |
| π£ | Your first agent β a step-by-step tutorial: one agent, from an empty file to a durable, tool-using, pinnable declaration, no Rust required |
| π§ | Concepts β the ideas the rest is built from |
| ποΈ | Architecture β the determinism boundary, the module layout, and where each mechanism lives |
| βοΈ | The effect protocol β at-most-once outward calls, unknown outcomes, sagas, transactional groups, stopping a run |
| π§Ύ | The journal β the hash chain, the signatures and the Merkle log, and the claims they refuse to make |
| πΊοΈ | Plans, cases and time β frozen authorization graphs, month-long cases, waits, timers, budgets, worklists |
| π | Models, agents and peers β everything this runtime calls that it does not own |
| π¦ | Publishing and pinning agents β the manifest as an artifact, and a registry that will not rewrite a version |
| π³ | Cookbook β task-shaped recipes, including wiring an MCP server beside typed tools |
| π | Manifest reference β every field, what enforces it, and what an absent value means; the published JSON Schema gives editors autocomplete and inline errors via one modeline |
| π§ͺ | Testing agents β the fake provider, fault injection, and proving a replay actually replayed |
| π¬ | How this is proven β model-checked specifications, mutation-tested specs, and every guarantee broken on purpose |
| π | Record format β the normative wire specification: canonical JSON, the chain, the Merkle log, the export file. Enough to verify a history without this crate |
| π | Security model β the trust boundary, and what it does not cover |
| ποΈ | Erasure and keys β erasure that reaches backups, key rotation and revocation, and how tenants are kept apart |
| βοΈ | Operations β deploying, HA, retention, observability |
| βοΈ | Regulation β EU AI Act obligation by obligation, and what is missing |
| π | Status β what is pre-alpha, what to pin, what is deliberately absent |
| β¬οΈ | Upgrading β what breaks between pre-alpha releases, and the shortest correct fix |
| π | Changelog β what changed and when, including every mechanism's reasoning as it landed |
| π€ | Contributing β the assurance ladder, and how to run it |
π§ͺ Assurance
Each layer answers a question the others structurally cannot.
Two are unusual enough to name:
π¬ Formal specs. TLA+ specifications are model-checked on every push β the effect protocol, effect groups, retry safety, sagas, fencing, authorization, delegation. And because a spec whose invariants cannot be violated proves nothing, each is re-checked against deliberately broken copies of itself; every mutant must be caught by the specific invariant written for it.
π A second reader of the record format. The
format specification is
normative prose, and tools/verify_export.py is written from it and reads none
of this crate's Rust β enforced by a guard, because a verifier that consulted
src/ would agree with the implementation by construction. just verify-golden
runs it: it re-derives all 27 record vectors from their parsed values with
its own canonicalizer and chain digest, verifies the sealed export end to end,
and then damages that export six ways and asserts each is reported. Vectors a
project generates and then checks are that project agreeing with itself; this
is the part that is not.
π An anchor from a party this plane does not control. The hash chain, the
signatures and the Merkle log all draw both halves of their comparison from
the store, so an operator who removes a run and recomputes the tree satisfies
every one of them β and agentplane audit says so rather than reporting a
clean history. RuntimeBuilder::witnesses(..) submits each checkpoint to
witnesses over C2SP tlog-witness on the periodic sweep, and agentplane audit --witness <prefix> --witness-key <name>=<key> reads back what they
hold. That second direction is the one that matters: the anchor reaches a
reader who did not get it from the operator, and two witnesses holding one
tree size with two different roots is a split view no single anchor exhibits.
π§Ύ Conformance by the protocol's own kit. just test-a2a-tck runs the
official a2a-tck against this crate's
A2A server on a live socket. Every other A2A test drives this server with this
crate's own client, which proves symmetry, not conformance β a client and
server written from the same misreading agree everywhere. The kit's first run
found five defects no in-repo test could reach.
π Tests against a real provider. just test-live runs the OpenAI and
Gemini drivers against the actual APIs. They are gated twice β an explicit AGENTPLANE_LIVE=1
and a key β because a credential being available is not a decision to spend
money with it, and they are never part of ci. They exist because a stubbed
provider is structurally unable to have the defects a real one finds: it never
rejects a malformed request and never returns a shape the driver mis-reads.
Writing them found two, both of which every offline test had passed. The Gemini
battery is the sharpest case: a thought signature is minted and validated by
Google, so a canned server accepts whatever a fixture tells it to and says
nothing about whether Gemini takes the signature back β the one check that
distinguishes a driver carrying the model's turn verbatim from one rebuilding
it, which is where the rest of the ecosystem has been losing this.
𧬠Mutation testing over the code. Every load-bearing guarantee is broken on purpose, and the test named for each one must fail. A mutation caught by some other test is reported weak, not passing β that usually means the guarantee has no test of its own and is being held up by one that could be rewritten without anyone noticing what it protected.
This is not decoration. The project shipped an unfalsifiable guarantee once: the refusal to replan on untrusted data was implemented, tested, and green β and deleting it would have failed no test, because the fixtures laundered the taint before it reached the check. It was found by accident. The sweep is so the next one is not.
It runs on every push, sharded ten ways. MUTANTS_SHARD=k/n takes a
contiguous slice of a list grouped by the feature set each mutation builds
under, cut on measured seconds rather than count β a mutation checked by a
library unit test costs six times one checked in an integration binary, and a
matrix finishes when its slowest job does. Each shard needs its own checkout:
the sweep rewrites source in place.
just anchors is the cheap half, and it checks text rather than types: a
mutation still matching the code it names does not prove its replacement still
compiles. That is what --verify is for.
π« Non-goals
| agentplane does not | Use instead |
|---|---|
| Ship a prompt library or IDE | Your prompts; agentplane pins the manifest that governs them by digest |
| Route or proxy model traffic | LiteLLM, Bifrost. The drivers themselves ship β what is out of scope is choosing between them at runtime |
| Implement a vector database | LanceDB / pgvector behind the SemanticRetriever seam; embedding is a journaled effect so the query vector is history rather than a recomputation |
| Ship a built-in tool catalogue | Write a typed Tool, or wire an MCP server. The tools other frameworks ship are mostly provider-hosted β they run during generation, so the call is never announced, authorized, metered or replayable |
| Replace a deterministic protocol engine | Keep it; agentplane sits beside it, never inside it |
| Require Kubernetes | One static binary |
| Train, fine-tune, or serve models | Permanently out of scope |
| Grade output quality | It emits replayable traces; grade them elsewhere |
| Interpret payload contents | Payloads are opaque, and labeled |
| Claim regulatory compliance | It provides technical means; compliance is the deployer's |
Who should not use this: a team running three agents against low-stakes data. The complexity is justified when agents touch money, meters, or regulated records.
π Status
Pre-alpha, pre-release, no API stability. Breaking changes land without deprecation. The journal record format and the storage schema will change.
Rust 1.94.1+. #![forbid(unsafe_code)]. One crate, feature-gated: an embedded
redb store by default β pure Rust, two crates
deep, no C toolchain β with everything else opt-in.
Honest framing on regulation: agentplane is not "compliant" and cannot be. Compliance attaches to a system in a context, assessed by its provider or deployer. What this gives you is the technical means to discharge EU AI Act Articles 12 and 14 β means that are already load-bearing for recovery and testing, and therefore cannot quietly rot. Regulation maps obligation to mechanism, names what is not built, and notes that the Digital Omnibus moved the high-risk dates to December 2027 without amending the articles.
π License
MIT OR Apache-2.0, at your option.