agentplane 0.34.0

Durable, replayable agent runtime β€” the journal is the plan of record
Documentation

agentplane

A durable, replayable, policy-governed runtime for AI agents β€” in Rust. πŸ¦€

crates.io API docs License Status MSRV

Documentation Β· Getting started Β· API reference Β· What is built

Not a prompt framework. Not an agent library. The layer beneath those β€” the thing that makes an agent's actions survivable, auditable, and governable when it is calling real systems that move real money.

// Performs its effects once, and journals everything.
let outcome = runtime.run("reconcile", Tainted::trusted(input)).await?;

// Replay re-executes the logic and reads every effect back from the journal.
// No tool is called again. No clock is read again. No invoice is issued twice.
runtime.replay(outcome.run_id, Mode::Strict).await?;

πŸ”₯ The problem

Production agents fail in ways a better model does not fix:

  • A 40-minute run dies at minute 38, and the retry re-issues every invoice.
  • "Why did the agent refund €4,200?" has no answer, because the reasoning was prose in a log line.
  • Untrusted tool output steers the next tool call.
  • A prompt change ships with no way to know what it broke.

These are runtime problems. agentplane is a runtime.

πŸ’‘ The idea

The journal is the plan of record. Orchestration is deterministic and replayable. Everything non-deterministic β€” model inference, tool calls, the clock, randomness β€” is an effect: performed at most once, written to an append-only hash-chained log, and read back on replay.

Get that right and six things fall out of one mechanism: crash recovery, audit, cost accounting, regression testing, tamper evidence, and regulatory record-keeping. They stop being six subsystems that can each rot independently.

And critically: the audit trail is also the recovery mechanism, so it cannot quietly stop working β€” the system would stop working with it. Logging that exists only to satisfy an auditor always rots.

πŸš€ Try it

cargo run --example hello_skill        # one skill, one run, one replay β€” start here
cargo run --example durable_pipeline   # crash, resume, divergence
cargo run --example clearing_case      # correlation, obligations, human tasks
cargo run --example plan_graph         # multi-step plans, contract, provenance
cargo run --example governed_transfer --features manifest
                                        # field provenance, protected arguments
cargo run --example saga_checkout      # reverse compensation, replay-safe unwind
cargo run --example effect_group       # calls that take together, or not at all
cargo run --example memory_run         # private/team memory, provenance, recall
cargo run --example batch_run          # one act, many items: a resume that
                                        # re-settles nothing, and partial failure
                                        # as a terminal state
cargo run --example budget_pause       # a ceiling pauses the run; a raise
                                        # resumes it, on the record, nothing repeated
cargo run --example answered_doubt     # a call nobody can account for: a person
                                        # supplies the fact, the runtime keeps the
                                        # verdict, and giving up leaves a finding
cargo run --example operator_stop      # cancel a run and it unwinds; halt a
                                        # tenant, an agent or one revision, and
                                        # nothing new starts
cargo run --example observability      # the last mile: latency without replays,
                                        # gauges from the census, one alert
                                        # predicate β€” and the OTLP wiring
cargo run --example recovered_run      # an instance dies mid-run; the survivor's
                                        # sweep finds it and finishes it
cargo run --example bedrock_live --features bedrock
                                        # env-gated Amazon Bedrock Converse call
cargo run --example openai_live --features providers
                                        # env-gated OpenAI Responses call
cargo run --example tool_loop --features redb,fake-model,manifest
                                        # a model choosing tools, and four refusals
cargo run --example approved_call --features redb,fake-model,manifest
                                        # a person approves the exact call β€”
                                        # suspend, worklist, approve or refuse
cargo run --example planned_run --features redb,fake-model,manifest
                                        # plan once, execute without the model β€”
                                        # a prompt injection with no reader, and
                                        # an invented recipient refused
cargo run --example camel_live --features redb,providers,manifest
                                        # the same, against two real models: a
                                        # privileged planner and a quarantined
                                        # extractor (env-gated)
cargo run --example sealed_run --features redb,testkit,keyring
                                        # erase a case: every copy unreadable,
                                        # and the chain still verifies

# Calls a model and replays without calling it again β€” no API key, no network.
cargo run --example model_run --features redb,fake-model

# Digest-only multimodal dispatch and zero-I/O replay β€” also fully offline.
cargo run --example media_run --features redb,fake-model,media

# An agent whose prompt, model, result shape and ceilings come from a file.
cargo run --example manifest_run --features redb,fake-model,manifest

# A real MCP server in this process beside a typed Rust tool β€” one agent
# reaching both, and a strict replay that calls neither.
cargo run --example mcp_tools --features redb,fake-model,manifest,mcp

# Four agents, one plane: a coded editor that dictates the sequence, and a
# YAML desk that consults the same specialists as tool://agent/... grants.
cargo run --example blog_room --features redb,fake-model,manifest

# This plane served as an A2A 1.0 agent, called the way a peer would call it:
# a public card, authenticated methods, and a message that arrives untrusted.
cargo run --example a2a_peer --features redb,a2a-server,manifest

# Two planes in one process: a served reviewer and a desk that consults it
# through `cx.call_peer` β€” the peer sees the run's chain plus one link, and a
# strict replay of the desk's run never reaches the reviewer.
cargo run --example peer_call --features redb,testkit,manifest,a2a,a2a-server

# Live tokens for a human, one journaled completion for the machine β€” and a
# replay that performs neither.
cargo run --example streaming_run --features redb,fake-model

# One customer's approved €500, spent across two separate runs, then revoked β€”
# with the terms still readable afterwards.
cargo run --example standing_authority --features redb,fake-model

Or skip Rust entirely β€” a file and a key are the whole agent, and a file may hold a whole room: several manifests separated by ---, the Kubernetes packaging convention. Each document keeps its own digest β€” the file is packaging, not identity β€” and a run starts at the room's declared orchestrator:

cargo install agentplane --features cli
agentplane run examples/summariser.yaml --input '{"ticket": "printer on fire"}'
agentplane run examples/room.yaml       --input '{"topic": "durable execution"}'

Or without a Rust toolchain at all β€” needing one to run a YAML file rather defeats the point of the file:

docker run --rm -v "$PWD/examples:/work:ro" ghcr.io/hupe1980/agentplane \
  run /work/summariser.yaml --input '{"ticket": "printer on fire"}'

Distroless, nonroot, no shell. It runs --read-only --network none because the default journal is genuinely in memory and the example's provider is the deterministic fake β€” so the first run needs neither a disk nor the internet. :slim (the default, and :latest) carries every model provider β€” Anthropic, OpenAI, Gemini, Bedrock, any OpenAI-compatible server; :full adds MCP, the A2A peer server, the operator HTTP API, Cedar, key rings, governed media and Postgres. Both are multi-arch, cosign-signed keylessly, and carry SLSA build provenance and an SBOM attached to the digest:

cosign verify ghcr.io/hupe1980/agentplane:slim \
  --certificate-identity-regexp 'https://github.com/hupe1980/agentplane/.*' \
  --certificate-oidc-issuer https://token.actions.githubusercontent.com
gh attestation verify oci://ghcr.io/hupe1980/agentplane:slim -R hupe1980/agentplane

durable_pipeline prints the whole claim in four steps: a live run, a strict replay that touches nothing, a crash that resumes without repeating work, and a changed build that is quarantined instead of quietly rewriting history.

A first program is one import and about forty lines:

use agentplane::prelude::*;

And when it goes wrong, the plane answers with what it does have rather than with a variant name β€” fn main() reports through Debug, so on the errors you hold, Debug is the message:

Error: no skill provides capability 'demo.greeet' β€” this plane provides:
demo.greet. `run` takes a capability, not a skill name; a skill declares its own
with `SkillDescriptor::new(..).provides(..)`

A declarative tool loop runs from a file too β€” the manifest grants tool://tickets/read, and which transport reaches tickets is deployment wiring rather than part of the reviewed declaration:

agentplane run examples/tool-calling.yaml --input '{"ticket": "T-1"}' \
  --mcp "tickets=python3 examples/mcp-server.py"

A grant naming a server nobody wired is refused at build, not on every run. Needs --features cli,mcp-stdio, or the :full image.

And an agent can be hosted from the same file β€” the A2A 1.0 server that passes the protocol project's own conformance kit, started without writing Rust:

agentplane serve examples/served.yaml \
  --url http://localhost:8080 \
  --policy examples/serve-policy.cedar \
  --tokens examples/serve-tokens.yaml \
  --store ./served.redb

Add --operator-addr 127.0.0.1:9090 and the operator surface is served too β€” the worklist and task decisions, plus the backlogs an on-call person asks for by question rather than by id: what is quarantined, what is escalated, which obligations were missed, which messages reached nobody, and which webhook receivers stopped accepting (the full table). Every backlog that is work has a verb that empties it, including the hard one: a run stopped on an effect nobody can account for names the call, takes a person's answer about what actually happened, and is then judged again by the runtime β€” or written off, in which case what it left standing becomes an audit finding rather than leaving with the status (answering a quarantine). A missed obligation is answered the same way, by an account that names who looked. The one listing with no verb is dead letters, and it is a diagnosis rather than a queue β€” the fix is a correlation key in somebody's emitter β€” so it is ordered newest-first, which is what a page onto a list nothing removes from has to be. On their own listener, off unless asked for, and separated from the peer surface by policy (peer reaches a2a:*, operator reaches api:*) rather than by the port. A served plane also sweeps deadlines, task expiry, dead letters, due timers and abandoned runs β€” a lease that expired while still naming an owner is an instance that died holding the run, and the sweep takes it over and resumes it β€” so a run that sleeps, waits or loses its instance actually finishes.

And it drains on SIGTERM: stop accepting, answer what is in hand, close admission, and give the runs still executing --drain-secs to reach a journaled resting point. A process killed mid-call leaves an effect nothing can decide the outcome of, and a rolling deploy should not be the ordinary way a plane produces those (stopping an instance).

Both --policy and --tokens are required and have no defaults. That is the design rather than an inconvenience: a permissive engine and no engine are the same behaviour, and a server that authenticates nobody has no actor to record a decision against. A token may carry its caller's own scope and not_after; every run that caller starts is then admitted under a chain rooted at the caller β€” checked against the plan, refused once expired β€” and the journal names the caller, never the plane, as who the run acted for. Needs --features cli,a2a-server,cedar, or the :full image.

New here? β†’ docs/getting-started.md

πŸ“¦ What you get

🧾 A journal you can audit β€” append-only, hash-chained, per-record signatures naming the workload that wrote them, and a per-plane Merkle log so deleting a whole run is detectable
⏱️ Durable execution β€” crash mid-run and resume from the last completed effect. Recovery is initiated, not merely possible: a sweep finds every run whose owner died holding it and takes it over, and a scheduled stop drains rather than becoming a crash
πŸ—‚οΈ Cases, not long-lived workflows β€” runs stay minutes, business processes span months, so a deploy never migrates an in-flight workflow. Admission claims an idempotency key in the transaction that writes the first record, so a redelivery is answered with the original run
πŸ›‘οΈ Policy before live dispatch β€” a total, I/O-free gate; denials are journaled, strict replay never re-judges history, and plan authority is checked before step 1
🏷️ Field-level information flow β€” outbound arguments carry hierarchical provenance, so an authority-bearing field can require a trusted or named source while ordinary content stays untrusted. Volume is the axis a label lacks, so the size crossing each sink is journaled and max_egress_bytes bounds it
πŸ’Έ Budgets and tenant quotas that bind β€” a failed model call is billed for what it burned, because the provider bills for it too, and a replayed run reaches the same tally at the same point β†’ budgets
🧬 Effects that take together, or not at all β€” each reversible member records the concrete call that undoes it, built from what that call actually returned; an irreversible send is deferred to commit, so an aborted group never sends it β†’ effects
πŸ‘€ Human oversight on the call, not a summary of it β€” a task carries the exact tool and arguments about to be dispatched, and a read-only preview puts four thousand records on the reviewer's screen instead of older_than: "2024-01-01" β†’ worklists
πŸ”‘ Erasure that reaches the backups β€” payload bytes are sealed under a per-case key the crate never holds, so erasing a case destroys the key and the backup taken an hour ago becomes unreadable too. The chain commits to the ciphertext, so an auditor holding no keys still verifies the run β†’ erasure and keys
πŸ“„ An agent that is only a file β€” agentplane run agent.yaml. No Rust, no main, no skill. The digest covers the agent in its entirety, and the run is journaled and deterministically replayable

Ten rows, not the inventory. The full surface β€” the export/audit/restore toolchain, a durable manifest registry with an enumerable inventory, typed release, standing authorities, effect groups that commit with the journal, batch runs over 10⁡ items with per-item journals and an item-granular resume, the scoped emergency stop, the audited sweeper, a scheduled recovery drill and a retention pass that says what it could not reach, model drivers and streaming, MCP and A2A on both sides, signed Agent Cards, governed media and memory, multi-tenancy, quotas, witnessing, break-glass, and why there is no AllowAll anywhere β€” is documented mechanism by mechanism on the site: what you get, in full.

What is deliberately not built, and what will move β†’ docs/status.md

πŸ“š Documentation

πŸš€ Getting started β€” first run, first skill, first replay
🐣 Your first agent β€” a step-by-step tutorial: one agent, from an empty file to a durable, tool-using, pinnable declaration, no Rust required
🧠 Concepts β€” the ideas the rest is built from
πŸ—οΈ Architecture β€” the determinism boundary, the module layout, and where each mechanism lives
βš›οΈ The effect protocol β€” at-most-once outward calls, unknown outcomes, sagas, transactional groups, stopping a run
🧾 The journal β€” the hash chain, the signatures and the Merkle log, and the claims they refuse to make
πŸ—ΊοΈ Plans, cases and time β€” frozen authorization graphs, month-long cases, waits, timers, budgets, worklists
πŸ”Œ Models, agents and peers β€” everything this runtime calls that it does not own
πŸ“¦ Publishing and pinning agents β€” the manifest as an artifact, and a registry that will not rewrite a version
🍳 Cookbook β€” task-shaped recipes, including wiring an MCP server beside typed tools
πŸ“„ Manifest reference β€” every field, what enforces it, and what an absent value means; the published JSON Schema gives editors autocomplete and inline errors via one modeline
πŸ§ͺ Testing agents β€” the fake provider, fault injection, and proving a replay actually replayed
πŸ”¬ How this is proven β€” model-checked specifications, mutation-tested specs, and every guarantee broken on purpose
πŸ“ Record format β€” the normative wire specification: canonical JSON, the chain, the Merkle log, the export file. Enough to verify a history without this crate
πŸ” Security model β€” the trust boundary, and what it does not cover
πŸ—οΈ Erasure and keys β€” erasure that reaches backups, key rotation and revocation, and how tenants are kept apart
βš™οΈ Operations β€” deploying, HA, retention, observability
βš–οΈ Regulation β€” EU AI Act obligation by obligation, and what is missing
πŸ“‹ Status β€” what is pre-alpha, what to pin, what is deliberately absent
⬆️ Upgrading β€” what breaks between pre-alpha releases, and the shortest correct fix
πŸ“œ Changelog β€” what changed and when, including every mechanism's reasoning as it landed
🀝 Contributing β€” the assurance ladder, and how to run it

πŸ§ͺ Assurance

Each layer answers a question the others structurally cannot.

just              # list every check
just ci           # lint Β· every feature alone Β· tests Β· examples Β· docs Β· packaging
just ci-full      # the above, plus TLA+ specs and the full mutation sweep

python3 tools/mutants.py <name> --verify   # break one guarantee, run its test

Two are unusual enough to name:

πŸ”¬ Formal specs. TLA+ specifications are model-checked on every push β€” the effect protocol, effect groups, retry safety, sagas, fencing, authorization, delegation. And because a spec whose invariants cannot be violated proves nothing, each is re-checked against deliberately broken copies of itself; every mutant must be caught by the specific invariant written for it.

πŸ“ A second reader of the record format. The format specification is normative prose, and tools/verify_export.py is written from it and reads none of this crate's Rust β€” enforced by a guard, because a verifier that consulted src/ would agree with the implementation by construction. just verify-golden runs it: it re-derives all 27 record vectors from their parsed values with its own canonicalizer and chain digest, verifies the sealed export end to end, and then damages that export six ways and asserts each is reported. Vectors a project generates and then checks are that project agreeing with itself; this is the part that is not.

πŸ”— An anchor from a party this plane does not control. The hash chain, the signatures and the Merkle log all draw both halves of their comparison from the store, so an operator who removes a run and recomputes the tree satisfies every one of them β€” and agentplane audit says so rather than reporting a clean history. RuntimeBuilder::witnesses(..) submits each checkpoint to witnesses over C2SP tlog-witness on the periodic sweep, and agentplane audit --witness <prefix> --witness-key <name>=<key> reads back what they hold. That second direction is the one that matters: the anchor reaches a reader who did not get it from the operator, and two witnesses holding one tree size with two different roots is a split view no single anchor exhibits.

🧾 Conformance by the protocol's own kit. just test-a2a-tck runs the official a2a-tck against this crate's A2A server on a live socket. Every other A2A test drives this server with this crate's own client, which proves symmetry, not conformance β€” a client and server written from the same misreading agree everywhere. The kit's first run found five defects no in-repo test could reach.

🌐 Tests against a real provider. just test-live runs the OpenAI, Gemini and OpenAI-compatible drivers, plus the embedding wire, against the actual APIs. They are gated twice β€” an explicit AGENTPLANE_LIVE=1 and a key β€” because a credential being available is not a decision to spend money with it, and they are never part of ci. They exist because a stubbed provider is structurally unable to have the defects a real one finds: it never rejects a malformed request and never returns a shape the driver mis-reads. What they catch had passed every offline test β€” including a plan format no provider with constrained decoding accepts, which left the dual-model execution kind unable to run for real at all. The Gemini battery is the sharpest case: a thought signature is minted and validated by Google, so a canned server accepts whatever a fixture tells it to and says nothing about whether Gemini takes the signature back β€” the one check that distinguishes a driver carrying the model's turn verbatim from one rebuilding it, which is where the rest of the ecosystem has been losing this.

🧬 Mutation testing over the code. Every load-bearing guarantee is broken on purpose, and the test named for each one must fail. A mutation caught by some other test is reported weak, not passing β€” that usually means the guarantee has no test of its own and is being held up by one that could be rewritten without anyone noticing what it protected.

This is not decoration. The project shipped an unfalsifiable guarantee once: the refusal to replan on untrusted data was implemented, tested, and green β€” and deleting it would have failed no test, because the fixtures laundered the taint before it reached the check. It was found by accident. The sweep is so the next one is not.

It runs on every push, sharded ten ways. MUTANTS_SHARD=k/n takes a contiguous slice of a list grouped by the feature set each mutation builds under, cut on measured seconds rather than count β€” a mutation checked by a library unit test costs six times one checked in an integration binary, and a matrix finishes when its slowest job does. Each shard needs its own checkout: the sweep rewrites source in place.

just anchors is the cheap half, and it checks text rather than types: a mutation still matching the code it names does not prove its replacement still compiles. That is what --verify is for.

🚫 Non-goals

agentplane does not Use instead
Ship a prompt library or IDE Your prompts; agentplane pins the manifest that governs them by digest
Route or proxy model traffic LiteLLM, Bifrost. The drivers themselves ship β€” what is out of scope is choosing between them at runtime
Implement a vector database LanceDB / pgvector behind the SemanticRetriever seam; embedding is a journaled effect so the query vector is history rather than a recomputation
Ship a built-in tool catalogue Write a typed Tool, or wire an MCP server. The tools other frameworks ship are mostly provider-hosted β€” they run during generation, so the call is never announced, authorized, metered or replayable
Replace a deterministic protocol engine Keep it; agentplane sits beside it, never inside it
Require Kubernetes One static binary
Train, fine-tune, or serve models Permanently out of scope
Grade output quality It emits replayable traces; grade them elsewhere
Interpret payload contents Payloads are opaque, and labeled
Claim regulatory compliance It provides technical means; compliance is the deployer's

Who should not use this: a team running three agents against low-stakes data. The complexity is justified when agents touch money, meters, or regulated records.

πŸ“Œ Status

Pre-alpha, pre-release, no API stability. Breaking changes land without deprecation. The journal record format and the storage schema will change.

Rust 1.94.1+. #![forbid(unsafe_code)]. One crate, feature-gated: an embedded redb store by default β€” pure Rust, two crates deep, no C toolchain β€” with everything else opt-in.

Honest framing on regulation: agentplane is not "compliant" and cannot be. Compliance attaches to a system in a context, assessed by its provider or deployer. What this gives you is the technical means to discharge EU AI Act Articles 12 and 14 β€” means that are already load-bearing for recovery and testing, and therefore cannot quietly rot. Regulation maps obligation to mechanism, names what is not built, and notes that the Digital Omnibus moved the high-risk dates to December 2027 without amending the articles.

πŸ“„ License

MIT OR Apache-2.0, at your option.