agentplane 0.1.0

Durable, replayable agent runtime โ€” the journal is the plan of record
Documentation

agentplane

A durable, replayable, policy-governed runtime for AI agents โ€” in Rust. ๐Ÿฆ€

License Status MSRV

Not a prompt framework. Not an agent library. The layer beneath those โ€” the thing that makes an agent's actions survivable, auditable, and governable when it is calling real systems that move real money.

// Performs its effects once, and journals everything.
let outcome = runtime.run("reconcile", input).await?;

// Replay re-executes the logic and reads every effect back from the journal.
// No tool is called again. No clock is read again. No invoice is issued twice.
runtime.replay(outcome.run_id, Mode::Strict).await?;

๐Ÿ”ฅ The problem

Production agents fail in ways a better model does not fix:

  • A 40-minute run dies at minute 38, and the retry re-issues every invoice.
  • "Why did the agent refund โ‚ฌ4,200?" has no answer, because the reasoning was prose in a log line.
  • Untrusted tool output steers the next tool call.
  • A prompt change ships with no way to know what it broke.

These are runtime problems. agentplane is a runtime.

๐Ÿ’ก The idea

The journal is the plan of record. Orchestration is deterministic and replayable. Everything non-deterministic โ€” model inference, tool calls, the clock, randomness โ€” is an effect: performed at most once, written to an append-only hash-chained log, and read back on replay.

Get that right and six things fall out of one mechanism: crash recovery, audit, cost accounting, regression testing, tamper evidence, and regulatory record-keeping. They stop being six subsystems that can each rot independently.

And critically: the audit trail is also the recovery mechanism, so it cannot quietly stop working โ€” the system would stop working with it. Logging that exists only to satisfy an auditor always rots.

๐Ÿš€ Try it

cargo run --example durable_pipeline   # crash, resume, divergence
cargo run --example clearing_case      # correlation, obligations, human tasks
cargo run --example plan_graph         # multi-step plans, contract, provenance

# Calls a model and replays without calling it again โ€” no API key, no network.
cargo run --example model_run --features redb,testkit

durable_pipeline prints the whole claim in four steps: a live run, a strict replay that touches nothing, a crash that resumes without repeating work, and a changed build that is quarantined instead of quietly rewriting history.

New here? โ†’ docs/getting-started.md

๐Ÿ“ฆ What you get

๐Ÿงพ A journal you can audit โ€” append-only, hash-chained, per-record signatures naming the workload that wrote them, and a per-plane Merkle log so deleting a whole run is detectable
โฑ๏ธ Durable execution โ€” crash mid-run and resume from the last completed effect; a suspended run costs a row on disk, not a task
๐Ÿ—‚๏ธ Cases, not long-lived workflows โ€” runs stay minutes, business processes span months, so a deploy never has to migrate an in-flight workflow
๐Ÿ›ก๏ธ Policy before every effect โ€” a total, I/O-free gate; a run denied at step 7 never starts at step 1
๐Ÿท๏ธ Information-flow labels โ€” may this principal act and may this value go there are different questions, and both are answerable
๐Ÿ’ธ Budgets that bind โ€” a failed model call is billed for what it burned, because the provider bills for it too
๐Ÿ‘ค Human oversight โ€” durable worklists with four-eyes, declared expiry behaviour, and an operator who can stop a run and have it unwind
๐Ÿ”Œ Real wires โ€” MCP tools, A2A peers, Anthropic and OpenAI drivers, each with a failure mapping that says whether the call landed

Full inventory, including what is not built โ†’ docs/status.md

๐Ÿ“š Documentation

๐Ÿš€ Getting started โ€” first run, first skill, first replay
๐Ÿง  Concepts โ€” the ideas the rest is built from
๐Ÿ—๏ธ Architecture โ€” how it actually works, mechanism by mechanism
๐Ÿณ Cookbook โ€” task-shaped recipes
๐Ÿ” Security model โ€” the trust boundary, and what it does not cover
โš™๏ธ Operations โ€” deploying, HA, retention, observability
โš–๏ธ Regulation โ€” EU AI Act obligation by obligation, and what is missing
๐Ÿ“‹ Status โ€” built vs designed-not-built
๐Ÿค Contributing โ€” the assurance ladder, and how to run it

๐Ÿงช Assurance

Each layer answers a question the others structurally cannot.

just              # list every check
just ci           # lint ยท 3 feature configs ยท examples ยท docs ยท packaging
just ci-full      # the above, plus TLA+ specs and the full mutation sweep

Two are unusual enough to name:

๐Ÿ”ฌ Formal specs. Six TLA+ specifications are model-checked on every push โ€” the effect protocol, retry safety, sagas, fencing, authorization, delegation. And because a spec whose invariants cannot be violated proves nothing, each is re-checked against 18 deliberately broken copies of itself; every mutant must be caught by the specific invariant written for it.

๐Ÿงฌ Mutation testing over the code. 106 guarantees are broken on purpose, and the test named for each one must fail. A mutation caught by some other test is reported weak, not passing โ€” that usually means the guarantee has no test of its own and is being held up by one that could be rewritten without anyone noticing what it protected.

This is not decoration. The project shipped an unfalsifiable guarantee once: the refusal to replan on untrusted data was implemented, tested, and green โ€” and deleting it would have failed no test, because the fixtures laundered the taint before it reached the check. It was found by accident. The sweep is so the next one is not.

๐Ÿšซ Non-goals

agentplane does not Use instead
Ship a prompt library or IDE Your manifests; agentplane hashes and versions them
Route or proxy model traffic LiteLLM, Bifrost, your own ModelProvider
Implement a vector database LanceDB / pgvector behind a seam
Replace a deterministic protocol engine Keep it; agentplane sits beside it, never inside it
Require Kubernetes One static binary
Train, fine-tune, or serve models Permanently out of scope
Grade output quality It emits replayable traces; grade them elsewhere
Interpret payload contents Payloads are opaque, and labeled
Claim regulatory compliance It provides technical means; compliance is the deployer's

Who should not use this: a team running three agents against low-stakes data. The complexity is justified when agents touch money, meters, or regulated records.

๐Ÿ“Œ Status

Pre-alpha, pre-release, no API stability. Breaking changes land without deprecation. The journal record format and the storage schema will change.

Rust 1.94+. #![forbid(unsafe_code)]. One crate, feature-gated: an embedded redb store by default โ€” pure Rust, two crates deep, no C toolchain โ€” with everything else opt-in.

Honest framing on regulation: agentplane is not "compliant" and cannot be. Compliance attaches to a system in a context, assessed by its provider or deployer. What this gives you is the technical means to discharge EU AI Act Articles 12 and 14 โ€” means that are already load-bearing for recovery and testing, and therefore cannot quietly rot. Regulation maps obligation to mechanism, names what is not built, and notes that the Digital Omnibus moved the high-risk dates to December 2027 without amending the articles.

๐Ÿ“„ License

MIT OR Apache-2.0, at your option.