lernie 0.0.2

A git-backed agent harness
Documentation

lernie

Release-plz

A git-backed agent harness. Design spec: docs/ARCHITECTURE.md. Principles catalog: docs/PRINCIPLES.md. Vocabulary reference: docs/TAXONOMY.md. Promise suite (the user stories 0.0.1 is evaluated against): docs/USER_STORIES.md.

CI runs make ci (fmt-check + lint + coverage with the 100% gate + test-install) on every push and pull request to main. The e2e tests exec the real provider adapter bz, which the test targets install themselves at the pinned version (see The pinned adapter under test). The Rust toolchain is pinned in rust-toolchain.toml — CI, the pre-commit gate, and every contributor build under the same rustc/rustfmt/clippy. That pin binds this git checkout only; it is excluded from the published crate, whose supported floor is the declared rust-version = "1.88" (the crate's let chains, not edition 2024's 1.85).

One command surface, two bindings

lernie is defined once as a command surface — the set of verbs, their arguments, and their products (ARCH §3.4). It is consumable two ways, and both are the same control plane:

  • Exec binding — run the lernie binary: exec("lernie", args) with env-var auth. This is what the CLI and every frontend use.
  • Linked binding — depend on the lernie crate and drive the same verb entries in-process. The crate's entire public API is lernie::cmd (the Cli/Command clap surface, one run entry per verb, the Fx/Outcome/Error binding seam, and the prelude binding preludes). The linked binding is pin-exact 0.x only — no semver stability, the posture brazen takes toward lernie.

Parity between the two is enforced mechanically, not by convention. tests/command_surface_parity/ asserts the bijection at three depths: it pairs each verb's Command variant with its module's entry as function values, so the compiler — not an assertion — proves the two share one argument type and one product type; it walks the crate's whole module graph (via syn) and asserts that every externally reachable declaration (item, field, enum variant, method, derive, trait impl) is exactly a verb's entry, its arguments, its products, or the binding preludes, with every src/**/*.rs proven reachable so nothing can hide in a file the walk never opened; and it asserts, per verb, that the CLI's introspected argument set (via clap) is exactly that verb's public Args fields — same names, same arity, same named-vs-positional form. It rides make check (hence the pre-commit hook and GitHub Actions), so a divergence between the linked surface and the CLI fails the build.

Quickstart

cargo install lernie --locked                    # or: make install, from a clone
cargo install brazen --version =0.0.4 --locked   # the provider adapter, always needed
lernie new ~/work/chat    # create a workspace (bare repo.git + config/default)
ANTHROPIC_API_KEY=... lernie prompt ~/work/chat 'hello'

Three install routes, not one — see Install for what each lays down. Every route needs bz; only make install installs it.

Install

There are three routes, and they do not lay down the same things. All three need a second binary — the provider adapter bz — which only the Makefile route installs for you.

cargo install lernie release tarball make install
binaries lernie lernie lernie, agent-eval, lernie-eval-agent
installs bz no no yes, at the pin
runs lernie prime no no yes
lands where cargo's bin dir wherever you unpack it $INSTALL_PREFIX/bin

From crates.io

cargo install lernie --locked
cargo install brazen --version =0.0.4 --locked   # the pinned provider adapter
lernie prime                                     # found the harness root

You get the lernie binary alone, in cargo's bin directory (~/.cargo/bin unless --root/CARGO_INSTALL_ROOT says otherwise) — no agent-eval, no lernie-eval-agent, no bz, and nothing runs after the build. The lernie prime line is optional but explicit: prime founds the harness root (below), and lernie new founds it too on its way to creating a workspace, so a user who skips prime is not stranded — only uninformed about where their state went.

From a GitHub release

Each v* release carries lernie-x86_64-unknown-linux-gnu.tar.gz: the lernie binary, this README, and the license. Unpack it, put lernie on your PATH, then run the cargo install brazen and lernie prime lines above — the tarball ships no adapter and runs nothing.

From a clone, with make

make install                                  # default: ~/.local/bin, XDG homes
make install INSTALL_PREFIX=/usr/local        # binaries -> /usr/local/bin/
make install LERNIE_HOME=/opt/lernie          # collapse both homes -> /opt/lernie/

make install runs a release build and then:

  1. Installs lernie, agent-eval, and lernie-eval-agent into $INSTALL_PREFIX/bin with install -m 0755 (atomic overwrite, no symlinks). Make sure that directory is on your PATH.
  2. Installs the provider adapter — brazen's bz — with cargo install brazen --version =<pin> --locked, where the pin is the brazen = "=<pin>" dependency in Cargo.toml — its one home; the Makefile and the load-time guard both derive from that line. One binary serves every provider (ARCH §4.4); the harness resolves bz on PATH, and a load-time guard rejects any bz whose version differs from the pin.
  3. Founds the harness root by invoking lernie prime — the single verb that seeds the installation substrate (ARCH §2.2), so the Makefile no longer duplicates the seeding. prime resolves the roots (XDG split, collapsed by LERNIE_HOME) and lays down the default models.yaml under the config root (model capabilities + context windows — no endpoints or auth, which are brazen's), the tools/ and skills/ pools and the workspaces/ tree under the data root, and the empty workflows/ templates dir. It is seed-if-absent throughout: a second run changes nothing, and a hand-edited models.yaml (or any operator-added pool entry) survives a re-install. The shipped assets are embedded in the binary, so prime needs no source tree — LERNIE_HOME=<dir> lernie prime seeds any fresh home. There is no frozen profile pool: the config a workspace runs under is its own config/default commit, authored from template/ at lernie new (fork is the freeze, ARCH §2.2).
  4. Smoke-tests the freshly installed binaries with lernie --version and a throwaway lernie new. Failure aborts the install with a non-zero exit.

Its closing banner prints what the other two routes leave you to find out: the install prefix, both harness roots, and the bz commands below.

The adapter is a second binary

Nothing prompts without bz. It is brazen's one stateless binary for every provider (ARCH §4.4), it is pinned exactly, and lernie refuses a bz at any other version rather than downgrading silently:

cargo install brazen --version =0.0.4 --locked

The pin is not folklore you have to read this file for — the installed binary carries it: lernie --version prints the linked pin beside its own version, lernie <version> (brazen 0.0.4). Its one home is the brazen = "=<pin>" line in Cargo.toml; the Makefile's BRAZEN_PIN, the load-time guard, lernie --version, and every pin printed in this file all derive from that line (a test holds them equal). With no bz at all, the first verb that drives a model call says so and hands you the command above.

Provider endpoints, auth, and wire dialects live entirely in brazen's own config (~/.config/brazen/config.toml; inspect with bz --dump-config, authenticate with bz --login --provider <id>). lernie references a provider row by name and never sees credential material (ARCH §4.1).

Where the state goes

The harness root is the installation-global substrate (ARCH §2.2), split by XDG lifetime: $XDG_CONFIG_HOME/lernie (hand-edited declarations — models.yaml, workflows/) and $XDG_DATA_HOME/lernie (machine- populated pools and the workspaces/ tree). LERNIE_HOME=<dir> collapses both to one directory, at install time and at runtime alike. lernie prime founds it, seed-if-absent throughout, so running it again — or after an upgrade — never clobbers a hand edit. Only make install runs prime for you; on the other two routes it is your first command, or lernie new's side effect.

make uninstall removes the three installed binaries; bz (installed via cargo) is removed with cargo uninstall brazen. The harness homes (the config and data roots, holding config and workspaces) stay put — clean them up manually if you want a true uninstall.

Configuration schemas

JSON Schemas for the harness-root and config-commit control files (per docs/ARCHITECTURE.md §2.2, §4.1) are generated from the Rust types under src/config/. make schemas writes them to schemas/ for editor integration and external validators. Generation is a golden test (config::schemas::write_to vs the checked-in schemas/): make schemas runs it with UPDATE_SCHEMAS=1 to rewrite the directory, and the same test under make check fails if schemas/ ever drifts from the source types — so the tree is always current, with no separate binary to run.

File Backed by Rust type Config-commit / on-disk file
schemas/version.json config::version::Version version (config commit)
schemas/manifest.json config::manifest::Manifest manifest.yaml (config commit)
schemas/workflow.json config::workflow::Workflow workflow.yaml (config commit)
schemas/providers.json config::per_repo_providers::PerRepoProviders providers.yaml (config commit, roles:)
schemas/models.json config::models::Models <config-root>/models.yaml

Layout: harness root and workspaces

The harness root is installation-global state, split by XDG lifetime into two homes (ARCH §2.2). LERNIE_HOME, if set and non-empty, collapses both to that one directory (test isolation, alternate installs). Three distinct on-disk locations:

  • Config root — hand-edited declarations, $XDG_CONFIG_HOME/lernie (default ~/.config/lernie). Holds the global models.yaml (model capabilities + context windows, plus an optional adapter: binary override — §4.2) and the workflows/ templates. Provider endpoints and auth live in brazen's config, not here (§4.1).
  • Data root — machine-populated pools, $XDG_DATA_HOME/lernie (default ~/.local/share/lernie). Holds the tools/ and skills/ pools plus the workspaces/ tree. Shared across every workspace.
  • Workspace — one git repository per workspace, at <data-root>/workspaces/<workspace>/ (ARCH §2.2): a bare repo.git holding config branches (config/<name>) and agent refs (agents/<agent-id>) — no main. The control files (providers.yaml roles: only — §4.3, manifest.yaml, workflow.yaml, version, souls/) live in the config commit, read from each agent's governing config commit (git merge-base against the config/* heads — derived from ancestry, never stored). Agent worktrees are siblings under agents/<agent-id>/; steps/ and inbox/ sit at the workspace root, outside every worktree. Workspace repositories are never pushed to a remote.

lernie new creates a workspace and authors its first config commit — an orphan root on config/default — from template/, the versioned skeleton embedded into the lernie binary at build time:

lernie new                     # auto-id under <data-root>/workspaces/
lernie new /path/to/my-workspace

Or via the Makefile wrapper:

make new-workspace DEST=/path/to/my-workspace

The binary founds the harness root first — it runs the same seed-if-absent routine lernie prime is (ARCH §2.2), so a data root nobody primed gains the tools/ and skills/ pools before they are read, and a primed install is untouched (nothing is clobbered, no flag is involved). That is what keeps the next step honest: the pools are an input to the config commit, so an unprimed root would otherwise author a commit with an empty descriptions/** and hand every agent forked off it an empty toolset. It then runs git init --bare -b config/default <dest>/repo.git, materializes a transient authoring checkout, extracts the template's control files into it, snapshots the data-root pools into descriptions/{tools,skills}/ (ARCH §3.3 descriptions-always), commits (config: init [config/default]), and tears the checkout down. The workspace is left with exactly one ref — the config commit every fresh root agent forks off (fork is the freeze, §2.2). The destination must either not exist or be an empty directory. With no path argument, the destination is <data-root>/workspaces/<auto-id>/; the created path is printed on stdout. goal.md and soul.md are intentionally not in the template — they are written per-branch at dispatch time (ARCH §2.3, §2.8), which also removes the control files from the agent's tree (§2.2: control is read from the config commit; worktrees hold only context).

Pre-v1 clean break (ARCH §10): the retired per-conversation layout (a root/ worktree with loose control files) is refused with an actionable error, not migrated — create a fresh workspace with lernie new.

First-run smoke test (required). lernie new authors the default providers.yaml with a concrete model id, but validates it against nothing — id validity is brazen's fact, and lernie runs no model-list reconciliation (ARCH §4.2, the settled stance). A wrong id surfaces only at the first live model call. The required next step after creating a workspace is therefore a live lernie prompt (see the quick start above): it is the cheapest — and, by that stance, only — check that the authored id actually resolves on the wire.

make smoke automates exactly this:

make smoke     # scaffold a throwaway workspace + one live 'lernie prompt'

It founds a throwaway harness root with lernie prime — from the assets embedded in the binary, the same front door make install uses, so the shipped install path is exercised too — scaffolds a workspace with lernie new, then runs one live lernie prompt against the shipped defaults — worker role, provider anthropic, model claude-sonnet-5 — through the real bz data plane. The verdict is read from observable state, never the agent's own claim: the lernie prompt exit code is 0, the agent ref (agents/<id>) carries a committed transcript entry, and the off-worktree step record (steps/<id>/001/) holds a response with no wire error and real assistant text. That last pair is the point: an auth-failed run still creates the branch and a step record whose response terminates in a clean end — the failure rides an error event ahead of it — so branch-exists and step-exists alone would pass a broken wire. make smoke requires exit 0 and no error event and an assistant content_delta.

By default make smoke runs the shipped default — provider anthropic, model claude-sonnet-5 — which needs a configured bz credential for the anthropic provider (bz --login --provider anthropic, or set ANTHROPIC_API_KEY / BRAZEN_API_KEY) and spends real money. To run the same live check against any other bz provider row, set both SMOKE_PROVIDER and SMOKE_MODEL (both-or-neither — one alone is a usage error; unset leaves the shipped default byte-for-byte):

make smoke SMOKE_PROVIDER=local SMOKE_MODEL=<a-pulled-ollama-model>
make smoke SMOKE_PROVIDER=codex SMOKE_MODEL=gpt-5.4

The override is laid into the throwaway config root through the same front doors a real install uses — a providers.yaml override under <config-root>/template/ (the config-root override) plus a models.yaml placed in the config root before lernie prime (its seed-if-absent contract, ARCH §4.2) — so there is no new lernie flag or verb. Local ollama (bz's local provider row) needs no credential, only a model that is actually pulled and served; the credential note above applies to the anthropic default alone.

What SMOKE_PROVIDER=local does and does not prove. bz's local row (protocol ollama_chat) rejects a canonical tool_result block: the second step of any tool-using run comes back as {"type":"error","kind":"parse_input","message":"user accepts only text content"}. So the local recipe validates the tool-free path only — one model call, assistant text, a committed transcript entry. It cannot exercise a tool step, a compactor (whose whole toolset is write_summary/mark_for_deletion), or any multi-step loop that runs a tool. This is a brazen-side gap in that provider row, not a lernie one, and is filed there as brazen bl-fba7; to smoke a tool-using path, point SMOKE_PROVIDER/SMOKE_MODEL at a row whose protocol carries tool results.

make smoke is deliberately not part of make check or the close gate: make check mocks the wire (httpmock Anthropic SSE), so it can never catch a shipped default that fails on the real provider — which is exactly how the fake id claude-sonnet-4-7 once shipped unnoticed. It runs only on demand.

Authoring config commits

lernie new authors a workspace's first config commit. Every later one — the general harness-assisted user act of ARCH §2.2 — is lernie config:

lernie config <workspace>                       # advance config/default
lernie config <workspace> <name>                # advance config/<name>
lernie config <workspace> <name> --from <src>   # fork config/<name> off config/<src>
lernie config <workspace> <name> --orphan       # fresh orphan lineage

The verb materializes a transient checkout of the target config lineage, refreshes the descriptions/** snapshot from the data-root pools (ARCH §3.3), opens the checkout in $EDITOR (falling back to vi) so you edit the control files (workflow.yaml, providers.yaml, manifest.yaml, souls/, version), commits, and tears the checkout down. <name> defaults to default. --from and --orphan are mutually exclusive and only apply when creating a new branch. A --from <src> naming a lineage the workspace does not have is resolved before the checkout is materialized, and declined by name:

lernie config: no config lineage "nosuch" in this workspace — existing lineages: default, strict

Declining is fine, and leaves nothing behind. Save no change and the pass is declined: there is nothing to commit, so no commit is authored, the branch does not move, and a --from / --orphan branch the pass would have created is not left behind. That is a success — lernie config exits 0 and prints the one line

config/default unchanged: the edit changed nothing, so no config commit was authored

so empty stdout means a commit landed. The transient checkout is torn down on every exit path (a decline, a git decline, an editor that fails), so the next lernie config always runs. Only a hard kill mid-pass can leave the checkout behind, and the next pass clears it before starting (ARCH §2.11 "the next touch heals") — at the cost of the killed pass's unsaved edit, which was never committed.

This is the only act that advances a config branch (ARCH §2.3); agents forked before it keep their governing config, and agents forked after it govern under the new head (fork is the freeze, §2.2).

Sending a prompt

lernie prompt /path/to/my-conversation 'hello'

lernie prompt is the root-agent path (ARCH §2.3, §2.6, §2.7, §2.8, §2.10). Each invocation spawns its own agents/<conv-id> branch off the default config branch's head (§2.2–§2.3 — there is no main), drives each step's model call through brazen's bz (§4.4), and steps until a terminal event. There is no terminal compaction stage (§2.7): compaction runs only at the checkpoints workflow.yaml declares, and a branch with no configured trigger never compacts. Merge-back is gone (§2.6): the root branch persists on its own ref (§2.4), and a child returns by depositing a result message into its parent's inbox (§2.6):

  1. Resolve the harness root (LERNIE_HOME, else XDG homes, ARCH §2.2) and guard the workspace layout (a non-workspace, or the retired per-conversation layout, is refused — §2.2, §10). Load <config-root>/models.yaml (capabilities + context windows + optional adapter: override — §4.2) and, from the config commit's tree (git show <config-commit>:providers.yaml, §2.2), providers.yaml (roles: block — §4.3); cross-validate roles.worker.{provider,model} against models.yaml. For a fresh root the config commit is config/default's head — the very commit the new agent forks off; lernie advance derives an existing agent's governing config commit from ancestry instead (git merge-base against the config/* heads — never stored).
  2. Run the load-time version guard: bz --version must equal the linked brazen crate version (§4.4). Under an adapter: override the guard is skipped and the in-band MessageStart.v handshake governs. Read the worker soul from the config commit's souls/worker.md (§2.2, §4.3).
  3. Spawn branch agents/<conv-id> (§2.3 — the id is the bare hyphenated descent; the agents/ prefix is the ref namespace) off config/default and allocate a worktree at <workspace>/agents/<conv-id>/ (§2.2). Write the branch goal to goal.md and the role soul to soul.md, remove the config commit's control files from the tree (§2.2 — the worktree holds only context), and commit — that commit's tree is step 1's read state (§2.10).
  4. Build a typed brazen::CanonicalRequest (linked crate — the fail-open extra map stays unreachable), mirror it to <workspace>/steps/<conv-id>/001/request.json (a diagnostic artifact, outside every worktree, never read at runtime, §2.3).
  5. Model call, harness-owned retry loop (§2.10, §4.4). Exec bz --json --provider <row> once per attempt, canonical request on stdin, appending each attempt's stdout verbatim to <workspace>/steps/<conv-id>/<NNN>/response.json as brazen v=1 NDJSON — one self-delimiting segment per attempt, each ending in a terminal end. On a retryable in-band Error (CanonicalError::retryable(), never re-derived) the harness re-invokes bz with the identical request, up to the workflow.yaml attempt cap with exponential backoff — floored by the failed attempt's Retry-After pacing hint (CanonicalError::retry_after_seconds) when it carries one, so the config schedule governs and the provider's hint can only lengthen it (§4.4). brazen never retries; auth and endpoints are entirely its own. The response.json fd is held open across every attempt and backoff sleep — its close is the §3.5 IN_CLOSE_WRITE completion signal. As the events stream, the harness tracks only their framing — the terminal end, an in-band Error, the handshake v — for retry/classification; meta.json carries {commit, started_at, ended_at}. The events' content streams into the transcript writer's (§2.3) staging file <workspace>/steps/<conv-id>/<NNN>/staging.json, appending each content block as it completes; segment authority (§4.4) truncates it on an Error attempt and the settling Finish seals it — one stream, two sinks (diagnostic response.json + transcript), never read back. When the model call completes, the sealed file is renamed into the worktree as messages/NNN-<model-id>.json — its origin token is the model that authored it (§2.3), the body a JSON array of canonical Content blocks — and committed. NNN is the branch's transcript counter, max-present-plus-one from the messages/ listing, evaluated at commit time. The initial user message now enters through the front door like any other (§2.11): the executor deposits it into the agent's own inbox, and the step-boundary drain delivers it as the first transcript entry messages/NNN-user.md (bl-1129) — no bespoke initial-message path beside the drain.
  6. Step loop (§2.5). At each step boundary the executor first drains the inbox (bl-1129, §2.11): after committing any renamed-but-uncommitted stray a prior death left in messages/, it moves each pending inbox/<agent-id>/<sender>-<NNN>.md into the worktree as messages/<counterNNN>-<sender>.md (a literal rename(2) — one home at every instant) and commits the move, in a deterministic (mtime, filename) order, ahead of the read-state capture so a delivered message is part of the commit the model call assembles from. Each step then re-assembles its model-facing history from the read-state commit's tree — readdir of messages/, sorted by the filename's NNN prefix, each entry composed by its origin token (NNN-<sender>.md → user text, NNN-<model-id>.json → the assistant message — any .json token but the reserved tool, NNN-tool.json → tool_result in the following user message), with consecutive same-side entries grouped into one alternating wire message. There is no in-memory history and no git-log walk; running, retry, and replay are one code path against one input, the commit's tree (§2.3, §5). If the settled model-output entry carries any tool_use block, run every one through the tool executor — the per-call records land under <workspace>/steps/<conv-id>/<NNN>/tools/<tool-id>/ (out of every worktree, §3.3; written but never read at runtime), and as each tool resolves the transcript writer commits messages/NNN-tool.json (its canonical tool_result block) — then loop into step <NNN+1>. A step with no tool_use block is terminal. Step ≥2 has no dispatch commit, but each step's transcript entries (assistant output, tool results) do advance the branch tip, which is that step's read state (§2.10). tool_use/tool_result pairing holds by construction: a tool result commits immediately after its emitting step's model-output entry, so it always lands in the immediately following user message. Closing each tool step, the executor reads the compaction checkpoint clock (§2.7, §6) — compaction.intermediate.trigger in workflow.yaml: every_n_commits, every_t_seconds, or the agent-elected on_flush, all derived from git (commits and elapsed seconds since the last compaction merge, or the branch root when none has landed — never a stored counter). When it is due, the worker_flush: dispatch(compactor) binding forks a compactor off the branch tip — the checkpoint commit C — and the branch keeps stepping straight through it; no quiescence is imposed. Omit the compaction: block and the branch never compacts.
  7. Terminal return (§2.6, §2.3 step 5). Every terminal event — normal completion (final-response), budget exhaustion (budget-exhausted, §6), and stop (stopped, §2.9 — the executor's SIGTERM handler deposits on its way out) — deposits a result message into the parent's inbox: an ordinary deposit whose frontmatter adds epitaph: and terminal_ref: (the branch tip) and whose body is the terminal response iff the agent spoke. For a root this is a structural no-op — a root has no parent inbox; its response answers the user (§2.4). The deposit is executor-side, never a model tool call ("Return is not a verb"). At delivery, a message carrying terminal_ref: applies the fork-point→terminal work-product transfer as one commit before its delivery commit, filtered to work products; a diff that fails to apply is declined at refs/lernie/conflicted/<agent-id> (§2.6).
  8. Exit protocol (§2.11). With the terminal deposit landed, the executor runs the branch's terminal workflow.yaml bindings (branch_stopped → mark_abandoned / notify_ui, §6), releases the executor lock, and only then spawns a driver at its own agent and — the deposit's own probe-and-launch — at the parent the deposit just revived. Both launches are fire-and-forget and both are decided by epitaph value: a final response launches, stopped and budget-exhausted never do. No terminal compactor is dispatched (§2.7): the v0.3 terminal-compaction stage is deleted, along with the Dispatcher re-entry that existed only to run it. Compaction is a checkpoint event (step 6), never an exit stage. Merge-back is gone (§2.6): the root branch persists on its own ref (§2.4); nothing merges back, and the agent's worktree is not torn down (quiescence, not teardown, §2.3 step 6).
  9. Print the agent id (the bare conv-id) on stdout.

After lernie prompt returns, inspect the agent against the bare workspace repository:

cd /path/to/my-workspace
git -C repo.git log --oneline --decorate agents/<conv-id> -4
git -C repo.git ls-tree --name-only agents/<conv-id> messages/
git -C repo.git show "agents/<conv-id>:messages/002-<model-id>.json"
ls steps/<conv-id>/

The log is the dispatch commit followed by one transcript NNN: commit per entry, its subject naming that entry's origin token — user, a sender's agent id, tool, or the authoring model's id:

f265de7 (agents/…) transcript 002: qwen3.5:9b […]
7ae527e transcript 001: user […]
f643a50 step 001: dispatch […]
6f4bd05 (config/default) config: init [config/default]

ls-tree lists the transcript itself (messages/001-user.md, messages/002-<model-id>.json, …) and show prints one entry — a model-output entry is a JSON array of canonical Content blocks, e.g. [{"type":"text","text":"pong"}]. ls steps/<conv-id>/ lists the off-worktree step records, one numbered directory per step, each holding request.json, response.json, and meta.json. There is no merge commit and no summary/ on a branch that never reached a compaction checkpoint (step 6) — those appear only once a compactor has returned.

The root branch persists unmerged by design (§2.4), so the health metric is no longer branch count but silent deaths and undelivered returns (ARCH §8) — read straight from git refs, the executor lock, and inbox listings, with no sidecar file.

Stopping a conversation

lernie stop /path/to/my-conversation <conv-id> [--stop-children]

Sends SIGTERM to the process group of the one executor driving <conv-id>, with a 5-second flush deadline before SIGKILL. This is the same cascade pattern adapter (§4.4) and tool (§3.3) cancellation use, applied to the harness itself (ARCH §2.9). The group signal reaches that executor's own bz and tool subprocesses — its limbs — and stops at the agent boundary: a dispatched child harness has taken its own process group, so a bare stop does not fell it. A running child outlives the stopped parent and revives it later by depositing its result (§2.11) — stopping a parent strands nothing.

--stop-children opts into the agent→agent cascade: it walks the id namespace — the descendants of <conv-id> are exactly the inbox directories prefixed <conv-id>- (§2.3), one prefix scan reaching every depth — and folds each descendant executor's group into the same sweep. The pid is discovered by scanning /proc/<pid>/fd/* for the process holding the agent's inbox-directory lock fd open — the executor lock (§2.11), held for the whole step loop, so a stop lands even during tool execution when no response.json is open — no sidecar pid file. Linux only.

The pgid that scan produces is vetted before anything is signalled, because a pid is discovered before its group has settled: between a driver's fork and the setpgid/setsid it runs at startup, /proc still reports the group it inherited from its spawner — your shell job. So a pgid is trusted only once it equals the holder's own pid (a group leader's does, and every driver becomes one), re-read a bounded number of times while it does not, and refused rather than signalled if it never settles; a stop that signals nothing is re-runnable, one that signals your shell is not. lernie stop additionally refuses any group it is itself standing in (§2.9).

The group signal reaches every member independently: bz installs no handler and dies at once (leaving the missing-end signature, §4.4), while the executor catches its own copy — SIGTERM is catchable — and, instead of dying on the spot, deposits its branch's stopped result on its way out (§2.9 step 3, executor-side, "Return is not a verb") and then exits cleanly. Catching shields nobody: the kernel already delivered to bz and the tools. For a root the deposit is a no-op (no parent inbox); the observable is the clean exit.

Because the model call is where the wall time goes, that is where a stop usually lands — so the clean exit is the ordinary case, not the rare one. The flag classifies, not the error's shape: a kill lands wherever the adapter was, leaving a half-stream, a torn JSON line, or a provider error depending on the instant, and with a stop pending each is read as the stop. With no stop pending the same faults still propagate non-zero, so a genuinely dying adapter is never hidden. The retry loop respects the flag too: a stop is never followed by another bz invocation.

Behavior:

  • Idempotent. A branch with no live writer (already stopped, or the harness exited cleanly) returns success without sending any signal.
  • Errors when the agent branch (agents/<conv-id>) doesn't exist. Surfaces as a non-zero exit with a lernie stop: prefix on stderr. (The old "already merged" refusal died with main: nothing merges, so there is no merged state to refuse — an already-terminal branch is simply the idempotent no-holder case above.)
  • No on-disk cancel marker. The §2.9 signature of a stopped branch is the latest step's response.json closed without a terminal end event — produced by bz dying mid-stream on its own SIGTERM (§4.4); the executor's stopped deposit is an independent write to the inbox tree and never touches that signature.

The frontend's stop button (per ARCH §3.5) exec's this exact subcommand; there is no second control surface.

Built-in tools (v0.3, +v0.4 Phase 2 dispatch)

The agent can call built-in tools that ship inside the lernie binary as lernie tool <name> subcommands (ARCH §3.3 / §12). The tool executor's resolution order — <data-root>/tools/lernie-tool-<name> → PATH → lernie tool <name> — falls through to this in-process route for tools not externalized.

Each built-in is the triple §3.3 pins:

  • Binary — the lernie tool <name> subcommand. Reads tool_use.input JSON from stdin, writes raw bytes to stdout, exits 0 on success or non-zero on failure (stderr is concatenated after stdout into tool_result.content when is_error is set).
  • JSON schema — at schemas/tools/<name>.json, seeded to <data-root>/tools/<name>.json by lernie prime (which make install invokes, ARCH §2.2). Sent verbatim as the input_schema of the tool's entry in the model call's tools: [...] array.
  • Skill — at skills/<name>/SKILL.md, seeded to <data-root>/skills/<name>/ by lernie prime. The frontmatter description is the tool's description in tools: [...]; the body explains when to reach for it.

The pool is discoverable from the CLI itself — lernie tool --help names it, and a name that is not in it is declined non-zero naming it too, the same way load_skill declines an unknown skill (ARCH §3.3):

$ lernie tool --help
Arguments:
  <NAME>  Built-in tool to run; one of: bash, dispatch, load_skill, message, read_file

$ echo '{}' | lernie tool nosuchtool
lernie tool nosuchtool: unknown built-in tool: "nosuchtool"; available: bash, dispatch, load_skill, message, read_file

Built-ins:

  • read_file — read the entire contents of a file at a given path. Rejects files larger than 1 MiB, reporting the file's true size (stat, not the capped read's length) so the agent can judge the magnitude it is up against; v0.4+ adds the oversized-output auto-dispatch shim (ARCH §3.3 / §12). Try it directly: echo '{"path":"README.md"}' | lernie tool read_file.
  • bash — runs a shell command via sh -c and returns its stdout. The shell runs in its own process group so a SIGTERM the harness sends is forwarded to the entire spawned tree (§2.9 cascade). Try it directly: echo '{"command":"ls"}' | lernie tool bash.
  • dispatch (v0.4 Phase 2) — spawns a subagent on a fresh branch with the supplied goal and returns {"status":"in_progress","handle":"<sub-branch>"} synchronously (ARCH §2.5). Input is {role, goal}; the role must resolve to souls/<role>.md and a roles: entry in providers.yaml — both read from the calling branch's governing config commit (§2.2). Reads the calling conversation's repo + branch from the harness-set LERNIE_CONV_REPO / LERNIE_CONV_BRANCH env vars (ARCH §3.3 env bullet); spawns through lernie dispatch <role> (§3.4). The handle it returns is the child's address — there is no polling tool to pair with it. The substrate redesign (ARCH §2.5 "Dispatch returns the child's address") dissolved the handle/await pair: the child's result comes back as a deposit into the parent's inbox carrying an epitaph (§2.6, §2.11), so await/check had nothing left to observe and are gone. The return path — the result-message deposit and the delivery-time work-product transfer — is built and live (bl-4ce8, bl-9f53, bl-c33b, §2.6), and children run full step loops: the dispatch's own front-door deposit finds the fresh child quiescent and launches the ordinary driver, lernie advance (§6) — there is no child-specific loop and no worker path — which steps the child to a terminal event, deposits its epitaph result (final-response, budget-exhausted, or stop) into the parent's inbox, and revives the parent, which delivers the result at its next step boundary.
  • message — deposits content into an existing agent's inbox (ARCH §2.11). Input is {agent, content}; the recipient is addressed by its agent id (its branch name / hyphenated descent). Unlike dispatch it starts no branch and returns no address — it deposits synchronously and returns {"status":"deposited"}. The sender is the calling agent's id, taken from the harness-set LERNIE_CONV_BRANCH (never model-supplied), so provenance cannot be forged. It goes through the front door — lernie message (below) — like dispatch goes through lernie dispatch, so it inherits the front door's recipient guards: an id that is not a single path component, or one with no agents/* ref, comes back as an is_error result naming the decline instead of a silently lost message. Shipped state: the deposit lands and the step-boundary drain delivers it (bl-1129) — the next driver to step the branch moves the inbox file into messages/ as a transcript entry at its next boundary. A deposit into a quiescent agent is self-delivering: the free-lease probe detach-spawns lernie advance (§6, below), which acquires the lease, delivers the deposit, and steps the branch.
  • load_skill — copies a pooled skill's body into the calling agent's worktree at skills/<name>/, where the next context assembly composes it (ARCH §3.3 Body-on-demand, §5.2). Input is {name}; the data-root pool + target worktree come from LERNIE_HOME/XDG and LERNIE_CONV_REPO / LERNIE_CONV_BRANCH. Returns {"status":"loaded","path":"skills/<name>"} on a fresh copy or already_loaded when the worktree already holds it (the loaded copy is the snapshot the branch is pinned to; rm and reload to refresh). An unknown or non-single-component name is declined (is_error, naming the available pool). Shipped state: the copy commits with the tool result — a tool commit now stages the whole worktree (git add -A, commit_tool), landing any tool's worktree side effects with its result entry (ARCH §2.3).

Messaging an existing agent directly

lernie message <workspace> <agent> <content> deposits a message into <agent>'s inbox and, finding the recipient quiescent, launches a driver to deliver it (ARCH §2.11, §3.4). The sender is read from LERNIE_CONV_BRANCH — the calling agent's id when the message tool re-enters the verb, else user for a bare invocation.

  • The recipient is guarded before anything is written. The id must be a single path component (ARCH §2.3) — .., a /, or an absolute path is declined, never sanitized, because Path::join would honour it and write outside the workspace — and an agents/<id> ref must exist for it: a message is addressed to an existing agent (§2.11), so a deposit no drain would ever come for is refused (lernie message: no agent "…" …, exit 1) rather than left in an inbox directory nothing will ever read. The id guard is the same rule at every verb taking an agent id from outside — message, advance, stop, dispatch, bundle — and literally the same code: one workspace-layout guard and one existence guard, each carrying the calling verb's own clause for why it needed an agent, so what differs between verbs is the cause, never the phrasing or the remedy.
  • The deposit is a create-only file at <workspace>/inbox/<agent>/ <sender>-<NNN>.md (temp-path + atomic rename), with from: / deposited_at: frontmatter and the content as its body. <NNN> is the sender's own sequence, derived as max-present-plus-one over its existing files in that inbox.
  • After depositing, the verb probes the executor lock (flock on the inbox directory): the same lease the shipped lernie prompt step loop holds for its whole run, releasing it on exit. A held lease means a driver is already stepping the branch (it will deliver at its next boundary); a free lease means the branch is quiescent.
  • On a free lease the verb launches a driver — lernie advance <workspace> <agent> (ARCH §6) — as a detached spawn (§2.11): setsid (its own session and process group), stdio to null, fire-and-forget. The driver outlives the lernie message process, so messaging is scriptable: the verb returns as soon as the deposit and spawn land, and delivery + stepping continue in the driver.
  • A failed branch is named, never refused. If the quiescent recipient's latest model call failed (its last response.json segment terminated in an error — retries exhausted or a non-retryable error, ARCH §2.10), the deposit and launch proceed unchanged — messaging is exactly how such a branch is retried once the cause is fixed — but the verb prints a stderr advisory naming the branch and pointing at steps/<agent>/ and lernie scan, so a silent death (ARCH §2.3, §8) is distinguishable from ordinary idleness at the verb that touches it. Exit code and stdout are untouched.

Driving a branch: lernie advance

lernie advance <workspace> <agent> is the §6 driver verb — the process every launch seam spawns, and the same verb an operator runs by hand. One invocation is one hop: guard the id (a single path component, and an agents/<id> ref must exist — a name that is no agent is refused with no agent "…" and exit 1 before any lease, so an operator typo neither drives anything nor leaves an inbox/<id>/ behind), take the lease (adopt the LERNIE_LOCK_FD fd published by a predecessor hop, else try-acquire the executor lock — losing it is a clean no-op), deliver pending inbox messages through the real drain (rematerializing a torn-down worktree first), derive warrant from the transcript tail (ends user-side → a model call is due; ends assistant-side without tool_use, or empty → exit silently; assistant tool_use with uncommitted results → decline loudly, the one non-replayable state), run one step, and hand off: a step that emitted tool_use runs its tools and exec's the successor lernie advance with the lock fd deliberately inherited (close-on- exec cleared just before exec; the successor fstat-validates the fd against the inbox directory and restores close-on-exec), while a terminal event ends the chain through the §2.11 exit protocol. Because the successor is exec'd in the same process, the pid, process group, and flock lease all survive the hop — lernie stop lands on whichever hop is current, and no rival driver can wedge between hops.

The exit protocol and the operator scan

Normal operation needs zero scanning (ARCH §2.11): lernie message deposits, probes the executor lock, and launches a driver if the agent is quiescent; the executor drains its inbox at every step boundary. The graceful-exit crack — a deposit landing after an executor's final drain but before its lock release — is closed by the exit protocol (§2.11, bl-5846): one terminal sequence, no agent kinds — deposit the result message (a structural no-op for a parentless agent) → release own lock → spawn a driver at own agent, fire-and-forget → probe-and- launch at the parent the deposit just landed in → exit. Two pins terminate the recursion: a driver that acquires and finds nothing to deliver exits silently (no step, no epitaph, no further launch — dispatch::driver::drive is that entry), and the launch is decided by epitaph value — a final response launches; stopped and budget-exhausted never do. The exit launch rides the same launcher seam as the writer probe, so it is the same detached lernie advance spawn (§6); the decision logic, ordering, driver entry, and the spawn itself are live and tested.

The parent-side step is what makes revival-on-deposit real (bl-4a6c): a child that returns to a quiescent — even torn-down — parent starts that parent's driver itself, through the same probe_and_launch the lernie message verb uses (one probe, no second copy), so the parent rematerializes, delivers the result, and steps with no lernie scan in the path. A parent whose lease is held gets nothing launched: its running executor delivers at its next step boundary. The epitaph decision governs this launch too, one level up: a stopped child would otherwise wake its parent to react to — perhaps re-dispatch around — the very branch the operator killed, and a budget-exhausted child's ceiling is the whole tree's (§6), so the woken parent would exhaust on its own next check and deposit again. In both cases the result still lands in the inbox and waits for the next explicit touch.

Crashes are accepted as a failure class (§2.11): everything is on disk, so a hard death strands results and messages late, never lost, and the next touch heals. That touch is a user reprompt — or the operator verb lernie scan <workspace> (§2.11, §8, bl-d148 + bl-5846): one workspace-wide pass, run by hand or by cron if you want a heartbeat, never wired into any driver hot path or default schedule (the events it compensates for happen at crash rate, not step rate). Two derived actions, no watcher (an idle workspace stays unswept until the next touch, by design):

  • Silent-death sweep. Every agent branch with no live executor (the §2.11 executor-lock probe) that either died mid-work — its latest step's model call never settled complete: response.json closed without a terminal end (killed/stopped, §2.9), or its final segment terminated in an error (retries exhausted or a non-retryable error, §2.10 — that segment closes with a clean end, so absence-of-end alone would misread the branch as idle) — or, for a child, never deposited a result message is a silent death (the §8 health count). Each one is named in the report (silent deaths: 1 (<agent-id>)): a dead root gets no deposit — it has no parent inbox — so its name here is how an operator learns which branch went quiet, and steps/<agent-id>/ is where to read why. For each hard-crashed child in that set, the sweep deposits a died-epitaph result message on the child's behalf (sender = the child — the sweep is the scribe, not the author), so the parent is revived rather than stalled. The "never deposited" test reads both the parent's inbox (undelivered) and its transcript (delivered), so a prior sweep's own deposit is seen on re-scan and never re-deposited — idempotent by construction.
  • Inbox flush. Every agent with pending inbox files and a free lock gets a driver launched — never drained: the scanner moves no files and commits nothing; only an agent's own lock-holding executor delivers. An agent whose lock is held is left alone. The inbox listing is intersected with the agents/* refs — the one registry of who exists — so an inbox directory with no matching ref is reported (inboxes with no agent branch: N) and left in place rather than driven: a driver launched for a name with no branch is refused by the existence guard (lernie advance: no agent "…", exit 1) on this pass and every pass after, writing nothing. The sweep's own deposits are picked up by the flush that follows in the same pass.

Shipped state. The scan (silent-death sweep + inbox flush) ships behind lernie scan and only there — driver startup (lernie prompt, lernie dispatch, lernie advance) runs no workspace scan. The flush and the exit launch reuse the same driver-launch seam as lernie message, and the spawn is real: each seam decides when a driver is needed and detach-spawns lernie advance (§6) for it. Children run full step loops (bl-c33b), so a died child is a state a real run reaches; the derivation is additionally exercised against constructed on-disk states, since a hard crash is not reproducible on demand.

Namespace note. The candidate enumeration is the agents/* ref namespace, exactly as ARCH §8 writes it (a root is agents/<conv-id>, a child agents/<parent>-<sub-id>); config branches are excluded structurally by the prefix — there is no main (§2.2).

Dispatching subagents directly

lernie dispatch <role> <repo> <branch> [--goal <text>] is the §3.4 re-entry point every child dispatch uses. It is writer-shaped, not an executor (ARCH §2.1): it forks the child branch, lands the dispatch commit, and deposits the dispatch message through the same front door every sender uses — the driver that deposit launches is the ordinary lernie advance (§6). The role name is positional and the role set is open (§4.3): a role is dispatchable iff the calling branch's governing config commit lists it under providers.yaml roles: and carries souls/<role>.md. The CLI enumerates no role names, so a verifier, a critic, or a role you author needs no CLI change; validity is checked before the fork, so a rejected role leaves no branch debris.

The id guard runs first, through the same two functions message, advance, stop and bundle call: the workspace layout, then the dispatching parent's agents/<id> ref. So all three refusals are the product's, never git's:

lernie dispatch worker <no-such-ws> someagent --goal hi
  → <path> is not a workspace (no repo.git) — create one with `lernie new` (ARCH §2.2)
lernie dispatch worker <ws> nosuchparent --goal hi
  → no agent "nosuchparent" in this workspace — a child forks off an existing parent (ARCH §2.5); …
lernie dispatch verifier <ws> <agent> --goal hi
  → role "verifier" is not defined in the providers.yaml governing agent "<agent>" — defined roles: compactor, worker

The role refusal names the pool that is defined — the same "name the pool" idiom load_skill and lernie tool decline with — and names the control file the user knows rather than the config commit's sha.

  • lernie dispatch compactor <workspace> <conv-id> forks a compactor-souled child off that agent's tip — exactly what a due compaction checkpoint does (§2.7), run by hand. The compactor is an ordinary child that makes a real model call through bz; it is not a stub, and it does not merge anything itself. Its goal is procedure-generated, so passing --goal is rejected. Its toolset is the deletion-only pair injected for the compactor role alone (never a providers.yaml tools: list): write_summary, which writes the next summary/<NNN>.md on the compactor's branch, and mark_for_deletion, a staged git rm that can remove but never write content — so the worst case is lost information, never corrupted information. Its request declares more than that pair: a compactor inherits the dispatching branch's transcript, so the model call also names whatever tools that transcript used — otherwise the provider refuses a request whose history mentions a tool it was not told about. Declaring is not permitting: a compactor reaching for one of those inherited tools gets an error tool result naming its own two, and nothing runs. The compaction merge lands later and elsewhere: when the compactor's result message is delivered, the dispatching agent's own executor interprets its compactor_return: compaction_merge binding (§6) and merges the compactor branch --no-ff — the one merge left in the system (§2.6). A compactor that ends on any other epitaph lands no merge; the branch simply continues uncompacted — enforced where the binding is interpreted: the delivered result's epitaph value gates compaction_merge, and a died/stopped/budget-exhausted compactor return is delivered like an ordinary child's result instead, so the parent sees the epitaph and nothing of the compactor's branch crosses (§2.6, §2.7).
  • lernie dispatch worker <workspace> <parent-id> --goal <text> spawns a worker child off the parent's tip. The new id is <parent>-<sub-id> (hyphenated descent, §2.2), its ref agents/<parent>-<sub-id> (§2.3), its worktree agents/<parent>-<sub-id>/; goal.md carries the supplied text and soul.md is read from the parent's governing config commit (souls/worker.md, §2.2), both committed as the dispatch commit (§2.3 step 2). The child then runs a full step loop under the lernie advance driver its dispatch deposit launched, and at its terminal event deposits a result message — epitaph, terminal ref, and the terminal response iff it spoke — into the parent's inbox, reviving the parent if it had gone quiescent (§2.6, §2.11). The v0.4 "Phase 1 stops at the dispatch commit" worker path (worker.rs) is deleted, not extended (bl-c33b).

Providers

Every model call goes through brazen — one small, stateless binary (bz) that adapts every provider and wire protocol behind a single pipe contract (see ARCH §4.4):

stdin (canonical request, JSON) → bz → stdout (v=1 event stream, NDJSON, one terminal `end`)

The harness execs bz --json --provider <row> once per attempt, pipes a typed brazen::CanonicalRequest on stdin, and appends bz's stdout verbatim to the step's response.json. lernie links the brazen crate (brazen = "=0.0.4") for the canonical types only — the data plane always crosses the subprocess boundary (§3.4). Two facts follow:

  • Retry is the harness's. brazen never retries — one bz process, one HTTP round-trip. On a retryable in-band Error (CanonicalError::retryable(), the linked crate's single home for the fact) the harness re-invokes bz up to the workflow.yaml attempt cap (§2.10). Each attempt appends one segment to response.json; the last is authoritative. Each attempt's bz stderr appends to the step's stderr.log beside it — empty on an ordinary run, because brazen speaks its failures in-band on stdout. A bz that dies before it can (a malformed brazen config) leaves an empty stream that reads exactly like a mid-stream kill, so the half-stream error quotes that capture's tail; with a stop pending it stays quiet, because the stop check point (§2.9) discards the outcome before anything is rendered.
  • Auth and endpoints are brazen's. Provider rows (endpoint, protocol, auth mode, model aliases) live in brazen's own config (~/.config/brazen/config.toml; bz --dump-config, bz --login). lernie references a row by name and never sees credential material (§4.1). A load-time guard (bz --version == the linked crate version) rejects a mismatched binary; make install installs the pin with cargo install brazen --version =0.0.4.

Adding a provider

  • A new provider on a supported protocol is a brazen config row — no code anywhere. Add the row (bz config), reference its name as a model's provider: in <config-root>/models.yaml, and point a role at that model in <repo>/providers.yaml.
  • A new wire protocol or auth mode is a contribution to brazen.
  • An alternate adapter binary that honors the same pipe contract slots in via the optional adapter: path in models.yaml (§4.2); the version guard is skipped for it and the in-band MessageStart.v handshake governs compatibility instead.

UI (v0.5)

The desktop frontend lives in its own repository, yog: an egui/eframe window that renders a workspace and issues user actions via lernie <subcommand>. It composes on lernie's public surfaces only — the CLI and the on-disk workspace layout (ARCH §3.5, §7.1) — and takes no Cargo dependency on this crate, so it builds, versions, and installs independently (make install there drops yog next to lernie). Keeping frontends out of this workspace is deliberate: lernie ships as a composable component, and anything that composes it (a GUI, a web view) lives outside it and meets it at those surfaces.

Evaluation: archival and the task suite (§9)

Archive a run. A "run" is an agent subtree, not a whole workspace (§9.2). lernie bundle <workspace> <agent> <out-dir> writes the subtree — the agents/<agent> branch and its agents/<agent>-* hyphen-descendants (§2.3), with all the ancestry those refs reach — plus the subtree's governing lineage: every config/* ref whose history reaches it (§2.2). Both go into one git bundle, and the matching steps/<id>* and inbox/<id>* diagnostic slices are copied beside it. One bundle plus two slices is the whole run.

The config refs are not decoration. An agent's control files are read from its governing config commit, which is derived — the nearest ancestor of the branch reachable from a config/* ref (§2.2). Ancestry alone carries that commit as an object but names no ref to take the merge-base against, so a replay of the agent refs alone yields a workspace no verb can drive. Carrying the refs (never a sidecar file — the refs are the single source) makes the replayed repo derive its governing config by the same computation, over the same candidate set, as the workspace it came from. "Every ref whose history reaches it" is broader than "every ancestor": a sibling config lineage that shares only a common root with the bundled subtree is still a merge-base candidate, so it rides too — carrying a ref that turns out not to be the nearest one is how the bundle stays a faithful copy of the computation, not a leak.

lernie bundle /path/to/workspace <agent-id> /path/to/archive

Replay a run. lernie replay <archive> reconstructs a scratch workspace under LERNIE_HOME's data root at replays/<primary-id>/ (the primary id is the subtree's root agent), fetches every branch out of the bundle into a fresh bare repo.git, materializes the primary's worktree under agents/, restores the slices, and prints the scratch path. Point the ordinary frontend at it — replay is not a mode (§2.3). Set LERNIE_HOME to an isolated directory to keep the replay sandboxed; the harness root it points at still supplies the machine-local pieces a config only names (models.yaml and the brazen provider rows, §4.2/§4.4).

A replayed workspace is an ordinary workspace: lernie prompt <scratch> "…" forks a fresh root off the config head that rode the bundle, and lernie message / lernie advance drive the replayed agent on its own governing config commit.

LERNIE_HOME=/tmp/replay lernie replay /path/to/archive

Task suite. The evaluation suite lives as data under tests/suite/ — 50 tasks with machine-checkable check scripts, tagged by the seven §9.1 failure categories (≥10 per category), format in tests/suite/README.md, well-formedness enforced by tests/suite.rs.

Run the suite. The agent-eval runner (a separate crate, crates/agent-eval, ARCH §9.3) executes an experiment against the suite N times per task and reports pass@1 (with 95% Wilson intervals) and pass@5, overall and per category:

agent-eval --config baseline --suite tests/suite --runs 5 --agent lernie-eval-agent

--config <name> names an experiment — a workflow.yaml variant under experiments/<name>/ (a config diff, no code changes; see experiments/README.md). baseline is the shipped default itself: its workflow.yaml is a symlink to template/workflow.yaml, because an experiment is a diff against the default and the baseline's diff is empty. Per run the runner seeds a fresh isolated LERNIE_HOME and working directory, runs the task setup, invokes the agent, then runs the task check — exit 0 is the sole pass signal (§9.1), so success is observable state, never the agent's own claim. --bundle-dir <dir> archives failing runs for triage via lernie bundle (§9.2). The runner is fully tested against a faked agent, so it needs no live model to validate.

The shipped driver is lernie-eval-agent (crates/lernie-eval-agent, workspace-internal like the runner; installed on PATH by make install). --agent <cmd> stays required with no default: which driver runs the agent under test is an experiment-defining input, so it is named explicitly. Per run the shipped driver seeds the run's isolated LERNIE_HOME from the machine's lernie config root (models.yaml plus the template/ config-root override — the wire is machine-local by design, §4.2/§9.2, and those two front doors are how a machine points evaluation runs at its own provider rows), then drives the harness exclusively through the front door, exec'ing lernie from PATH: lernie new, lernie config (applying the experiment — below), and one lernie prompt carrying the task prompt grounded in the shared working directory. The contract any driver must honour, per run:

Given How
the task prompt argv[1]
the isolated harness root for this run LERNIE_HOME in the env
the experiment's workflow.yaml LERNIE_EXPERIMENT in the env — an absolute path
where to report back LERNIE_EVAL_REPORT in the env — a file path
the working directory cwd (shared with the task's setup and check)

LERNIE_EXPERIMENT is a hand-off, not a hook: nothing in the harness reads that variable. The harness takes its workflow.yaml from the workspace's config commit (§2.2), never from the environment, so applying the experiment is the driver's job. The shipped driver does it through lernie config, with $EDITOR set to copy the experiment over the authoring checkout's workflow.yaml — the experiment lands as an ordinary config commit, exactly the "config diff, no code changes" §9.3 promises (for baseline the diff is empty and the authoring pass declines: the default is already in force).

LERNIE_EVAL_REPORT names a file the driver may write with exactly two lines — the workspace path, then the agent id — which is what lernie bundle needs to archive the run if it fails (§9.2). It is the driver's only channel back to the runner. Writing nothing, or anything malformed, only makes a failing run un-bundleable; it is never an error, and it never affects pass/fail, which is the task check alone. The driver's own exit code is likewise ignored. Failure to spawn the driver, by contrast, is a hard error naming the program.

Contributing

The instructions below are for contributors building lernie from source. Users installing a release don't need any of this — Install covers the three user-facing routes, only one of which involves a clone.

Contributor setup

make install-hooks

Sets core.hooksPath to .githooks. Required on every fresh clone — git does not track .git/config, so the hooks are not active until installed. That arms both the pre-commit gate and the auto-push hook.

The Rust toolchain is pinned in rust-toolchain.toml (channel 1.95.0, with rustfmt, clippy, and llvm-tools-preview). rustup reads it automatically for every cargo command in the tree and installs the pinned toolchain on first use — no manual rustup step. This is what keeps fmt-check and lint from drifting between your machine, another agent's, and CI.

Build targets

Target What it does
make build cargo build
make release cargo build --release
make test cargo test, with the pinned bz first on PATH (below)
make test-install cargo test --test install — the install contract end-to-end, uninstrumented (it is cfg_attr(tarpaulin, ignore), so coverage skips it); ~45s warm, and it re-installs bz at the brazen pin
make coverage cargo tarpaulin --fail-under 100 (llvm engine), same pinned PATH (below); hard-gated on tarpaulin 0.35.2 exactly (TARPAULIN_PIN in the Makefile — its one home; any other version aborts with the cargo install cargo-tarpaulin --version 0.35.2 --locked fix-it line)
make lint cargo clippy --all-targets -- -D warnings
make fmt cargo fmt
make fmt-check cargo fmt --check
make schemas Regenerate schemas/*.json from the Rust types
make new-workspace DEST=<path> Create a workspace (bare repo.git + first config commit from template/)
make eval CONFIG=<exp> SUITE=<dir> RUNS=<n> AGENT=<driver-cmd> Run the evaluation runner (ARCH §9.3): experiment × suite × N (see Task suite above). AGENT is required and has no default — the shipped driver is lernie-eval-agent (see "Run the suite"), and naming it is deliberate: the driver is an experiment-defining input
make check fmt-check + lint + coverage + test-install
make ci Alias for check
make smoke Live-wire smoke test: one real lernie prompt against the shipped defaults (override with SMOKE_PROVIDER/SMOKE_MODEL); the default needs a bz anthropic credential and spends money; NOT part of check
make install-hooks Point git at .githooks/
make install-bz Install the provider adapter bz on your PATH at the version Cargo.toml pins (ARCH §4.4); a no-op when the bz there already matches. For running lernie — the tests feed themselves (below)
make brazen-pin Print that pinned version and nothing else — CI keys its bz cache on it so no workflow file names a version
make install [INSTALL_PREFIX=<p> LERNIE_HOME=<h>] Release-build; drop lernie/agent-eval into $INSTALL_PREFIX/bin (default: ~/.local/bin); install the provider adapter bz via make install-bz at the version Cargo.toml pins (the ARCH §4.4 version pin — the number's one home); then invoke lernie prime to found the harness root — config root (default ~/.config/lernie) with a default models.yaml and an empty workflows/ templates dir, data root (default ~/.local/share/lernie) with the tools//skills/ pools and the workspaces/ tree — seed-if-absent (ARCH §2.2); LERNIE_HOME collapses both
make uninstall [INSTALL_PREFIX=<p> LERNIE_HOME=<h>] Remove the installed binaries; leaves the harness homes (config + data roots) in place

The pinned adapter under test

The e2e tests exec the real bz (against a mock HTTP endpoint, not a provider), and lernie's load-time version guard (ARCH §4.4) demands the pinned version exactly. The pin's one home is the brazen = "=<version>" line in Cargo.toml.

The trap. bz normally resolves from PATH — that is ~/.cargo/bin/bz, machine-global mutable state shared by every checkout and every agent on the box. Anyone running make install rewrites that binary at their tree's pin. If your tree pins a different version, your next test run dies in five-plus e2e tests with

bz version "0.0.3" does not match the linked brazen crate "0.0.4"

which looks nothing like "someone else installed a binary" and everything like a regression you just wrote.

The cure. make test and make coverage do not use the PATH bz at all. They depend on $XDG_CACHE_HOME/lernie/bz/<pin>/bin/bz — installed from crates.io on first use — and put that directory first on PATH for the run, so the tests always exercise the pin this tree names, whatever the machine's bz happens to be. The version comes from BRAZEN_PIN in the Makefile, derived from Cargo.toml; the cache directory is named after it, so bumping the pin is a cache miss and nothing else, and a stale entry is never overwritten in place. Cost: one cargo install (~25s) per pin per machine — sibling worktrees share the cache — and nothing at all when warm, since it is an ordinary make file prerequisite.

Two consequences worth knowing:

  • Bare cargo test is still exposed. It inherits your PATH and so runs whatever bz is installed there. Use make test; if you must run cargo test directly, make install-bz first to line the global binary up with the tree's pin.
  • No test writes the global bz. make install does — that is its job — but the install test that runs it (tests/install.rs) points CARGO_INSTALL_ROOT at a per-worktree root under target/, so the pinned bz lands there and ~/.cargo/bin/bz is never touched by a test run.
  • Runtime resolution is unchanged. This is test determinism only — lernie itself still resolves the adapter per ARCH §4.4 (the models.yaml adapter: override, else a binding-injected target, else bz on PATH), and make install still puts the pinned bz on your PATH for real use.

the_makefile_derives_the_same_pin (src/prompt/tests/pin.rs) keeps the two readers of that one line honest: the Makefile's BRAZEN_PIN (which names the cached binary) and the crate's brazen_pin() (which the version guard compares against) must agree, or the tests would fail the guard against a binary the Makefile itself installed.

Workflow

All changes land on main via bl squash-merges. Direct commits to main are rejected by the pre-commit hook, and every landing on main is pushed to origin automatically (see Auto-push hook).

bl prime --as <you>
bl claim <task-id>              # creates a worktree; cd into it
# ...edit, test, commit...
bl close  <task-id> -m "..."    # squash-merges into main; run from the repo root

See bl skill for the full guide.

What gets published

cargo package ships the crate, not the repo. Cargo.toml's exclude keeps out everything that serves this git checkout only — docs/, tests/, experiments/, scripts/, .github/, .githooks/, .balls/, Makefile, tarpaulin.toml, release-plz.toml, AGENTS.md, CLAUDE.md, and rust-toolchain.toml (which would otherwise force a source builder onto this repo's exact pinned toolchain). What remains is src/, README.md, LICENSE, Cargo.lock, and the embedded asset trees template/, schemas/, skills/, install/models.yaml — those four are include_dir!/include_str! inputs, so excluding any of them is a build failure, not a smaller tarball. Verify a change to the list with cargo package --list and then cargo package, which compiles the extracted tarball.

crates/agent-eval is publish = false: it is workspace-internal and is not part of the published crate at all.

Pre-commit hook

.githooks/pre-commit enforces three rules on every commit:

  1. No direct commits to mainline. main and master are rejected unless the commit is the tail of a merge (MERGE_MSG/SQUASH_MSG present), which is how bl close lands squash-merges.
  2. 300-line cap on code files. The cap is a repo invariant, not a per-commit property, so the hook sweeps every tracked code file in the tree (git ls-files), not just the staged set — a file that crosses the cap in one commit and is untouched afterward is still caught. Docs (*.md, *.txt), config (*.toml, *.yaml, *.yml, *.json, *.lock), Makefile, .gitignore, LICENSE, and anything under .githooks/ are exempt.
  3. make check on every commit that touches a Cargo project: fmt-check (formatting), lint (clippy -D warnings), coverage (cargo tarpaulin --fail-under 100), and test-install (cargo test --test install). The hook invokes make check rather than re-listing the commands, so the close gate is always exactly what make check is — the Makefile is the single source. Formatting and lint drift therefore cannot land invisibly. test-install is a separate step because the install test shells out to a release build and cargo install brazen, which contend with tarpaulin's target/ lock; it is cfg_attr(tarpaulin, ignore), so without its own uninstrumented step the install contract — the first thing every user touches — would never run at the gate at all. It costs ~45s warm and leaves the machine-global ~/.cargo/bin/bz alone: the test redirects make install's cargo install brazen into a per-worktree root under target/ with CARGO_INSTALL_ROOT, so a sibling worktree at another pin is never rolled over. The toolchain is pinned in rust-toolchain.toml and the tarpaulin version in tarpaulin.toml (also .github/workflows/ci.yml) so fmt-check, lint, and the coverage denominator mean the same thing locally and on CI — newer tarpaulin releases have silently dropped inline #[cfg(test)] mod tests; files from the count, weakening the floor. make coverage aborts with an install hint if the local tarpaulin version drifts.

A floor of exactly 100% only holds if every line's coverage is caused by the code's own structure and not by winning a race, so no line may be reachable only while a clock has not yet run out. With several agents measuring coverage at once, whichever side of such a race the machine happens to pick that minute decides the verdict, and the gate reports an uncovered line on a diff that touched nothing. Two shapes to write around:

  • A retry budget is a count of attempts, never a wall-clock deadline. PROBE_RETRIES (src/prompt/tests/exit_launch.rs) is the one budget every executor-lock probe shares; a deadline expires on load rather than on evidence, so under load the give-up arm can be taken on the first pass and the retry arm never runs at all.
  • A poll loop waits because its child is still running, not because a flag has yet to land. wait_with_cascade (.../builtin/bash/mod.rs) and wait_with_stop (.../tool/subprocess.rs) therefore sleep between the reap and the flag read: the interval is entered for as long as the child lives, instead of only while a stop scheduled milliseconds out has not arrived yet.

The same objection reaches past coverage to the verdict, and the end-to-end tests answer it the same way: a poll waiting on a detached driver is bounded by consecutive probes that saw no change in the workspace tree, never by wall time (src/e2e/poll.rs, docs/ARCHITECTURE.md §9). A live driver writes continuously and a wedged one writes nothing, so a loaded box only makes the pass path slower — where a stopwatch would have turned a slow success red, and (as bl-2bf0 found) hid a real defect behind a timeout that read like machine load.

There is no --no-verify escape hatch in the workflow. If the hook rejects a commit, fix the underlying issue rather than skipping.

Auto-push hook

.githooks/reference-transaction pushes main to origin the moment local main advances. Landing and publishing are one act: a bl close reaches GitHub and the push triggers the Release-plz workflow, which contains CI as a called job (needs: ci) and only publishes once it is green. origin/main cannot silently fall months behind local main again.

Why a reference-transaction hook and not post-commit. Nothing lands on this repo's main through git commit. bl close delivers by plumbing — git commit-tree, then git update-ref refs/heads/main — which fires no commit hook and no merge hook at all — every commit bl has landed on main arrived that way. Git's reference-transaction hook is the one event every landing path shares: the plumbing delivery, a git merge --no-ff, and a plain commit alike all end in an update of refs/heads/main.

The hook acts only on the committed state of a transaction that moves refs/heads/main to a new value, and only when an origin remote exists. Everything else — side branches, refs/remotes/* (including the ones its own push writes, so it cannot recurse), no-op rewrites like git pack-refs, and a deletion of main — falls through untouched.

It cannot block or hang a landing. Git aborts a ref transaction when this hook exits non-zero in the prepared state, so every path in it exits 0 — which is also why it does not set -e. A push that fails prints one warning line on stderr and nothing else, and timeout 30 bounds an offline push rather than stalling the commit behind a TCP timeout. Git runs reference-transaction hooks from 2.28 onward; on anything older the file is simply never invoked and main has to be pushed by hand.

tests/hooks.rs exercises the shipped hook file itself against a local bare repository as origin — never the real remote — and covers all six behaviours above: a commit on main pushes, a commit-tree + update-ref delivery pushes, a --no-ff merge pushes, a side-branch commit pushes nothing, an unreachable origin warns without failing the commit, and a repo with no origin is silent.

Commit-identity guard (opt-in, per machine)

main's history carries exactly one human identity, mudbungie <mudbungie@gmail.com>, and no Co-Authored-By trailers — it was normalized to that on 2026-07-26. tests/commit_hygiene.rs keeps it that way, but only on a machine that asks for it: the test arms itself on the presence of $XDG_CONFIG_HOME/lernie/enforce-commit-identity (default ~/.config/lernie/enforce-commit-identity), an empty marker file outside the repo. Absent — the default in public CI and in every clone — the test returns without asserting anything.

Armed, it walks all of refs/heads/main and fails on any commit whose author or committer is neither mudbungie <mudbungie@gmail.com> nor github-actions[bot] (the bot stays allowed: release-plz authors the release commit as it), on any Co-Authored-By trailer, and on any mention of a throwaway or personal address in an identity or a message. The policy lives in the marker, not in the code: rm it and the guard is off, with no code edit and no flag. Create it with touch ~/.config/lernie/enforce-commit-identity.

License

MIT. See LICENSE.