# lernie
[](https://github.com/mudbungie/lernie/actions/workflows/release-plz.yml)
A git-backed agent harness. Design spec: [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md).
Principles catalog: [`docs/PRINCIPLES.md`](docs/PRINCIPLES.md).
Vocabulary reference: [`docs/TAXONOMY.md`](docs/TAXONOMY.md).
Promise suite (the user stories 0.0.1 is evaluated against): [`docs/USER_STORIES.md`](docs/USER_STORIES.md).
CI runs `make ci` (`fmt-check` + `lint` + `coverage` with the 100% gate + `test-install`) on every push and pull request to `main`. The e2e tests exec the real provider adapter `bz`, which the test targets install themselves at the pinned version (see **[The pinned adapter under test](#the-pinned-adapter-under-test)**). The Rust toolchain is pinned in `rust-toolchain.toml` — CI, the pre-commit gate, and every contributor build under the same `rustc`/`rustfmt`/`clippy`. That pin binds this git checkout only; it is excluded from the published crate, whose supported floor is the declared `rust-version = "1.88"` (the crate's `let` chains, not edition 2024's 1.85).
## One command surface, two bindings
lernie is defined once as a **command surface** — the set of verbs, their arguments, and their products (ARCH §3.4). It is consumable two ways, and both are the *same* control plane:
- **Exec binding** — run the `lernie` binary: `exec("lernie", args)` with env-var auth. This is what the CLI and every frontend use.
- **Linked binding** — depend on the `lernie` crate and drive the same verb entries in-process. The crate's entire public API is `lernie::cmd` (the `Cli`/`Command` clap surface, one `run` entry per verb, the `Fx`/`Outcome`/`Error` binding seam, and the `prelude` binding preludes). The linked binding is **pin-exact 0.x only** — no semver stability, the posture brazen takes toward lernie.
Parity between the two is enforced mechanically, not by convention. `tests/command_surface_parity/` asserts the bijection at three depths: it pairs each verb's `Command` variant with its module's entry *as function values*, so the compiler — not an assertion — proves the two share one argument type and one product type; it walks the crate's whole module graph (via `syn`) and asserts that every externally reachable declaration (item, field, enum variant, method, derive, trait impl) is exactly a verb's entry, its arguments, its products, or the binding preludes, with every `src/**/*.rs` proven reachable so nothing can hide in a file the walk never opened; and it asserts, per verb, that the CLI's introspected argument set (via clap) is exactly that verb's public `Args` fields — same names, same arity, same named-vs-positional form. It rides `make check` (hence the pre-commit hook and GitHub Actions), so a divergence between the linked surface and the CLI fails the build.
## Quickstart
```
cargo install lernie --locked # or: make install, from a clone
cargo install brazen --version =0.0.4 --locked # the provider adapter, always needed
lernie new ~/work/chat # create a workspace (bare repo.git + config/default)
ANTHROPIC_API_KEY=... lernie prompt ~/work/chat 'hello'
```
Three install routes, not one — see **[Install](#install)** for what each
lays down. Every route needs `bz`; only `make install` installs it.
## Install
There are three routes, and they do not lay down the same things. All
three need a second binary — the provider adapter `bz` — which only the
Makefile route installs for you.
| | `cargo install lernie` | release tarball | `make install` |
|---|---|---|---|
| binaries | `lernie` | `lernie` | `lernie`, `agent-eval`, `lernie-eval-agent` |
| installs `bz` | no | no | yes, at the pin |
| runs `lernie prime` | no | no | yes |
| lands where | cargo's bin dir | wherever you unpack it | `$INSTALL_PREFIX/bin` |
### From crates.io
```
cargo install lernie --locked
cargo install brazen --version =0.0.4 --locked # the pinned provider adapter
lernie prime # found the harness root
```
You get the `lernie` binary alone, in cargo's bin directory
(`~/.cargo/bin` unless `--root`/`CARGO_INSTALL_ROOT` says otherwise) —
no `agent-eval`, no `lernie-eval-agent`, no `bz`, and nothing runs after
the build. The `lernie prime` line is optional but explicit: `prime`
founds the harness root (below), and `lernie new` founds it too on its
way to creating a workspace, so a user who skips `prime` is not stranded
— only uninformed about where their state went.
### From a GitHub release
Each `v*` release carries `lernie-x86_64-unknown-linux-gnu.tar.gz`: the
`lernie` binary, this README, and the license. Unpack it, put `lernie`
on your `PATH`, then run the `cargo install brazen` and `lernie prime`
lines above — the tarball ships no adapter and runs nothing.
### From a clone, with make
```
make install # default: ~/.local/bin, XDG homes
make install INSTALL_PREFIX=/usr/local # binaries -> /usr/local/bin/
make install LERNIE_HOME=/opt/lernie # collapse both homes -> /opt/lernie/
```
`make install` runs a release build and then:
1. Installs `lernie`, `agent-eval`, and `lernie-eval-agent` into
`$INSTALL_PREFIX/bin` with `install -m 0755` (atomic overwrite, no
symlinks). Make sure that directory is on your `PATH`.
2. Installs the provider adapter — brazen's `bz` — with
`cargo install brazen --version =<pin> --locked`, where the pin is
the `brazen = "=<pin>"` dependency in `Cargo.toml` — its one home;
the Makefile and the load-time guard both derive from that line.
One binary serves every provider
(ARCH §4.4); the harness resolves `bz` on `PATH`, and a load-time
guard rejects any `bz` whose version differs from the pin.
3. **Founds the harness root by invoking `lernie prime`** — the single
verb that seeds the installation substrate (ARCH §2.2), so the
Makefile no longer duplicates the seeding. `prime` resolves the roots
(XDG split, collapsed by `LERNIE_HOME`) and lays down the default
`models.yaml` under the **config root** (model capabilities + context
windows — no endpoints or auth, which are brazen's), the `tools/` and
`skills/` pools and the `workspaces/` tree under the **data root**,
and the empty `workflows/` templates dir. It is **seed-if-absent
throughout**: a second run changes nothing, and a hand-edited
`models.yaml` (or any operator-added pool entry) survives a re-install.
The shipped assets are embedded in the binary, so `prime` needs no
source tree — `LERNIE_HOME=<dir> lernie prime` seeds any fresh home.
There is no frozen profile pool: the config a workspace runs under is
its own `config/default` commit, authored from
[`template/`](template/) at `lernie new` (fork is the freeze, ARCH §2.2).
4. Smoke-tests the freshly installed binaries with `lernie --version`
and a throwaway `lernie new`. Failure aborts the install with a
non-zero exit.
Its closing banner prints what the other two routes leave you to find
out: the install prefix, both harness roots, and the `bz` commands
below.
### The adapter is a second binary
Nothing prompts without `bz`. It is brazen's one stateless binary for
every provider (ARCH §4.4), it is **pinned exactly**, and lernie refuses
a `bz` at any other version rather than downgrading silently:
```
cargo install brazen --version =0.0.4 --locked
```
The pin is not folklore you have to read this file for — the installed
binary carries it: `lernie --version` prints the linked pin beside its
own version, `lernie <version> (brazen 0.0.4)`. Its one home is the
`brazen = "=<pin>"` line in `Cargo.toml`;
the Makefile's `BRAZEN_PIN`, the load-time guard, `lernie --version`,
and every pin printed in this file all derive from that line (a test
holds them equal). With no `bz` at all, the first verb that drives a
model call says so and hands you the command above.
Provider endpoints, auth, and wire dialects live entirely in brazen's
own config (`~/.config/brazen/config.toml`; inspect with
`bz --dump-config`, authenticate with `bz --login --provider <id>`).
lernie references a provider *row* by name and never sees credential
material (ARCH §4.1).
### Where the state goes
The harness root is the installation-global substrate (ARCH §2.2), split
by XDG lifetime: `$XDG_CONFIG_HOME/lernie` (hand-edited declarations —
`models.yaml`, `workflows/`) and `$XDG_DATA_HOME/lernie` (machine-
populated pools and the `workspaces/` tree). `LERNIE_HOME=<dir>`
collapses both to one directory, at install time and at runtime alike.
`lernie prime` founds it, seed-if-absent throughout, so running it again
— or after an upgrade — never clobbers a hand edit. Only `make install`
runs `prime` for you; on the other two routes it is your first command,
or `lernie new`'s side effect.
`make uninstall` removes the three installed binaries; `bz`
(installed via cargo) is removed with `cargo uninstall brazen`. The
harness homes (the config and data roots, holding config and
workspaces) stay put — clean them up manually if you want a true
uninstall.
## Configuration schemas
JSON Schemas for the harness-root and config-commit control files (per
[`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md) §2.2, §4.1) are generated
from the Rust types under `src/config/`. `make schemas` writes them to
`schemas/` for editor integration and external validators. Generation is a
golden test (`config::schemas::write_to` vs the checked-in `schemas/`):
`make schemas` runs it with `UPDATE_SCHEMAS=1` to rewrite the directory,
and the same test under `make check` fails if `schemas/` ever drifts from
the source types — so the tree is always current, with no separate binary
to run.
| File | Backed by Rust type | Config-commit / on-disk file |
|-------------------------------|--------------------------------------------|---------------------------------------|
| `schemas/version.json` | `config::version::Version` | `version` (config commit) |
| `schemas/manifest.json` | `config::manifest::Manifest` | `manifest.yaml` (config commit) |
| `schemas/workflow.json` | `config::workflow::Workflow` | `workflow.yaml` (config commit) |
| `schemas/providers.json` | `config::per_repo_providers::PerRepoProviders` | `providers.yaml` (config commit, `roles:`) |
| `schemas/models.json` | `config::models::Models` | `<config-root>/models.yaml` |
## Layout: harness root and workspaces
The harness root is installation-global state, split by XDG lifetime
into two homes (ARCH §2.2). `LERNIE_HOME`, if set and non-empty,
collapses both to that one directory (test isolation, alternate
installs). Three distinct on-disk locations:
- **Config root** — hand-edited declarations, `$XDG_CONFIG_HOME/lernie`
(default `~/.config/lernie`). Holds the global
[`models.yaml`](docs/ARCHITECTURE.md#42-model-abstraction) (model
capabilities + context windows, plus an optional `adapter:` binary
override — §4.2) and the `workflows/` templates. Provider endpoints
and auth live in brazen's config, not here (§4.1).
- **Data root** — machine-populated pools, `$XDG_DATA_HOME/lernie`
(default `~/.local/share/lernie`). Holds the `tools/` and `skills/`
pools plus the `workspaces/` tree. Shared across every workspace.
- **Workspace** — one git repository per workspace, at
`<data-root>/workspaces/<workspace>/` (ARCH §2.2): a bare `repo.git`
holding config branches (`config/<name>`) and agent refs
(`agents/<agent-id>`) — **no `main`**. The control files
(`providers.yaml` `roles:` only — §4.3, `manifest.yaml`,
`workflow.yaml`, `version`, `souls/`) live in the **config commit**,
read from each agent's governing config commit (`git merge-base`
against the `config/*` heads — derived from ancestry, never stored).
Agent worktrees are siblings under `agents/<agent-id>/`; `steps/` and
`inbox/` sit at the workspace root, outside every worktree.
Workspace repositories are never pushed to a remote.
`lernie new` creates a workspace and authors its **first config commit**
— an orphan root on `config/default` — from [`template/`](template/),
the versioned skeleton embedded into the `lernie` binary at build time:
```
lernie new # auto-id under <data-root>/workspaces/
lernie new /path/to/my-workspace
```
Or via the Makefile wrapper:
```
make new-workspace DEST=/path/to/my-workspace
```
The binary **founds the harness root first** — it runs the same
seed-if-absent routine `lernie prime` is (ARCH §2.2), so a data root
nobody primed gains the `tools/` and `skills/` pools before they are
read, and a primed install is untouched (nothing is clobbered, no flag
is involved). That is what keeps the next step honest: the pools are an
input to the config commit, so an unprimed root would otherwise author a
commit with an empty `descriptions/**` and hand every agent forked off
it an empty toolset. It then runs
`git init --bare -b config/default <dest>/repo.git`,
materializes a transient authoring checkout, extracts the template's
control files into it, snapshots the data-root pools into
`descriptions/{tools,skills}/` (ARCH §3.3 descriptions-always), commits
(`config: init [config/default]`), and tears the checkout down. The
workspace is left with exactly one ref — the config commit every fresh
root agent forks off (fork is the freeze, §2.2). The destination must
either not exist or be an empty directory. With no path argument, the
destination is `<data-root>/workspaces/<auto-id>/`; the created path is
printed on stdout. `goal.md` and `soul.md` are intentionally not in the
template — they are written per-branch at dispatch time (ARCH §2.3,
§2.8), which also **removes the control files from the agent's tree**
(§2.2: control is read from the config commit; worktrees hold only
context).
**Pre-v1 clean break (ARCH §10):** the retired per-conversation layout
(a `root/` worktree with loose control files) is refused with an
actionable error, not migrated — create a fresh workspace with
`lernie new`.
**First-run smoke test (required).** `lernie new` authors the default
`providers.yaml` with a concrete model id, but validates it against
nothing — id validity is brazen's fact, and lernie runs no model-list
reconciliation (ARCH §4.2, the settled stance). A wrong id surfaces only
at the first live model call. The required next step after creating a
workspace is therefore a live `lernie prompt` (see the quick start
above): it is the cheapest — and, by that stance, only — check that the
authored id actually resolves on the wire.
`make smoke` automates exactly this:
```
make smoke # scaffold a throwaway workspace + one live 'lernie prompt'
```
It founds a throwaway harness root with `lernie prime` — from the assets
**embedded in the binary**, the same front door `make install` uses, so
the shipped install path is exercised too — scaffolds a workspace with
`lernie new`, then runs one live `lernie prompt` against the **shipped
defaults** — worker role, provider `anthropic`, model `claude-sonnet-5` —
through the real `bz` data plane.
The verdict is read from **observable state, never the agent's own
claim**: the `lernie prompt` exit code is 0, the agent ref
(`agents/<id>`) carries a committed transcript entry, and the off-worktree
step record (`steps/<id>/001/`) holds a response with **no wire error and
real assistant text**. That last pair is the point: an auth-failed run
still creates the branch and a step record whose response terminates in a
clean `end` — the failure rides an `error` event ahead of it — so
branch-exists and step-exists alone would pass a broken wire. `make smoke`
requires exit 0 **and** no `error` event **and** an assistant
`content_delta`.
By default `make smoke` runs the **shipped default** — provider
`anthropic`, model `claude-sonnet-5` — which needs a configured `bz`
credential for the `anthropic` provider (`bz --login --provider
anthropic`, or set `ANTHROPIC_API_KEY` / `BRAZEN_API_KEY`) and spends real
money. To run the same live check against **any other `bz` provider row**,
set **both** `SMOKE_PROVIDER` and `SMOKE_MODEL` (both-or-neither — one
alone is a usage error; unset leaves the shipped default byte-for-byte):
```
make smoke SMOKE_PROVIDER=local SMOKE_MODEL=<a-pulled-ollama-model>
make smoke SMOKE_PROVIDER=codex SMOKE_MODEL=gpt-5.4
```
The override is laid into the throwaway config root through the same front
doors a real install uses — a `providers.yaml` override under
`<config-root>/template/` (the config-root override) plus a `models.yaml`
placed in the config root before `lernie prime` (its seed-if-absent
contract, ARCH §4.2) — so there is no new `lernie` flag or verb. Local
`ollama` (bz's `local` provider row) needs no credential, only a model
that is actually pulled and served; the credential note above applies to
the `anthropic` default alone.
**What `SMOKE_PROVIDER=local` does and does not prove.** bz's `local`
row (protocol `ollama_chat`) rejects a canonical `tool_result` block:
the second step of any tool-using run comes back as
`{"type":"error","kind":"parse_input","message":"user accepts only text
content"}`. So the local recipe validates the **tool-free path only** —
one model call, assistant text, a committed transcript entry. It cannot
exercise a tool step, a compactor (whose whole toolset is
`write_summary`/`mark_for_deletion`), or any multi-step loop that runs a
tool. This is a brazen-side gap in that provider row, not a lernie one,
and is filed there as brazen `bl-fba7`; to smoke a tool-using path,
point `SMOKE_PROVIDER`/`SMOKE_MODEL` at a row whose protocol carries
tool results.
`make smoke` is deliberately **not** part of `make check` or the close
gate: `make check` mocks the wire (httpmock Anthropic SSE), so it can
never catch a shipped default that fails on the real provider — which is
exactly how the fake id `claude-sonnet-4-7` once shipped unnoticed. It
runs only on demand.
## Authoring config commits
`lernie new` authors a workspace's *first* config commit. Every later
one — the general harness-assisted user act of ARCH §2.2 — is
`lernie config`:
```
lernie config <workspace> # advance config/default
lernie config <workspace> <name> # advance config/<name>
lernie config <workspace> <name> --from <src> # fork config/<name> off config/<src>
lernie config <workspace> <name> --orphan # fresh orphan lineage
```
The verb materializes a transient checkout of the target config lineage,
refreshes the `descriptions/**` snapshot from the data-root pools (ARCH
§3.3), opens the checkout in `$EDITOR` (falling back to `vi`) so you edit
the control files (`workflow.yaml`, `providers.yaml`, `manifest.yaml`,
`souls/`, `version`), commits, and tears the checkout down. `<name>`
defaults to `default`. `--from` and `--orphan` are mutually exclusive and
only apply when creating a new branch. A `--from <src>` naming a lineage
the workspace does not have is resolved *before* the checkout is
materialized, and declined by name:
```
lernie config: no config lineage "nosuch" in this workspace — existing lineages: default, strict
```
**Declining is fine, and leaves nothing behind.** Save no change and the
pass is *declined*: there is nothing to commit, so no commit is authored,
the branch does not move, and a `--from` / `--orphan` branch the pass
would have created is not left behind. That is a success — `lernie
config` exits 0 and prints the one line
```
config/default unchanged: the edit changed nothing, so no config commit was authored
```
so empty stdout means a commit landed. The transient checkout is torn
down on every exit path (a decline, a git decline, an editor that fails),
so the next `lernie config` always runs. Only a hard kill mid-pass can
leave the checkout behind, and the next pass clears it before starting
(ARCH §2.11 "the next touch heals") — at the cost of the killed pass's
unsaved edit, which was never committed.
This is the **only** act that advances a config branch (ARCH §2.3);
agents forked before it keep their governing config, and agents forked
after it govern under the new head (fork is the freeze, §2.2).
## Sending a prompt
```
lernie prompt /path/to/my-conversation 'hello'
```
`lernie prompt` is the root-agent path (ARCH §2.3, §2.6, §2.7,
§2.8, §2.10). Each invocation spawns its own `agents/<conv-id>` branch
off the default config branch's head (§2.2–§2.3 — there is no `main`),
drives each step's model call through brazen's `bz` (§4.4), and steps
until a terminal event. **There is no terminal compaction stage** (§2.7):
compaction runs only at the checkpoints `workflow.yaml` declares, and a
branch with no configured trigger never compacts. Merge-back is gone
(§2.6): the root branch persists on its own ref (§2.4), and a child
returns by depositing a result message into its parent's inbox (§2.6):
1. Resolve the harness root (`LERNIE_HOME`, else XDG homes, ARCH
§2.2) and guard the workspace layout (a non-workspace, or the
retired per-conversation layout, is refused — §2.2, §10). Load
`<config-root>/models.yaml` (capabilities + context windows +
optional `adapter:` override — §4.2) and, from the config commit's
tree (`git show <config-commit>:providers.yaml`, §2.2),
`providers.yaml` (`roles:` block — §4.3); cross-validate
`roles.worker.{provider,model}` against `models.yaml`. For a fresh
root the config commit is `config/default`'s head — the very commit
the new agent forks off; `lernie advance` derives an existing
agent's **governing config commit** from ancestry instead
(`git merge-base` against the `config/*` heads — never stored).
2. Run the load-time version guard: `bz --version` must equal the
linked brazen crate version (§4.4). Under an `adapter:` override the
guard is skipped and the in-band `MessageStart.v` handshake governs.
Read the worker soul from the config commit's `souls/worker.md`
(§2.2, §4.3).
3. Spawn branch `agents/<conv-id>` (§2.3 — the id is the bare
hyphenated descent; the `agents/` prefix is the ref namespace) off
`config/default` and allocate a worktree at
`<workspace>/agents/<conv-id>/` (§2.2). Write the branch goal to
`goal.md` and the role soul to `soul.md`, remove the config commit's
control files from the tree (§2.2 — the worktree holds only
context), and commit — that commit's tree is step 1's read state
(§2.10).
4. Build a typed `brazen::CanonicalRequest` (linked crate — the
fail-open `extra` map stays unreachable), mirror it to
`<workspace>/steps/<conv-id>/001/request.json` (a diagnostic
artifact, outside every worktree, never read at runtime, §2.3).
5. **Model call, harness-owned retry loop (§2.10, §4.4).** Exec
`bz --json --provider <row>` once per *attempt*, canonical request
on stdin, appending each attempt's stdout verbatim to
`<workspace>/steps/<conv-id>/<NNN>/response.json` as brazen `v=1`
NDJSON — one self-delimiting segment per attempt, each ending in a
terminal `end`. On a retryable in-band `Error`
(`CanonicalError::retryable()`, never re-derived) the harness
re-invokes `bz` with the identical request, up to the `workflow.yaml`
attempt cap with exponential backoff — floored by the failed
attempt's `Retry-After` pacing hint
(`CanonicalError::retry_after_seconds`) when it carries one, so the
config schedule governs and the provider's hint can only lengthen it
(§4.4). brazen never retries; auth and
endpoints are entirely its own. The `response.json` fd is held open
across every attempt and backoff sleep — its close is the §3.5
IN_CLOSE_WRITE completion signal. As the events stream, the harness
tracks only their *framing* — the terminal `end`, an in-band `Error`,
the handshake `v` — for retry/classification; `meta.json` carries
`{commit, started_at, ended_at}`. The events' *content* streams into
the **transcript writer**'s (§2.3) staging file
`<workspace>/steps/<conv-id>/<NNN>/staging.json`,
appending each content block as it completes; segment authority
(§4.4) truncates it on an `Error` attempt and the settling `Finish`
seals it — one stream, two sinks (diagnostic `response.json` +
transcript), never read back. When the model call completes, the
sealed file is renamed into the worktree as
`messages/NNN-<model-id>.json` — its origin token is the model that
authored it (§2.3), the body a JSON array of canonical `Content`
blocks — and committed. `NNN` is the branch's transcript counter,
max-present-plus-one from the `messages/` listing, evaluated at
commit time. The initial user message now enters through the front
door like any other (§2.11): the executor deposits it into the agent's
own inbox, and the step-boundary drain delivers it as the first
transcript entry `messages/NNN-user.md` (bl-1129) — no bespoke
initial-message path beside the drain.
6. **Step loop (§2.5).** At each step boundary the executor first
**drains the inbox** (bl-1129, §2.11): after committing any
renamed-but-uncommitted stray a prior death left in `messages/`, it
moves each pending `inbox/<agent-id>/<sender>-<NNN>.md` into the
worktree as `messages/<counterNNN>-<sender>.md` (a literal `rename(2)`
— one home at every instant) and commits the move, in a deterministic
`(mtime, filename)` order, ahead of the read-state capture so a
delivered message is part of the commit the model call assembles from.
Each step then re-assembles its model-facing history
from the read-state commit's tree — `readdir` of `messages/`, sorted
by the filename's `NNN` prefix, each entry composed by its origin
token (`NNN-<sender>.md` → user text, `NNN-<model-id>.json` → the
assistant message — any `.json` token but the reserved `tool`,
`NNN-tool.json` → `tool_result` in the following
user message), with consecutive same-side entries grouped into one
alternating wire message. There is no in-memory history and no
git-log walk; running, retry, and replay are one code path against
one input, the commit's tree (§2.3, §5). If the settled model-output
entry carries any `tool_use` block, run every one through the tool
executor — the per-call records land under
`<workspace>/steps/<conv-id>/<NNN>/tools/<tool-id>/` (out of every
worktree, §3.3; written but never read at runtime), and as each tool
resolves the transcript writer commits `messages/NNN-tool.json` (its
canonical `tool_result` block) — then loop into step `<NNN+1>`. A
step with no `tool_use` block is terminal. Step ≥2 has no *dispatch*
commit, but each step's transcript entries (assistant output, tool
results) do advance the branch tip, which is that step's read state
(§2.10). `tool_use`/`tool_result` pairing holds by construction: a
tool result commits immediately after its emitting step's model-output
entry, so it always lands in the immediately following user message.
Closing each tool step, the executor reads the **compaction
checkpoint clock** (§2.7, §6) — `compaction.intermediate.trigger` in
`workflow.yaml`: `every_n_commits`, `every_t_seconds`, or the
agent-elected `on_flush`, all derived from git (commits and elapsed
seconds since the last compaction merge, or the branch root when none
has landed — never a stored counter). When it is due, the
`worker_flush: dispatch(compactor)` binding forks a compactor off the
branch tip — the checkpoint commit `C` — and the branch keeps
stepping straight through it; no quiescence is imposed. Omit the
`compaction:` block and the branch never compacts.
7. **Terminal return (§2.6, §2.3 step 5).** Every terminal event —
normal completion (`final-response`), budget exhaustion
(`budget-exhausted`, §6), and stop (`stopped`, §2.9 — the executor's
SIGTERM handler deposits on its way out) — deposits a **result
message** into the parent's inbox: an ordinary deposit whose
frontmatter adds `epitaph:`
and `terminal_ref:` (the branch tip) and whose body is the terminal
response iff the agent spoke. For a root this is a structural no-op —
a root has no parent inbox; its response answers the user (§2.4). The
deposit is executor-side, never a model tool call ("Return is not a
verb"). At delivery, a message carrying `terminal_ref:` applies the
fork-point→terminal **work-product transfer** as one commit before its
delivery commit, filtered to work products; a diff that fails to apply
is declined at `refs/lernie/conflicted/<agent-id>` (§2.6).
8. **Exit protocol (§2.11).** With the terminal deposit landed, the
executor runs the branch's terminal `workflow.yaml` bindings
(`branch_stopped` → `mark_abandoned` / `notify_ui`, §6), releases the
executor lock, and only then spawns a driver at its own agent and —
the deposit's own probe-and-launch — at the parent the deposit just
revived. Both launches are fire-and-forget and both are decided by
epitaph *value*: a final response launches, `stopped` and
`budget-exhausted` never do. **No terminal compactor is dispatched**
(§2.7): the v0.3 terminal-compaction stage is deleted, along with the
`Dispatcher` re-entry that existed only to run it. Compaction is a
checkpoint event (step 6), never an exit stage. **Merge-back is gone
(§2.6):** the root branch persists on its own ref (§2.4); nothing
merges back, and the agent's worktree is not torn down (quiescence,
not teardown, §2.3 step 6).
9. Print the agent id (the bare conv-id) on stdout.
After `lernie prompt` returns, inspect the agent against the bare
workspace repository:
```
cd /path/to/my-workspace
git -C repo.git log --oneline --decorate agents/<conv-id> -4
git -C repo.git ls-tree --name-only agents/<conv-id> messages/
git -C repo.git show "agents/<conv-id>:messages/002-<model-id>.json"
ls steps/<conv-id>/
```
The log is the dispatch commit followed by one `transcript NNN:` commit
per entry, its subject naming that entry's **origin token** — `user`, a
sender's agent id, `tool`, or the authoring model's id:
```
f265de7 (agents/…) transcript 002: qwen3.5:9b […]
7ae527e transcript 001: user […]
f643a50 step 001: dispatch […]
6f4bd05 (config/default) config: init [config/default]
```
`ls-tree` lists the transcript itself (`messages/001-user.md`,
`messages/002-<model-id>.json`, …) and `show` prints one entry — a
model-output entry is a JSON array of canonical `Content` blocks, e.g.
`[{"type":"text","text":"pong"}]`. `ls steps/<conv-id>/` lists the
off-worktree step records, one numbered directory per step, each holding
`request.json`, `response.json`, and `meta.json`. There is **no merge
commit and no `summary/`** on a branch that never reached a compaction
checkpoint (step 6) — those appear only once a compactor has returned.
The root branch persists unmerged by design (§2.4), so the health metric
is no longer branch count but silent deaths and undelivered returns
(ARCH §8) — read straight from git refs, the executor lock, and inbox
listings, with no sidecar file.
## Stopping a conversation
```
lernie stop /path/to/my-conversation <conv-id> [--stop-children]
```
Sends `SIGTERM` to the process group of the **one executor** driving
`<conv-id>`, with a 5-second flush deadline before `SIGKILL`. This is the
same cascade pattern adapter (§4.4) and tool (§3.3) cancellation use,
applied to the harness itself
([ARCH §2.9](docs/ARCHITECTURE.md#29-stopped-branches)). The group signal
reaches that executor's own `bz` and tool subprocesses — its limbs — and
**stops at the agent boundary**: a dispatched child harness has taken its
own process group, so a bare stop does not fell it. A running child
outlives the stopped parent and revives it later by depositing its result
(§2.11) — stopping a parent strands nothing.
`--stop-children` opts into the agent→agent cascade: it walks the id
namespace — the descendants of `<conv-id>` are exactly the inbox
directories prefixed `<conv-id>-` (§2.3), one prefix scan reaching every
depth — and folds each descendant executor's group into the same sweep.
The pid is discovered by scanning `/proc/<pid>/fd/*` for the process
holding the agent's inbox-directory lock fd open — the executor lock
(§2.11), held for the whole step loop, so a stop lands even during tool
execution when no `response.json` is open — no sidecar pid file. Linux
only.
The pgid that scan produces is **vetted before anything is signalled**,
because a pid is discovered before its group has settled: between a
driver's fork and the `setpgid`/`setsid` it runs at startup, `/proc`
still reports the group it inherited from its spawner — your shell job.
So a pgid is trusted only once it equals the holder's own pid (a group
leader's does, and every driver becomes one), re-read a bounded number of
times while it does not, and refused rather than signalled if it never
settles; a stop that signals nothing is re-runnable, one that signals
your shell is not. `lernie stop` additionally refuses any group it is
itself standing in (§2.9).
The group signal reaches every member independently: `bz` installs no
handler and dies at once (leaving the missing-`end` signature, §4.4),
while the **executor catches its own copy** — SIGTERM is catchable — and,
instead of dying on the spot, deposits its branch's `stopped` result on
its way out (§2.9 step 3, executor-side, "Return is not a verb") and then
exits cleanly. Catching shields nobody: the kernel already delivered to
`bz` and the tools. For a root the deposit is a no-op (no parent inbox);
the observable is the clean exit.
Because the model call is where the wall time goes, that is where a stop
usually lands — so the clean exit is the ordinary case, not the rare one.
The **flag classifies, not the error's shape**: a kill lands wherever the
adapter was, leaving a half-stream, a torn JSON line, or a provider error
depending on the instant, and with a stop pending each is read as the
stop. With no stop pending the same faults still propagate non-zero, so a
genuinely dying adapter is never hidden. The retry loop respects the flag
too: a stop is never followed by another `bz` invocation.
Behavior:
- **Idempotent.** A branch with no live writer (already stopped, or the
harness exited cleanly) returns success without sending any signal.
- **Errors when** the agent branch (`agents/<conv-id>`) doesn't exist.
Surfaces as a non-zero exit with a `lernie stop:` prefix on stderr.
(The old "already merged" refusal died with `main`: nothing merges,
so there is no merged state to refuse — an already-terminal branch is
simply the idempotent no-holder case above.)
- **No on-disk cancel marker.** The §2.9 signature of a stopped branch
is the latest step's `response.json` closed without a terminal `end`
event — produced by `bz` dying mid-stream on its own SIGTERM (§4.4);
the executor's `stopped` deposit is an independent write to the inbox
tree and never touches that signature.
The frontend's stop button (per [ARCH §3.5](docs/ARCHITECTURE.md#35-ui-contract))
exec's this exact subcommand; there is no second control surface.
## Built-in tools (v0.3, +v0.4 Phase 2 dispatch)
The agent can call **built-in tools** that ship inside the `lernie`
binary as `lernie tool <name>` subcommands (ARCH §3.3 / §12). The tool
executor's resolution order — `<data-root>/tools/lernie-tool-<name>`
→ `PATH` → `lernie tool <name>` — falls through to this in-process
route for tools not externalized.
Each built-in is the triple §3.3 pins:
- **Binary** — the `lernie tool <name>` subcommand. Reads
`tool_use.input` JSON from stdin, writes raw bytes to stdout, exits
0 on success or non-zero on failure (stderr is concatenated after
stdout into `tool_result.content` when `is_error` is set).
- **JSON schema** — at [`schemas/tools/<name>.json`](schemas/tools/),
seeded to `<data-root>/tools/<name>.json` by `lernie prime` (which
`make install` invokes, ARCH §2.2). Sent verbatim as the
`input_schema` of the tool's entry in the model call's `tools: [...]`
array.
- **Skill** — at [`skills/<name>/SKILL.md`](skills/), seeded to
`<data-root>/skills/<name>/` by `lernie prime`. The frontmatter `description` is
the tool's description in `tools: [...]`; the body explains when to
reach for it.
The pool is discoverable from the CLI itself — `lernie tool --help`
names it, and a name that is not in it is declined non-zero naming it
too, the same way `load_skill` declines an unknown skill (ARCH §3.3):
```
$ lernie tool --help
Arguments:
<NAME> Built-in tool to run; one of: bash, dispatch, load_skill, message, read_file
$ echo '{}' | lernie tool nosuchtool
lernie tool nosuchtool: unknown built-in tool: "nosuchtool"; available: bash, dispatch, load_skill, message, read_file
```
Built-ins:
- **`read_file`** — read the entire contents of a file at a given
path. Rejects files larger than 1 MiB, reporting the file's **true**
size (`stat`, not the capped read's length) so the agent can judge
the magnitude it is up against; v0.4+ adds the oversized-output
auto-dispatch shim (ARCH §3.3 / §12). Try it directly:
`echo '{"path":"README.md"}' | lernie tool read_file`.
- **`bash`** — runs a shell command via `sh -c` and returns its
stdout. The shell runs in its own process group so a SIGTERM the
harness sends is forwarded to the entire spawned tree (§2.9
cascade). Try it directly: `echo '{"command":"ls"}' | lernie tool bash`.
- **`dispatch`** (v0.4 Phase 2) — spawns a subagent on a fresh
branch with the supplied goal and returns
`{"status":"in_progress","handle":"<sub-branch>"}` synchronously
(ARCH §2.5). Input is `{role, goal}`;
the role must resolve to `souls/<role>.md` and a `roles:` entry in
`providers.yaml` — both read from the calling branch's governing
config commit (§2.2). Reads the calling
conversation's repo + branch from the harness-set
`LERNIE_CONV_REPO` / `LERNIE_CONV_BRANCH` env vars (ARCH §3.3 env
bullet); spawns through `lernie dispatch <role>` (§3.4). The handle
it returns is the child's *address* — there is no polling tool to
pair with it. The substrate redesign (ARCH §2.5 "Dispatch returns
the child's address") dissolved the handle/`await` pair: the child's
result comes back as a **deposit into the parent's inbox** carrying
an epitaph (§2.6, §2.11), so `await`/`check` had nothing left to
observe and are gone. The return path — the result-message deposit
and the delivery-time work-product transfer — is built and live
(bl-4ce8, bl-9f53, bl-c33b, §2.6), and **children run full step
loops**: the dispatch's own front-door deposit finds the fresh child
quiescent and launches the ordinary driver, `lernie advance` (§6) —
there is no child-specific loop and no worker path — which steps the
child to a terminal event, deposits its epitaph result (final-response,
budget-exhausted, or stop) into the parent's inbox, and revives the
parent, which delivers the result at its next step boundary.
- **`message`** — deposits content into an *existing* agent's inbox
(ARCH §2.11). Input is `{agent, content}`; the recipient is addressed
by its agent id (its branch name / hyphenated descent). Unlike
`dispatch` it starts no branch and returns no address — it deposits
synchronously and returns `{"status":"deposited"}`. The sender is the
calling agent's id, taken from the harness-set `LERNIE_CONV_BRANCH`
(never model-supplied), so provenance cannot be forged. It goes
through the front door — `lernie message` (below) — like `dispatch`
goes through `lernie dispatch`, so it inherits the front door's
recipient guards: an id that is not a single path component, or one
with no `agents/*` ref, comes back as an `is_error` result naming the
decline instead of a silently lost message. **Shipped state:** the deposit lands
and the step-boundary drain delivers it (bl-1129) — the next driver to
step the branch moves the inbox file into `messages/` as a transcript
entry at its next boundary. A deposit into a *quiescent* agent is
self-delivering: the free-lease probe detach-spawns `lernie advance`
(§6, below), which acquires the lease, delivers the deposit, and steps
the branch.
- **`load_skill`** — copies a pooled skill's body into the calling
agent's worktree at `skills/<name>/`, where the next context assembly
composes it (ARCH §3.3 *Body-on-demand*, §5.2). Input is `{name}`; the
data-root pool + target worktree come from `LERNIE_HOME`/XDG and
`LERNIE_CONV_REPO` / `LERNIE_CONV_BRANCH`. Returns
`{"status":"loaded","path":"skills/<name>"}` on a fresh copy or
`already_loaded` when the worktree already holds it (the loaded copy is
the snapshot the branch is pinned to; `rm` and reload to refresh). An
unknown or non-single-component name is declined (`is_error`, naming
the available pool). **Shipped state:** the copy commits with the tool
result — a tool commit now stages the whole worktree (`git add -A`,
`commit_tool`), landing any tool's worktree side effects with its
result entry (ARCH §2.3).
## Messaging an existing agent directly
`lernie message <workspace> <agent> <content>` deposits a message into
`<agent>`'s inbox and, finding the recipient quiescent, launches a
driver to deliver it (ARCH §2.11, §3.4). The sender is read from
`LERNIE_CONV_BRANCH` — the calling agent's id when the `message` tool
re-enters the verb, else `user` for a bare invocation.
- **The recipient is guarded before anything is written.** The id must
be a single path component (ARCH §2.3) — `..`, a `/`, or an absolute
path is declined, never sanitized, because `Path::join` would honour
it and write outside the workspace — and an `agents/<id>` ref must
exist for it: a message is addressed to an *existing* agent (§2.11),
so a deposit no drain would ever come for is refused (`lernie
message: no agent "…" …`, exit 1) rather than left in an inbox
directory nothing will ever read. The id guard is the same rule at every verb taking an agent id
from outside — `message`, `advance`, `stop`, `dispatch`, `bundle` —
and literally the same code: one workspace-layout guard and one
existence guard, each carrying the calling verb's own clause for *why*
it needed an agent, so what differs between verbs is the cause, never
the phrasing or the remedy.
- The deposit is a create-only file at `<workspace>/inbox/<agent>/
<sender>-<NNN>.md` (temp-path + atomic rename), with `from:` /
`deposited_at:` frontmatter and the content as its body. `<NNN>` is
the sender's own sequence, derived as max-present-plus-one over its
existing files in that inbox.
- After depositing, the verb probes the **executor lock** (`flock` on
the inbox directory): the same lease the shipped `lernie prompt` step
loop holds for its whole run, releasing it on exit. A held lease means
a driver is already stepping the branch (it will deliver at its next
boundary); a free lease means the branch is quiescent.
- On a free lease the verb launches a driver — `lernie advance
<workspace> <agent>` (ARCH §6) — as a **detached spawn** (§2.11):
`setsid` (its own session and process group), stdio to null,
fire-and-forget. The driver outlives the `lernie message` process, so
messaging is scriptable: the verb returns as soon as the deposit and
spawn land, and delivery + stepping continue in the driver.
- **A failed branch is named, never refused.** If the quiescent
recipient's latest model call failed (its last `response.json` segment
terminated in an `error` — retries exhausted or a non-retryable error,
ARCH §2.10), the deposit and launch proceed unchanged — messaging is
exactly how such a branch is retried once the cause is fixed — but the
verb prints a stderr advisory naming the branch and pointing at
`steps/<agent>/` and `lernie scan`, so a silent death (ARCH §2.3, §8)
is distinguishable from ordinary idleness at the verb that touches
it. Exit
code and stdout are untouched.
## Driving a branch: `lernie advance`
`lernie advance <workspace> <agent>` is the §6 driver verb — the
process every launch seam spawns, and the same verb an operator runs by
hand. One invocation is one **hop**: guard the id (a single path
component, and an `agents/<id>` ref must exist — a name that is no
agent is refused with `no agent "…"` and exit 1 before any lease, so
an operator typo neither drives anything nor leaves an `inbox/<id>/`
behind), take the lease (adopt the
`LERNIE_LOCK_FD` fd published by a predecessor hop, else try-acquire
the executor lock — losing it is a clean no-op), deliver pending inbox
messages through the real drain (rematerializing a torn-down worktree
first), derive warrant from the transcript tail (ends user-side → a
model call is due; ends assistant-side without `tool_use`, or empty →
exit silently; assistant `tool_use` with uncommitted results → decline
loudly, the one non-replayable state), run one step, and hand off: a
step that emitted `tool_use` runs its tools and **exec's the successor
`lernie advance`** with the lock fd deliberately inherited (close-on-
exec cleared just before exec; the successor fstat-validates the fd
against the inbox directory and restores close-on-exec), while a
terminal event ends the chain through the §2.11 exit protocol. Because
the successor is `exec`'d in the same process, the pid, process group,
and flock lease all survive the hop — `lernie stop` lands on whichever
hop is current, and no rival driver can wedge between hops.
## The exit protocol and the operator scan
Normal operation needs zero scanning (ARCH §2.11): `lernie message`
deposits, probes the executor lock, and launches a driver if the agent
is quiescent; the executor drains its inbox at every step boundary. The
graceful-exit crack — a deposit landing after an executor's final drain
but before its lock release — is closed by the **exit protocol**
(§2.11, bl-5846): one terminal sequence, no agent kinds — deposit the
result message (a structural no-op for a parentless agent) → release
own lock → spawn a driver at own agent, fire-and-forget → probe-and-
launch at the parent the deposit just landed in → exit. Two
pins terminate the recursion: a driver that acquires and finds nothing
to deliver exits silently (no step, no epitaph, no further launch —
`dispatch::driver::drive` is that entry), and the launch is decided by
epitaph value — a final response launches; `stopped` and
`budget-exhausted` never do. The exit launch rides the same launcher
seam as the writer probe, so it is the same detached `lernie advance`
spawn (§6); the decision logic, ordering, driver entry, and the spawn
itself are live and tested.
The parent-side step is what makes **revival-on-deposit** real
(bl-4a6c): a child that returns to a quiescent — even torn-down —
parent starts that parent's driver itself, through the *same*
`probe_and_launch` the `lernie message` verb uses (one probe, no
second copy), so the parent rematerializes, delivers the result, and
steps with no `lernie scan` in the path. A parent whose lease is held
gets nothing launched: its running executor delivers at its next step
boundary. The epitaph decision governs this launch too, one level up:
a `stopped` child would otherwise wake its parent to react to — perhaps
re-dispatch around — the very branch the operator killed, and a
`budget-exhausted` child's ceiling is the whole tree's (§6), so the
woken parent would exhaust on its own next check and deposit again. In
both cases the result still lands in the inbox and waits for the next
explicit touch.
Crashes are accepted as a failure class (§2.11): everything is on disk,
so a hard death strands results and messages *late*, never lost, and
the next touch heals. That touch is a user reprompt — or the operator
verb **`lernie scan <workspace>`** (§2.11, §8, bl-d148 + bl-5846): one
workspace-wide pass, run by hand or by cron if you want a heartbeat,
never wired into any driver hot path or default schedule (the events it
compensates for happen at crash rate, not step rate). Two derived
actions, no watcher (an idle workspace stays unswept until the next
touch, by design):
- **Silent-death sweep.** Every agent branch with no live executor (the
§2.11 executor-lock probe) that either died mid-work — its latest
step's model call never settled complete: `response.json` closed
without a terminal `end` (killed/stopped, §2.9), *or* its final
segment terminated in an `error` (retries exhausted or a non-retryable
error, §2.10 — that segment closes with a clean `end`, so
absence-of-`end` alone would misread the branch as idle) — or, for a
child, never deposited a result message is a *silent death* (the §8
health count). Each one is **named** in the report
(`silent deaths: 1 (<agent-id>)`): a dead **root** gets no deposit —
it has no parent inbox — so its name here is how an operator learns
which branch went quiet, and `steps/<agent-id>/` is where to read
why. For each hard-crashed **child** in that set, the sweep deposits
a `died`-epitaph result message *on the child's behalf* (sender = the
child — the sweep is the scribe, not the author), so the parent is
revived rather than stalled. The "never deposited" test reads both the
parent's inbox (undelivered) and its transcript (delivered), so a prior
sweep's own deposit is seen on re-scan and never re-deposited —
idempotent by construction.
- **Inbox flush.** Every agent with pending inbox files and a free lock
gets a driver **launched** — never drained: the scanner moves no files
and commits nothing; only an agent's own lock-holding executor
delivers. An agent whose lock is held is left alone. The inbox listing
is intersected with the `agents/*` refs — the one registry of who
exists — so an inbox directory with no matching ref is reported
(`inboxes with no agent branch: N`) and left in place rather than
driven: a driver launched for a name with no branch is refused by the
existence guard (`lernie advance: no agent "…"`, exit 1) on this pass
and every pass after, writing nothing. The sweep's own deposits
are picked up by the flush that follows in the same pass.
**Shipped state.** The scan (silent-death sweep + inbox flush) ships
behind `lernie scan` and *only* there — driver startup (`lernie prompt`,
`lernie dispatch`, `lernie advance`) runs no workspace scan. The flush
and the exit launch reuse the same driver-launch seam as `lernie
message`, and the spawn is real: each seam decides *when* a driver is
needed and detach-spawns `lernie advance` (§6) for it. Children run full
step loops (bl-c33b), so a `died` child is a state a real run reaches; the
derivation is additionally exercised against constructed on-disk states,
since a hard crash is not reproducible on demand.
**Namespace note.** The candidate enumeration is the `agents/*` ref
namespace, exactly as ARCH §8 writes it (a root is `agents/<conv-id>`,
a child `agents/<parent>-<sub-id>`); config branches are excluded
structurally by the prefix — there is no `main` (§2.2).
## Dispatching subagents directly
`lernie dispatch <role> <repo> <branch> [--goal <text>]` is the §3.4
re-entry point every child dispatch uses. It is **writer-shaped, not an
executor** (ARCH §2.1): it forks the child branch, lands the dispatch
commit, and deposits the dispatch message through the same front door
every sender uses — the driver that deposit launches is the ordinary
`lernie advance` (§6). The role name is positional and the role set is
**open** (§4.3): a role is dispatchable iff the calling branch's
governing config commit lists it under `providers.yaml` `roles:` and
carries `souls/<role>.md`. The CLI enumerates no role names, so a
verifier, a critic, or a role you author needs no CLI change; validity
is checked *before* the fork, so a rejected role leaves no branch debris.
The **id guard runs first**, through the same two functions `message`,
`advance`, `stop` and `bundle` call: the workspace layout, then the
dispatching parent's `agents/<id>` ref. So all three refusals are the
product's, never git's:
```
lernie dispatch worker <no-such-ws> someagent --goal hi
→ <path> is not a workspace (no repo.git) — create one with `lernie new` (ARCH §2.2)
lernie dispatch worker <ws> nosuchparent --goal hi
→ no agent "nosuchparent" in this workspace — a child forks off an existing parent (ARCH §2.5); …
lernie dispatch verifier <ws> <agent> --goal hi
→ role "verifier" is not defined in the providers.yaml governing agent "<agent>" — defined roles: compactor, worker
```
The role refusal names the pool that *is* defined — the same "name the
pool" idiom `load_skill` and `lernie tool` decline with — and names the
control file the user knows rather than the config commit's sha.
- `lernie dispatch compactor <workspace> <conv-id>` forks a
compactor-souled child off that agent's tip — exactly what a due
compaction checkpoint does (§2.7), run by hand. The compactor is an
**ordinary child that makes a real model call** through `bz`; it is not
a stub, and it does not merge anything itself. Its goal is
procedure-generated, so passing `--goal` is rejected. Its toolset is
the deletion-only pair injected for the compactor role alone (never a
`providers.yaml` `tools:` list): `write_summary`, which writes the next
`summary/<NNN>.md` on the compactor's branch, and `mark_for_deletion`,
a staged `git rm` that can remove but never write content — so the
worst case is lost information, never corrupted information. Its
request *declares* more than that pair: a compactor inherits the
dispatching branch's transcript, so the model call also names whatever
tools that transcript used — otherwise the provider refuses a request
whose history mentions a tool it was not told about. Declaring is not
permitting: a compactor reaching for one of those inherited tools gets
an error tool result naming its own two, and nothing runs. The
**compaction merge** lands later and elsewhere: when the compactor's
result message is delivered, the dispatching agent's own executor
interprets its `compactor_return: compaction_merge` binding (§6) and
merges the compactor branch `--no-ff` — the one merge left in the
system (§2.6). A compactor that ends on any other epitaph lands no
merge; the branch simply continues uncompacted — enforced where the
binding is interpreted: the delivered result's **epitaph value** gates
`compaction_merge`, and a `died`/`stopped`/`budget-exhausted`
compactor return is delivered like an ordinary child's result instead,
so the parent sees the epitaph and nothing of the compactor's branch
crosses (§2.6, §2.7).
- `lernie dispatch worker <workspace> <parent-id> --goal <text>`
spawns a worker child off the parent's tip. The new id is
`<parent>-<sub-id>` (hyphenated descent, §2.2), its ref
`agents/<parent>-<sub-id>` (§2.3), its worktree
`agents/<parent>-<sub-id>/`; `goal.md` carries the supplied text and
`soul.md` is read from the parent's governing config commit
(`souls/worker.md`, §2.2), both committed as the dispatch commit
(§2.3 step 2). The child then **runs a full step loop** under the
`lernie advance` driver its dispatch deposit launched, and at its
terminal event deposits a result message — epitaph, terminal ref, and
the terminal response iff it spoke — into the parent's inbox, reviving
the parent if it had gone quiescent (§2.6, §2.11). The v0.4 "Phase 1
stops at the dispatch commit" worker path (`worker.rs`) is **deleted**,
not extended (bl-c33b).
## Providers
Every model call goes through **brazen** — one small, stateless binary
(`bz`) that adapts every provider and wire protocol behind a single pipe
contract (see [ARCH §4.4](docs/ARCHITECTURE.md#44-the-provider-adapter-brazen)):
```
stdin (canonical request, JSON) → bz → stdout (v=1 event stream, NDJSON, one terminal `end`)
```
The harness execs `bz --json --provider <row>` once per attempt, pipes a
typed `brazen::CanonicalRequest` on stdin, and appends bz's stdout
verbatim to the step's `response.json`. lernie links the `brazen` crate
(`brazen = "=0.0.4"`) for the canonical *types* only — the data plane
always crosses the subprocess boundary (§3.4). Two facts follow:
- **Retry is the harness's.** brazen never retries — one `bz` process,
one HTTP round-trip. On a retryable in-band `Error`
(`CanonicalError::retryable()`, the linked crate's single home for the
fact) the harness re-invokes `bz` up to the `workflow.yaml` attempt cap
(§2.10). Each attempt appends one segment to `response.json`; the last
is authoritative. Each attempt's `bz` stderr appends to the step's
`stderr.log` beside it — empty on an ordinary run, because brazen
speaks its failures in-band on stdout. A `bz` that dies *before* it can
(a malformed brazen config) leaves an empty stream that reads exactly
like a mid-stream kill, so the half-stream error quotes that capture's
tail; with a stop pending it stays quiet, because the stop check point
(§2.9) discards the outcome before anything is rendered.
- **Auth and endpoints are brazen's.** Provider *rows* (endpoint,
protocol, auth mode, model aliases) live in brazen's own config
(`~/.config/brazen/config.toml`; `bz --dump-config`, `bz --login`).
lernie references a row by name and never sees credential material
(§4.1). A load-time guard (`bz --version` == the linked crate version)
rejects a mismatched binary; `make install` installs the pin with
`cargo install brazen --version =0.0.4`.
### Adding a provider
- **A new provider on a supported protocol** is a brazen config row — no
code anywhere. Add the row (`bz` config), reference its name as a
model's `provider:` in `<config-root>/models.yaml`, and point a role
at that model in `<repo>/providers.yaml`.
- **A new wire protocol or auth mode** is a contribution to brazen.
- **An alternate adapter binary** that honors the same pipe contract
slots in via the optional `adapter:` path in `models.yaml` (§4.2); the
version guard is skipped for it and the in-band `MessageStart.v`
handshake governs compatibility instead.
## UI (v0.5)
The desktop frontend lives in its own repository, `yog`: an
egui/eframe window that renders a workspace and issues user actions via
`lernie <subcommand>`. It composes on lernie's public surfaces only —
the CLI and the on-disk workspace layout (ARCH §3.5, §7.1) — and takes
no Cargo dependency on this crate, so it builds, versions, and installs
independently (`make install` there drops `yog` next to
`lernie`). Keeping frontends out of this workspace is deliberate:
lernie ships as a composable component, and anything that composes it
(a GUI, a web view) lives outside it and meets it at those surfaces.
## Evaluation: archival and the task suite (§9)
**Archive a run.** A "run" is an agent subtree, not a whole workspace (§9.2).
`lernie bundle <workspace> <agent> <out-dir>` writes the subtree — the
`agents/<agent>` branch and its `agents/<agent>-*` hyphen-descendants (§2.3),
with all the ancestry those refs reach — plus the subtree's **governing
lineage**: every `config/*` ref whose history reaches it (§2.2). Both go into
one `git bundle`, and the matching `steps/<id>*` and `inbox/<id>*` diagnostic
slices are copied beside it. One bundle plus two slices is the whole run.
The config refs are not decoration. An agent's control files are read from its
*governing config commit*, which is derived — the nearest ancestor of the
branch reachable from a `config/*` ref (§2.2). Ancestry alone carries that
commit as an object but names no ref to take the merge-base against, so a
replay of the agent refs alone yields a workspace no verb can drive. Carrying
the refs (never a sidecar file — the refs are the single source) makes the
replayed repo derive its governing config by the same computation, over the
same candidate set, as the workspace it came from. "Every ref whose history
reaches it" is broader than "every ancestor": a **sibling** config lineage
that shares only a common root with the bundled subtree is still a
merge-base *candidate*, so it rides too — carrying a ref that turns out not
to be the nearest one is how the bundle stays a faithful copy of the
computation, not a leak.
```
lernie bundle /path/to/workspace <agent-id> /path/to/archive
```
**Replay a run.** `lernie replay <archive>` reconstructs a scratch workspace
under `LERNIE_HOME`'s data root at `replays/<primary-id>/` (the primary id is
the subtree's root agent), fetches every branch out of the bundle into a
fresh bare `repo.git`, materializes the primary's worktree under `agents/`,
restores the slices, and prints the scratch path. Point the ordinary frontend
at it — replay is not a mode (§2.3). Set `LERNIE_HOME` to an isolated
directory to keep the replay sandboxed; the harness root it points at still
supplies the machine-local pieces a config only *names* (`models.yaml` and the
brazen provider rows, §4.2/§4.4).
A replayed workspace is an ordinary workspace: `lernie prompt <scratch> "…"`
forks a fresh root off the config head that rode the bundle, and `lernie
message` / `lernie advance` drive the replayed agent on its own governing
config commit.
```
LERNIE_HOME=/tmp/replay lernie replay /path/to/archive
```
**Task suite.** The evaluation suite lives as data under `tests/suite/` — 50
tasks with machine-checkable `check` scripts, tagged by the seven §9.1 failure
categories (≥10 per category), format in `tests/suite/README.md`,
well-formedness enforced by `tests/suite.rs`.
**Run the suite.** The `agent-eval` runner (a separate crate, `crates/agent-eval`,
ARCH §9.3) executes an experiment against the suite N times per task and reports
pass@1 (with 95% Wilson intervals) and pass@5, overall and per category:
```
agent-eval --config baseline --suite tests/suite --runs 5 --agent lernie-eval-agent
```
`--config <name>` names an experiment — a `workflow.yaml` variant under
`experiments/<name>/` (a config diff, no code changes; see `experiments/README.md`).
`baseline` is the shipped default itself: its `workflow.yaml` is a symlink to
`template/workflow.yaml`, because an experiment is a diff against the default and
the baseline's diff is empty.
Per run the runner seeds a fresh isolated `LERNIE_HOME` and working directory,
runs the task `setup`, invokes the agent, then runs the task `check` — **exit 0
is the sole pass signal** (§9.1), so success is observable state, never the
agent's own claim. `--bundle-dir <dir>` archives failing runs for triage via
`lernie bundle` (§9.2). The runner is fully tested against a faked agent, so it
needs no live model to validate.
**The shipped driver is `lernie-eval-agent`** (`crates/lernie-eval-agent`,
workspace-internal like the runner; installed on `PATH` by `make install`).
`--agent <cmd>` stays required with no default: which driver runs the agent
under test is an experiment-defining input, so it is named explicitly. Per run
the shipped driver seeds the run's isolated `LERNIE_HOME` from the machine's
lernie config root (`models.yaml` plus the `template/` config-root override —
the wire is machine-local by design, §4.2/§9.2, and those two front doors are
how a machine points evaluation runs at its own provider rows), then drives
the harness exclusively through the front door, exec'ing `lernie` from `PATH`:
`lernie new`, `lernie config` (applying the experiment — below), and one
`lernie prompt` carrying the task prompt grounded in the shared working
directory. The contract any driver must honour, per run:
| Given | How |
|---|---|
| the task prompt | argv[1] |
| the isolated harness root for this run | `LERNIE_HOME` in the env |
| the experiment's `workflow.yaml` | `LERNIE_EXPERIMENT` in the env — an absolute path |
| where to report back | `LERNIE_EVAL_REPORT` in the env — a file path |
| the working directory | cwd (shared with the task's `setup` and `check`) |
`LERNIE_EXPERIMENT` is a hand-off, not a hook: **nothing in the harness reads
that variable.** The harness takes its `workflow.yaml` from the workspace's
config commit (§2.2), never from the environment, so *applying* the experiment
is the driver's job. The shipped driver does it through `lernie config`, with
`$EDITOR` set to copy the experiment over the authoring checkout's
`workflow.yaml` — the experiment lands as an ordinary config commit, exactly
the "config diff, no code changes" §9.3 promises (for `baseline` the diff is
empty and the authoring pass declines: the default is already in force).
`LERNIE_EVAL_REPORT` names a file the driver **may** write with exactly two
lines — the workspace path, then the agent id — which is what `lernie bundle`
needs to archive the run if it fails (§9.2). It is the driver's only channel
back to the runner. Writing nothing, or anything malformed, only makes a failing
run un-bundleable; it is never an error, and it never affects pass/fail, which
is the task `check` alone. The driver's own exit code is likewise ignored.
Failure to *spawn* the driver, by contrast, is a hard error naming the program.
## Contributing
The instructions below are for contributors building lernie from source.
Users installing a release don't need any of this — **[Install](#install)**
covers the three user-facing routes, only one of which involves a clone.
### Contributor setup
```
make install-hooks
```
Sets `core.hooksPath` to `.githooks`. Required on every fresh clone — git
does not track `.git/config`, so the hooks are not active until installed.
That arms both the [pre-commit gate](#pre-commit-hook) and the
[auto-push hook](#auto-push-hook).
The Rust toolchain is pinned in `rust-toolchain.toml` (channel `1.95.0`, with
`rustfmt`, `clippy`, and `llvm-tools-preview`). rustup reads it automatically
for every `cargo` command in the tree and installs the pinned toolchain on
first use — no manual `rustup` step. This is what keeps `fmt-check` and
`lint` from drifting between your machine, another agent's, and CI.
### Build targets
| Target | What it does |
|-----------------------|-------------------------------------------------------|
| `make build` | `cargo build` |
| `make release` | `cargo build --release` |
| `make test` | `cargo test`, with the pinned `bz` first on `PATH` (below) |
| `make test-install` | `cargo test --test install` — the install contract end-to-end, uninstrumented (it is `cfg_attr(tarpaulin, ignore)`, so `coverage` skips it); ~45s warm, and it re-installs `bz` at the `brazen` pin |
| `make coverage` | `cargo tarpaulin --fail-under 100` (llvm engine), same pinned `PATH` (below); hard-gated on tarpaulin **0.35.2** exactly (`TARPAULIN_PIN` in the `Makefile` — its one home; any other version aborts with the `cargo install cargo-tarpaulin --version 0.35.2 --locked` fix-it line) |
| `make lint` | `cargo clippy --all-targets -- -D warnings` |
| `make fmt` | `cargo fmt` |
| `make fmt-check` | `cargo fmt --check` |
| `make schemas` | Regenerate `schemas/*.json` from the Rust types |
| `make new-workspace DEST=<path>` | Create a workspace (bare repo.git + first config commit from `template/`) |
| `make eval CONFIG=<exp> SUITE=<dir> RUNS=<n> AGENT=<driver-cmd>` | Run the evaluation runner (ARCH §9.3): experiment × suite × N (see **Task suite** above). `AGENT` is required and has no default — the shipped driver is `lernie-eval-agent` (see "Run the suite"), and naming it is deliberate: the driver is an experiment-defining input |
| `make check` | `fmt-check` + `lint` + `coverage` + `test-install` |
| `make ci` | Alias for `check` |
| `make smoke` | Live-wire smoke test: one real `lernie prompt` against the shipped defaults (override with `SMOKE_PROVIDER`/`SMOKE_MODEL`); the default needs a `bz` anthropic credential and spends money; NOT part of `check` |
| `make install-hooks` | Point git at `.githooks/` |
| `make install-bz` | Install the provider adapter `bz` on your `PATH` at the version Cargo.toml pins (ARCH §4.4); a no-op when the `bz` there already matches. For *running* lernie — the tests feed themselves (below) |
| `make brazen-pin` | Print that pinned version and nothing else — CI keys its `bz` cache on it so no workflow file names a version |
| `make install` [`INSTALL_PREFIX=<p>` `LERNIE_HOME=<h>`] | Release-build; drop `lernie`/`agent-eval` into `$INSTALL_PREFIX/bin` (default: `~/.local/bin`); install the provider adapter `bz` via `make install-bz` at the version Cargo.toml pins (the ARCH §4.4 version pin — the number's one home); then invoke `lernie prime` to found the harness root — config root (default `~/.config/lernie`) with a default `models.yaml` and an empty `workflows/` templates dir, data root (default `~/.local/share/lernie`) with the `tools/`/`skills/` pools and the `workspaces/` tree — seed-if-absent (ARCH §2.2); `LERNIE_HOME` collapses both |
| `make uninstall` [`INSTALL_PREFIX=<p>` `LERNIE_HOME=<h>`] | Remove the installed binaries; leaves the harness homes (config + data roots) in place |
### The pinned adapter under test
The e2e tests exec the **real** `bz` (against a mock HTTP endpoint, not a
provider), and lernie's load-time version guard (ARCH §4.4) demands the
pinned version *exactly*. The pin's one home is the `brazen = "=<version>"`
line in `Cargo.toml`.
**The trap.** `bz` normally resolves from `PATH` — that is
`~/.cargo/bin/bz`, machine-global mutable state shared by every checkout and
every agent on the box. Anyone running `make install` rewrites that binary at
*their* tree's pin. If your tree pins a different version, your next test run
dies in five-plus e2e tests with
```
bz version "0.0.3" does not match the linked brazen crate "0.0.4"
```
which looks nothing like "someone else installed a binary" and everything
like a regression you just wrote.
**The cure.** `make test` and `make coverage` do not use the `PATH` `bz` at
all. They depend on `$XDG_CACHE_HOME/lernie/bz/<pin>/bin/bz` — installed
from crates.io on first use — and put that directory *first* on `PATH` for
the run, so the tests always exercise the pin **this** tree names, whatever
the machine's `bz` happens to be. The version comes from `BRAZEN_PIN` in the
`Makefile`, derived from `Cargo.toml`; the cache directory is named after
it, so bumping the pin is a cache miss and nothing else, and a stale entry is
never overwritten in place. Cost: one `cargo install` (~25s) per pin per
machine — sibling worktrees share the cache — and nothing at all when warm,
since it is an ordinary make file prerequisite.
Two consequences worth knowing:
- **Bare `cargo test` is still exposed.** It inherits your `PATH` and so
runs whatever `bz` is installed there. Use `make test`; if you must run
`cargo test` directly, `make install-bz` first to line the global binary up
with the tree's pin.
- **No test writes the global `bz`.** `make install` does — that is its job —
but the install test that runs it (`tests/install.rs`) points
`CARGO_INSTALL_ROOT` at a per-worktree root under `target/`, so the pinned
`bz` lands there and `~/.cargo/bin/bz` is never touched by a test run.
- **Runtime resolution is unchanged.** This is test determinism only —
`lernie` itself still resolves the adapter per ARCH §4.4 (the `models.yaml`
`adapter:` override, else a binding-injected target, else `bz` on `PATH`),
and `make install` still puts the pinned `bz` on your `PATH` for real use.
`the_makefile_derives_the_same_pin` (`src/prompt/tests/pin.rs`) keeps the two
readers of that one line honest: the Makefile's `BRAZEN_PIN` (which names the
cached binary) and the crate's `brazen_pin()` (which the version guard
compares against) must agree, or the tests would fail the guard against a
binary the Makefile itself installed.
### Workflow
All changes land on `main` via `bl` squash-merges. Direct commits to `main` are
rejected by the pre-commit hook, and every landing on `main` is pushed to
`origin` automatically (see **[Auto-push hook](#auto-push-hook)**).
```
bl prime --as <you>
bl claim <task-id> # creates a worktree; cd into it
# ...edit, test, commit...
bl close <task-id> -m "..." # squash-merges into main; run from the repo root
```
See `bl skill` for the full guide.
### What gets published
`cargo package` ships the crate, not the repo. `Cargo.toml`'s `exclude` keeps
out everything that serves this git checkout only — `docs/`, `tests/`,
`experiments/`, `scripts/`, `.github/`, `.githooks/`, `.balls/`, `Makefile`,
`tarpaulin.toml`, `release-plz.toml`, `AGENTS.md`, `CLAUDE.md`, and
`rust-toolchain.toml` (which would otherwise force a source builder onto this
repo's exact pinned toolchain). What remains is `src/`, `README.md`, `LICENSE`,
`Cargo.lock`, and the embedded asset trees `template/`, `schemas/`, `skills/`,
`install/models.yaml` — those four are `include_dir!`/`include_str!` inputs, so
excluding any of them is a build failure, not a smaller tarball. Verify a change
to the list with `cargo package --list` and then `cargo package`, which
compiles the extracted tarball.
`crates/agent-eval` is `publish = false`: it is workspace-internal and is not
part of the published crate at all.
### Pre-commit hook
`.githooks/pre-commit` enforces three rules on every commit:
1. **No direct commits to mainline.** `main` and `master` are rejected unless
the commit is the tail of a merge (`MERGE_MSG`/`SQUASH_MSG` present), which
is how `bl close` lands squash-merges.
2. **300-line cap on code files.** The cap is a repo *invariant*, not a
per-commit property, so the hook sweeps **every tracked code file in the
tree** (`git ls-files`), not just the staged set — a file that crosses the
cap in one commit and is untouched afterward is still caught. Docs (`*.md`,
`*.txt`), config (`*.toml`, `*.yaml`, `*.yml`, `*.json`, `*.lock`),
`Makefile`, `.gitignore`, `LICENSE`, and anything under `.githooks/` are
exempt.
3. **`make check`** on every commit that touches a Cargo project: `fmt-check`
(formatting), `lint` (`clippy -D warnings`), `coverage` (`cargo tarpaulin
--fail-under 100`), and `test-install` (`cargo test --test install`). The
hook invokes `make check` rather than re-listing the commands, so the close
gate is always exactly what `make check` is — the Makefile is the single
source. Formatting and lint drift therefore cannot land invisibly.
`test-install` is a separate step because the install test shells out to a
release build and `cargo install brazen`, which contend with tarpaulin's
`target/` lock; it is `cfg_attr(tarpaulin, ignore)`, so without its own
uninstrumented step the install contract — the first thing every user
touches — would never run at the gate at all. It costs ~45s warm and leaves
the machine-global `~/.cargo/bin/bz` alone: the test redirects `make
install`'s `cargo install brazen` into a per-worktree root under `target/`
with `CARGO_INSTALL_ROOT`, so a sibling worktree at another pin is never
rolled over. The
toolchain is pinned in `rust-toolchain.toml` and the tarpaulin version in
`tarpaulin.toml` (also
`.github/workflows/ci.yml`) so `fmt-check`, `lint`, and the coverage
denominator mean the same thing locally and on CI — newer tarpaulin
releases have silently dropped inline `#[cfg(test)] mod tests;` files from
the count, weakening the floor. `make coverage` aborts with an install
hint if the local tarpaulin version drifts.
A floor of exactly 100% only holds if every line's coverage is caused by the
code's own structure and not by winning a race, so **no line may be reachable
only while a clock has not yet run out.** With several agents measuring
coverage at once, whichever side of such a race the machine happens to pick
that minute decides the verdict, and the gate reports an uncovered line on a
diff that touched nothing. Two shapes to write around:
- *A retry budget is a count of attempts, never a wall-clock deadline.*
`PROBE_RETRIES` (`src/prompt/tests/exit_launch.rs`) is the one budget every
executor-lock probe shares; a deadline expires on load rather than on
evidence, so under load the give-up arm can be taken on the first pass and
the retry arm never runs at all.
- *A poll loop waits because its child is still running, not because a flag
has yet to land.* `wait_with_cascade` (`.../builtin/bash/mod.rs`) and
`wait_with_stop` (`.../tool/subprocess.rs`) therefore sleep between the
reap and the flag read: the interval is entered for as long as the child
lives, instead of only while a stop scheduled milliseconds out has not
arrived yet.
The same objection reaches past coverage to the **verdict**, and the
end-to-end tests answer it the same way: a poll waiting on a detached driver
is bounded by *consecutive probes that saw no change in the workspace tree*,
never by wall time (`src/e2e/poll.rs`, `docs/ARCHITECTURE.md` §9). A live
driver writes continuously and a wedged one writes nothing, so a loaded box
only makes the pass path slower — where a stopwatch would have turned a slow
success red, and (as bl-2bf0 found) hid a real defect behind a timeout that
read like machine load.
There is no `--no-verify` escape hatch in the workflow. If the hook rejects a
commit, fix the underlying issue rather than skipping.
### Auto-push hook
`.githooks/reference-transaction` pushes `main` to `origin` the moment local
`main` advances. Landing and publishing are one act: a `bl close` reaches
GitHub and the push triggers the Release-plz workflow, which contains CI as a
called job (`needs: ci`) and only publishes once it is green. `origin/main`
cannot silently fall months behind local `main` again.
**Why a `reference-transaction` hook and not `post-commit`.** Nothing lands on
this repo's `main` through `git commit`. `bl close` delivers by plumbing —
`git commit-tree`, then `git update-ref refs/heads/main` — which fires no
commit hook and no merge hook at all — every commit `bl` has landed on `main`
arrived that way. Git's `reference-transaction` hook is the one event every
landing path shares: the plumbing delivery, a `git merge --no-ff`, and a plain
commit alike all end in an update of `refs/heads/main`.
The hook acts only on the `committed` state of a transaction that moves
`refs/heads/main` to a new value, and only when an `origin` remote exists.
Everything else — side branches, `refs/remotes/*` (including the ones its own
push writes, so it cannot recurse), no-op rewrites like `git pack-refs`, and a
deletion of `main` — falls through untouched.
It cannot block or hang a landing. Git aborts a ref transaction when this hook
exits non-zero in the `prepared` state, so every path in it exits 0 — which is
also why it does not `set -e`. A push that fails prints one warning line on
stderr and nothing else, and `timeout 30` bounds an offline push rather than
stalling the commit behind a TCP timeout. Git runs `reference-transaction`
hooks from 2.28 onward; on anything older the file is simply never invoked and
`main` has to be pushed by hand.
`tests/hooks.rs` exercises the shipped hook file itself against a local bare
repository as `origin` — never the real remote — and covers all six behaviours
above: a commit on `main` pushes, a `commit-tree` + `update-ref` delivery
pushes, a `--no-ff` merge pushes, a side-branch commit pushes nothing, an
unreachable `origin` warns without failing the commit, and a repo with no
`origin` is silent.
### Commit-identity guard (opt-in, per machine)
`main`'s history carries exactly one human identity, `mudbungie
<mudbungie@gmail.com>`, and no `Co-Authored-By` trailers — it was normalized to
that on 2026-07-26. `tests/commit_hygiene.rs` keeps it that way, but only on a
machine that asks for it: the test arms itself on the presence of
`$XDG_CONFIG_HOME/lernie/enforce-commit-identity` (default
`~/.config/lernie/enforce-commit-identity`), an empty marker file **outside** the
repo. Absent — the default in public CI and in every clone — the test returns
without asserting anything.
Armed, it walks all of `refs/heads/main` and fails on any commit whose author or
committer is neither `mudbungie <mudbungie@gmail.com>` nor
`github-actions[bot]` (the bot stays allowed: release-plz authors the release
commit as it), on any `Co-Authored-By` trailer, and on any mention of a
throwaway or personal address in an identity or a message. The policy lives in
the marker, not in the code: `rm` it and the guard is off, with no code edit and
no flag. Create it with `touch ~/.config/lernie/enforce-commit-identity`.
## License
MIT. See [`LICENSE`](LICENSE).