lernie
A git-backed agent harness. Design spec: docs/ARCHITECTURE.md.
Principles catalog: docs/PRINCIPLES.md.
Vocabulary reference: docs/TAXONOMY.md.
Promise suite (the user stories 0.0.1 is evaluated against): docs/USER_STORIES.md.
CI runs make ci (fmt-check + lint + coverage with the 100% gate + test-install) on every push and pull request to main. The e2e tests exec the real provider adapter bz, which the test targets install themselves at the pinned version (see The pinned adapter under test). The Rust toolchain is pinned in rust-toolchain.toml — CI, the pre-commit gate, and every contributor build under the same rustc/rustfmt/clippy. That pin binds this git checkout only; it is excluded from the published crate, whose supported floor is the declared rust-version = "1.88" (the crate's let chains, not edition 2024's 1.85).
One command surface, two bindings
lernie is defined once as a command surface — the set of verbs, their arguments, and their products (ARCH §3.4). It is consumable two ways, and both are the same control plane:
- Exec binding — run the
lerniebinary:exec("lernie", args)with env-var auth. This is what the CLI and every frontend use. - Linked binding — depend on the
lerniecrate and drive the same verb entries in-process. The crate's entire public API islernie::cmd(theCli/Commandclap surface, onerunentry per verb, theFx/Outcome/Errorbinding seam, and thepreludebinding preludes). The linked binding is pin-exact 0.x only — no semver stability, the posture brazen takes toward lernie.
Parity between the two is enforced mechanically, not by convention. tests/command_surface_parity/ asserts the bijection at three depths: it pairs each verb's Command variant with its module's entry as function values, so the compiler — not an assertion — proves the two share one argument type and one product type; it walks the crate's whole module graph (via syn) and asserts that every externally reachable declaration (item, field, enum variant, method, derive, trait impl) is exactly a verb's entry, its arguments, its products, or the binding preludes, with every src/**/*.rs proven reachable so nothing can hide in a file the walk never opened; and it asserts, per verb, that the CLI's introspected argument set (via clap) is exactly that verb's public Args fields — same names, same arity, same named-vs-positional form. It rides make check (hence the pre-commit hook and GitHub Actions), so a divergence between the linked surface and the CLI fails the build.
Quickstart
cargo install lernie --locked # or: make install, from a clone
cargo install brazen --version =0.0.4 --locked # the provider adapter, always needed
lernie new ~/work/chat # create a workspace (bare repo.git + config/default)
ANTHROPIC_API_KEY=... lernie prompt ~/work/chat 'hello'
Three install routes, not one — see Install for what each
lays down. Every route needs bz; only make install installs it.
Install
There are three routes, and they do not lay down the same things. All
three need a second binary — the provider adapter bz — which only the
Makefile route installs for you.
cargo install lernie |
release tarball | make install |
|
|---|---|---|---|
| binaries | lernie |
lernie |
lernie, agent-eval, lernie-eval-agent |
installs bz |
no | no | yes, at the pin |
runs lernie prime |
no | no | yes |
| lands where | cargo's bin dir | wherever you unpack it | $INSTALL_PREFIX/bin |
From crates.io
cargo install lernie --locked
cargo install brazen --version =0.0.4 --locked # the pinned provider adapter
lernie prime # found the harness root
You get the lernie binary alone, in cargo's bin directory
(~/.cargo/bin unless --root/CARGO_INSTALL_ROOT says otherwise) —
no agent-eval, no lernie-eval-agent, no bz, and nothing runs after
the build. The lernie prime line is optional but explicit: prime
founds the harness root (below), and lernie new founds it too on its
way to creating a workspace, so a user who skips prime is not stranded
— only uninformed about where their state went.
From a GitHub release
Each v* release carries lernie-x86_64-unknown-linux-gnu.tar.gz: the
lernie binary, this README, and the license. Unpack it, put lernie
on your PATH, then run the cargo install brazen and lernie prime
lines above — the tarball ships no adapter and runs nothing.
From a clone, with make
make install # default: ~/.local/bin, XDG homes
make install INSTALL_PREFIX=/usr/local # binaries -> /usr/local/bin/
make install LERNIE_HOME=/opt/lernie # collapse both homes -> /opt/lernie/
make install runs a release build and then:
- Installs
lernie,agent-eval, andlernie-eval-agentinto$INSTALL_PREFIX/binwithinstall -m 0755(atomic overwrite, no symlinks). Make sure that directory is on yourPATH. - Installs the provider adapter — brazen's
bz— withcargo install brazen --version =<pin> --locked, where the pin is thebrazen = "=<pin>"dependency inCargo.toml— its one home; the Makefile and the load-time guard both derive from that line. One binary serves every provider (ARCH §4.4); the harness resolvesbzonPATH, and a load-time guard rejects anybzwhose version differs from the pin. - Founds the harness root by invoking
lernie prime— the single verb that seeds the installation substrate (ARCH §2.2), so the Makefile no longer duplicates the seeding.primeresolves the roots (XDG split, collapsed byLERNIE_HOME) and lays down the defaultmodels.yamlunder the config root (model capabilities + context windows — no endpoints or auth, which are brazen's), thetools/andskills/pools and theworkspaces/tree under the data root, and the emptyworkflows/templates dir. It is seed-if-absent throughout: a second run changes nothing, and a hand-editedmodels.yaml(or any operator-added pool entry) survives a re-install. The shipped assets are embedded in the binary, soprimeneeds no source tree —LERNIE_HOME=<dir> lernie primeseeds any fresh home. There is no frozen profile pool: the config a workspace runs under is its ownconfig/defaultcommit, authored fromtemplate/atlernie new(fork is the freeze, ARCH §2.2). - Smoke-tests the freshly installed binaries with
lernie --versionand a throwawaylernie new. Failure aborts the install with a non-zero exit.
Its closing banner prints what the other two routes leave you to find
out: the install prefix, both harness roots, and the bz commands
below.
The adapter is a second binary
Nothing prompts without bz. It is brazen's one stateless binary for
every provider (ARCH §4.4), it is pinned exactly, and lernie refuses
a bz at any other version rather than downgrading silently:
cargo install brazen --version =0.0.4 --locked
The pin is not folklore you have to read this file for — the installed
binary carries it: lernie --version prints the linked pin beside its
own version, lernie <version> (brazen 0.0.4). Its one home is the
brazen = "=<pin>" line in Cargo.toml;
the Makefile's BRAZEN_PIN, the load-time guard, lernie --version,
and every pin printed in this file all derive from that line (a test
holds them equal). With no bz at all, the first verb that drives a
model call says so and hands you the command above.
Provider endpoints, auth, and wire dialects live entirely in brazen's
own config (~/.config/brazen/config.toml; inspect with
bz --dump-config, authenticate with bz --login --provider <id>).
lernie references a provider row by name and never sees credential
material (ARCH §4.1).
Where the state goes
The harness root is the installation-global substrate (ARCH §2.2), split
by XDG lifetime: $XDG_CONFIG_HOME/lernie (hand-edited declarations —
models.yaml, workflows/) and $XDG_DATA_HOME/lernie (machine-
populated pools and the workspaces/ tree). LERNIE_HOME=<dir>
collapses both to one directory, at install time and at runtime alike.
lernie prime founds it, seed-if-absent throughout, so running it again
— or after an upgrade — never clobbers a hand edit. Only make install
runs prime for you; on the other two routes it is your first command,
or lernie new's side effect.
make uninstall removes the three installed binaries; bz
(installed via cargo) is removed with cargo uninstall brazen. The
harness homes (the config and data roots, holding config and
workspaces) stay put — clean them up manually if you want a true
uninstall.
Configuration schemas
JSON Schemas for the harness-root and config-commit control files (per
docs/ARCHITECTURE.md §2.2, §4.1) are generated
from the Rust types under src/config/. make schemas writes them to
schemas/ for editor integration and external validators. Generation is a
golden test (config::schemas::write_to vs the checked-in schemas/):
make schemas runs it with UPDATE_SCHEMAS=1 to rewrite the directory,
and the same test under make check fails if schemas/ ever drifts from
the source types — so the tree is always current, with no separate binary
to run.
| File | Backed by Rust type | Config-commit / on-disk file |
|---|---|---|
schemas/version.json |
config::version::Version |
version (config commit) |
schemas/manifest.json |
config::manifest::Manifest |
manifest.yaml (config commit) |
schemas/workflow.json |
config::workflow::Workflow |
workflow.yaml (config commit) |
schemas/providers.json |
config::per_repo_providers::PerRepoProviders |
providers.yaml (config commit, roles:) |
schemas/models.json |
config::models::Models |
<config-root>/models.yaml |
Layout: harness root and workspaces
The harness root is installation-global state, split by XDG lifetime
into two homes (ARCH §2.2). LERNIE_HOME, if set and non-empty,
collapses both to that one directory (test isolation, alternate
installs). Three distinct on-disk locations:
- Config root — hand-edited declarations,
$XDG_CONFIG_HOME/lernie(default~/.config/lernie). Holds the globalmodels.yaml(model capabilities + context windows, plus an optionaladapter:binary override — §4.2) and theworkflows/templates. Provider endpoints and auth live in brazen's config, not here (§4.1). - Data root — machine-populated pools,
$XDG_DATA_HOME/lernie(default~/.local/share/lernie). Holds thetools/andskills/pools plus theworkspaces/tree. Shared across every workspace. - Workspace — one git repository per workspace, at
<data-root>/workspaces/<workspace>/(ARCH §2.2): a barerepo.githolding config branches (config/<name>) and agent refs (agents/<agent-id>) — nomain. The control files (providers.yamlroles:only — §4.3,manifest.yaml,workflow.yaml,version,souls/) live in the config commit, read from each agent's governing config commit (git merge-baseagainst theconfig/*heads — derived from ancestry, never stored). Agent worktrees are siblings underagents/<agent-id>/;steps/andinbox/sit at the workspace root, outside every worktree. Workspace repositories are never pushed to a remote.
lernie new creates a workspace and authors its first config commit
— an orphan root on config/default — from template/,
the versioned skeleton embedded into the lernie binary at build time:
lernie new # auto-id under <data-root>/workspaces/
lernie new /path/to/my-workspace
Or via the Makefile wrapper:
make new-workspace DEST=/path/to/my-workspace
The binary founds the harness root first — it runs the same
seed-if-absent routine lernie prime is (ARCH §2.2), so a data root
nobody primed gains the tools/ and skills/ pools before they are
read, and a primed install is untouched (nothing is clobbered, no flag
is involved). That is what keeps the next step honest: the pools are an
input to the config commit, so an unprimed root would otherwise author a
commit with an empty descriptions/** and hand every agent forked off
it an empty toolset. It then runs
git init --bare -b config/default <dest>/repo.git,
materializes a transient authoring checkout, extracts the template's
control files into it, snapshots the data-root pools into
descriptions/{tools,skills}/ (ARCH §3.3 descriptions-always), commits
(config: init [config/default]), and tears the checkout down. The
workspace is left with exactly one ref — the config commit every fresh
root agent forks off (fork is the freeze, §2.2). The destination must
either not exist or be an empty directory. With no path argument, the
destination is <data-root>/workspaces/<auto-id>/; the created path is
printed on stdout. goal.md and soul.md are intentionally not in the
template — they are written per-branch at dispatch time (ARCH §2.3,
§2.8), which also removes the control files from the agent's tree
(§2.2: control is read from the config commit; worktrees hold only
context).
Pre-v1 clean break (ARCH §10): the retired per-conversation layout
(a root/ worktree with loose control files) is refused with an
actionable error, not migrated — create a fresh workspace with
lernie new.
First-run smoke test (required). lernie new authors the default
providers.yaml with a concrete model id, but validates it against
nothing — id validity is brazen's fact, and lernie runs no model-list
reconciliation (ARCH §4.2, the settled stance). A wrong id surfaces only
at the first live model call. The required next step after creating a
workspace is therefore a live lernie prompt (see the quick start
above): it is the cheapest — and, by that stance, only — check that the
authored id actually resolves on the wire.
make smoke automates exactly this:
make smoke # scaffold a throwaway workspace + one live 'lernie prompt'
It founds a throwaway harness root with lernie prime — from the assets
embedded in the binary, the same front door make install uses, so
the shipped install path is exercised too — scaffolds a workspace with
lernie new, then runs one live lernie prompt against the shipped
defaults — worker role, provider anthropic, model claude-sonnet-5 —
through the real bz data plane.
The verdict is read from observable state, never the agent's own
claim: the lernie prompt exit code is 0, the agent ref
(agents/<id>) carries a committed transcript entry, and the off-worktree
step record (steps/<id>/001/) holds a response with no wire error and
real assistant text. That last pair is the point: an auth-failed run
still creates the branch and a step record whose response terminates in a
clean end — the failure rides an error event ahead of it — so
branch-exists and step-exists alone would pass a broken wire. make smoke
requires exit 0 and no error event and an assistant
content_delta.
By default make smoke runs the shipped default — provider
anthropic, model claude-sonnet-5 — which needs a configured bz
credential for the anthropic provider (bz --login --provider anthropic, or set ANTHROPIC_API_KEY / BRAZEN_API_KEY) and spends real
money. To run the same live check against any other bz provider row,
set both SMOKE_PROVIDER and SMOKE_MODEL (both-or-neither — one
alone is a usage error; unset leaves the shipped default byte-for-byte):
make smoke SMOKE_PROVIDER=local SMOKE_MODEL=<a-pulled-ollama-model>
make smoke SMOKE_PROVIDER=codex SMOKE_MODEL=gpt-5.4
The override is laid into the throwaway config root through the same front
doors a real install uses — a providers.yaml override under
<config-root>/template/ (the config-root override) plus a models.yaml
placed in the config root before lernie prime (its seed-if-absent
contract, ARCH §4.2) — so there is no new lernie flag or verb. Local
ollama (bz's local provider row) needs no credential, only a model
that is actually pulled and served; the credential note above applies to
the anthropic default alone.
What SMOKE_PROVIDER=local does and does not prove. bz's local
row (protocol ollama_chat) rejects a canonical tool_result block:
the second step of any tool-using run comes back as
{"type":"error","kind":"parse_input","message":"user accepts only text content"}. So the local recipe validates the tool-free path only —
one model call, assistant text, a committed transcript entry. It cannot
exercise a tool step, a compactor (whose whole toolset is
write_summary/mark_for_deletion), or any multi-step loop that runs a
tool. This is a brazen-side gap in that provider row, not a lernie one,
and is filed there as brazen bl-fba7; to smoke a tool-using path,
point SMOKE_PROVIDER/SMOKE_MODEL at a row whose protocol carries
tool results.
make smoke is deliberately not part of make check or the close
gate: make check mocks the wire (httpmock Anthropic SSE), so it can
never catch a shipped default that fails on the real provider — which is
exactly how the fake id claude-sonnet-4-7 once shipped unnoticed. It
runs only on demand.
Authoring config commits
lernie new authors a workspace's first config commit. Every later
one — the general harness-assisted user act of ARCH §2.2 — is
lernie config:
lernie config <workspace> # advance config/default
lernie config <workspace> <name> # advance config/<name>
lernie config <workspace> <name> --from <src> # fork config/<name> off config/<src>
lernie config <workspace> <name> --orphan # fresh orphan lineage
The verb materializes a transient checkout of the target config lineage,
refreshes the descriptions/** snapshot from the data-root pools (ARCH
§3.3), opens the checkout in $EDITOR (falling back to vi) so you edit
the control files (workflow.yaml, providers.yaml, manifest.yaml,
souls/, version), commits, and tears the checkout down. <name>
defaults to default. --from and --orphan are mutually exclusive and
only apply when creating a new branch. A --from <src> naming a lineage
the workspace does not have is resolved before the checkout is
materialized, and declined by name:
lernie config: no config lineage "nosuch" in this workspace — existing lineages: default, strict
Declining is fine, and leaves nothing behind. Save no change and the
pass is declined: there is nothing to commit, so no commit is authored,
the branch does not move, and a --from / --orphan branch the pass
would have created is not left behind. That is a success — lernie config exits 0 and prints the one line
config/default unchanged: the edit changed nothing, so no config commit was authored
so empty stdout means a commit landed. The transient checkout is torn
down on every exit path (a decline, a git decline, an editor that fails),
so the next lernie config always runs. Only a hard kill mid-pass can
leave the checkout behind, and the next pass clears it before starting
(ARCH §2.11 "the next touch heals") — at the cost of the killed pass's
unsaved edit, which was never committed.
This is the only act that advances a config branch (ARCH §2.3); agents forked before it keep their governing config, and agents forked after it govern under the new head (fork is the freeze, §2.2).
Sending a prompt
lernie prompt /path/to/my-conversation 'hello'
lernie prompt is the root-agent path (ARCH §2.3, §2.6, §2.7,
§2.8, §2.10). Each invocation spawns its own agents/<conv-id> branch
off the default config branch's head (§2.2–§2.3 — there is no main),
drives each step's model call through brazen's bz (§4.4), and steps
until a terminal event. There is no terminal compaction stage (§2.7):
compaction runs only at the checkpoints workflow.yaml declares, and a
branch with no configured trigger never compacts. Merge-back is gone
(§2.6): the root branch persists on its own ref (§2.4), and a child
returns by depositing a result message into its parent's inbox (§2.6):
- Resolve the harness root (
LERNIE_HOME, else XDG homes, ARCH §2.2) and guard the workspace layout (a non-workspace, or the retired per-conversation layout, is refused — §2.2, §10). Load<config-root>/models.yaml(capabilities + context windows + optionaladapter:override — §4.2) and, from the config commit's tree (git show <config-commit>:providers.yaml, §2.2),providers.yaml(roles:block — §4.3); cross-validateroles.worker.{provider,model}againstmodels.yaml. For a fresh root the config commit isconfig/default's head — the very commit the new agent forks off;lernie advancederives an existing agent's governing config commit from ancestry instead (git merge-baseagainst theconfig/*heads — never stored). - Run the load-time version guard:
bz --versionmust equal the linked brazen crate version (§4.4). Under anadapter:override the guard is skipped and the in-bandMessageStart.vhandshake governs. Read the worker soul from the config commit'ssouls/worker.md(§2.2, §4.3). - Spawn branch
agents/<conv-id>(§2.3 — the id is the bare hyphenated descent; theagents/prefix is the ref namespace) offconfig/defaultand allocate a worktree at<workspace>/agents/<conv-id>/(§2.2). Write the branch goal togoal.mdand the role soul tosoul.md, remove the config commit's control files from the tree (§2.2 — the worktree holds only context), and commit — that commit's tree is step 1's read state (§2.10). - Build a typed
brazen::CanonicalRequest(linked crate — the fail-openextramap stays unreachable), mirror it to<workspace>/steps/<conv-id>/001/request.json(a diagnostic artifact, outside every worktree, never read at runtime, §2.3). - Model call, harness-owned retry loop (§2.10, §4.4). Exec
bz --json --provider <row>once per attempt, canonical request on stdin, appending each attempt's stdout verbatim to<workspace>/steps/<conv-id>/<NNN>/response.jsonas brazenv=1NDJSON — one self-delimiting segment per attempt, each ending in a terminalend. On a retryable in-bandError(CanonicalError::retryable(), never re-derived) the harness re-invokesbzwith the identical request, up to theworkflow.yamlattempt cap with exponential backoff — floored by the failed attempt'sRetry-Afterpacing hint (CanonicalError::retry_after_seconds) when it carries one, so the config schedule governs and the provider's hint can only lengthen it (§4.4). brazen never retries; auth and endpoints are entirely its own. Theresponse.jsonfd is held open across every attempt and backoff sleep — its close is the §3.5 IN_CLOSE_WRITE completion signal. As the events stream, the harness tracks only their framing — the terminalend, an in-bandError, the handshakev— for retry/classification;meta.jsoncarries{commit, started_at, ended_at}. The events' content streams into the transcript writer's (§2.3) staging file<workspace>/steps/<conv-id>/<NNN>/staging.json, appending each content block as it completes; segment authority (§4.4) truncates it on anErrorattempt and the settlingFinishseals it — one stream, two sinks (diagnosticresponse.json+ transcript), never read back. When the model call completes, the sealed file is renamed into the worktree asmessages/NNN-<model-id>.json— its origin token is the model that authored it (§2.3), the body a JSON array of canonicalContentblocks — and committed.NNNis the branch's transcript counter, max-present-plus-one from themessages/listing, evaluated at commit time. The initial user message now enters through the front door like any other (§2.11): the executor deposits it into the agent's own inbox, and the step-boundary drain delivers it as the first transcript entrymessages/NNN-user.md(bl-1129) — no bespoke initial-message path beside the drain. - Step loop (§2.5). At each step boundary the executor first
drains the inbox (bl-1129, §2.11): after committing any
renamed-but-uncommitted stray a prior death left in
messages/, it moves each pendinginbox/<agent-id>/<sender>-<NNN>.mdinto the worktree asmessages/<counterNNN>-<sender>.md(a literalrename(2)— one home at every instant) and commits the move, in a deterministic(mtime, filename)order, ahead of the read-state capture so a delivered message is part of the commit the model call assembles from. Each step then re-assembles its model-facing history from the read-state commit's tree —readdirofmessages/, sorted by the filename'sNNNprefix, each entry composed by its origin token (NNN-<sender>.md→ user text,NNN-<model-id>.json→ the assistant message — any.jsontoken but the reservedtool,NNN-tool.json→tool_resultin the following user message), with consecutive same-side entries grouped into one alternating wire message. There is no in-memory history and no git-log walk; running, retry, and replay are one code path against one input, the commit's tree (§2.3, §5). If the settled model-output entry carries anytool_useblock, run every one through the tool executor — the per-call records land under<workspace>/steps/<conv-id>/<NNN>/tools/<tool-id>/(out of every worktree, §3.3; written but never read at runtime), and as each tool resolves the transcript writer commitsmessages/NNN-tool.json(its canonicaltool_resultblock) — then loop into step<NNN+1>. A step with notool_useblock is terminal. Step ≥2 has no dispatch commit, but each step's transcript entries (assistant output, tool results) do advance the branch tip, which is that step's read state (§2.10).tool_use/tool_resultpairing holds by construction: a tool result commits immediately after its emitting step's model-output entry, so it always lands in the immediately following user message. Closing each tool step, the executor reads the compaction checkpoint clock (§2.7, §6) —compaction.intermediate.triggerinworkflow.yaml:every_n_commits,every_t_seconds, or the agent-electedon_flush, all derived from git (commits and elapsed seconds since the last compaction merge, or the branch root when none has landed — never a stored counter). When it is due, theworker_flush: dispatch(compactor)binding forks a compactor off the branch tip — the checkpoint commitC— and the branch keeps stepping straight through it; no quiescence is imposed. Omit thecompaction:block and the branch never compacts. - Terminal return (§2.6, §2.3 step 5). Every terminal event —
normal completion (
final-response), budget exhaustion (budget-exhausted, §6), and stop (stopped, §2.9 — the executor's SIGTERM handler deposits on its way out) — deposits a result message into the parent's inbox: an ordinary deposit whose frontmatter addsepitaph:andterminal_ref:(the branch tip) and whose body is the terminal response iff the agent spoke. For a root this is a structural no-op — a root has no parent inbox; its response answers the user (§2.4). The deposit is executor-side, never a model tool call ("Return is not a verb"). At delivery, a message carryingterminal_ref:applies the fork-point→terminal work-product transfer as one commit before its delivery commit, filtered to work products; a diff that fails to apply is declined atrefs/lernie/conflicted/<agent-id>(§2.6). - Exit protocol (§2.11). With the terminal deposit landed, the
executor runs the branch's terminal
workflow.yamlbindings (branch_stopped→mark_abandoned/notify_ui, §6), releases the executor lock, and only then spawns a driver at its own agent and — the deposit's own probe-and-launch — at the parent the deposit just revived. Both launches are fire-and-forget and both are decided by epitaph value: a final response launches,stoppedandbudget-exhaustednever do. No terminal compactor is dispatched (§2.7): the v0.3 terminal-compaction stage is deleted, along with theDispatcherre-entry that existed only to run it. Compaction is a checkpoint event (step 6), never an exit stage. Merge-back is gone (§2.6): the root branch persists on its own ref (§2.4); nothing merges back, and the agent's worktree is not torn down (quiescence, not teardown, §2.3 step 6). - Print the agent id (the bare conv-id) on stdout.
After lernie prompt returns, inspect the agent against the bare
workspace repository:
cd /path/to/my-workspace
git -C repo.git log --oneline --decorate agents/<conv-id> -4
git -C repo.git ls-tree --name-only agents/<conv-id> messages/
git -C repo.git show "agents/<conv-id>:messages/002-<model-id>.json"
ls steps/<conv-id>/
The log is the dispatch commit followed by one transcript NNN: commit
per entry, its subject naming that entry's origin token — user, a
sender's agent id, tool, or the authoring model's id:
f265de7 (agents/…) transcript 002: qwen3.5:9b […]
7ae527e transcript 001: user […]
f643a50 step 001: dispatch […]
6f4bd05 (config/default) config: init [config/default]
ls-tree lists the transcript itself (messages/001-user.md,
messages/002-<model-id>.json, …) and show prints one entry — a
model-output entry is a JSON array of canonical Content blocks, e.g.
[{"type":"text","text":"pong"}]. ls steps/<conv-id>/ lists the
off-worktree step records, one numbered directory per step, each holding
request.json, response.json, and meta.json. There is no merge
commit and no summary/ on a branch that never reached a compaction
checkpoint (step 6) — those appear only once a compactor has returned.
The root branch persists unmerged by design (§2.4), so the health metric is no longer branch count but silent deaths and undelivered returns (ARCH §8) — read straight from git refs, the executor lock, and inbox listings, with no sidecar file.
Stopping a conversation
lernie stop /path/to/my-conversation <conv-id> [--stop-children]
Sends SIGTERM to the process group of the one executor driving
<conv-id>, with a 5-second flush deadline before SIGKILL. This is the
same cascade pattern adapter (§4.4) and tool (§3.3) cancellation use,
applied to the harness itself
(ARCH §2.9). The group signal
reaches that executor's own bz and tool subprocesses — its limbs — and
stops at the agent boundary: a dispatched child harness has taken its
own process group, so a bare stop does not fell it. A running child
outlives the stopped parent and revives it later by depositing its result
(§2.11) — stopping a parent strands nothing.
--stop-children opts into the agent→agent cascade: it walks the id
namespace — the descendants of <conv-id> are exactly the inbox
directories prefixed <conv-id>- (§2.3), one prefix scan reaching every
depth — and folds each descendant executor's group into the same sweep.
The pid is discovered by scanning /proc/<pid>/fd/* for the process
holding the agent's inbox-directory lock fd open — the executor lock
(§2.11), held for the whole step loop, so a stop lands even during tool
execution when no response.json is open — no sidecar pid file. Linux
only.
The pgid that scan produces is vetted before anything is signalled,
because a pid is discovered before its group has settled: between a
driver's fork and the setpgid/setsid it runs at startup, /proc
still reports the group it inherited from its spawner — your shell job.
So a pgid is trusted only once it equals the holder's own pid (a group
leader's does, and every driver becomes one), re-read a bounded number of
times while it does not, and refused rather than signalled if it never
settles; a stop that signals nothing is re-runnable, one that signals
your shell is not. lernie stop additionally refuses any group it is
itself standing in (§2.9).
The group signal reaches every member independently: bz installs no
handler and dies at once (leaving the missing-end signature, §4.4),
while the executor catches its own copy — SIGTERM is catchable — and,
instead of dying on the spot, deposits its branch's stopped result on
its way out (§2.9 step 3, executor-side, "Return is not a verb") and then
exits cleanly. Catching shields nobody: the kernel already delivered to
bz and the tools. For a root the deposit is a no-op (no parent inbox);
the observable is the clean exit.
Because the model call is where the wall time goes, that is where a stop
usually lands — so the clean exit is the ordinary case, not the rare one.
The flag classifies, not the error's shape: a kill lands wherever the
adapter was, leaving a half-stream, a torn JSON line, or a provider error
depending on the instant, and with a stop pending each is read as the
stop. With no stop pending the same faults still propagate non-zero, so a
genuinely dying adapter is never hidden. The retry loop respects the flag
too: a stop is never followed by another bz invocation.
Behavior:
- Idempotent. A branch with no live writer (already stopped, or the harness exited cleanly) returns success without sending any signal.
- Errors when the agent branch (
agents/<conv-id>) doesn't exist. Surfaces as a non-zero exit with alernie stop:prefix on stderr. (The old "already merged" refusal died withmain: nothing merges, so there is no merged state to refuse — an already-terminal branch is simply the idempotent no-holder case above.) - No on-disk cancel marker. The §2.9 signature of a stopped branch
is the latest step's
response.jsonclosed without a terminalendevent — produced bybzdying mid-stream on its own SIGTERM (§4.4); the executor'sstoppeddeposit is an independent write to the inbox tree and never touches that signature.
The frontend's stop button (per ARCH §3.5) exec's this exact subcommand; there is no second control surface.
Built-in tools (v0.3, +v0.4 Phase 2 dispatch)
The agent can call built-in tools that ship inside the lernie
binary as lernie tool <name> subcommands (ARCH §3.3 / §12). The tool
executor's resolution order — <data-root>/tools/lernie-tool-<name>
→ PATH → lernie tool <name> — falls through to this in-process
route for tools not externalized.
Each built-in is the triple §3.3 pins:
- Binary — the
lernie tool <name>subcommand. Readstool_use.inputJSON from stdin, writes raw bytes to stdout, exits 0 on success or non-zero on failure (stderr is concatenated after stdout intotool_result.contentwhenis_erroris set). - JSON schema — at
schemas/tools/<name>.json, seeded to<data-root>/tools/<name>.jsonbylernie prime(whichmake installinvokes, ARCH §2.2). Sent verbatim as theinput_schemaof the tool's entry in the model call'stools: [...]array. - Skill — at
skills/<name>/SKILL.md, seeded to<data-root>/skills/<name>/bylernie prime. The frontmatterdescriptionis the tool's description intools: [...]; the body explains when to reach for it.
The pool is discoverable from the CLI itself — lernie tool --help
names it, and a name that is not in it is declined non-zero naming it
too, the same way load_skill declines an unknown skill (ARCH §3.3):
$ lernie tool --help
Arguments:
<NAME> Built-in tool to run; one of: bash, dispatch, load_skill, message, read_file
$ echo '{}' | lernie tool nosuchtool
lernie tool nosuchtool: unknown built-in tool: "nosuchtool"; available: bash, dispatch, load_skill, message, read_file
Built-ins:
read_file— read the entire contents of a file at a given path. Rejects files larger than 1 MiB, reporting the file's true size (stat, not the capped read's length) so the agent can judge the magnitude it is up against; v0.4+ adds the oversized-output auto-dispatch shim (ARCH §3.3 / §12). Try it directly:echo '{"path":"README.md"}' | lernie tool read_file.bash— runs a shell command viash -cand returns its stdout. The shell runs in its own process group so a SIGTERM the harness sends is forwarded to the entire spawned tree (§2.9 cascade). Try it directly:echo '{"command":"ls"}' | lernie tool bash.dispatch(v0.4 Phase 2) — spawns a subagent on a fresh branch with the supplied goal and returns{"status":"in_progress","handle":"<sub-branch>"}synchronously (ARCH §2.5). Input is{role, goal}; the role must resolve tosouls/<role>.mdand aroles:entry inproviders.yaml— both read from the calling branch's governing config commit (§2.2). Reads the calling conversation's repo + branch from the harness-setLERNIE_CONV_REPO/LERNIE_CONV_BRANCHenv vars (ARCH §3.3 env bullet); spawns throughlernie dispatch <role>(§3.4). The handle it returns is the child's address — there is no polling tool to pair with it. The substrate redesign (ARCH §2.5 "Dispatch returns the child's address") dissolved the handle/awaitpair: the child's result comes back as a deposit into the parent's inbox carrying an epitaph (§2.6, §2.11), soawait/checkhad nothing left to observe and are gone. The return path — the result-message deposit and the delivery-time work-product transfer — is built and live (bl-4ce8, bl-9f53, bl-c33b, §2.6), and children run full step loops: the dispatch's own front-door deposit finds the fresh child quiescent and launches the ordinary driver,lernie advance(§6) — there is no child-specific loop and no worker path — which steps the child to a terminal event, deposits its epitaph result (final-response, budget-exhausted, or stop) into the parent's inbox, and revives the parent, which delivers the result at its next step boundary.message— deposits content into an existing agent's inbox (ARCH §2.11). Input is{agent, content}; the recipient is addressed by its agent id (its branch name / hyphenated descent). Unlikedispatchit starts no branch and returns no address — it deposits synchronously and returns{"status":"deposited"}. The sender is the calling agent's id, taken from the harness-setLERNIE_CONV_BRANCH(never model-supplied), so provenance cannot be forged. It goes through the front door —lernie message(below) — likedispatchgoes throughlernie dispatch, so it inherits the front door's recipient guards: an id that is not a single path component, or one with noagents/*ref, comes back as anis_errorresult naming the decline instead of a silently lost message. Shipped state: the deposit lands and the step-boundary drain delivers it (bl-1129) — the next driver to step the branch moves the inbox file intomessages/as a transcript entry at its next boundary. A deposit into a quiescent agent is self-delivering: the free-lease probe detach-spawnslernie advance(§6, below), which acquires the lease, delivers the deposit, and steps the branch.load_skill— copies a pooled skill's body into the calling agent's worktree atskills/<name>/, where the next context assembly composes it (ARCH §3.3 Body-on-demand, §5.2). Input is{name}; the data-root pool + target worktree come fromLERNIE_HOME/XDG andLERNIE_CONV_REPO/LERNIE_CONV_BRANCH. Returns{"status":"loaded","path":"skills/<name>"}on a fresh copy oralready_loadedwhen the worktree already holds it (the loaded copy is the snapshot the branch is pinned to;rmand reload to refresh). An unknown or non-single-component name is declined (is_error, naming the available pool). Shipped state: the copy commits with the tool result — a tool commit now stages the whole worktree (git add -A,commit_tool), landing any tool's worktree side effects with its result entry (ARCH §2.3).
Messaging an existing agent directly
lernie message <workspace> <agent> <content> deposits a message into
<agent>'s inbox and, finding the recipient quiescent, launches a
driver to deliver it (ARCH §2.11, §3.4). The sender is read from
LERNIE_CONV_BRANCH — the calling agent's id when the message tool
re-enters the verb, else user for a bare invocation.
- The recipient is guarded before anything is written. The id must
be a single path component (ARCH §2.3) —
.., a/, or an absolute path is declined, never sanitized, becausePath::joinwould honour it and write outside the workspace — and anagents/<id>ref must exist for it: a message is addressed to an existing agent (§2.11), so a deposit no drain would ever come for is refused (lernie message: no agent "…" …, exit 1) rather than left in an inbox directory nothing will ever read. The id guard is the same rule at every verb taking an agent id from outside —message,advance,stop,dispatch,bundle— and literally the same code: one workspace-layout guard and one existence guard, each carrying the calling verb's own clause for why it needed an agent, so what differs between verbs is the cause, never the phrasing or the remedy. - The deposit is a create-only file at
<workspace>/inbox/<agent>/ <sender>-<NNN>.md(temp-path + atomic rename), withfrom:/deposited_at:frontmatter and the content as its body.<NNN>is the sender's own sequence, derived as max-present-plus-one over its existing files in that inbox. - After depositing, the verb probes the executor lock (
flockon the inbox directory): the same lease the shippedlernie promptstep loop holds for its whole run, releasing it on exit. A held lease means a driver is already stepping the branch (it will deliver at its next boundary); a free lease means the branch is quiescent. - On a free lease the verb launches a driver —
lernie advance <workspace> <agent>(ARCH §6) — as a detached spawn (§2.11):setsid(its own session and process group), stdio to null, fire-and-forget. The driver outlives thelernie messageprocess, so messaging is scriptable: the verb returns as soon as the deposit and spawn land, and delivery + stepping continue in the driver. - A failed branch is named, never refused. If the quiescent
recipient's latest model call failed (its last
response.jsonsegment terminated in anerror— retries exhausted or a non-retryable error, ARCH §2.10), the deposit and launch proceed unchanged — messaging is exactly how such a branch is retried once the cause is fixed — but the verb prints a stderr advisory naming the branch and pointing atsteps/<agent>/andlernie scan, so a silent death (ARCH §2.3, §8) is distinguishable from ordinary idleness at the verb that touches it. Exit code and stdout are untouched.
Driving a branch: lernie advance
lernie advance <workspace> <agent> is the §6 driver verb — the
process every launch seam spawns, and the same verb an operator runs by
hand. One invocation is one hop: guard the id (a single path
component, and an agents/<id> ref must exist — a name that is no
agent is refused with no agent "…" and exit 1 before any lease, so
an operator typo neither drives anything nor leaves an inbox/<id>/
behind), take the lease (adopt the
LERNIE_LOCK_FD fd published by a predecessor hop, else try-acquire
the executor lock — losing it is a clean no-op), deliver pending inbox
messages through the real drain (rematerializing a torn-down worktree
first), derive warrant from the transcript tail (ends user-side → a
model call is due; ends assistant-side without tool_use, or empty →
exit silently; assistant tool_use with uncommitted results → decline
loudly, the one non-replayable state), run one step, and hand off: a
step that emitted tool_use runs its tools and exec's the successor
lernie advance with the lock fd deliberately inherited (close-on-
exec cleared just before exec; the successor fstat-validates the fd
against the inbox directory and restores close-on-exec), while a
terminal event ends the chain through the §2.11 exit protocol. Because
the successor is exec'd in the same process, the pid, process group,
and flock lease all survive the hop — lernie stop lands on whichever
hop is current, and no rival driver can wedge between hops.
The exit protocol and the operator scan
Normal operation needs zero scanning (ARCH §2.11): lernie message
deposits, probes the executor lock, and launches a driver if the agent
is quiescent; the executor drains its inbox at every step boundary. The
graceful-exit crack — a deposit landing after an executor's final drain
but before its lock release — is closed by the exit protocol
(§2.11, bl-5846): one terminal sequence, no agent kinds — deposit the
result message (a structural no-op for a parentless agent) → release
own lock → spawn a driver at own agent, fire-and-forget → probe-and-
launch at the parent the deposit just landed in → exit. Two
pins terminate the recursion: a driver that acquires and finds nothing
to deliver exits silently (no step, no epitaph, no further launch —
dispatch::driver::drive is that entry), and the launch is decided by
epitaph value — a final response launches; stopped and
budget-exhausted never do. The exit launch rides the same launcher
seam as the writer probe, so it is the same detached lernie advance
spawn (§6); the decision logic, ordering, driver entry, and the spawn
itself are live and tested.
The parent-side step is what makes revival-on-deposit real
(bl-4a6c): a child that returns to a quiescent — even torn-down —
parent starts that parent's driver itself, through the same
probe_and_launch the lernie message verb uses (one probe, no
second copy), so the parent rematerializes, delivers the result, and
steps with no lernie scan in the path. A parent whose lease is held
gets nothing launched: its running executor delivers at its next step
boundary. The epitaph decision governs this launch too, one level up:
a stopped child would otherwise wake its parent to react to — perhaps
re-dispatch around — the very branch the operator killed, and a
budget-exhausted child's ceiling is the whole tree's (§6), so the
woken parent would exhaust on its own next check and deposit again. In
both cases the result still lands in the inbox and waits for the next
explicit touch.
Crashes are accepted as a failure class (§2.11): everything is on disk,
so a hard death strands results and messages late, never lost, and
the next touch heals. That touch is a user reprompt — or the operator
verb lernie scan <workspace> (§2.11, §8, bl-d148 + bl-5846): one
workspace-wide pass, run by hand or by cron if you want a heartbeat,
never wired into any driver hot path or default schedule (the events it
compensates for happen at crash rate, not step rate). Two derived
actions, no watcher (an idle workspace stays unswept until the next
touch, by design):
- Silent-death sweep. Every agent branch with no live executor (the
§2.11 executor-lock probe) that either died mid-work — its latest
step's model call never settled complete:
response.jsonclosed without a terminalend(killed/stopped, §2.9), or its final segment terminated in anerror(retries exhausted or a non-retryable error, §2.10 — that segment closes with a cleanend, so absence-of-endalone would misread the branch as idle) — or, for a child, never deposited a result message is a silent death (the §8 health count). Each one is named in the report (silent deaths: 1 (<agent-id>)): a dead root gets no deposit — it has no parent inbox — so its name here is how an operator learns which branch went quiet, andsteps/<agent-id>/is where to read why. For each hard-crashed child in that set, the sweep deposits adied-epitaph result message on the child's behalf (sender = the child — the sweep is the scribe, not the author), so the parent is revived rather than stalled. The "never deposited" test reads both the parent's inbox (undelivered) and its transcript (delivered), so a prior sweep's own deposit is seen on re-scan and never re-deposited — idempotent by construction. - Inbox flush. Every agent with pending inbox files and a free lock
gets a driver launched — never drained: the scanner moves no files
and commits nothing; only an agent's own lock-holding executor
delivers. An agent whose lock is held is left alone. The inbox listing
is intersected with the
agents/*refs — the one registry of who exists — so an inbox directory with no matching ref is reported (inboxes with no agent branch: N) and left in place rather than driven: a driver launched for a name with no branch is refused by the existence guard (lernie advance: no agent "…", exit 1) on this pass and every pass after, writing nothing. The sweep's own deposits are picked up by the flush that follows in the same pass.
Shipped state. The scan (silent-death sweep + inbox flush) ships
behind lernie scan and only there — driver startup (lernie prompt,
lernie dispatch, lernie advance) runs no workspace scan. The flush
and the exit launch reuse the same driver-launch seam as lernie message, and the spawn is real: each seam decides when a driver is
needed and detach-spawns lernie advance (§6) for it. Children run full
step loops (bl-c33b), so a died child is a state a real run reaches; the
derivation is additionally exercised against constructed on-disk states,
since a hard crash is not reproducible on demand.
Namespace note. The candidate enumeration is the agents/* ref
namespace, exactly as ARCH §8 writes it (a root is agents/<conv-id>,
a child agents/<parent>-<sub-id>); config branches are excluded
structurally by the prefix — there is no main (§2.2).
Dispatching subagents directly
lernie dispatch <role> <repo> <branch> [--goal <text>] is the §3.4
re-entry point every child dispatch uses. It is writer-shaped, not an
executor (ARCH §2.1): it forks the child branch, lands the dispatch
commit, and deposits the dispatch message through the same front door
every sender uses — the driver that deposit launches is the ordinary
lernie advance (§6). The role name is positional and the role set is
open (§4.3): a role is dispatchable iff the calling branch's
governing config commit lists it under providers.yaml roles: and
carries souls/<role>.md. The CLI enumerates no role names, so a
verifier, a critic, or a role you author needs no CLI change; validity
is checked before the fork, so a rejected role leaves no branch debris.
The id guard runs first, through the same two functions message,
advance, stop and bundle call: the workspace layout, then the
dispatching parent's agents/<id> ref. So all three refusals are the
product's, never git's:
lernie dispatch worker <no-such-ws> someagent --goal hi
→ <path> is not a workspace (no repo.git) — create one with `lernie new` (ARCH §2.2)
lernie dispatch worker <ws> nosuchparent --goal hi
→ no agent "nosuchparent" in this workspace — a child forks off an existing parent (ARCH §2.5); …
lernie dispatch verifier <ws> <agent> --goal hi
→ role "verifier" is not defined in the providers.yaml governing agent "<agent>" — defined roles: compactor, worker
The role refusal names the pool that is defined — the same "name the
pool" idiom load_skill and lernie tool decline with — and names the
control file the user knows rather than the config commit's sha.
lernie dispatch compactor <workspace> <conv-id>forks a compactor-souled child off that agent's tip — exactly what a due compaction checkpoint does (§2.7), run by hand. The compactor is an ordinary child that makes a real model call throughbz; it is not a stub, and it does not merge anything itself. Its goal is procedure-generated, so passing--goalis rejected. Its toolset is the deletion-only pair injected for the compactor role alone (never aproviders.yamltools:list):write_summary, which writes the nextsummary/<NNN>.mdon the compactor's branch, andmark_for_deletion, a stagedgit rmthat can remove but never write content — so the worst case is lost information, never corrupted information. Its request declares more than that pair: a compactor inherits the dispatching branch's transcript, so the model call also names whatever tools that transcript used — otherwise the provider refuses a request whose history mentions a tool it was not told about. Declaring is not permitting: a compactor reaching for one of those inherited tools gets an error tool result naming its own two, and nothing runs. The compaction merge lands later and elsewhere: when the compactor's result message is delivered, the dispatching agent's own executor interprets itscompactor_return: compaction_mergebinding (§6) and merges the compactor branch--no-ff— the one merge left in the system (§2.6). A compactor that ends on any other epitaph lands no merge; the branch simply continues uncompacted — enforced where the binding is interpreted: the delivered result's epitaph value gatescompaction_merge, and adied/stopped/budget-exhaustedcompactor return is delivered like an ordinary child's result instead, so the parent sees the epitaph and nothing of the compactor's branch crosses (§2.6, §2.7).lernie dispatch worker <workspace> <parent-id> --goal <text>spawns a worker child off the parent's tip. The new id is<parent>-<sub-id>(hyphenated descent, §2.2), its refagents/<parent>-<sub-id>(§2.3), its worktreeagents/<parent>-<sub-id>/;goal.mdcarries the supplied text andsoul.mdis read from the parent's governing config commit (souls/worker.md, §2.2), both committed as the dispatch commit (§2.3 step 2). The child then runs a full step loop under thelernie advancedriver its dispatch deposit launched, and at its terminal event deposits a result message — epitaph, terminal ref, and the terminal response iff it spoke — into the parent's inbox, reviving the parent if it had gone quiescent (§2.6, §2.11). The v0.4 "Phase 1 stops at the dispatch commit" worker path (worker.rs) is deleted, not extended (bl-c33b).
Providers
Every model call goes through brazen — one small, stateless binary
(bz) that adapts every provider and wire protocol behind a single pipe
contract (see ARCH §4.4):
stdin (canonical request, JSON) → bz → stdout (v=1 event stream, NDJSON, one terminal `end`)
The harness execs bz --json --provider <row> once per attempt, pipes a
typed brazen::CanonicalRequest on stdin, and appends bz's stdout
verbatim to the step's response.json. lernie links the brazen crate
(brazen = "=0.0.4") for the canonical types only — the data plane
always crosses the subprocess boundary (§3.4). Two facts follow:
- Retry is the harness's. brazen never retries — one
bzprocess, one HTTP round-trip. On a retryable in-bandError(CanonicalError::retryable(), the linked crate's single home for the fact) the harness re-invokesbzup to theworkflow.yamlattempt cap (§2.10). Each attempt appends one segment toresponse.json; the last is authoritative. Each attempt'sbzstderr appends to the step'sstderr.logbeside it — empty on an ordinary run, because brazen speaks its failures in-band on stdout. Abzthat dies before it can (a malformed brazen config) leaves an empty stream that reads exactly like a mid-stream kill, so the half-stream error quotes that capture's tail; with a stop pending it stays quiet, because the stop check point (§2.9) discards the outcome before anything is rendered. - Auth and endpoints are brazen's. Provider rows (endpoint,
protocol, auth mode, model aliases) live in brazen's own config
(
~/.config/brazen/config.toml;bz --dump-config,bz --login). lernie references a row by name and never sees credential material (§4.1). A load-time guard (bz --version== the linked crate version) rejects a mismatched binary;make installinstalls the pin withcargo install brazen --version =0.0.4.
Adding a provider
- A new provider on a supported protocol is a brazen config row — no
code anywhere. Add the row (
bzconfig), reference its name as a model'sprovider:in<config-root>/models.yaml, and point a role at that model in<repo>/providers.yaml. - A new wire protocol or auth mode is a contribution to brazen.
- An alternate adapter binary that honors the same pipe contract
slots in via the optional
adapter:path inmodels.yaml(§4.2); the version guard is skipped for it and the in-bandMessageStart.vhandshake governs compatibility instead.
UI (v0.5)
The desktop frontend lives in its own repository, yog: an
egui/eframe window that renders a workspace and issues user actions via
lernie <subcommand>. It composes on lernie's public surfaces only —
the CLI and the on-disk workspace layout (ARCH §3.5, §7.1) — and takes
no Cargo dependency on this crate, so it builds, versions, and installs
independently (make install there drops yog next to
lernie). Keeping frontends out of this workspace is deliberate:
lernie ships as a composable component, and anything that composes it
(a GUI, a web view) lives outside it and meets it at those surfaces.
Evaluation: archival and the task suite (§9)
Archive a run. A "run" is an agent subtree, not a whole workspace (§9.2).
lernie bundle <workspace> <agent> <out-dir> writes the subtree — the
agents/<agent> branch and its agents/<agent>-* hyphen-descendants (§2.3),
with all the ancestry those refs reach — plus the subtree's governing
lineage: every config/* ref whose history reaches it (§2.2). Both go into
one git bundle, and the matching steps/<id>* and inbox/<id>* diagnostic
slices are copied beside it. One bundle plus two slices is the whole run.
The config refs are not decoration. An agent's control files are read from its
governing config commit, which is derived — the nearest ancestor of the
branch reachable from a config/* ref (§2.2). Ancestry alone carries that
commit as an object but names no ref to take the merge-base against, so a
replay of the agent refs alone yields a workspace no verb can drive. Carrying
the refs (never a sidecar file — the refs are the single source) makes the
replayed repo derive its governing config by the same computation, over the
same candidate set, as the workspace it came from. "Every ref whose history
reaches it" is broader than "every ancestor": a sibling config lineage
that shares only a common root with the bundled subtree is still a
merge-base candidate, so it rides too — carrying a ref that turns out not
to be the nearest one is how the bundle stays a faithful copy of the
computation, not a leak.
lernie bundle /path/to/workspace <agent-id> /path/to/archive
Replay a run. lernie replay <archive> reconstructs a scratch workspace
under LERNIE_HOME's data root at replays/<primary-id>/ (the primary id is
the subtree's root agent), fetches every branch out of the bundle into a
fresh bare repo.git, materializes the primary's worktree under agents/,
restores the slices, and prints the scratch path. Point the ordinary frontend
at it — replay is not a mode (§2.3). Set LERNIE_HOME to an isolated
directory to keep the replay sandboxed; the harness root it points at still
supplies the machine-local pieces a config only names (models.yaml and the
brazen provider rows, §4.2/§4.4).
A replayed workspace is an ordinary workspace: lernie prompt <scratch> "…"
forks a fresh root off the config head that rode the bundle, and lernie message / lernie advance drive the replayed agent on its own governing
config commit.
LERNIE_HOME=/tmp/replay lernie replay /path/to/archive
Task suite. The evaluation suite lives as data under tests/suite/ — 50
tasks with machine-checkable check scripts, tagged by the seven §9.1 failure
categories (≥10 per category), format in tests/suite/README.md,
well-formedness enforced by tests/suite.rs.
Run the suite. The agent-eval runner (a separate crate, crates/agent-eval,
ARCH §9.3) executes an experiment against the suite N times per task and reports
pass@1 (with 95% Wilson intervals) and pass@5, overall and per category:
agent-eval --config baseline --suite tests/suite --runs 5 --agent lernie-eval-agent
--config <name> names an experiment — a workflow.yaml variant under
experiments/<name>/ (a config diff, no code changes; see experiments/README.md).
baseline is the shipped default itself: its workflow.yaml is a symlink to
template/workflow.yaml, because an experiment is a diff against the default and
the baseline's diff is empty.
Per run the runner seeds a fresh isolated LERNIE_HOME and working directory,
runs the task setup, invokes the agent, then runs the task check — exit 0
is the sole pass signal (§9.1), so success is observable state, never the
agent's own claim. --bundle-dir <dir> archives failing runs for triage via
lernie bundle (§9.2). The runner is fully tested against a faked agent, so it
needs no live model to validate.
The shipped driver is lernie-eval-agent (crates/lernie-eval-agent,
workspace-internal like the runner; installed on PATH by make install).
--agent <cmd> stays required with no default: which driver runs the agent
under test is an experiment-defining input, so it is named explicitly. Per run
the shipped driver seeds the run's isolated LERNIE_HOME from the machine's
lernie config root (models.yaml plus the template/ config-root override —
the wire is machine-local by design, §4.2/§9.2, and those two front doors are
how a machine points evaluation runs at its own provider rows), then drives
the harness exclusively through the front door, exec'ing lernie from PATH:
lernie new, lernie config (applying the experiment — below), and one
lernie prompt carrying the task prompt grounded in the shared working
directory. The contract any driver must honour, per run:
| Given | How |
|---|---|
| the task prompt | argv[1] |
| the isolated harness root for this run | LERNIE_HOME in the env |
the experiment's workflow.yaml |
LERNIE_EXPERIMENT in the env — an absolute path |
| where to report back | LERNIE_EVAL_REPORT in the env — a file path |
| the working directory | cwd (shared with the task's setup and check) |
LERNIE_EXPERIMENT is a hand-off, not a hook: nothing in the harness reads
that variable. The harness takes its workflow.yaml from the workspace's
config commit (§2.2), never from the environment, so applying the experiment
is the driver's job. The shipped driver does it through lernie config, with
$EDITOR set to copy the experiment over the authoring checkout's
workflow.yaml — the experiment lands as an ordinary config commit, exactly
the "config diff, no code changes" §9.3 promises (for baseline the diff is
empty and the authoring pass declines: the default is already in force).
LERNIE_EVAL_REPORT names a file the driver may write with exactly two
lines — the workspace path, then the agent id — which is what lernie bundle
needs to archive the run if it fails (§9.2). It is the driver's only channel
back to the runner. Writing nothing, or anything malformed, only makes a failing
run un-bundleable; it is never an error, and it never affects pass/fail, which
is the task check alone. The driver's own exit code is likewise ignored.
Failure to spawn the driver, by contrast, is a hard error naming the program.
Contributing
The instructions below are for contributors building lernie from source. Users installing a release don't need any of this — Install covers the three user-facing routes, only one of which involves a clone.
Contributor setup
make install-hooks
Sets core.hooksPath to .githooks. Required on every fresh clone — git
does not track .git/config, so the hooks are not active until installed.
That arms both the pre-commit gate and the
auto-push hook.
The Rust toolchain is pinned in rust-toolchain.toml (channel 1.95.0, with
rustfmt, clippy, and llvm-tools-preview). rustup reads it automatically
for every cargo command in the tree and installs the pinned toolchain on
first use — no manual rustup step. This is what keeps fmt-check and
lint from drifting between your machine, another agent's, and CI.
Build targets
| Target | What it does |
|---|---|
make build |
cargo build |
make release |
cargo build --release |
make test |
cargo test, with the pinned bz first on PATH (below) |
make test-install |
cargo test --test install — the install contract end-to-end, uninstrumented (it is cfg_attr(tarpaulin, ignore), so coverage skips it); ~45s warm, and it re-installs bz at the brazen pin |
make coverage |
cargo tarpaulin --fail-under 100 (llvm engine), same pinned PATH (below); hard-gated on tarpaulin 0.35.2 exactly (TARPAULIN_PIN in the Makefile — its one home; any other version aborts with the cargo install cargo-tarpaulin --version 0.35.2 --locked fix-it line) |
make lint |
cargo clippy --all-targets -- -D warnings |
make fmt |
cargo fmt |
make fmt-check |
cargo fmt --check |
make schemas |
Regenerate schemas/*.json from the Rust types |
make new-workspace DEST=<path> |
Create a workspace (bare repo.git + first config commit from template/) |
make eval CONFIG=<exp> SUITE=<dir> RUNS=<n> AGENT=<driver-cmd> |
Run the evaluation runner (ARCH §9.3): experiment × suite × N (see Task suite above). AGENT is required and has no default — the shipped driver is lernie-eval-agent (see "Run the suite"), and naming it is deliberate: the driver is an experiment-defining input |
make check |
fmt-check + lint + coverage + test-install |
make ci |
Alias for check |
make smoke |
Live-wire smoke test: one real lernie prompt against the shipped defaults (override with SMOKE_PROVIDER/SMOKE_MODEL); the default needs a bz anthropic credential and spends money; NOT part of check |
make install-hooks |
Point git at .githooks/ |
make install-bz |
Install the provider adapter bz on your PATH at the version Cargo.toml pins (ARCH §4.4); a no-op when the bz there already matches. For running lernie — the tests feed themselves (below) |
make brazen-pin |
Print that pinned version and nothing else — CI keys its bz cache on it so no workflow file names a version |
make install [INSTALL_PREFIX=<p> LERNIE_HOME=<h>] |
Release-build; drop lernie/agent-eval into $INSTALL_PREFIX/bin (default: ~/.local/bin); install the provider adapter bz via make install-bz at the version Cargo.toml pins (the ARCH §4.4 version pin — the number's one home); then invoke lernie prime to found the harness root — config root (default ~/.config/lernie) with a default models.yaml and an empty workflows/ templates dir, data root (default ~/.local/share/lernie) with the tools//skills/ pools and the workspaces/ tree — seed-if-absent (ARCH §2.2); LERNIE_HOME collapses both |
make uninstall [INSTALL_PREFIX=<p> LERNIE_HOME=<h>] |
Remove the installed binaries; leaves the harness homes (config + data roots) in place |
The pinned adapter under test
The e2e tests exec the real bz (against a mock HTTP endpoint, not a
provider), and lernie's load-time version guard (ARCH §4.4) demands the
pinned version exactly. The pin's one home is the brazen = "=<version>"
line in Cargo.toml.
The trap. bz normally resolves from PATH — that is
~/.cargo/bin/bz, machine-global mutable state shared by every checkout and
every agent on the box. Anyone running make install rewrites that binary at
their tree's pin. If your tree pins a different version, your next test run
dies in five-plus e2e tests with
bz version "0.0.3" does not match the linked brazen crate "0.0.4"
which looks nothing like "someone else installed a binary" and everything like a regression you just wrote.
The cure. make test and make coverage do not use the PATH bz at
all. They depend on $XDG_CACHE_HOME/lernie/bz/<pin>/bin/bz — installed
from crates.io on first use — and put that directory first on PATH for
the run, so the tests always exercise the pin this tree names, whatever
the machine's bz happens to be. The version comes from BRAZEN_PIN in the
Makefile, derived from Cargo.toml; the cache directory is named after
it, so bumping the pin is a cache miss and nothing else, and a stale entry is
never overwritten in place. Cost: one cargo install (~25s) per pin per
machine — sibling worktrees share the cache — and nothing at all when warm,
since it is an ordinary make file prerequisite.
Two consequences worth knowing:
- Bare
cargo testis still exposed. It inherits yourPATHand so runs whateverbzis installed there. Usemake test; if you must runcargo testdirectly,make install-bzfirst to line the global binary up with the tree's pin. - No test writes the global
bz.make installdoes — that is its job — but the install test that runs it (tests/install.rs) pointsCARGO_INSTALL_ROOTat a per-worktree root undertarget/, so the pinnedbzlands there and~/.cargo/bin/bzis never touched by a test run. - Runtime resolution is unchanged. This is test determinism only —
lernieitself still resolves the adapter per ARCH §4.4 (themodels.yamladapter:override, else a binding-injected target, elsebzonPATH), andmake installstill puts the pinnedbzon yourPATHfor real use.
the_makefile_derives_the_same_pin (src/prompt/tests/pin.rs) keeps the two
readers of that one line honest: the Makefile's BRAZEN_PIN (which names the
cached binary) and the crate's brazen_pin() (which the version guard
compares against) must agree, or the tests would fail the guard against a
binary the Makefile itself installed.
Workflow
All changes land on main via bl squash-merges. Direct commits to main are
rejected by the pre-commit hook, and every landing on main is pushed to
origin automatically (see Auto-push hook).
bl prime --as <you>
bl claim <task-id> # creates a worktree; cd into it
# ...edit, test, commit...
bl close <task-id> -m "..." # squash-merges into main; run from the repo root
See bl skill for the full guide.
What gets published
cargo package ships the crate, not the repo. Cargo.toml's exclude keeps
out everything that serves this git checkout only — docs/, tests/,
experiments/, scripts/, .github/, .githooks/, .balls/, Makefile,
tarpaulin.toml, release-plz.toml, AGENTS.md, CLAUDE.md, and
rust-toolchain.toml (which would otherwise force a source builder onto this
repo's exact pinned toolchain). What remains is src/, README.md, LICENSE,
Cargo.lock, and the embedded asset trees template/, schemas/, skills/,
install/models.yaml — those four are include_dir!/include_str! inputs, so
excluding any of them is a build failure, not a smaller tarball. Verify a change
to the list with cargo package --list and then cargo package, which
compiles the extracted tarball.
crates/agent-eval is publish = false: it is workspace-internal and is not
part of the published crate at all.
Pre-commit hook
.githooks/pre-commit enforces three rules on every commit:
- No direct commits to mainline.
mainandmasterare rejected unless the commit is the tail of a merge (MERGE_MSG/SQUASH_MSGpresent), which is howbl closelands squash-merges. - 300-line cap on code files. The cap is a repo invariant, not a
per-commit property, so the hook sweeps every tracked code file in the
tree (
git ls-files), not just the staged set — a file that crosses the cap in one commit and is untouched afterward is still caught. Docs (*.md,*.txt), config (*.toml,*.yaml,*.yml,*.json,*.lock),Makefile,.gitignore,LICENSE, and anything under.githooks/are exempt. make checkon every commit that touches a Cargo project:fmt-check(formatting),lint(clippy -D warnings),coverage(cargo tarpaulin --fail-under 100), andtest-install(cargo test --test install). The hook invokesmake checkrather than re-listing the commands, so the close gate is always exactly whatmake checkis — the Makefile is the single source. Formatting and lint drift therefore cannot land invisibly.test-installis a separate step because the install test shells out to a release build andcargo install brazen, which contend with tarpaulin'starget/lock; it iscfg_attr(tarpaulin, ignore), so without its own uninstrumented step the install contract — the first thing every user touches — would never run at the gate at all. It costs ~45s warm and leaves the machine-global~/.cargo/bin/bzalone: the test redirectsmake install'scargo install brazeninto a per-worktree root undertarget/withCARGO_INSTALL_ROOT, so a sibling worktree at another pin is never rolled over. The toolchain is pinned inrust-toolchain.tomland the tarpaulin version intarpaulin.toml(also.github/workflows/ci.yml) sofmt-check,lint, and the coverage denominator mean the same thing locally and on CI — newer tarpaulin releases have silently dropped inline#[cfg(test)] mod tests;files from the count, weakening the floor.make coverageaborts with an install hint if the local tarpaulin version drifts.
A floor of exactly 100% only holds if every line's coverage is caused by the code's own structure and not by winning a race, so no line may be reachable only while a clock has not yet run out. With several agents measuring coverage at once, whichever side of such a race the machine happens to pick that minute decides the verdict, and the gate reports an uncovered line on a diff that touched nothing. Two shapes to write around:
- A retry budget is a count of attempts, never a wall-clock deadline.
PROBE_RETRIES(src/prompt/tests/exit_launch.rs) is the one budget every executor-lock probe shares; a deadline expires on load rather than on evidence, so under load the give-up arm can be taken on the first pass and the retry arm never runs at all. - A poll loop waits because its child is still running, not because a flag
has yet to land.
wait_with_cascade(.../builtin/bash/mod.rs) andwait_with_stop(.../tool/subprocess.rs) therefore sleep between the reap and the flag read: the interval is entered for as long as the child lives, instead of only while a stop scheduled milliseconds out has not arrived yet.
The same objection reaches past coverage to the verdict, and the
end-to-end tests answer it the same way: a poll waiting on a detached driver
is bounded by consecutive probes that saw no change in the workspace tree,
never by wall time (src/e2e/poll.rs, docs/ARCHITECTURE.md §9). A live
driver writes continuously and a wedged one writes nothing, so a loaded box
only makes the pass path slower — where a stopwatch would have turned a slow
success red, and (as bl-2bf0 found) hid a real defect behind a timeout that
read like machine load.
There is no --no-verify escape hatch in the workflow. If the hook rejects a
commit, fix the underlying issue rather than skipping.
Auto-push hook
.githooks/reference-transaction pushes main to origin the moment local
main advances. Landing and publishing are one act: a bl close reaches
GitHub and the push triggers the Release-plz workflow, which contains CI as a
called job (needs: ci) and only publishes once it is green. origin/main
cannot silently fall months behind local main again.
Why a reference-transaction hook and not post-commit. Nothing lands on
this repo's main through git commit. bl close delivers by plumbing —
git commit-tree, then git update-ref refs/heads/main — which fires no
commit hook and no merge hook at all — every commit bl has landed on main
arrived that way. Git's reference-transaction hook is the one event every
landing path shares: the plumbing delivery, a git merge --no-ff, and a plain
commit alike all end in an update of refs/heads/main.
The hook acts only on the committed state of a transaction that moves
refs/heads/main to a new value, and only when an origin remote exists.
Everything else — side branches, refs/remotes/* (including the ones its own
push writes, so it cannot recurse), no-op rewrites like git pack-refs, and a
deletion of main — falls through untouched.
It cannot block or hang a landing. Git aborts a ref transaction when this hook
exits non-zero in the prepared state, so every path in it exits 0 — which is
also why it does not set -e. A push that fails prints one warning line on
stderr and nothing else, and timeout 30 bounds an offline push rather than
stalling the commit behind a TCP timeout. Git runs reference-transaction
hooks from 2.28 onward; on anything older the file is simply never invoked and
main has to be pushed by hand.
tests/hooks.rs exercises the shipped hook file itself against a local bare
repository as origin — never the real remote — and covers all six behaviours
above: a commit on main pushes, a commit-tree + update-ref delivery
pushes, a --no-ff merge pushes, a side-branch commit pushes nothing, an
unreachable origin warns without failing the commit, and a repo with no
origin is silent.
Commit-identity guard (opt-in, per machine)
main's history carries exactly one human identity, mudbungie <mudbungie@gmail.com>, and no Co-Authored-By trailers — it was normalized to
that on 2026-07-26. tests/commit_hygiene.rs keeps it that way, but only on a
machine that asks for it: the test arms itself on the presence of
$XDG_CONFIG_HOME/lernie/enforce-commit-identity (default
~/.config/lernie/enforce-commit-identity), an empty marker file outside the
repo. Absent — the default in public CI and in every clone — the test returns
without asserting anything.
Armed, it walks all of refs/heads/main and fails on any commit whose author or
committer is neither mudbungie <mudbungie@gmail.com> nor
github-actions[bot] (the bot stays allowed: release-plz authors the release
commit as it), on any Co-Authored-By trailer, and on any mention of a
throwaway or personal address in an identity or a message. The policy lives in
the marker, not in the code: rm it and the guard is off, with no code edit and
no flag. Create it with touch ~/.config/lernie/enforce-commit-identity.
License
MIT. See LICENSE.