litany 0.0.13

A git-backed agent harness
Documentation
events:
  user_message:
    - dispatch(worker)
  worker_return:
    - deliver_result
  worker_flush:
    - dispatch(compactor)
  compactor_return:
    - land_compaction
  branch_stopped:
    - mark_abandoned
    - notify_ui

# Intermediate compaction checkpoints (ARCH §2.6–§2.7, §6). The executor
# reads this at each step boundary: when the trigger fires it dispatches a
# compactor off the compaction point — the branch tip, or `HEAD~keep_recent`
# when `keep_recent` is set — and the compactor's return lands by
# rebase-forward (the `compactor_return: land_compaction` binding above):
# the span before the point squashes into a compaction base and the live
# tail replays on top, zero downtime. `trigger` is one of
# `every_n_commits`, `every_t_seconds` (both take `n`), `window_percent`
# (`n` = the percent of the model's context window the branch's last usage
# fills, 1..=100 — the provider's own numbers, both of them: the prompt
# side of the newest model entry's `usage` report against the
# `context_window` the same report carries), or `on_flush` (the
# agent-elected `flush`, no `n`). `keep_recent` (optional, default 0) keeps
# the most recent commits out of the span; it must stay below `n` under
# `every_n_commits`. `keep_recent_tokens` states the same retained tail in
# the provider's own unit instead — keep the longest stretch of the
# transcript that cost at most that many prompt tokens to append, read off
# successive model entries' `usage` reports, no tokenizer and no stored
# counter. One tail or the other, never both: declaring both is refused at
# load. `extract_bytes` (optional) caps the extract the landing itself
# derives — `summary/<NNN>.refs.md`, the verbatim user messages, error
# strings, pull-request numbers, commit shas and paths the compaction
# takes out of context, written by code beside the compactor's prose
# (docs/DESIGN_CONTEXT_ECONOMY.md §5.3); omit it and no extract is
# written. Omit the whole block and the branch never compacts.
#
# The window is on the wire (brazen 0.0.9 stamps `context_window` onto
# every usage event), but the shipped trigger stays `every_n_commits`: a
# model whose provider row states no window is *declined* under
# `window_percent` rather than left silently never-due, so defaulting to
# it would refuse those workspaces at their first boundary. That is not a
# hypothetical any more — the survey ran (bl-4c64) and refused the flip.
# Of brazen 0.0.9's eight built-in rows exactly ONE can state a window at
# all (`google`, whose protocol names `inputTokenLimit`); the rest name no
# key to lift one from, `anthropic` and `claude-code` among them. And even
# that one states nothing until `bz --list-models` has run for it, because
# the window is read off the local model cache and the data plane's own
# writer stores a bare id with no metadata. So the variant is available,
# not defaulted: turn it on here when your row states a window, and see
# docs/ARCHITECTURE.md's bl-4c64 shipped-state note for what would make it
# the default (a change in brazen's rows, not in litany).
#
# `keep_recent_tokens` is unset here for the same reason and one more: a
# token tail under a COMMIT clock leaves a stretch where the clock is over
# threshold and the span keeps coming back empty — the branch re-walks its
# transcript at every step boundary and compacts nothing until the tail
# outgrows the budget. That is the shape `keep_recent >= n` is refused for,
# in units no load-time check can compare. The token tail belongs with the
# window trigger; they flip together or not at all.
#
# `n: 60` and `extract_bytes: 8192` are chosen together with the
# `tool_output:` bound below, because the clock counts commits and what
# a commit COSTS is that bound (bl-ce09). A step writes two commits, so
# 60 is about thirty steps. At the shipped 4 KiB-per-result bound a
# tool-heavy step appends a few thousand prompt tokens, so a span of
# thirty steps is tens of thousands — inside every shipped model's
# window, with the retained tail and the pinned head on top. The
# previous `n: 20` fired every ten steps whatever the context held: on
# the measured goals an ordinary conversation compacted six times, and
# each compaction is a model dispatch carrying the whole inherited
# transcript, so the clock cost more than the context it was reclaiming.
# The rule the pair is picked under is: an ORDINARY conversation should
# finish without compacting once, and a long one should compact rarely
# rather than continuously.
#
# `extract_bytes` moved with it for a reason that is not symmetry. Since
# bl-2071 the landing sweeps the span's transcript entries itself, so
# the extract is now derived from the whole span rather than from
# whatever a model happened to nominate — it will actually fill toward
# its cap, and unlike a tool result it stays in context until its
# summary is shed. 8 KiB is about two thousand tokens of references per
# compaction; 32 KiB was a cap nothing reached before and would now be
# paid every time.
compaction:
  intermediate:
    trigger: every_n_commits
    n: 60
    extract_bytes: 8192

# Harness-owned retry policy for a step's model call (ARCH §2.10, §4.4):
# brazen never retries — the harness re-invokes `bz` on a retryable
# in-band Error, up to max_attempts, with exponential backoff.
retry:
  max_attempts: 3
  backoff: exponential

# Bounded transcript projection of tool output (ARCH §3.3, §6). Each
# stream of a tool result (stdout and stderr independently) is bounded
# to its first head_bytes and last tail_bytes before the result envelope
# is rendered; the omitted middle is replaced by a marker stating the
# original byte/line counts and where the full record lives
# (steps/<agent-id>/<NNN>/tools/<tool-id>/output.json — always complete).
# Counts are bytes, never tokens. Omit the block and tool output reaches
# the transcript unbounded.
#
# 2 KiB + 2 KiB is the shipped bound, and it is small on purpose
# (bl-ce09). The number that ships is the one an ordinary conversation
# pays on EVERY step for the rest of its life, so it is chosen against
# the ordinary case and not against the rare one that wants the whole
# capture. At 4 KiB a result is roughly a thousand tokens: a `--help`,
# an `ls -la`, a `git status`, a test summary all land whole or land
# with their two useful ends and a marker between them. The previous
# 16 KiB + 16 KiB made ONE result worth about eight thousand tokens —
# measured, three ordinary calls filled a context, one `find` over a
# home tree put 32,985 bytes into a transcript essentially whole, and a
# conversation reading a repository reached 122,000 prompt tokens by its
# eighth step on `cat` output alone.
#
# What it costs is real and is priced here: a source file read whole is
# cut in the middle, and the model gets the marker instead. That is the
# intended trade — the full capture is on disk, the marker names its
# path and the byte and line counts, and re-reading a named range costs
# one cheap tool call, where carrying every whole file forever costs
# every later step. The range is named for the model rather than
# guessed by it: since bl-cbe0 `read_file` takes `offset`/`limit` in
# lines and every result says which lines it returned, how many the
# file has, and the offset to continue at. Raise both numbers on a
# workspace whose work really is reading long files end to end; that is
# what a severable policy block is for.
tool_output:
  head_bytes: 2048
  tail_bytes: 2048

# Context files (ARCH §3.3 *Context files ride the next tool result*,
# `docs/DESIGN_CONTEXT_ECONOMY.md` §6). File NAMES, looked for in every
# directory on the path from the enclosing repository's top level down to
# the agent's working directory. Each one the agent has not been shown
# yet is appended to its next tool result, framed <file path="..."> and
# bounded by tool_output above as its own stream; "already shown" is read
# off the transcript, so a compaction that drops the entry shows the file
# again. Omit the block and nothing is discovered.
context_files: [AGENTS.md, CLAUDE.md]

# Tool control (ARCH §3.3 *Tool control*, §6): an adjudicator binary
# consulted before every granted tool invocation executes — it answers
# pass, refuse, or hold (park for out-of-band review). Deliberately not
# configured here: no control ships, and omitting the block leaves the
# tool window unchanged. To wire one:
# tool_control:
#   command: /path/to/control

# Whole-tree spend limits (ARCH §6 "Budgets (v0.7)"). One frozen ceiling
# for the whole agent tree, not a per-agent allowance: every driver in
# the tree — root or subagent — checks the tree's total against these
# same numbers, and a dispatch inherits no fresh budget. Checked at
# every model-call boundary before the adapter is invoked; spend, wall,
# and depth are derived from disk each check — no stored counter. Omit a
# limit (or the whole block) to leave that axis unbounded.
#
# What ships bounded is DEPTH, and only depth (bl-c701).
#
# The two SPEND ceilings stay off, for the reason the 2026-08-16
# operator ruling gave: a whole-tree ceiling binds far earlier than its
# number reads, because a root and every agent below it spend one shared
# allowance — an hour of accumulated wall across a tree ends a
# conversation that is working, and raising the number moves that cliff
# rather than removing it. Nothing here bounds tokens or wall; an
# operator who wants a spend ceiling declares one. To wire one:
# budgets:
#   max_total_tokens: 2000000
#   max_wall_seconds: 3600
#
# max_depth was swept out with them rather than judged on its own, and
# it is a different kind of limit. It is not a spend judgement: it is
# the tree's only prohibition on GROWTH (ARCH §6 "The depth ceiling is
# the tree's only prohibition on growth"). Every agent may dispatch
# children, a parent never blocks, and stopping a parent does not
# cascade downward on its own — so with no depth ceiling a tree that
# re-dispatches itself has nothing structural to stop it. That is not
# hypothetical: the runaway reproduced in yog bl-d023 re-dispatched
# every ~3 seconds per generation and was ended by an operator at
# depth 4, not by the harness. Unlike a token or wall cap, a depth
# ceiling cannot cut off legitimate long work; it refuses only a shape.
#
# Why 5. Depth counts dispatches from the root, which is depth 0, and
# max_depth is the deepest ALLOWED depth: a dispatch that would land a
# child deeper is refused before the fork, so the deepest agent a
# dispatch can create sits at exactly max_depth. Five is therefore a
# root plus five levels of delegation. Observed practice in real fleets
# is two or three levels; five leaves two clear levels of headroom above
# it, while a runaway recursion crosses it in seconds.
#
# The cost, stated rather than hidden: the fork gate makes no
# distinction between a dispatch the model asked for and one the
# workflow ran, so an agent sitting AT max_depth cannot fork a compactor
# either — its compaction checkpoint is refused and the branch steps on
# uncompacted (ARCH §6 "One gate, every dispatch"). That is the second
# reason the number sits above ordinary practice rather than at it: the
# band that cannot compact should be a band nothing ordinarily reaches.
#
# Severable, and already severed for one consumer: delete the two lines
# below and the tree is unbounded on every axis again — config deleted,
# no code touched. yog does exactly that, stripping any top-level
# budgets: block from every workspace it manages at every start (yog
# bl-56af), because a seat holds a dollar ceiling at a conversation's
# birth and an operator watching the board, and two ceilings over one
# concern is the second representation that drifts. So this default
# binds the plain-litany operator — whose fleet nobody is watching,
# which is exactly whose fleet it was filed about.
budgets:
  max_depth: 5