yog 0.0.1

yog: a balls-oriented session manager for lernie loops (egui frontend)
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
# yog — The paved-path story ladder

**Benchmark: OpenAI Codex.** Zero-to-running is a login flow plus Enter in a
text box, and it *ends with the reply streaming* — a spawned process is not
the payoff. This document is the acceptance ladder for that path: user
stories escalating in operator skill, each rung adding features **without
burdening the rung below** — the S0 user never meets a concept from S3.
DESIGN.md remains the architecture authority; this file binds *experience*
to it and enumerates the integration tests that prove each story. A story is
done when its tests pass against the fake substrate **and** the flow works
against the real one.

Vocabulary is DESIGN §1's (no "session"): a conversation is a root agent in a
workspace; starting one is *prompt into the focused workspace, created if
none exists* (§3.4).

Phase-1 premise, owned: yog shells to host binaries, so "yog + current
substrate binaries installed" is the entry bar; the toolchain pane names any
gap and its exact install/upgrade command (W5), and phase 2's exact-pinned
embedded crates dissolve the install story entirely (§16.4) — S9 is that rung,
the one that escalates *downward* by deleting this premise.

## Test harness (all stories)

Integration tests live in `tests/` and drive the dispatch layer — the same
`pub` functions the shell's click-glue calls (`start::*`, `actions::*`,
`AppModel`, the view-model modules) — never egui widgets (§11's split: glue
is thin and excluded; everything a click *calls* is covered). Substrates are
fake recorder script binaries injected as `Cli::new(path)` / `Deps{…}` at
the dispatch API (the `tests/editor_roundtrip.rs` idiom): each records
argv+env+cwd to a log and plays a canned stdout/exit per verb. (The
`LERNIE_BINARY`/`BL_BINARY`/`BZ_BINARY` env vars are production wiring,
covered by the existing `resolve_with` unit tests — tests never mutate
process-global env under the parallel runner.) Every story test asserts up
to three surfaces: the recorded spawns, the `ops.jsonl` trail, and the
derived view-model the shell would paint.

**Test→enabling-task map** (M6 §15; a test lands red-first *inside the
worktree of the task that turns it green* — written first, red in-worktree,
green at close; the precommit gate forbids landing red on main):

| Tests | Enabling task |
|---|---|
| S1-T2, S1-T3, INV-3, S0-T2 (seed-skip half) | Z7 (fixture; green today) |
| S0-T1, S0-T3 (abort half), S1-T1, S2-T1, S3-T1..T4, S3-T6, INV-1 | Z3 (with Z2 under it) |
| S0-T4 | Z6 |
| S0-T3 (ops-row + view-model halves), INV-2 | Z5 |
| S0-T5, S0-T6 | Z8 |
| S1-T4 | Z5 (view-model) |
| S3-T5, S4-T1..T3 | Z4 |
| S3-T7, S4-T4..T7 | Z10 (to file) — the board: `nav::{tabs,convs,group}` |
| S5-T1..T6 | Z11 (to file) — toolgate (Z6) + the three §9 editors |
| S6-T1..T5 | Z12 (to file) — `attention` (Y10) + the activity accessory |
| S7-T1..T5 | Z13 (to file) — the inspector altitudes (Y9/Y12/Y13/Y20) |
| S8-T1..T4 | Z14 (to file) — `world::{mod,hatch,marks}` (W1/W2/W4/W6) |
| S9-T1..T4 | W8–W11 (phase 2; each gated on its upstream, §16.7) |

Z10–Z14 are the ladder's remaining enabling tasks. Their rungs are **already
built** — the board landed with Z9/bl-de16, the editors with Y18–Y21, the
world with W1–W6 — so each task's deliverable is its row of tests, not a
feature: written red against the built surface, green at close, the same rule
as every other row. Their bl ids land in DESIGN §15 M6 as each is filed,
exactly as Z1–Z9's did.

## Real-substrate drive (the second done-bar half)

The intro's done-bar has two halves: a story is done when its tests pass
**against the fake substrate** (the `tests/` harness above) *and* **the flow
works against the real one**. The fake harness proves the dispatch layer in
isolation; nothing in it exercises the installed `yog` binary, a live egui
window, or a real model wire. This section is the second half's harness — a
graduated, repeatable drive of the *real* `yog` against the *real* `lernie`
wire, so "works against the real one" is a run you can re-run, not a claim.

**Scripts** (`scripts/drive/`, bash, no repo deps):

- `yogdrive.sh` — the seat primitive. It drives a real `yog` on an **isolated
  Xvfb display (`:99`)**, never the user's live seat: synthetic `xdotool`
  input on a live seat leaks keystrokes into the operator's own apps, so input
  is confined to `:99` (no compositor there — XTEST works natively).
  `launch <scratch>` spawns `yog` under `XDG_DATA_HOME=<scratch>` (with
  `WAYLAND_DISPLAY` unset, `DISPLAY=:99`), finds the window by pid, and prints
  `PID WID`; `shot`/`type`/`key`/`click`/`stop` are the verbs. Capture is
  `ffmpeg -f x11grab` to a PNG. The driven `yog` resolves from `PATH`, so a
  worktree build is driven by prefixing its `target/release` — the drive
  proves the build in hand, not whatever is installed.
- `stories.sh` — the story runner. `seed <scratch>` lays the world seed;
  `run <scratch> <out>` fires the S0/S1 beats and asserts on the three real
  surfaces (`ops.jsonl`, the workspace tree, the on-disk `messages/`), one
  PASS/FAIL line per beat, screenshots to `<out>` for visual review.

**The world seed (DESIGN §16.6 W3).** Before launch the runner copies the
ambient world's `world/lernie/models.yaml` (which carries a `gpt-5.4` codex
model entry) and `world/lernie/template/providers.yaml` (codex in both the
worker and compactor roles) into the scratch world, and `yogdrive.sh` symlinks
the scratch `brazen/credentials` to the ambient one (brazen config/creds stay
shared, §16.2). The `models.yaml` *is* lernie's seeded marker, so a seeded
world skips `lernie prime` — the S0-T2 general path with the seed present
(§3.4), not a bootstrap branch.

**Beats driven** (the P0 rungs — S0+S1, the Codex bar):

- **S0 bare start** — launch to composer, type a goal, Enter: assert the argv
  trail is `lernie new` then a detached `lernie prompt` (no `prime`), the wire
  reply lands on disk, and the transcript renders it in the focused view.
- **S1 message-to-agent** — type into the focused conversation, Enter → a
  `lernie message` verb, its reply rendered.
- **S1 restart-equivalence** — kill and relaunch on the same world: the
  workspace, conversation, and ops trail re-derive from disk with **no spawn
  at idle** (INV-1 / I1: restart is re-read).
- **S1 prompt-into-existing** — Enter in the focused workspace's composer →
  `lernie prompt` only, a new root agent, **no re-mint** (no `new`/`prime`).

**Selection must not ride on pixels.** A hardcoded click coordinate is a
standing regression source against a live UI: a row inserted above a target
silently retargets every click below it — the 2026-07-24 run's S1 message beat
fired a *prompt* because the conversation list had grown its `recent | by ball`
toggle where the beat clicked. So the runner's steering rule is: **drive the
DESIGN §11 keyboard binding wherever one exists.** ↓/↑ step the flattened
conversation roster and land through the `focus_agent` path, setting the
focused workspace *and* the selected agent in one key — which is why the
runner's whole selection gesture is a single `xdotool key Down` where it was
once a workspace-tab click plus a conversation-row click; digits 1–5 pick an
inspector tab the same way. The clicks that remain are only the ones no
binding covers — giving a **text box** egui keyboard focus, since synthetic
typing goes nowhere unless a widget holds focus — and each is tagged
`CLICK (no binding)` in `stories.sh` with the widget it targets.

**Drivable next.** The harness asserts S0/S1 today. The rungs below reach the
real substrate with no new machinery beyond selection: S3's close (a ball in
the world, `bl close` stamped with the workspace name), S4's second
conversation and the grouped-by-ball toggle, S6's attention strip clearing on
focus and the activity chip's counts. S5's editors, S7's inspector tabs and
S8's hatches are drivable but need world fixtures the runner does not lay yet
(a primed project, a config branch). S9 is not drivable until phase 2 exists.

**Run logs** live in `docs/drive-logs/` — one file per real run, quoting the
wire replies verbatim and recording pass/fail per beat plus any UI observation
that becomes follow-up work.

## S0 — Stranger: first launch to first reply

Machine state: yog + substrate binaries installed; nothing configured — no
world, no workspaces, no projects, possibly no credentials.

1. Launch `yog`: the window opens to the composer, focused, with the wordmark
   and a one-line invitation. Greyed above the box: `You are <name>.` — the
   identity preview, pre-minted as a pure read (§3.3 as amended; nothing
   spawns before Enter except the read-only capability probes, W5). No
   wizard, no empty dead end, no setup screen.
2. Type a goal, hit Enter — **one box, one Enter**. On this one explicit
   action (I7) the world materializes: `lernie prime` seeds the nested home
   (skipped forever after — the seeded world is the general path with the
   seed present, §3.4/W3), the
   previewed name mints (re-derived at fire; the stamp is the truth, §3.3),
   `lernie new <root>/<name>` creates the workspace, and the goal — identity
   preamble harness-stamped, payload exactly as typed — fires as a detached
   `lernie prompt` with `YOG_NAME=<name>`, cwd `~`.
3. The start renders in flight (busy indicator per step), then **the reply
   streams into the focused view** — the root appears within a watch tick
   and its streaming tail paints as it grows (§5.1 #10). That is the Codex
   bar: text in, text out, nothing between.
4. Any step failure is a **rendered fact**: the ops line (argv, cwd, exit,
   stderr — §4.2 as amended) expands in the ops pane, and the originating
   surface shows ichor red with argv + stderr tail. That holds for a *driver*
   that dies too, not only a step yog itself ran: the conversation renders the
   §7.3 **no-response wound** — "driver produced no response" in ichor beside
   the step and bannered on the conversation — instead of a `0 attempts · 0 tok`
   row that reads as a quiet step, while the driver's own stderr rides its
   per-spawn sink into the ops trail (§8.1/§13.3). The typed goal survives
   in the composer (the draft is RAM until *sent* — a failed start has not
   sent it). A substrate failure aborts *before* any `bl` mutation (§8.1
   order) — no half-committed state, ever.
5. If the model call needs credentials, the failure surfaces as derived
   agent state (§13.3 — the detached driver has no captured stderr): the
   auth-failed step renders with a **Login** affordance one click away,
   which runs `bz --login --provider <row>` streamed-piped (§8.3 as
   amended) and paints the device code/URL lines live. Credentials stay
   bz's; yog renders, never stores.
6. A stale-but-present substrate is caught **up front**: the capability gate
   (W5 as amended) probes every driven verb before first render; a Mismatch
   names the missing verb *and its remediation command*, and every mutating
   dispatch refuses with that same verdict while reading/rendering
   continues.

Burden check: the user never sees "world", "ball", "worktree", "claim", or a
name picker. One box, one Enter. The identity preview is one grey line.

Tests:
- **S0-T1 bootstrap-bare-start**: empty world, submit → argv sequence
  `lernie prime`, `lernie new <names-root>/<a-b>`, detached `lernie prompt`
  carrying the §3.3 stamped identity preamble + the typed text, env
  `YOG_NAME=<a-b>`, cwd `~`; the pre-submit view-model already carried the
  `You are <a-b>.` preview; ops trail complete.
- **S0-T2 seeded-skip**: seeded world (marker present) → no `prime` spawn.
- **S0-T3 seed-failure-surfaces**: fake `lernie prime` exits 2 → **no `bl`
  spawn recorded**; the error carries argv + stderr; the ops row holds
  stderr; the failure view-model renders argv + stderr tail; the draft
  remains.
- **S0-T4 capability-gate**: fake lernie whose `prime --help` exits 2 → the
  `pub` verdict names `prime` + remediation; `start::prepare` refuses with
  it; read derivation is unaffected.
- **S0-T5 login-stream**: fake `bz --login` prints a device code → the
  streamed-piped runner's view-model carries the lines verbatim; outcome
  line lands in ops at exit.
- **S0-T6 login-detection**: fixture workspace with an auth-failed step on
  disk → the step's view-model carries the Login affordance.

## S1 — Returner: the conversation continues

1. Relaunch yog: the workspace, its agents, and every transcript render from
   disk (I1 — restart is re-read; nothing to restore, nothing to resume).
   Quitting was always safe: drivers are setsid-detached (§7.3) — closing
   the window never kills a running agent, and the next launch re-derives
   them mid-flight. A second instance alongside converges identically (I0).
2. Enter in the composer → a new root in the focused workspace. Re-opening
   is the same gesture as opening (§3.4) — no resume concept.
3. Select an agent and type → `lernie message` (the resume gesture, §8.2).
   Stop (± children) and Scan work per §8.2. (Platform caveat, §10: `lernie
   stop` is /proc-based — on macOS the Stop failure renders verbatim; the
   fix is upstream lernie's, tracked there.)
4. The streaming tail keeps painting for any in-flight agent (S0 step 3's
   surface, same derivation).

Burden check: identical gesture set to S0; the second launch adds zero steps.

Tests:
- **S1-T1 prompt-into-existing**: focused workspace, submit → `lernie
  prompt` only (no mint, no `new`, no `prime`).
- **S1-T2 restart-equivalence**: two AppModels over the same disk derive
  identical view-models — workspace tabs and per-workspace trees (I0).
- **S1-T3 agent-verbs**: message/stop/scan argv per §8.2, outcomes in ops.
- **S1-T4 reply-streams**: fixture workspace with an open `response.json` →
  the transcript view-model carries the live tail, visually distinct.

## S2 — Director: point the conversation at a directory

1. The composer grows one optional affordance: a work target (a directory).
   Submitting with a path appends the §3.3 target preamble verbatim and sets
   driver cwd to the path. The directory need not be a bl project (§3.4).

Burden check: the affordance is ignorable; empty target = S0/S1 exactly.

Tests:
- **S2-T1 path-rung**: submit with dir → goal carries the target preamble,
  spawn cwd = dir, no `bl` spawns.

## S3 — Tracker: work rides a ball

Projects enter yog's world by being primed *in the world* (shared store
branch, §16.3). v1 keeps `bl prime` out of the UI (§8.3): the paved interim
is `yog exec bl prime` in the repo — and the left panel's empty balls
section renders exactly that hint (owned by Z4; tested by S3-T5).

1. A ready ball's ▶ Start claims it `--as <workspace-name>` (§3.2) — *after*
   the workspace exists (§8.1 order) — and the composer prefills ball title,
   body verbatim, and the worktree preamble; driver cwd = the work worktree.
2. A new ball from the composer is `bl create` then the existing-ball path —
   the new→existing transition is the convergence (§8.1).
3. A ball already claimed by a local workspace name re-plans as a prompt
   into that workspace — resume, never a second mint (§8.1).
4. The §3.5 join table renders every row state; `bl close` surfaces gate
   failures verbatim.
5. Shipping is one verb: Close stamps `bl close <id> --as <the ball's bound
   workspace name>` (§8.2's identity rider — the claimant delivers its own
   ball; the operator's `$USER` never appears). The ball file is then gone, so
   the row re-derives as **delivered** from the closed listing's claimant
   (§3.4) — grouped under the same workspace, nothing stored to remember it.

Burden check: with zero projects in the world, no ball UI exists at all —
the §3.5 "unassigned workspace" row is the general case (no ball column).

Tests:
- **S3-T1 ball-rung**: ready ball → `bl claim <id> --as <name>` after
  `lernie new`, goal carries title/body/worktree preamble, cwd = worktree.
- **S3-T2 new-ball-converges**: create → re-planned as existing; one claim.
- **S3-T3 resume-not-remint**: ball claimed by local name → prompt into
  that workspace; no mint, no claim.
- **S3-T4 close-gate-verbatim**: fake `bl close` fails its gate → stderr
  verbatim in ops + view-model.
- **S3-T5 empty-project-hint**: zero projects in the world → the balls-
  section view-model carries the `yog exec bl prime` hint.
- **S3-T6 abort-before-claim**: ball rung with a fake `lernie prime` (or
  `lernie new`) failing → **no `bl create`/`bl claim` recorded** — the §8.1
  load-bearing order proven on the rung that has a `bl` mutation to abort
  (S0-T3's no-`bl` assertion is vacuous on the bare rung).
- **S3-T7 close-stamps-and-delivers**: Close on a bound ball → `bl close <id>
  --as <bound workspace name>`, cwd = the project; with the ball then absent
  from the live set, the join re-derives it delivered under that same
  workspace and the conversation's badge turns ash (§3.5).

## S4 — Organizer: spheres, corrals, and the board

1. + New workspace (the deliberate sphere-wall verb, §3.4/§11) mints a
   second name — the first time the user ever meets one.
2. Assign / move / release balls (§8.2) with enablement from the join state;
   bound balls render in the focused workspace's balls section (§11 — the
   full per-project ball views return in the ball-views wave).
3. Claimed-elsewhere renders the claimant verbatim; delivered balls group
   under their claimant on demand (§3.4). Workspace retirement is lernie's
   retention (30-day default, §3.4) — yog deletes nothing and grows no
   delete verb; clutter ages out where the data lives.
4. **The conversation list is the board.** With many roots in flight, each row
   carries its subtree's aggregated state, its age, and — when the
   conversation was started from a ball — that ball's badge, coloured by the
   ball's own §3.5 row state: bound green, delivered ash, blocked or
   claimed-elsewhere brazen, orphaned ichor. Sort is §6's: attention >
   running > recency.
5. **The badge is honest about what it cannot know** (§3.2's two altitudes).
   A conversation started bare or by path shows *no* badge; a stamped id this
   machine's join does not know shows the id **uncoloured** — the stamp is
   truth, the colour is the join's when it has one. And a ball an agent
   claimed mid-conversation shows in the *workspace's* bound-ball rows only:
   no fact records which conversation picked it up, so no badge is invented.
6. **One toggle re-organizes the same rows**: `by ball` partitions the sorted
   list so each stamped ball heads its conversations, the stamp-less ones
   trailing in a single unassociated group; `recent` is the flat default. It
   is a re-ordering of rows already on screen — no row appears, disappears, or
   changes meaning — and the toggle itself is viewport ephemera (§13.1),
   stored nowhere.
7. **The tab strip is the sphere wall**: one tab per named workspace, pinned
   first then name order, each badged with its own attention count; foreign
   and replay workspaces live in the ⋯ overflow (real, but not regimes) and
   the overflow button carries their aggregate. Pinning hoists one out.

Burden check: S0–S3 never minted explicitly; nothing here changes their path.
A one-conversation workspace renders identically under both orderings, and a
world with no balls renders no badge column at all.

Tests:
- **S4-T1 new-workspace-verb**: explicit mint + `lernie new`; occupied set
  respected (names root + claimants).
- **S4-T2 assign-move-release**: argv per §8.2; enablement predicates refuse
  what balls would refuse.
- **S4-T3 join-rows**: one fixture per §3.5 row state; the balls section
  groups bound balls under their claimant workspace.
- **S4-T4 conversation-badges**: four conversations in one workspace — goal
  stamped with a bound ball, stamped with a delivered ball, stamped with an id
  the join does not know, unstamped → badge hues green / ash / **uncoloured
  id** / **none**; plus a ball claimed by the workspace with no stamp anywhere,
  asserted present in the workspace's bound rows and absent from every
  conversation row (§3.2's honest limit, never a fabricated row).
- **S4-T5 group-partition**: grouping is a pure stable partition of the sorted
  rows — group order = first appearance, within-group order preserved, the
  unassociated group emitted last and only when non-empty; flattening the
  groups returns the input rows unchanged.
- **S4-T6 board-order**: subtree aggregation (InFlight > Live > the root's
  settled state, with the §10 "?" suffix) and the §6 sort, over one fixture
  carrying an attention-flagged idle root, a running root, and two settled
  roots of different ages.
- **S4-T7 tab-strip**: pinned tabs hoist in pin order ahead of name order,
  each tab's badge is its own workspace's rollup, foreign and replay
  workspaces fall to the overflow, and the overflow button carries their
  aggregate attention.

## S5 — Operator: the tools and their configuration

1. Toolchain pane: per-tool capability verdicts against the normative
   driven-verb list (W5). A Mismatch **names the missing verb and the exact
   command that fixes it**; a capable tuple renders no remediation text at all.
   Login per provider row (§8.3) streams `bz --login`'s device code and URL
   live, with the exact-command fallback if the piped flow exits non-zero.
2. **The gate has teeth.** On a Mismatch every mutating dispatch refuses with
   that same verdict — the start flow and the per-agent verbs alike, because
   the gate is consulted in the dispatch layer, not in click-glue — while
   reading and rendering continue untouched. A gate only some verbs honor is
   not a gate.
3. Config editing per §9, three surfaces, one gesture: brazen's `config.toml`
   as raw TOML, lernie's global `models.yaml`/`workflows/*` as raw text, and a
   workspace's config branches — the last written by `lernie config` driven
   with yog itself as `$EDITOR` (§9.3), because that verb is the only lawful
   writer of `config/*`.
4. **One discipline for all three**: load → RAM buffer → Apply = stage →
   validate where a validator exists → hash-guard → atomic rename. A malformed
   brazen config cannot land — `bz` itself rejects the staged file and the
   draft survives in the box. A file changed underneath (the other instance,
   or vi) refuses the Apply and says so rather than overwriting.
5. Credentials are never yog's: presence renders, contents never do, and the
   only write path is bz's own login.

Burden check: nothing here sits on the S0 path — the toolchain pane and Config
are left-panel entries the stranger never opens, and a healthy machine renders
neither a verdict nor a remediation string.

Tests:
- **S5-T1 toolchain-verdicts**: fake `lernie` whose `scan --help` exits 2 →
  the verdict names `scan` and carries its remediation command; the
  fully-capable tuple yields a green row per tool with an empty remediation.
- **S5-T2 gate-refuses-mutations**: with that same Mismatch, `start::prepare`
  and every `actions::verbs` dispatcher refuse carrying the verdict, the
  refusal lands as a `["yog-step",…]` ops row (never a silent no-op), and the
  derived view-models still build from disk unchanged.
- **S5-T3 brazen-validate-rejects**: a staged buffer whose `bz --config <temp>
  --dump-config` exits non-zero → the destination file is byte-identical
  afterwards, no temp survives, stderr renders, the draft is kept.
- **S5-T4 hash-guard**: the file changes on disk after load → Apply refuses
  and names the drift; after a reload the same Apply lands. Asserted once per
  editor (brazen, lernie-global, config branch) — one discipline, three doors.
- **S5-T5 config-branch-shim**: `lernie config <ws> <name>` is spawned with
  `EDITOR=<yog> --editor-apply` and `YOG_EDIT_SRC=<stage>`; the shim copies
  **only** the drafted files (a `descriptions/` file already in the checkout
  is untouched), an empty diff surfaces as "no change", and the staging dir is
  gone at exit (`tests/editor_roundtrip.rs` is this test's seed).
- **S5-T6 login-rows**: provider rows derive from the §5.1 #20/#21 config
  reads; the streamed runner's view-model carries the device lines verbatim
  (shared with S0-T5) and a non-zero exit falls back to the exact command.

## S6 — Triager: what needs you, and what went wrong

Machine state: several workspaces, many conversations, some finished, one
dead.

1. The attention strip answers "does anything need me?" with no click: totals
   per signal kind across every workspace, each tab badged with its own count,
   and one jump-to-next control that walks them in derived order.
2. Attention is a predicate over disk (§6), never a stored flag: unacked
   notify, stopped-without-abandoned, budget, conflicted — and pending mail
   nobody is driving. That last one is deliberately **not** silenceable: a
   stall you can dismiss is a stall you will miss.
3. **Acknowledging is focusing.** Landing on a conversation records the
   current evidence oids as seen. The acknowledgement *converges* — the second
   instance stops flagging it too — while each instance's own focus stays its
   own (§13.1). The mark is lernie's; the acknowledgement is yog's.
4. **Acknowledging clears the signal, not the fact.** A dead conversation goes
   quiet in the strip and keeps its state badge, and an auth-shaped death
   keeps its inline Login one click away.
5. The other half is the activity accessory: one chip, `activity · N ops ·
   M ⚠` (ichor when M > 0), expanding to the ops tail; any row expands to
   argv, cwd, exit and stderr verbatim — including actions that never
   spawned. A trail that hides *why* is not a trail. yog's own plumbing lives
   there and never between conversation messages.

Burden check: at S0 the strip reads "nothing stirs" and the chip reads
`activity · 0 ops`. The rung adds no gesture — it adds meaning to two surfaces
already on screen.

Tests:
- **S6-T1 attention-predicates**: one fixture per §6 rule 1–5, each asserted
  true, then false once the matching watermark is written — except rule 5,
  which stays true across every watermark and self-clears only when a driver
  takes the lock.
- **S6-T2 ack-converges**: two AppModels over one `ui.json`; acknowledging in
  A clears the signal in B after the adopt, and B's focus and scroll are
  unmoved (§13.1's line, tested from both sides).
- **S6-T3 failure-stirs-then-settles**: a fixture whose latest step failed →
  the conversation stirs the strip through rule 2; after the ack the strip is
  quiet while the state badge and the Login affordance still render.
- **S6-T4 rollups-and-jump**: workspace rollup = max over its agents, strip
  totals sum across workspaces, and jump-to-next is total over the derived
  order — it wraps and never sticks on the current row.
- **S6-T5 activity-chip**: an ops fixture with k failures → the chip label
  carries `N ops · k ⚠` in ichor; an expanded row yields argv/cwd/exit/stderr
  verbatim, and a line clipped by the §4.2 cap renders its `truncated` marker
  rather than silently short bytes. (INV-2 is the mechanized half: *every*
  dispatch error has a row.)

## S7 — Forensic: every byte inspectable

Machine state: a conversation that did something surprising.

1. The center is transcript-first. Only a conversation **with children** grows
   the compact descent tree — one selectable row per member, and selecting one
   is the same acknowledgement gesture as anywhere else (§6).
2. Five tabs answer five questions: what was said (Transcript), what ran
   (Steps), what is waiting (Inbox), what was written (Files), what policy
   governed it (Config — "policy frozen at `<short-oid>`", derived from git
   ancestry, not from a stored pointer).
3. Steps drill into meta / request / response / staging and per-tool input and
   output through one collapsible JSON widget, and **every tab has a Raw
   toggle** showing the verbatim bytes. Nothing is summarized away; a file yog
   cannot parse renders as an error row rather than vanishing.
4. Budget is a fold over the subtree's usage events. Limits render as the raw
   `workflow.yaml` text, because yog owns no YAML parser and will not become a
   second authority on one.
5. Pending mail explains `✉n`, and Flush is `lernie scan`. A committed
   `tool_use` with no `tool_result` renders "tool in progress" — a fact read
   off the transcript, not a guess about a running process.

Burden check: five tabs behind one selection; the S0 stranger lands on
Transcript and never leaves it. The digit keys 1–5 are those same five tabs —
one concept, two ways in.

Tests:
- **S7-T1 tabs-dispatch**: one fixture per tab; each builds from disk alone,
  and the Raw toggle yields the underlying file's bytes unaltered.
- **S7-T2 step-drilldown**: a step carrying request/response/staging and a
  tool call → the jsonview row tree matches the parsed value; a malformed step
  file renders an error row and the sibling tabs still build.
- **S7-T3 budget-fold**: usage across a root and two children folds to the
  subtree total shown in the header; the limit line is raw text, never parsed.
- **S7-T4 inbox-and-progress**: deposits parse from/deposited_at/epitaph,
  `✉n` equals the count, Flush dispatches `lernie scan <ws>`; a committed
  tool_use with no result renders "tool in progress".
- **S7-T5 tree-only-with-children**: a single-agent conversation renders no
  descent tree; adding a child grows one row per member with aggregated
  badges, and selecting a member retargets the inspector.

## S8 — Neighbour: yog beside your own shell

Machine state: an operator who also drives `bl` and `lernie` by hand.

1. **yog's substrate state is yog's.** One nested world under
   `$XDG_DATA_HOME/yog` (§16.2) holds the nested lernie home, the nested balls
   state and yog's own two artifacts. `rm -rf` that directory and the world is
   gone, with the ambient lernie, balls and brazen untouched — severability
   you can actually run.
2. **The overlap is chosen, not accidental.** The balls store branch is shared
   by default, so a ball yog claims shows up in the operator's own `bl list`
   and vice versa; brazen's config, credentials and model cache are the
   ambient ones — one `bz`, one login, no second copy to keep in step.
3. `eval "$(yog env)"` drops a shell *into* the world; `yog exec <cmd…>` runs
   one command there. That is how a project is primed into the world in v1 —
   and the empty balls section renders exactly that command rather than
   growing a button (§8.3).
4. An operator who wants no marks in the shared store flips a per-project knob
   that **drives `bl conf`** (`task-remote none`, or a custom task-branch).
   The pane renders the trade: stealth makes yog's claims invisible to the
   ambient `bl list`. There is no yog config file — the policy lives in balls'
   capability, so removing the knob deletes config, not code.
5. Two yogs side by side render the same data, acknowledgements included (I0);
   only focus, scroll and unsent drafts stay per-instance.

Burden check: an operator who never opens a shell meets none of this. The
world is *composed* at startup, which is pure; it is *materialized* only by
the same first Start that S0 already describes (I7).

Tests:
- **S8-T1 world-compose**: the composed env overrides exactly `LERNIE_HOME`
  and `XDG_STATE_HOME`, leaves `XDG_DATA_HOME`/`XDG_CACHE_HOME`/
  `BRAZEN_CONFIG` ambient, and re-deriving the anchor through the world env is
  a fixed point (no bootstrap special case).
- **S8-T2 every-spawn-nested**: every dispatched verb — the start steps, the
  per-agent verbs, the detached prompt, the `bz` runners — carries the world
  overrides, asserted through the world `Cli` so a new call site inherits them
  by construction; the dir yog watches and the dir a spawned `bl` writes are
  one path.
- **S8-T3 hatches**: `yog env` prints exactly the export lines the world
  composes, and `yog exec <cmd…>` runs its argv under them with the requested
  cwd — both pure entrypoints of the yog binary, neither a substrate spawn.
- **S8-T4 marks-knob**: stealth dispatches `bl conf task-remote none` in the
  project and the custom-branch setting its `bl conf` counterpart; **no
  yog-owned file is written by either**, and the rendered trade text is
  derived from the resulting `bl conf` state.

## S9 — Settler: the only thing you installed is yog

**This rung escalates downward.** Every rung above adds skill; this one
subtracts a prerequisite, and its acceptance is that S0's entry bar
disappears. It lands with phase 2 (§16.7) and only as fast as its upstreams.

Machine state: a machine with yog and nothing else — no `lernie`, no `bl`, no
`bz` on `PATH`.

1. Launch yog and S0 runs **exactly as written**. Nothing is probed for
   absence because nothing is shelled to for the substrate work yog owns:
   balls, brazen and lernie are exact-pinned crates and the pin *is* the
   version (§16.5).
2. The toolchain pane loses its remediation column. There is no install
   command to show, so the phase-1 capability gate is deleted with the phase
   it belonged to — S0-T4 and S5-T1/T2 are **removed, not skipped** (§16.4).
3. Agents working on yog's world get `lernie-tool-bl` — yog re-execing itself
   against the embedded balls crate and the nested roots — so an agent's
   `bl claim` computes the world's paths and shares yog's own implementation.
4. **Identity stops being an instruction.** The shim stamps `--as $YOG_NAME`
   on any verb the caller left unstamped, and §3.3's phase-1 preamble sentence
   is deleted with it: the agent cannot forget what it never had to remember.
5. The driver is yog re-execed (W11), and nothing about the concurrency model
   moves: drivers are still flock-holding processes, plugin dispatch is still
   a subprocess.

Burden check (inverted): the intro's phase-1 premise is retired. S0's bar
falls from "yog + current substrate binaries" to "yog"; no rung above gains a
gesture, and every story test above must pass unchanged against the embedded
substrates — that equality is this rung's real assertion.

Tests:
- **S9-T1 no-host-binaries**: the story suite runs with `lernie`/`bl`/`bz`
  absent from `PATH` → S0-T1 and S1-T1 pass unchanged; the gate tests are
  gone from the suite rather than ignored.
- **S9-T2 tool-shim-argv**: `lernie-tool-bl <verb> …` honors bl's argv
  contract and resolves the **nested** clone and worktree roots, never the
  ambient ones — the §16.4 correctness argument, asserted as paths.
- **S9-T3 identity-injected**: an unstamped verb through the shim gains
  `--as $YOG_NAME`; a verb that carries its own `--as` is passed through
  untouched; the composed preamble no longer contains the stamp instruction.
- **S9-T4 pin-is-the-version**: the toolchain pane renders the embedded crate
  versions as plain facts from the lockfile, with no remediation affordance
  anywhere — version skew is unrepresentable, not merely unlikely.

## Invariant tests (every rung)

- **INV-1 idle-is-pure**: constructing AppModel + N ticks with no user
  action performs **no mutating spawn and no substrate write** (I7). The
  permitted read spawns are enumerated: the W5 capability probes (`--help`/
  `--version`) and the §7.2 fetch cadence (`bl list --json` and kin). The
  fake binaries record everything; the test asserts nothing outside that
  read set.
- **INV-2 no-swallowed-errors**: mechanized two ways — every dispatch-layer
  `Err` lands in `ops.jsonl` per the amended §4.2 (spawn failures included,
  non-spawn steps as `["yog-step",…]` rows), and `grep -r eprintln
  src/shell/` finds **zero** occurrences. *(bl-a649 closed the one error the
  dispatch layer never saw: a detached `lernie prompt` that launches cleanly and
  then dies has no `Err` to log — its stderr goes to a per-spawn sink file
  (§8.1/§13.3) which the ops sweep folds into the `-2` row, making the death a
  rendered failure instead of a prompt that "does nothing".)*
- **INV-3 convergence**: re-running any start plan after a mid-plan kill
  converges (§8.1 idempotent-or-convergent steps).

## Priority (this epic)

P0 = S0+S1 (the Codex bar), P1 = S2+S3, P2 = S4+S6 (the many-things regime:
you cannot organize a board you cannot triage), P3 = S5+S7+S8, P4 = S9 (phase
2, each beat gated on its own upstream, §16.7). Implementation
rides §15 M6: Z2 (binding), Z3 (start rungs — `start::plan` is the one
planner, `prepare()` executes it), Z4 (verbs + hints), Z5 (error surfacing),
Z6 (capability gate), Z7 (test fixture), Z8 (login flow), Z9 (tabs +
conversation-first) — bl-tracked; ids recorded in §15 M6 as each is filed.
Z10–Z14 (the test map above) carry the S3–S8 story tests over surfaces that
are already built; S9's carriers are W8–W11.