cloud/lib.rs
1//! @yah:relay(R040, "yah-cloud Phase 1: mirror bootstrap (Hetzner driver, Podman runtime, yah-yubaba)")
2//! @yah:status(review)
3//! @yah:parent(Q058)
4//! @yah:next("yah-A track (8 tickets) — substrate for noisetable cloud + future managed camps; gates noisetable C6 (E2E) and D1–D5 (corpus migration)")
5//! @yah:next("Adjacent relays: R019 (SSH-RPC remote camps), R032 (yah-agentd), R034 (identity registry — hostkeys overlap)")
6//! @yah:next("Reparented under Q058 cloud quest 2026-05-07. Largely superseded by R091 (yubaba orchestration + integration testing) and R092 (managed-camps CLI + .yah/cloud/ schema) — net new work should land on those relays. Remaining unique R040 scope (Cloudflare Tunnel cookbook F15, pg-on-mesh recipe F16, mesh-promote phasing F18-F19) stays here as historical/operational reference; consider archiving once R091+R092 reach review.")
7//! @yah:next("Refinement open questions: subcommand framework alignment with yah arch/board, where hostkey fingerprints live, reverse proxy choice. Crate-layout question soft-answered by this stub being a sibling crate; A1 may still refactor into a sub-module of an existing crate if desired (move the //! block + delete this stub).")
8//! @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md)
9//! @arch:see(architecture/yah-managed-camps-topology.md)
10//!
11//! @yah:ticket(R040-F1, "A1: .yah/cloud/ schema + parser (MachineConfig, MirrorConfig, ServiceConfig)")
12//! @yah:at(2026-05-05T00:29:11Z)
13//! @yah:assignee(agent:claude)
14//! @yah:status(review)
15//! @yah:phase(P1)
16//! @yah:parent(R040)
17//! @yah:next("Per-file TOML layout: machines/<name>.toml, mirrors/<camp>.toml, services/<name>.toml — diffability + parallel editing")
18//! @yah:next("Sketch types in arch doc; refine to fit actual yah crate structure (this stub crate is one option; folding into an existing crate is the other)")
19//! @yah:verify("cargo test -p cloud config::tests round-trips all three config types")
20//! @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md)
21//! @yah:handoff("Schema + parser complete: MachineConfig, MirrorConfig, ServiceConfig, BucketSpec, PortMapping in crates/yah/cloud/src/config.rs. Per-file TOML layout via CloudConfig::load. MachineConfig::save for write-back (A4 hostkey). 6 tests pass (round-trips + tempdir integration + save/reload). tempfile dev-dep added. No warnings.")
22//! @yah:next("A2: yah cloud subcommand stub — wire into yah CLI using same clap/subcommand pattern as yah board/arch. Commands: machine {status,provision,destroy}, mirror {show,status}, service {deploy,status}, agent {ping,services,logs}. All no-op except status/show which dump parsed .yah/cloud/ config. Verify: yah cloud mirror show noisetable prints declared regions/services.")
23//!
24//! @yah:ticket(R040-F3, "A3: Hetzner driver — MachineProvider trait + hcloud impl (server, bucket, status, destroy)")
25//! @yah:at(2026-05-05T00:29:11Z)
26//! @yah:assignee(agent:claude)
27//! @yah:status(review)
28//! @yah:phase(P1)
29//! @yah:parent(R040)
30//! @yah:next("API token from HETZNER_API_TOKEN env; document in .yah/cloud/SECRETS.md")
31//! @yah:next("Gotcha: Hetzner has TWO APIs (Cloud and Robot) — Phase 1 is Cloud only; Object Storage is S3-compat, separate surface")
32//! @yah:next("Verify before A6: Hetzner Hillsboro + Ashburn Object Storage both GA; CPX-22 available in PDX/IAD/FSN")
33//! @yah:verify("cargo test -p cloud provider::hetzner -- --ignored (gated; needs token)")
34//! @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md)
35//! @yah:handoff("MachineProvider trait + HetznerDriver impl complete in crates/yah/cloud/src/provider/. Cloud API (create_server, server_status, destroy_server) uses HETZNER_API_TOKEN via reqwest+bearer. create_bucket uses Hetzner Object Storage S3-compat with AWS Sig V4 signing (HETZNER_S3_ACCESS_KEY + HETZNER_S3_SECRET_KEY). Location enum maps pdx/iad/fsn to hetzner cloud IDs and S3 endpoints. 9 unit tests pass; 2 integration tests behind #[ignore] gate. .yah/cloud/SECRETS.md created documenting all three token types.")
36//! @yah:next("A2: wire yah cloud subcommand stub into CLI (yah board/arch pattern). All commands no-op except machine status + mirror show which dump parsed .yah/cloud/ config. Verify: yah cloud mirror show noisetable prints declared regions/services.")
37//!
38//! @yah:ticket(R040-F4, "A4: cloud-init YAML + yah cloud machine provision (Debian + Podman + Tailscale + yah-yubaba)")
39//! @yah:at(2026-05-05T00:29:11Z)
40//! @yah:status(review)
41//! @yah:phase(P1)
42//! @yah:parent(R040)
43//! @yah:next("Template at .yah/cloud/cloud-init/mirror.yml; renderer substitutes YAH_WARDEN_BASE64 + HEADSCALE_PREAUTH_KEY + TAGS per machine")
44//! @yah:next("Hostkey generated ON the machine (ssh-keygen in cloud-init); private half never touches dev machines")
45//! @yah:next("Provision command polls cloud-init done, reads back hostkey fingerprint via yah-yubaba /identity, writes back to machine.toml")
46//! @yah:next("Surface /var/log/cloud-init.log and /var/log/cloud-init-output.log on failure")
47//! @yah:verify("yah cloud machine provision noisetable-pdx-1 --dry-run prints rendered cloud-init YAML")
48//! @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md)
49//! @yah:handoff("Cloud-init template + renderer + provision command landed. Template at .yah/cloud/cloud-init/mirror.yml (canonical) with embedded copy at crates/yah/cloud/templates/mirror.yml (drift-test in cloud_init::tests::embedded_template_matches_workspace_canonical). Renderer at crates/yah/cloud/src/cloud_init.rs substitutes {{KEY}} (no-spaces) for MACHINE_NAME, YAH_WARDEN_BASE64, HEADSCALE_PREAUTH_KEY, TAGS; doc comments use {{ KEY }} (with spaces) so they survive rendering. find_unsubstituted() trips on any unsubstituted real placeholder. Provision orchestrator at crates/yah/cloud/src/provision.rs decoupled from concrete provider (takes &dyn MachineProvider). CLI wired in app/yah/cli/src/cloud.rs: yah cloud machine provision <name> [--dry-run] [--yubaba <path>]; --dry-run with no flags uses placeholders for the yubaba binary + preauth key (with NOTE lines explaining); live path requires --yubaba + HEADSCALE_PREAUTH_KEY + HETZNER_API_TOKEN. 20 unit tests pass (8 new for cloud_init + provision). Vocab: per-node infra daemon is yah-yubaba (was yah-mirror-agent — renamed 2026-05-01 because mirror is a deployment, not a node).")
50//! @yah:next("A8 yah-yubaba /identity endpoint: poll cloud-init done after Hetzner accepts the create_server call, GET /identity for hostkey fingerprint, write back to .yah/cloud/machines/<name>.toml via MachineConfig::save(). Wiring stub already in handle_provision after rt.block_on — replace the next: print with the real poll loop.")
51//! @yah:next("Optional polish: surface /var/log/cloud-init.log + /var/log/cloud-init-output.log on cloud-init failure (probably via the yah-yubaba /diagnostics endpoint when A8 lands, since direct SSH from the dev machine isn't part of Phase 1).")
52//!
53//! @yah:ticket(R040-T5, "A5: yah cloud machine status — drift report (declared vs Hetzner vs agent health)")
54//! @yah:at(2026-05-05T00:29:11Z)
55//! @yah:status(review)
56//! @yah:phase(P2)
57//! @yah:parent(R040)
58//! @yah:next("Drift categories: missing-server, wrong-machine-type, missing-bucket, hostkey-fingerprint-mismatch, missing-co-tenant, agent-unreachable, service-unhealthy")
59//! @yah:next("Exit non-zero on drift unless --quiet; bare 'yah cloud machine status' dumps all declared machines")
60//! @yah:next("Can land any time after A4 (parallel-safe with A6/A7/A8)")
61//! @yah:verify("Pre-A6: reports 'not provisioned' for each declared machine")
62//! @yah:verify("Post-A6+A7+A8: all three report 'in sync'")
63//! @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md)
64//! @yah:handoff("Drift module landed at crates/yah/cloud/src/status.rs (DriftFinding + MachineReport + collect_machine_report). MachineProvider trait gained find_server_by_name (GET /servers?name=) + bucket_exists (S3 HEAD) — Hetzner impls share a refactored sign_s3_empty_body helper for PUT/HEAD signing. CLI: `yah cloud machine status [name] [--quiet]` queries Hetzner, prints declared+live state with per-finding [drift]/[note ] tags, and bails non-zero on real drift unless --quiet. Soft findings (missing creds, A8-not-landed, bucket-uncheckable) are surfaced as notes only. Pre-A6 verify: smoke-tested against this workspace's noisetable-pdx-1 declaration with no token — reports 'unknown' + provider-error note, exit 0. 7 new drift unit tests pass via in-memory FakeProvider; 27 total pass in crates/yah/cloud.")
65//! @yah:next("A8 will plug in real agent-ping (yah-yubaba /health + /identity). The HostkeyFingerprintMismatch + AgentUnreachable variants are wired through the report rendering — A8 just needs to flip the AgentUnreachable emit-site in collect_machine_report() from the always-emit pre-A8 stub to a live probe.")
66//!
67//! @yah:ticket(R040-T6, "A6: provision noisetable-{pdx,iad,fsn}-1 (operational; yah dogfoods on noisetable's mirror, no yah-cloud SaaS)")
68//! @yah:at(2026-05-05T00:29:11Z)
69//! @yah:status(review)
70//! @yah:phase(P2)
71//! @yah:parent(R040)
72//! @yah:verify("yah cloud machine status --path /Users/user/ss/noisetable — all three 'in sync'")
73//! @yah:verify("aws s3 ls --endpoint-url <hetzner-region-endpoint> s3://noisetable-assets-<region>-1 — all three list")
74//! @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md)
75//! @yah:handoff("Code path complete across R040-F1..F8 + cloud-init runcmd alignment. Gated on operator runbook (release tag → live provision against Hetzner). Open repo-config decision (noisetable side, not yah): noisetable/.yah/.gitignore line 3 ignores '/cloud' wholesale — blocks tracking machines/, mirrors/, SECRETS.md. 30 cloud-crate + 17 yah cloud:: tests pass.")
76//! @yah:next("Operator runbook lives in events.jsonl history; the strategic next is the live mirror-status verify (`yah cloud machine status --path /Users/user/ss/noisetable` → all three 'in sync') which is gated on the release-tag + provision steps.")
77//!
78//! @yah:ticket(R040-F7, "A7: yah cloud service deploy — Podman compose generation + Cloudflare front + tier isolation")
79//! @yah:at(2026-05-05T00:29:11Z)
80//! @yah:assignee(agent:claude)
81//! @yah:status(review)
82//! @yah:phase(P3)
83//! @yah:parent(R040)
84//! @yah:next("Resolves service configs + mirror declarations → per-machine compose.yml → push via yah-yubaba (A8) → systemctl restart yah-cloud-services")
85//! @yah:next("Tier isolation at SOFTWARE layer: separate PG roles per service set, distinct Headscale tags (tier:t0 yah meta vs tier:t2 noisetable assets), mesh_only flag controls Cloudflare exposure")
86//! @yah:next("Cloudflare DNS records (manual, in SECRETS.md): pdx.cloud.noisetable.example etc. → public IPs via orange-cloud")
87//! @yah:next("Gotcha: Cloudflare free tier proxies HTTP/HTTPS but NOT raw TCP/UDP — Headscale mesh uses direct WireGuard between machines")
88//! @yah:next("Assumes: Cloudflare account exists + noisetable.example DNS delegated to it (else spawns Cloudflare-bootstrap sub-ticket)")
89//! @yah:next("Open question (refinement): Caddy vs nginx vs direct-to-Cloudflare for per-machine reverse proxy")
90//! @yah:verify("curl https://pdx.cloud.noisetable.example/healthz (asset-registry) returns 200")
91//! @yah:verify("curl https://pdx.cloud.yah.example/healthz (yah meta-directory) returns 200")
92//! @yah:verify("yah cloud service status --all — all green on all 3 machines")
93//! @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md)
94//! @yah:handoff("Compose generation + yubaba service deploy + CLI wiring landed. New file: crates/yah/cloud/src/compose.rs (generate_compose_bundle: tier-named networks, Caddy service for mesh_only:false, Caddyfile with hostname when cloud_domain is set in mirrors/*.toml). MirrorConfig gained cloud_domain: Option<String>. Yubaba: replaced 501 stubs with real /services (podman compose ps --format json; empty array when no compose.yml) and /compose (write compose.yml + Caddyfile + write yah-cloud-services.service unit + systemctl enable --now). ServerState gained compose_dir: PathBuf (default /etc/yah-cloud) + with_compose_dir() builder. cloud-client: added ComposeDeployRequest / ComposeDeployResponse wire types + deploy_compose() method. CLI: yah cloud service deploy <name> resolves machines via mirror→service lookup, generates per-machine bundle (using location+cloud_domain as public hostname when available), and POSTs to each machine's yubaba. yah cloud service status [name|--all] queries /services on each machine and prints [+]/[-] per container. All tests pass: 18 yubaba, 47 cloud, 8 cloud-client, 120 yah bin.")
95//! @yah:next("Caddy in compose needs a Cloudflare origin cert or Let's Encrypt config once real domains land — Caddyfile currently uses :port placeholders which work for testing. Operator action: set cloud_domain in mirrors/<camp>.toml then re-run yah cloud service deploy to regenerate with real hostnames.")
96//! @yah:next("Cloudflare Tunnel integration (R040-F15): add a cloudflared service to the compose stack — currently services are exposed via Caddy on public ports 80/443 (orange-cloud model). F15 will replace with no-public-port cloudflared tunnel approach.")
97//! @yah:next("podman compose ps --format json output format varies by podman-compose version — current implementation passes raw JSON through. If operators hit parse issues, add a normalization layer in get_services() that extracts Name+State from both podman-compose and Podman 4.x native formats.")
98//! @yah:next("yah cloud service rolling <name> not yet implemented — stub left in place. Compose rolling restarts are just podman compose up --no-deps <svc>; wire it when rolling deploys are needed.")
99//! @yah:handoff("STALE ENTRY CORRECTED 2026-09-11 by R876-B12, not by this ticket's owner: the `@yah:next` above saying \"yah cloud service rolling <name> not yet implemented - stub left in place\" is no longer true. The stub at app/yah/cli/src/cloud.rs (`ServiceCommands::Rolling { .. } => stub(...)`) is retired and the verb now means something different from what that bullet imagined - it is NOT a compose rolling restart. It is the no-build redeploy path for a service-mirror `[providers.bundle]` slot: re-issue the deploy RPC for the digest already recorded on the node, which makes yubaba re-run `admit_bundle` and restores a lost service record (the R876-B11 recovery). If a compose-tier rolling restart is still wanted, it needs a different verb or an explicit tier switch; do not re-implement it under this name.")
100//!
101//! @yah:ticket(R040-F8, "A8: yah-yubaba (Rust binary on machine) + yah-cloud-client + desktop Cloud panel")
102//! @yah:at(2026-05-05T00:29:11Z)
103//! @yah:status(review)
104//! @yah:phase(P3)
105//! @yah:parent(R040)
106//! @yah:verify("cargo test -p yubaba -p cloud -p cloud-client + cargo test -p yah --bin yah cloud:: — 18 + 47 + 9 + 17 tests pass")
107//! @yah:verify("cargo run -p yubaba -- serve --bind 127.0.0.1:7450 --state /tmp/i.json then curl POST /register-hostkey with a real ssh-ed25519 pubkey, then `yah cloud machine attach <machine> --host 127.0.0.1:7450 --wait 5 --path <camp>` prints `attached: <machine> (fingerprint recorded: SHA256:…)` and writes the line into .yah/cloud/machines/<name>.toml")
108//! @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md)
109//! @arch:see(architecture/yah-managed-camps-topology.md)
110//! @yah:handoff("Session 3 landed `yah cloud machine attach <name>` (the hostkey write-back follow-on). app/yah/cli/src/cloud.rs::handle_attach loads MachineConfig, polls cloud-client GET /identity until success or --wait seconds elapse (default 180s, retry every 3s on NotRegistered or Transport errors), then calls MachineConfig::save() with the fingerprint. Decision logic in `decide_attach_action` covers all four states: Write (no fingerprint on file), Idempotent (match — no-op write), Mismatch (bail with --force hint), Overwrite (mismatch + --force). Subcommand args harmonized with agent ping: --host (default http://<machine>:7443), --wait <seconds>, --force, --path. handle_provision's stale stub print (`next: hostkey fingerprint write-back lands with A8`) replaced with `next: once cloud-init finishes, run yah cloud machine attach <name>`. End-to-end smoke verified all four AttachAction paths against a live yubaba on 127.0.0.1:7450 — pre-registration polling + bail, fresh-write, idempotent re-run, mismatch refusal, --force overwrite. 5 new unit tests (decide_attach_action_{writes_when_no_declared,idempotent_on_match,bails_on_mismatch_without_force,overwrites_with_force,force_is_idempotent_when_match}) — total 9 cloud:: tests in the yah bin. Cumulative R040-F8 deliverables across sessions 1–3: yubaba binary (12 tests), cloud-client crate (6 tests + doctest), AgentProbe trait + drift wire-up (4 new in cloud, 29 total), `yah cloud agent {ping,services,logs}` CLI, `yah cloud machine attach`, provision next-step hint correctly points at attach. Out of scope still: Tailscale-mesh discovery, desktop Cloud panel, mTLS hardening.")
111//! @yah:next("Tailscale-mesh discovery for --host: `yah cloud agent ping/services/attach` defaults to http://<machine>:7443 which only resolves inside the mesh. Smart discovery would query `tailscale status --json` for the machine's mesh IP. Until then, operators pass --host explicitly. Add a TS_DEVICES env shortcut or `yah cloud agent ip <machine>` helper if the manual path stays painful.")
112//! @yah:next("Desktop Cloud panel (out of session, larger): packages/yah/ui/src/cloud/ — machines list, per-machine service list, log tail. Tauri command bridges in app/yah/desktop/src/. Pulls on the cloud-client crate (already in workspace). Defer until at least Hostkey write-back lands so the panel has something useful to display per-machine.")
113//! @yah:next("Hardening parking lot (post-Phase-1): mTLS (server cert from machine hostkey, client cert from desktop user identity); /metrics endpoint (Prometheus); streaming /logs from journald with heartbeat. None of these gate Phase 1 verify; cloud-client's CloudClient::with_timeout exists so log streaming has a hook.")
114//! @yah:gotcha("yubaba binds 0.0.0.0:7443 by default; cloud-init template adds ufw rules to restrict to tailscale0. If a future deployment skips ufw, public IP exposure is a real risk. Better long-term: bind to tailscale0 IP directly (cloud-init systemd unit can compute it via `ip -4 addr show tailscale0`).")
115//! @yah:gotcha("Re-running register-hostkey with a different pubkey replaces the stored identity (idempotent for same key, swaps for different). Probably what you want, but worth flagging if someone designs an audit log around 'first registration wins'.")
116//! @yah:gotcha("Default URL resolution (`http://<machine.name>:7443`) is shared between WardenHealthProbe (status drift), agent ping/services, and machine attach — outside Tailscale mesh every default-host call against a real Hetzner machine will fail. Status emits AgentUnreachable as a soft note; agent/attach surface the transport error directly. Once mesh discovery lands, all three call sites should resolve via Tailscale before falling back to plain hostname.")
117//!
118//! @yah:ticket(R040-F18, "Phase 1a — yah mesh start: camp bootstrap with stable URL from day 1")
119//! @yah:at(2026-05-05T00:29:11Z)
120//! @yah:status(review)
121//! @yah:assignee(agent:claude)
122//! @yah:parent(R040)
123//! @yah:phase(P1)
124//! @yah:next("Implements Phase 1a from .yah/docs/architecture/A041-yah-mesh-bootstrap.md — design committed 2026-05-04. Phase 1b lands as R040-F19, Phase 2 (HA) as R040-F20+F21.")
125//! @yah:next("Subcommand surface: `yah mesh start` brings up Headscale locally (yah-desktop or camp daemon); auto-detects direct vs. cloudflared reachability; configures DNS for mesh.<your-domain>; stores URL in vault as `mesh-url` so provision picks it up.")
126//! @yah:next("Cloudflared bundling: yah-desktop should manage the cloudflared binary so the operator doesn't need a separate install. Same daemon can carry yah serve ingress + Headscale tunnel under one install — that's the user-acceptable story.")
127//! @yah:next("Stable URL contract: `mesh.<your-domain>` set ONCE here and never changes across Phase 1b promotion or Phase 2 leader changes. Nodes embed this URL in their tailscale config; downstream phases re-point DNS, not nodes.")
128//! @yah:next("Headscale binary provisioning: download + manage like cloudflared (precedent: yah already shells out to cloudflared/tailscale/cargo). Config template per arch doc 'Headscale config' section. Initial ACL writes a permissive `tag:*` policy.")
129//! @yah:next("Provision integration: when `mesh-url` is set, `yah cloud machine provision` calls Headscale API for a single-use preauth key and embeds {MESH_URL, PREAUTH_KEY, TAGS} into cloud-init's `tailscale up` line. Static `headscale-preauth-key` in vault becomes optional.")
130//! @yah:next("Also lands `yah mesh status` (coordinator location, connected nodes, preauth count). `yah mesh backup`/`restore` can be Phase-1a stubs; full litestream comes with R040-F21.")
131//! @yah:verify("yah mesh start on a fresh camp brings up Headscale + DNS, prints stable URL, stores in vault")
132//! @yah:verify("yah cloud machine provision <name> with mesh-url set generates preauth key and joins the new machine to the camp's Headscale")
133//! @yah:verify("tailscale status on the new machine shows it connected via the camp's tunnel; node-to-node ping works")
134//! @arch:see(.yah/docs/architecture/A041-yah-mesh-bootstrap.md)
135//! @yah:handoff("Phase 1a landed. New files: crates/yah/cloud/src/mesh.rs (HeadscaleClient + generate_headscale_config + headscale_download_url + DEFAULT_ACL_POLICY), app/yah/cli/src/mesh.rs (yah mesh start/status/stop/backup-stub/restore-stub). cloud_init::RenderInput gained mesh_url: Option<String>; renderer substitutes {{MESH_LOGIN_SERVER_ARG}} → ' --login-server <url>' or '' depending on field. provision::build_request now takes mesh_url: Option<String>. template mirror.yml updated with MESH_LOGIN_SERVER_ARG placeholder. cloud.rs::handle_provision now reads mesh-url from vault/HEADSCALE_URL env; when set, calls HeadscaleClient::from_vault_or_env() to auto-generate a single-use preauth key (falls back to static headscale-preauth-key when no API key is available). yah mesh start downloads headscale binary from GitHub (HEADSCALE_VERSION=0.23.0), writes config.yaml + acls.yaml, creates default user, spawns headscale serve in background, saves PID, stores mesh-url in vault, prints DNS/tunnel operator instructions. yah mesh status shows PID, mesh-url, node count + per-node online status. yah mesh stop does SIGTERM+SIGKILL with 5s grace. Cloud crate: 36 tests pass (+4 new: render_with_mesh_url_adds_login_server, render_without_mesh_url_no_login_server, provision::build_request_with_mesh_url_adds_login_server, mesh::config_contains_server_url, mesh::config_all_paths_in_data_dir, mesh::download_url_current_platform). yah lib: 116 unit tests pass (+4 new mesh:: tests). arch dogfood integration tests are pre-existing failures unrelated to this work.")
136//! @yah:next("Operator runbook for Phase 1a: (1) run yah mesh start --url https://mesh.<domain> on the camp; (2) set up DNS + cloudflared (or direct port-forward) for that URL; (3) headscale apikeys create to get an API key; (4) yah keys set headscale-api-key; (5) yah cloud machine provision <name> --yubaba-url <URL> --yubaba <PATH> --path /Users/user/ss/noisetable — provision now auto-generates a preauth key from Headscale.")
137//! @yah:next("Phase 1b (R040-F19): yah mesh promote <machine> — migrates Headscale from camp to first cluster machine. The stable-URL contract from this ticket means no node reconfiguration is needed; only DNS re-points.")
138//! @yah:next("Cloudflared bundling (R040-F15): cloud-init template needs cloudflared install + token. yah mesh start should detect reachability and auto-configure cloudflared when --url points through a Cloudflare Tunnel. The operator-instruction path this session is the MVP fallback.")
139//! @yah:next("Headscale API key auto-generation: currently the operator must manually run headscale apikeys create. A future improvement: yah mesh start can use the local UNIX socket (headscale.sock) to generate the API key and store it in the vault automatically without needing the binary in PATH.")
140//!
141//! @yah:ticket(R040-F19, "Phase 1b — yah mesh promote: migrate Headscale from camp to first cluster machine")
142//! @yah:at(2026-05-05T00:29:11Z)
143//! @yah:assignee(agent:claude)
144//! @yah:status(review)
145//! @yah:phase(P2)
146//! @yah:parent(R040)
147//! @yah:next("Implements Phase 1b from .yah/docs/architecture/A041-yah-mesh-bootstrap.md. Depends on R040-F18 (Phase 1a) shipping the stable-URL contract and camp Headscale.")
148//! @yah:next("Subcommand: yah mesh promote <machine-name>. Refuses if target machine isn't healthy (yubaba /health) or hostkey not registered.")
149//! @yah:next("Migration sequence: (1) stop Headscale on camp, (2) copy SQLite DB + config + private keys to target via yubaba, (3) start Headscale on target as a systemd unit yubaba manages, (4) update Cloudflare DNS A/AAAA for mesh.<your-domain> to point at target, (5) tear down camp Headscale + (if used) cloudflared.")
150//! @yah:next("Single-member raft is the degenerate case at this point — yubaba runs on the target, no quorum partners yet. Acceptable; HA shows up in Phase 2. Mark cluster as 'single-machine' in mesh status output so the operator knows the SPOF is real.")
151//! @yah:next("Atomicity / failure recovery: keep a snapshot of camp Headscale state until target is verified serving traffic. yah mesh promote --abort restores the camp coordinator if the new target proves unhealthy mid-migration. Dry-run mode prints the plan without executing.")
152//! @yah:next("Brief control-plane outage (~10s) is expected during DNS cutover. Data plane (existing WireGuard tunnels) is unaffected — verify by leaving a node-to-node ping running across the migration.")
153//! @yah:next("Out of scope: litestream replication, openraft, multi-member promotion. All of that lands in R040-F20/F21.")
154//! @yah:verify("yah mesh promote noisetable-pdx-1 migrates the coordinator; yah mesh status shows the new location; yah cloud machine provision <new-machine> against the migrated coordinator joins successfully")
155//! @yah:verify("Existing nodes' tailnet connectivity uninterrupted across the migration (continuous ping from one node to another stays green)")
156//! @arch:see(.yah/docs/architecture/A041-yah-mesh-bootstrap.md)
157//! @yah:handoff("Phase 1b landed. New command: yah mesh promote <machine> [--dry-run] [--abort] [--host <URL>] [--path <dir>]. Migration sequence: (1) verify yubaba health + hostkey registered, (2) stop camp headscale, (3) POST SQLite DB + WireGuard keys + ACL policy to yubaba POST /headscale/deploy (yubaba downloads headscale binary from GitHub, writes files to configurable headscale_dir, starts systemd unit), (4) poll GET /headscale/health until running (90s timeout), (5) update Cloudflare DNS A record via CF API (CLOUDFLARE_API_TOKEN + CLOUDFLARE_ZONE_ID) or print manual instructions, (6) persist coordinator type to vault (mesh-coordinator-type=cluster, mesh-coordinator-machine=<name>). --abort restarts camp headscale from existing local state. Rollback on mid-migration failure auto-restarts camp headscale. yah mesh status updated: shows 'cluster [<machine>] (single-machine — SPOF)' after promote, camp PID status before. New yubaba endpoints: POST /headscale/deploy (accepts base64-encoded state files, writes to configurable headscale_dir, curl-downloads binary, systemctl enable --now), GET /headscale/health (systemctl is-active + localhost:8080 HTTP probe). ServerSummary gained public_ipv4: Option<String> (parsed from Hetzner public_net.ipv4.ip). Cloudflare DNS helper: update_cloudflare_dns() + cloudflare_credentials() in cloud::mesh. cloud-client: deploy_headscale() + headscale_health() methods. Tests: +8 yubaba (15 total), +2 cloud-client (9 total), +4 yah mesh:: (120 total). All pass.")
158//! @yah:next("Operator runbook: (1) yah mesh start --url https://mesh.<domain> on camp, (2) provision a machine, (3) yah cloud machine attach <name>, (4) yah mesh promote <name> [--host <yubaba-ip>:7443 pre-mesh]. DNS update needs CLOUDFLARE_API_TOKEN + CLOUDFLARE_ZONE_ID or manual update.")
159//! @yah:next("Mesh URL connectivity: before Tailscale mesh is up, --host must be the machine's public IP (pre-mesh access). Once mesh is established, the default http://<machine>:7443 works. Consider yah cloud agent ip <machine> helper or mesh-discovery to auto-resolve (same gap as R040-F8 Tailscale discovery).")
160//! @yah:next("HEADSCALE_VERSION is pinned to 0.23.0. If the operator's local headscale.db was created by a different version, the remote headscale may reject the database. Add a version-mismatch warning to the deploy step.")
161//!
162//! @yah:relay(R085, "Infra-info MCP — yubaba / machine / mirror status read surface")
163//! @yah:assignee(agent:claude)
164//! @yah:status(review)
165//! @yah:parent(Q082)
166//! @yah:handoff("Four read-only cloud.* MCP tools landed in crates/yah/agent-tools/src/cloud_tools.rs. Registration: KgToolRegistry::with_cloud() in tools.rs (same builder pattern as with_board/with_scryer/with_task). cloud-client dep added to agent-tools Cargo.toml. Local-config tools (cloud.machines, cloud.mirror_state) read .yah/cloud/{machines,mirrors,services}/*.toml from ctx.camp_root with lightweight serde structs — no cloud crate dep, no network. Network tools (cloud.warden_status, cloud.service_ports) probe yubaba HTTP API via cloud_client::CloudClient and return reachable:false rather than error when unreachable. 14 unit tests, all pass. Total agent-tools tests: 140.")
167//! @yah:next("Wire with_cloud() into the agent session startup (camp.rs or agent.rs in app/yah/cli) alongside with_board() and with_scryer() so the tools appear in the agent's tool list automatically.")
168//! @yah:next("Feed cloud.warden_status into the drift detection surface (status.rs AgentProbe) so yubaba reachability populates machine status reports rather than being a separate probe path.")
169//! @yah:verify("cargo test -p agent-tools -- cloud 2>&1 | grep 'test result' shows 14 passed; 0 failed")
170//! @yah:verify("KgToolRegistry::standard_read_only(ctx).with_cloud().schemas() returns 4 tools all named cloud.*")
171//! @arch:see(.yah/quests.md)
172//!
173//! @yah:relay(R168, "yah-cloud TOML config schema + validation")
174//! @yah:assignee(agent:claude)
175//! @yah:at(2026-05-13T19:07:07Z)
176//! @yah:status(review)
177//! @yah:parent(Q156)
178//! @yah:handoff("Schema doc written at .yah/docs/architecture/A031-yah-cloud-config-shape.md. MirrorConfig.camp renamed to .camp with serde alias 'camp' for migration compat (R137 rename). load_mirrors() added to handle both flat mirrors/<id>.toml and folder mirrors/<id>/mirror.toml layouts — yah-com/mirror.toml now loads. cloud_tools.rs local stub updated to match. All 90 cloud-crate tests pass + 4 new layout tests. agent-tools + yah bin check clean.")
179//! @yah:next("Remove serde alias 'camp' from MirrorConfig once all known mirrors are migrated (tracked in R137)")
180//! @yah:verify("cargo test -p cloud — 90 tests pass including mirror_folder_layout_loads + mirror_malformed_fails_with_field_path")
181//! @yah:verify("Schema doc links back to yah-public-site.md deployment-shape section")
182//! @arch:see(.yah/docs/architecture/A031-yah-cloud-config-shape.md)
183//! @arch:see(.yah/docs/architecture/A045-yah-public-site.md)
184//!
185//! @yah:relay(R260, "OpenRouter model almanac — track top-weekly models for backend selection")
186//! @yah:assignee(agent:claude)
187//! @yah:at(2026-05-20T22:23:52Z)
188//! @yah:kind(spike)
189//! @yah:status(in-progress)
190//! @arch:see(architecture/yah-cloud-config-shape.md)
191//! @yah:next("Unique signal: popularity + freshness + capability tags. Pricing is already covered by `crates/yah/probe/data/pricing.json` (LiteLLM snapshot) — almanac does NOT duplicate $/token; it adds rank, `:free` flag, modality, capability tags (code / vision / tool-use / long-context), and a last_seen_at so consumers can tell when an entry has gone stale.")
192//! @yah:next("Primary consumer: the OpenRouter character bundle at `crates/yah/kg/preroll/openrouter/{agents,subclasses}.json`. Currently a single preroll agent (\"Quick Hand\") pinned to `qwen/qwen3-coder:free` — this WILL silently break when OpenRouter rotates the free-tier list. Almanac's first job is to keep that subclass pointed at a currently-available free model (and let the bundle expand into a roster of currently-hot free models, not just one).")
193//! @yah:next("Free-model rotation is the load-bearing constraint. Refresh cadence has to match how often OpenRouter shuffles the `:free` set (empirically weekly-ish). Storage options ranked by how well they handle rotation: (1) scheduled daemon refresh writing back into `crates/yah/kg/preroll/openrouter/` — auto-heals; (2) on-demand fetch+cache in the runner — never stale at session start but offline-fragile; (3) build-time embedded snapshot like `probe/data/pricing.json` — simplest, goes stale between releases (probably wrong shape for free-tier).")
194//! @yah:next("Two approaches to evaluate side-by-side: (A) straight call to `https://openrouter.ai/api/v1/models` (free, no key, structured JSON — already returns `context_length`, `pricing.prompt`/`pricing.completion`, `architecture.modality`; `crates/yah/runner/src/resolver/openrouter.rs` already talks to OpenRouter so a `fetch_models()` sibling to its existing `GET /api/v1/credits` is the natural home); (B) inference-based tool that distills `https://openrouter.ai/models?order=top-weekly` HTML into our schema. Confirm (A) first — if the JSON exposes the top-weekly ordering and `:free` flag, (B) is unnecessary; if not, (B) becomes the rotation-detector.")
195//! @yah:next("Query-resolvable subclass (the real primitive): `AgentSubclass.model` at `crates/yah/kg/src/party.rs:520` is a hard `provider:model` string today. Promote it to a sum type — literal `provider:model` OR `ModelQuery { tier, capability_tags, min_context_tokens, max_price_per_token, … }` — resolved by the almanac at session start. Slots into the existing `fallbacks: Vec<FallbackRule>` machinery at party.rs:540: emit `ConfigSwitch{ModelResolved}` (sibling of `FallbackTriggered`) when the query lands, so the UI can show \"Quick Hand → qwen3-coder:free (current pick, refreshed 2h ago)\". The preroll bundle then declares intent (\"Thief-shaped, free, code-tuned\") instead of a brittle pin.")
196//! @yah:next("Constraint vocabulary — the hard numbers behind \"Thief-shaped, free, code-tuned\", layered by cost-to-implement so we ship something useful at T0 and can grow. **T0 (already free from /models):** `context_length`, `pricing.prompt`/`pricing.completion`, `architecture.modality` come back in the same fetch. MVP heuristic = Pareto frontier of (context_window, price_per_token); at `:free` tier price collapses to 0 so it cleanly reduces to \"max context with matching capability tag\". This is enough to pick a non-broken Quick Hand. **T1 (external quality scores):** join on canonical model id against artificialanalysis.ai's quality index (general + code subscore), Aider's code leaderboard, or LiveBench. Cheap to add once the T0 pipe exists; lets queries express `min_quality` or `prefer_code_subscore`. **T2 (almanac-as-evaluator):** spend a few credits running a small fixed corpus (lint-fix, summarize, ack-and-route) through each new free candidate, score against a reference, persist scores into the almanac. The dogfood option — most signal, most expensive, most interesting. Out of scope for the spike; file as a follow-up child once T0/T1 are in flight.")
197//! @yah:next("Recommendation surface: popularity-ranked picker entries in `packages/yah/ui/src/components/agent/Picker/toPickerAgents.ts` + `AgentProvidersPanel.tsx` so adding an OpenRouter backend surfaces the currently-hot models first; query-resolved subclasses display their resolved pick + a freshness chip; stale literal-pinned `:free` references show a warning.")
198//! @yah:next("Spike output: a design note at `.yah/docs/working/W091-openrouter-almanac.md` + a working (A) prototype printing the top 10 models in our T0 schema (id, context_length, price, `:free` flag, capability tags, popularity rank). From there, refine into child tickets for: (i) refresh mechanism, (ii) `ModelSpec::Literal | ModelSpec::Query` sum type + T0 resolver, (iii) preroll-bundle migration from pinned-model to ModelQuery, (iv) UI recommendation + freshness chips, (v) T1 external-benchmark join, (vi) T2 self-eval as a stretch/follow-up.")
199//! @yah:handoff("Spike output delivered: working prototype + design note + 6 child tickets. Prototype lives in crates/yah/runner/src/resolver/openrouter.rs (AlmanacEntry struct + fetch_models() against /api/frontend/models/find?order=top-weekly + derive_capability_tags) plus app/yah/cli/src/agent.rs (yah agent almanac --top N --free-only --tag code --json). Design note at .yah/docs/working/W091-openrouter-almanac.md. Six follow-ups filed: R260-F1 (refresh scheduler), R260-F2 (ModelSpec sum type + T0 resolver), R260-T3 (preroll bundle migration → ModelSpec::Query), R260-F4 (UI freshness chips + stale-pin warning), R260-F5 (T1 benchmark join), R260-S6 (T2 self-eval, stretch). T0 prototype confirms the rotation problem: `yah agent almanac --top 10 --free-only --tag code` shows qwen/qwen3-coder:free (Quick Hand's current pin) is NO LONGER in the top-10 free code models — current leaders are openrouter/owl-alpha, nvidia/nemotron-3-super, poolside/laguna-m.1.")
200//! @yah:verify("cargo run -p yah -- agent almanac --top 10 — prints OpenRouter top-10 weekly models in T0 schema")
201//! @yah:verify("cargo run -p yah -- agent almanac --top 10 --free-only --tag code — shows the currently-hot free-tier code-tuned roster (the Quick Hand replacement set)")
202//! @yah:verify("cargo check -p runner -p yah — clean (note: cargo test -p runner --lib is blocked by pre-existing R258-F3 test breakage on AgentSession; tracked as gotcha)")
203//! @yah:gotcha("Pre-existing R258-F3 test breakage in crates/yah/runner/src/{anthropic.rs:833, sessions.rs} — `cargo test -p runner --lib` fails with E0063 (AgentSession initialiser missing last_text + slot_id fields). R258-F3 is currently in review. R260's new resolver::openrouter::tests entries are byte-perfect but un-runnable until R258-F3's test fixtures are updated.")
204//! @yah:gotcha("OpenRouter's /api/frontend/models/find?order=top-weekly is undocumented. The almanac treats it as best-effort and the design (see openrouter-almanac.md) calls for graceful degradation to public /api/v1/models alone when the join fails. ?order= is silently ignored on the documented endpoint (it always returns newest-first); approach (B) HTML scraping is unnecessary and was ruled out.")
205//!
206//! @yah:relay(R414, "yah-cloud pill sync + publish + Pills panel")
207//! @yah:assignee(bundle-anthropic-miravel)
208//! @yah:at(2026-06-03T07:07:03Z)
209//! @yah:status(review)
210//! @yah:phase(P2)
211//! @yah:parent(Q410)
212//! @arch:see(.yah/docs/working/W141-pill-rings.md)
213//! @arch:see(.yah/docs/architecture/A031-yah-cloud-config-shape.md)
214//! @yah:depends_on(R411)
215//! @yah:handoff("P2 delivered: resolve_pill_catalog (chips crate) + party_resolve_pill_catalog Tauri command + env wiring + PillsPanel in FullCharacterEditor. Pills panel shows all categorized pills grouped by category with per-character tri-state toggle (always-on/discoverable/off); edits agent.persona.chips.pill_state map via patchAgent.")
216//! @yah:next("P3: local pills.toml loader — add ~/.yah/pills.toml and .yah/pills.toml as new resolution layers (separate from chips.toml) per W141 layer stack")
217//! @yah:next("P3: MCP URL trust display — when a content-pill with mcp[] is toggled always-on, show one-time approval dialog listing the URLs before activating")
218//! @yah:next("P4: cloud sync stub — 'Share to my account' button in PillsPanel (disabled until account system exists)")
219//! @yah:verify("cargo test -p chips — 84 passed (3 new resolve_pill_catalog tests)")
220//! @yah:verify("bun run typecheck — no new errors (6 pre-existing unchanged)")
221//!
222//! @arch:see(.yah/docs/working/W193-asset-dependency-status-surface.md)
223//!
224//! @yah:ticket(R743-T7, "yah-cloud: 5 test binaries to 1 (or 2 — pond_smoke may stay isolated)")
225//! @yah:at(2026-08-11T01:18:56Z)
226//! @yah:status(review)
227//! @yah:phase(P2)
228//! @yah:parent(R743)
229//! @yah:next("tests/main.rs mod'ing the siblings + autotests = false and [[test]] name = \"main\" in oss/yubaba/crates/cloud/Cargo.toml.")
230//! @yah:verify("cargo test -p yah-cloud -- --list count unchanged; three green runs. One commit — oss subtree.")
231//! @yah:gotcha("tests/pond_smoke.rs:198 is the only real fixed-port bind in the whole in-scope set — MinIO on http://127.0.0.1:9000, an external dependency it does not own. It is a legitimate deliberate-isolation candidate: if merging it makes the suite flaky or order-dependent, keep it as its own [[test]] target and say so on the ticket rather than force-merging.")
232//! @yah:gotcha("pond_smoke.rs and mesofact_static_e2e.rs both host live @yah: annotations.")
233//! @yah:handoff("LANDED: 5 test binaries → 2. New tests/main.rs mods live_workspace_smoke, mesofact_static_e2e, pg_driver_live and whisper_derive_e2e; Cargo.toml gains autotests = false plus explicit [[test]] main (tests/main.rs) and [[test]] pond_smoke (tests/pond_smoke.rs). Files were mod'd, not concatenated, so mesofact_static_e2e.rs's live R441-B2 annotation block stays where the harvester expects it and CARGO_MANIFEST_DIR (which three of the four walk up from) is unchanged. autotests = false only gates [[test]] discovery — examples/load_probe.rs is still autodiscovered. Test names gained a module prefix (e.g. whisper_derive_e2e::derive_pipeline_upload_skip_prune_reproducibility); substring filters still match, but `--test <file-stem>` for the four merged files is now `--test main -- <module>::`.")
234//! @yah:handoff("ISOLATION DECISION: pond_smoke STAYS its own [[test]] target — 2 targets, not 1, and deliberately so. Three reasons, strongest first. (1) DECISIVE: .yah/qed/pond-smoke.toml's `pond-spinup-budget` step shells out to `cargo test --release --locked -p cloud --test pond_smoke -- --nocapture`. Folding pond_smoke into `main` turns that committed pipeline step into a hard cargo error, and that file is outside this crate (peer-owned path, not mine to edit). (2) It is the only test in the crate that reaches a FIXED external address — MinIO at http://127.0.0.1:9000 (pond_smoke.rs:199) — and the only one that creates/`docker rm -f`s named containers in a Drop guard, so its blast radius on failure is outside the process and a standalone runnable handle is worth keeping. xtask/src/lib.rs's @yah:assumes census independently found this to be the ONLY real port bind across all ten in-scope crates. (3) It is benchmark-shaped and libtest parallelises within one binary: both its tests assert wall-clock budgets (WARM_RESTART_BUDGET 3s, WARDEN_COLD_BUDGET 5s, COLD_START_BUDGET 15s). Merging would drop whisper_derive_e2e (two in-process axum servers + BLAKE3 + tempdir IO, ungated) and live_workspace_smoke (a full CloudConfig::load walk of the real workspace, ungated) onto sibling threads inside that measurement window, and the pipeline runs it --nocapture for the timing report, which merging would interleave with every other test's output. The four that DID merge have no such conflict: read-only fs, 127.0.0.1:0 ephemeral ports, tempdirs, and no process-global state (no set_var / set_current_dir / top-level statics anywhere in the set).")
235//! @yah:verify("BASELINE (HEAD layout, cargo test -p yah-cloud -- --list, run with the change temporarily reverted): 6 test binaries — lib 924 tests; live_workspace_smoke 1; mesofact_static_e2e 1; pg_driver_live 1; pond_smoke 2; whisper_derive_e2e 1; Doc-tests cloud 1. Integration total 6 across 5 binaries.")
236//! @yah:verify("AFTER (same command, exit 0): lib 924 tests; tests/main.rs 4 tests; tests/pond_smoke.rs 2 tests; Doc-tests cloud 1. Integration total 6 across 2 binaries — COUNT UNCHANGED, every one of the 6 names accounted for, now module-prefixed inside main.")
237//! @yah:verify("THREE GREEN RUNS: cargo test -p yah-cloud --test main --test pond_smoke, three consecutive times — main 3 passed / 0 failed / 1 ignored, pond_smoke 2 passed / 0 failed, all three runs identical. No order-dependence observed.")
238//! @yah:verify("HONEST SCOPE OF WHAT ACTUALLY EXECUTED (confirmed with -- --nocapture, container has no docker / no live pg / no mesofact-dev binary): REALLY RAN — live_workspace_smoke::live_yah_workspace_loads_cleanly (real CloudConfig::load of the checked-in .yah/ tree, real assertions) and whisper_derive_e2e::derive_pipeline_upload_skip_prune_reproducibility (in-process axum fake-S3 + fake upstream on 127.0.0.1:0, all three runs). SELF-SKIPPED at runtime — mesofact_static_e2e (prints 'skipping: set YAH_RECONCILER_E2E_BIN'), pond_smoke::pond_spinup_budget and pond_smoke::warden_container_spinup_budget (both print 'SKIP: set YAH_LOCAL_SIM_E2E=1', need orbstack/colima/docker). IGNORED — pg_driver_live (#[ignore], needs a built yah-pg-dev binary + network). So the live/e2e legs were compiled and linked but NOT exercised here; the pipeline steps that do exercise them are unchanged for pond_smoke and reachable as `--test main -- mesofact_static_e2e::` for the merged one.")
239//! @yah:gotcha("PRE-EXISTING, NOT CAUSED BY THIS TICKET: .yah/qed/local-sim-smoke.toml's only step runs `cargo test -p cloud --test local_sim_smoke`, but tests/local_sim_smoke.rs does not exist and `git log --all` finds it at neither oss/yubaba/crates/cloud/tests/ nor the pre-OSS crates/yah/cloud/tests/ path. That pipeline was already dangling before this change (it most likely wants pond_smoke). Left alone: .yah/qed/ is outside this crate. Flagging it because autotests = false makes a target-name typo fail identically to this, and the next person to hit it should not blame R743-T7.")
240//! @yah:gotcha("Historical @yah:verify strings on already-landed tickets still cite the retired per-file target names — `--test whisper_derive_e2e` (reconciler/static_asset.rs:46,118), `--test mesofact_static_e2e` (the R441-B2 block inside tests/mesofact_static_e2e.rs itself). Those are landed records, not live invocations, so they were deliberately NOT rewritten; the working form is now `cargo test -p yah-cloud --test main -- <module>::`.")
241//! @yah:gotcha("CONTAINER, not code: the first attempt at the post-change --list died with `error: linking with cc failed … ld terminated with signal 7 [Bus error]` on both the merged `main` target and the untouched `lib test` target. Root cause was the container's root filesystem at 100% (8.0K free) — oss/yubaba/target alone was 23G with no cargo-orphan-gc installed here to reclaim it. Cleared by deleting the regenerable incremental caches (oss/yubaba/target/debug/incremental, target/debug/incremental) while no cargo was running; the listing then exited 0. Worth knowing because the failure mode reads exactly like the R770 orphan-gc symptom described in CLAUDE.md but is plain ENOSPC.")
242//! @yah:tier(Warrior)
243//!
244//! @yah:ticket(R895-T2, "Delete the LegacyServiceConfig compose/Caddy generation once nothing renders it")
245//! @yah:at(2026-09-13T21:02:37Z)
246//! @yah:status(review)
247//! @yah:assignee(agent:bundle-anthropic-ashguard)
248//! @yah:parent(R895)
249//! @yah:next("Tier: Cleric — mostly deletion, but the liveness check must be real. This module (podman-compose + Caddyfile generation over LegacyServiceConfig, R040-F7) and mesh_service.rs's compose-era host-network pattern are the pre-workload-spec generation of the data plane. Before deleting: establish by call-graph, not by name, whether any live path still renders a ComposeBundle (POST /compose consumers, cloud-init templates, provision.rs) — if one does, migrate that consumer to the workload-spec path first and delete in the same relay, per the below-1.0 break-don't-tape rule. The one thing worth salvaging is documented intent: the tenant network-split rules (W206/R558-T2, NetworkPlan) express a policy the W343 model must also answer; carry the policy statement into W343's doc or the container_net module docs before the code that encodes it goes.")
250//! @yah:gotcha("Liveness check done (sniffer, session:77442b98) and the ticket's premise needs correcting: compose.rs is NOT dead by call graph. There is one real live chain — clap `Commands::Cloud` (app/yah/cli/src/cli.rs:1081) -> handle_cloud_command (cloud.rs:2472) -> handle_service (cloud.rs:7161) -> handle_service_deploy (cloud.rs:7202) -> generate_compose_bundle (cloud.rs:7204) -> ComposeDeployRequest -> cloud_client::deploy_compose (crates/yah/cloud-client/src/lib.rs:1373) -> POST /compose. The route is live at oss/yubaba/crates/yubaba/src/lib.rs:2673 with handler deploy_compose at :6566-6649 (auth class OPERATOR), and `ServiceCommands::Deploy` (cloud.rs:1506) carries no deprecation or hidden attribute. So this ticket deletes a SHIPPED CLI verb and an HTTP route, not just an unreferenced module. NOTE (R895-T2 courier, session:49bbf176): compose.rs and mesh_service.rs are now DELETED (this ticket) — the file:line references above (compose.rs, cloud.rs pre-deletion line numbers, lib.rs:6566-6649) are historical, describing the call chain as it existed before this deletion.")
251//! @yah:gotcha("The path is nonetheless dead by DATA here and non-functional on the fleet, which is why deletion is still the right call. (a) `.yah/cloud/` in this camp holds only machines/ and status.jsonl — no services/ — so CloudConfig::load (config.rs:1575) yields empty legacy_services and handle_service_deploy bails at cloud.rs:7190 before reaching the renderer. (b) deploy_compose writes /etc/yah-cloud/compose.yml (lib.rs:6578) but app/yah/cli/resources/yubaba.service:245-246 sets ProtectSystem=strict with ReadWritePaths covering only /var/lib/yah/yubaba, /run/yubaba, /var/lib/yah/qed — no /etc — so the write EROFSes and propagates as a 500. The same silent-EROFS site is already recorded at .yah/infra/machines/us-west-001.toml:13. (c) the unit it generates runs `podman compose up` (lib.rs:6662) and the provisioning template stopped installing podman in R092-F2 (oss/yubaba/crates/cloud/templates/mirror.yml:41-43). Deleting a verb that already 500s is not a regression.")
252//! @yah:gotcha("SALVAGE is larger than the ticket assumed, and is the load-bearing part of this work. Deleting compose.rs removes the ONLY executable encoding of tenant network isolation in the tree. NetworkPlan (compose.rs:112, derive :118, network_for :129, declared :139) is what actually splits co-resident tenants onto `<tenant>-<base>` bridge networks per W206/R558-T2. WorkloadSpec.tenant survives (oss/yah-base/crates/workload-spec/src/lib.rs:2715) but NOTHING consumes it for networking: container_net.rs is per-node by construction (\"One bridge per node, one veth pair per workload, one routed /24 per node\", :15-16) and has no tenant concept, and W343 never states a tenant-separation rule at all. After this deletion WorkloadSpec.tenant is declarative-only and W206's \"cross-tenant traffic is denied by default\" has no enforcer anywhere. No capability is actually lost today (the encoder is unreachable per the gotchas above), but the design commitment loses its last home in code — so the policy MUST be transplanted into W343 plus the container_net module docs, marked explicitly as required-and-unimplemented, before the code goes. DONE (R895-T2 courier, session:49bbf176): transplanted into W343 \"Open, deliberately\" section and oss/kamaji/crates/kamaji/src/container_net.rs module docs, both citing R895-F3.")
253//! @yah:gotcha("Co-deletion target established: oss/yubaba/crates/cloud/src/mesh_service.rs (158 lines, declared at cloud/src/lib.rs:263) is reachable ONLY through compose.rs — compose.rs:210 uses MESH_IP_ENV_FILE, compose.rs:259 uses ufw_rules_for_mesh_port, and its other four exports (pg_hba_snippet, mesh_ip_env_runcmd, MESH_SUBNET, TAILSCALE_IFACE) have zero callers outside their own cfg(test) block at :108-158. It shares no types with compose.rs, so it becomes unreferenced in the same commit. Every leg of the host-network pattern its module doc describes is already false: cloud_init.rs no longer writes /etc/yah-cloud/mesh-ip.env (zero compose/caddy hits in that file), and compose.rs:215 points at `yah cloud service recipe postgres`, a subcommand that does not exist (ServiceCommands at cloud.rs:1505 has only Deploy/Rolling/Status/Prune). Successor documented at oss/yubaba/crates/yubaba/src/service_records.rs:4-8. DELETED along with compose.rs (R895-T2 courier, session:49bbf176) — `git show HEAD:oss/yubaba/crates/cloud/src/mesh_service.rs` carried no @yah: annotation of its own, so nothing else was re-homed from it. Its committed-HEAD content matched this description; the file was additionally dirty (uncommitted) at deletion time and that uncommitted diff was not read before `rm` — if it carried an annotation beyond HEAD's, it was not recovered.")
254//! @yah:gotcha("Three near-miss names that are LIVE and must not be swept into this deletion. (1) The caddy-container provider for the local-sim tier (.yah/qed/local-sim-smoke.toml:28-40, A031, A046:40) is owned by oss/yubaba/crates/cloud/src/reconciler/pond.rs, not by compose.rs. (2) crates/yah/runner/src/compose.rs is prompt-block composition, unrelated. (3) workload-spec's compose_import (oss/yah-base/crates/workload-spec/src/lib.rs:489) is `yah workload import docker-compose.yml`, a one-way CLI importer per A054:437. Also confirmed clean and out of scope: provision.rs and cloud_init.rs have zero compose/caddy hits, and the sibling GET /services route (lib.rs:2589) stopped shelling out to podman in R556-F7-T3 — test services_is_compose_independent_post_r556_f7_t3 at lib.rs:11338 pins that — so `yah cloud service status` survives; only `deploy` dies.")
255//! @yah:verify("Orphan-unit risk CHECKED AND CLEARED on the live fleet, 2026-09-13 (read-only sweep, session:1c2ba0c4). All 9 nodes declared in .yah/infra/machines/*.toml were reachable over the mesh — us-east-001, us-south-001, us-west-001, us-west-002, us-west-003, us-west-011, us-west-013, us-west-014, us-west-015 — zero unreachable, so this is a complete answer rather than a sample. On every node: no yah-cloud-services.service unit (systemctl is-enabled reports not-found), no /etc/yah-cloud/compose.yml, no Caddyfile, and no podman binary. The /etc/yah-cloud/ directories that do exist hold only current cert-store.env / litestream.env / durability env files with Sep 8-12 mtimes — none are artifacts deploy_compose ever wrote. us-west-015 is a macOS box with no systemd, so the unit path is architecturally moot there. CONCLUSION: deleting the POST /compose route and the ServiceCommands::Deploy verb orphans nothing on the running fleet; no pre-deploy node action is required.")
256//! @yah:handoff("SALVAGE (phase 1): DONE. Tenant-isolation policy from NetworkPlan/W206/R558-T2 transplanted into TWO places, both citing R895-F3 as required-and-unimplemented: (a) .yah/docs/working/W343-per-workload-mesh-addressing.md, new bullet added to the 'Open, deliberately' section ('Tenant network isolation (W206/R558-T2) — required, currently unimplemented'); (b) oss/kamaji/crates/kamaji/src/container_net.rs module doc, new '## Tenant isolation (W206/R558-T2) — required, currently unimplemented' section inserted before the @yah:ticket(R605-F22,...) annotation block. Both describe the rule (no cross-tenant L2/3 sharing, same-tenant cross-namespace OK, deny-by-default with opt-in, single-tenant stays a no-op) and note the likely implementation shape (tenant-keyed bridge or packet filter, since the addressing model is already per-node not per-tenant). This is the one part of the ticket the compiler can't verify — read it directly if in doubt.")
257//! @yah:handoff("DELETIONS, file by file, each verified by the grep shown (all return empty/only-annotation-text unless noted): oss/yubaba/crates/cloud/src/compose.rs (765 lines) — DELETED via rm; its own @yah:ticket(R895-T2,...) annotation (5 gotchas + verify) was read in full BEFORE deletion and re-homed into oss/yubaba/crates/cloud/src/lib.rs's module header (appended after the pre-existing @yah:tier(Warrior) line at :242, own block starts :244) — `board.show R895-T2` and `board.show R743-T7` (the ticket that owned the tier line) both resolve cleanly, confirmed no interleaving.")
258//! @yah:handoff("oss/yubaba/crates/cloud/src/mesh_service.rs (158 lines) — DELETED via rm. `git show HEAD:oss/yubaba/crates/cloud/src/mesh_service.rs | grep '@yah:'` returns nothing, so the last COMMITTED version carried no annotation. BUT this file showed 'M' (dirty vs HEAD) in the git status at session start, meaning it had uncommitted edits from an earlier session, and I ran `rm` before reading the working-tree content — that uncommitted diff is not recoverable via git (never staged/committed). @Ashguard:blade has a forensics session running (session:b49fa21f) trying to recover it from claude-cli transcripts; that recovery is not mine to chase further.")
259//! @yah:handoff("oss/yubaba/crates/cloud/src/lib.rs — `pub mod compose;`, `pub mod mesh_service;`, `pub use compose::{generate_compose_bundle, ComposeBundle};`, and `LegacyServiceConfig` dropped from the `pub use config::{...}` re-export block. Verified: `grep -nE 'LegacyServiceConfig|generate_compose_bundle|ComposeBundle|mod compose|mod mesh_service' oss/yubaba/crates/cloud/src/lib.rs` returns only the R895-T2 annotation text itself (expected, it's prose).")
260//! @yah:handoff("oss/yubaba/crates/cloud/src/config.rs — LegacyServiceConfig struct deleted (was ~:1391, doc said 'Deprecated... replaced by workloads/ in R092-F1'); `legacy_services: Vec<LegacyServiceConfig>` field removed from CloudConfig + doc comment on the struct updated to note removal; the `load_dir::<LegacyServiceConfig>(...)` call and its slot in the load-tuple removed from CloudConfig::load; both `legacy_services: vec![]` default-construction sites removed (load_from_config_dir, and the make_empty_cfg test helper); 4 now-impossible tests deleted (round_trip_service_legacy, service_bind_interface_round_trips, service_bind_interface_absent_is_none, service_bind_interface_skipped_when_none); 2 remaining `cfg.legacy_services` assertions removed (cloud_config_load_and_lookup — also removed its now-dead 'legacy services/ dir' fixture setup — and one more standalone assertion); unused top-level imports HashMap and TenantId dropped from the `use` block (TenantId is still imported locally inside several `#[cfg(test)] mod` blocks elsewhere in the file — untouched, still needed there). Verified: `grep -n 'LegacyServiceConfig|legacy_services' config.rs` returns only the doc-comment mention explaining the removal.")
261//! @yah:handoff("app/yah/cli/src/cloud.rs — ServiceCommands::Deploy variant removed from the enum; the `ServiceCommands::Deploy { .. } => handle_service_deploy(...)` match arm and the handle_service_deploy() function body removed entirely; collect_machine_services() and derive_public_hostname() removed (both became fully unreferenced once handle_service_deploy was gone — confirmed via grep, zero other call sites); `compose::generate_compose_bundle` and `LegacyServiceConfig` imports dropped from the `use cloud::{...}` block; machines_for_service() KEPT (still called from handle_service_status). Also fixed two now-stale strings inside handle_service_status: a comment claiming '/services' shells to 'podman compose ps' (false since R556-F7-T3) and a 'no services running — deploy first with `yah cloud service deploy <name>`' hint pointing at the just-deleted verb.")
262//! @yah:handoff("crates/yah/cloud-client/src/lib.rs — ComposeDeployRequest/ComposeDeployResponse structs and the deploy_compose() client method removed; the spawn_yubaba() test helper's compose_dir tempdir creation and .with_compose_dir(...) call removed; deploy_compose_round_trips_via_yubaba test deleted; services_returns_empty_array_when_no_compose_deployed renamed to services_returns_empty_array_when_no_workloads_deployed and simplified (no longer creates a compose_dir). Verified via grep only — NOT yet compiled/tested by me (see 'not done' below).")
263//! @yah:handoff("oss/yubaba/crates/cloud/src/reconciler/mesofact_bundle.rs — its cfg_with() test helper (a downstream CloudConfig struct-literal builder) had `legacy_services: vec![]`; removed, this was a real compile error caught by the yah-cloud build.")
264//! @yah:handoff("legacy_mirrors / LegacyMirrorConfig verdict: KEEP, do not delete. Confirmed reachable and heavily live: machines_for_service (cloud.rs), the mirror-status/mirror-show CLI path (build_mirror_rows, print_mirror_show_table), handle_mirror-family functions, mesofact_bundle.rs and lan_tunnel.rs test fixtures. This is unrelated to the compose/service-deploy system being deleted here.")
265//! @yah:handoff("BUILD STATUS: `cargo check --manifest-path oss/yubaba/Cargo.toml -p yah-cloud --lib --tests` is GREEN — Finished, 0 errors. Remaining warnings in that output are pre-existing and unrelated to this diff (dead_code in oss/yah-base/crates/object-store/src/r2.rs, unused_mut in cloud/src/reconciler/mesofact_static.rs, an unused struct field in cloud/src/app_manifest.rs, an unused fn in cloud/src/reconciler/pond_door.rs, one non_snake_case test name in cloud/src/reconciler/mod.rs) — none touch anything this ticket edited.")
266//! @yah:handoff("Tree-wide stale-reference sweep run (grep for LegacyServiceConfig, generate_compose_bundle, ComposeBundle, mesh_service, compose::, MESH_IP_ENV_FILE, ufw_rules_for_mesh_port, pg_hba_snippet, deploy_compose, ComposeDeployRequest/Response, with_compose_dir, DEFAULT_COMPOSE_DIR, COMPOSE_UNIT across oss/ app/ crates/): the only hits outside this diff's own annotation text are (a) oss/yubaba/crates/yubaba/src/service_records.rs:4-8, a doc comment historically describing why service_records.rs exists instead of 'cloud::mesh_service' — cosmetically references a deleted module name but is prose, not code, does not block compile, left as-is; (b) crates/yah/runner/src/compose.rs and its ~9 call sites — confirmed unrelated (prompt-block composition), explicitly out of scope per the ticket's own gotchas.")
267//! @yah:handoff("Tree anchor at handoff: e0530813af8f7d86f5eb7ea9a6b5a57d386bf30b — the shared tree as I left it. Diff against it (`git diff e0530813af8f7d86f5eb7ea9a6b5a57d386bf30b..HEAD`) to see what landed under you, and quote this SHA rather than 'HEAD' in any revert/restore instruction.")
268//! @yah:next("NOT DONE, in priority order for the next courier (dispatched at Warrior per the leader): (1) cargo check on the yah CLI crate --all-targets (confirm the actual package name first, per the ticket's own instruction — do not guess) to verify app/yah/cli/src/cloud.rs compiles after the ServiceCommands::Deploy removal; I never got a green confirmation on this file myself. (2) cargo test -p <cloud-client's actual package name> to verify crates/yah/cloud-client/src/lib.rs's test-helper edits compile and pass. (3) cargo test --manifest-path oss/yubaba/Cargo.toml --workspace (full suite, not just yah-cloud lib+tests) and specifically confirm services_is_compose_independent_post_r556_f7_t3 passes. (4) Measure a real baseline vs after test-count comparison as the ticket's own Verify section asks — I did not do this under wind-down time pressure.")
269//! @yah:next("If (1)/(2)/(3) turn up more compile errors from downstream CloudConfig/ServiceCommands consumers I did not find in my sweep, fix them the same way as mesofact_bundle.rs was fixed (they are mechanical — remove a `legacy_services:` field initializer or a ServiceCommands::Deploy match arm), then re-run the full-workspace check.")
270//! @yah:next("Once (1)-(3) are green: this ticket is functionally complete except for the mesh_service.rs uncommitted-diff loss, which is @Ashguard:blade's forensics track (session:b49fa21f), not blocking sign-off of THIS ticket's own scope.")
271//! @yah:verify("TEST-COUNT BASELINE MEASURED (R895-T2 courier, session:218a3724). Method: `#[test]`/`#[tokio::test]` attribute count per file, anchor commit e0530813af8f7d86f5eb7ea9a6b5a57d386bf30b (== HEAD at measurement time) vs working tree, then reconciled BY TEST-FUNCTION NAME rather than by count. Deliberate removals total 33 and every one is accounted for: compose.rs -22 (whole file deleted), mesh_service.rs -6 (whole file deleted), config.rs -4 (exactly round_trip_service_legacy, service_bind_interface_round_trips, service_bind_interface_absent_is_none, service_bind_interface_skipped_when_none), cloud-client -1 (deploy_compose_round_trips_via_yubaba) plus 1 rename (services_returns_empty_array_when_no_compose_deployed -> services_returns_empty_array_when_no_workloads_deployed, both names confirmed present on their respective sides). app/yah/cli/src/cloud.rs removed ZERO tests (handle_service_deploy/collect_machine_services/derive_public_hostname carried none). Nothing else went missing: the removed-name lists are exactly the deliberate lists, with no extras.")
272//! @yah:gotcha("A NAIVE before/after test COUNT on this ticket reads backwards, and will mislead the next person who tries it. Three of the touched files gained tests from concurrent peer sessions between the anchor commit e0530813af8f7d86f5eb7ea9a6b5a57d386bf30b and this measurement, all uncommitted and unrelated to R895: oss/yubaba/crates/cloud/src/config.rs 220->218 (-4 mine, +2 peer: a_declared_health_path_survives_the_loader, a_relative_health_path_is_refused_at_load), app/yah/cli/src/cloud.rs 203->206 (0 mine, +3 peer: a_unit_off_the_allowlist_is_refused_before_any_config_load, unit_log_script_carries_the_sudo_prefix, unit_log_script_is_read_only_journalctl), crates/yah/cloud-client/src/lib.rs 43->46 (-2 mine, +5 peer: the logs_* family and the_client_allow_list_matches_the_nodes). So cloud.rs and cloud-client's raw counts go UP across a deletion-only ticket. Reconcile by test-function NAME (comm against `git show <anchor>:<path>`), never by count, on this shared tree.")
273//! @yah:verify("MESH_SERVICE.RS LOSS INVESTIGATED AND CLOSED — NOTHING WAS LOST (forensics session:b49fa21f, 2026-09-13). This supersedes the worry recorded in the co-deletion gotcha above. Method: searched all 3,486 files in .yah/sessions/ (44 mention mesh_service at all) for any edit/write/bash tool call touching that file's content, across this camp's entire history. THERE IS NONE — no session log shows content ever being written to oss/yubaba/crates/cloud/src/mesh_service.rs. The only prior touch was a read by the liveness sniffer (session:77442b98) at 12:36:58 PDT, 39 minutes before deletion, reporting 6242 bytes / 159 lines; `git show HEAD:<path>` is also exactly 6242 bytes (158 by wc -l, the 158-vs-159 being the missing-trailing-newline convention, not a content difference). So the working copy was byte-for-byte identical to HEAD shortly before the rm, and nothing touched it afterward. git log shows exactly one commit ever touched the file (3f7c44ad, the months-old warden->yubaba rename). CONSEQUENCE: the deleted content is fully recoverable from HEAD at any time; no peer work was destroyed. UNEXPLAINED RESIDUE, stated rather than smoothed over: the file did show `M` in git status at the deleting session's start, and forensics could not attribute that flag to any edit — the sniffer was read-only and the deleting courier never read the file. If the flag was real its origin is outside anything this camp's session logs capture (plausibly a mode or line-ending touch). Not chased further because the content question is settled.")
274//! @yah:verify("COMPOSE.RS uncommitted delta also fully accounted for by the same forensics pass, with no residue. Its working copy was 772 lines against HEAD's 765, and the entire 7-line delta was this ticket's own live-growing @yah: annotation block, appended by the relay leader (session:7ca0970b) via board.update between 12:38 and 13:17 PDT. That text is preserved twice over — in the board's own store (board.show R895-T2 resolved throughout, including after deletion) and in the block re-homed into oss/yubaba/crates/cloud/src/lib.rs at 13:15:31, one edit before the rm. Every other edit/write against cloud/src/compose.rs in the session logs predates this relay by weeks (Jul-Aug 2026). Nothing beyond the annotation was uncommitted in that file either.")
275//! @yah:handoff("DOWNSTREAM COMPILE BREAKAGE FOUND AND FIXED (R895-T2 courier, session:218a3724). The previous courier's sweep missed four CloudConfig struct-literal sites that still initialized the deleted `legacy_services` field, so `cargo check -p yah --all-targets` was RED with E0560 (no field named `legacy_services`): 1 error in `yah` (lib) + 5 in `yah` (lib test). Sites removed, same mechanical fix as reconciler/mesofact_bundle.rs: app/yah/cli/src/cloud.rs three occurrences of `legacy_services: vec![],` (formerly :17636, :17781, :18504, all in test CloudConfig builders) and app/yah/cli/src/lan_tunnel.rs:1001 `legacy_services: Vec::new(),`. NOTE the earlier sweep could not have caught these by its own grep method: a tree-wide `rg legacy_services` DOES return them, but the raw output is swamped by this ticket's own multi-KB @yah: annotation prose in oss/yubaba/crates/cloud/src/lib.rs, which is what made them easy to miss. Filter with `rg -n 'legacy_services' --glob '*.rs' . | rg -v '^\\S+:[0-9]+://[/!]'` to drop doc-comment lines and the four code hits stand out immediately. Post-fix the only remaining tree-wide references to LegacyServiceConfig / ServiceCommands::Deploy in CODE (as opposed to annotation prose) are zero.")
276//! @yah:verify("CHECK 1 GREEN, independently corroborated. `cargo check -p yah --all-targets` EXIT=0, zero errors, 3m40s (session:218a3724, 2026-09-13). This is the wide gate: --all-targets covers lib + lib-test + all bins and tests, which is where 5 of the original 6 E0560 errors lived, so the deletion is verified across the full target set rather than the library alone. Corroborated from outside this relay by @Glimmerstone's R365 courier running `cargo check -p yah --lib` against the same tree and reporting zero errors attributable to this ticket. Confirmed by both runs: no LegacyServiceConfig was reintroduced anywhere and nothing was repointed at LegacyMirrorConfig — every fix was a removal, which is the only correct shape here.")
277//! @yah:gotcha("ROOT CAUSE OF THE ~50-MINUTE CAMP-WIDE OUTAGE ON 2026-09-13, recorded so this ticket is not read as a clean deletion. `cargo check -p yah` was red camp-wide for roughly 50 minutes, blocking R365-B27, R365-F22, R856 and R898-T4, each retrying against the shared cargo lock and amplifying the queue. Three compounding causes, in order of blame: (1) LEADER ERROR — the relay lead briefed the first courier that a red tree between deletion phases was expected and not to worry about it. That is true on a private branch and false on a shared working tree, where a broken `cloud` crate takes out every session depending on it. Standing rule since: keep the tree compiling at every stopping point, never end a turn red. (2) The stale-reference sweep that should have caught the remaining sites returned a truncated clean-looking result because this ticket's own multi-KB @yah: annotation prose blew ripgrep's 32KB output cap in the same file — filed as R901-B1. (3) The courier's background build reported \"exit code 0\" in the harness notification while cargo itself reported EXIT=101, because the notification carries the wrapper's status — filed as R901-B2. The four sites that actually broke it were E0560 `legacy_services:` initializers at app/yah/cli/src/cloud.rs x3 and app/yah/cli/src/lan_tunnel.rs:1001. Two peers independently misread the diff as a LegacyServiceConfig -> LegacyMirrorConfig RENAME; it is a deletion, LegacyMirrorConfig is a separate live type, and any fix that repoints a call site at it is wrong.")
278//! @yah:verify("CHECKS 1-4 MEASURED (R895-T2 courier, session:218a3724), pass/fail counts not logs. (1) `cargo check -p yah --all-targets` = EXIT 0, zero errors, 3m40s, after the four legacy_services removals; corroborated independently by a later desktop run whose dependency leg compiled `yah` (lib) with 26 warnings / 0 errors. (2) `cargo test -p cloud-client` = GREEN, 51 passed + 1 passed (doc/second binary), 0 failed, 0 ignored; the renamed test `services_returns_empty_array_when_no_workloads_deployed` is present and passes, confirming the cloud-client test-helper edits both compile and run. (3) `cargo test --manifest-path oss/yubaba/Cargo.toml --workspace` = 11 test binaries, 10 fully green totalling 2204 passed / 0 failed / 11 ignored; THE PIN ASKED FOR, `services_is_compose_independent_post_r556_f7_t3`, PASSES (yubaba-ws.log:2357) — so GET /services is confirmed compose-independent and `yah cloud service status` survives the deletion of `deploy`. The 11th binary reported 61 passed / 33 failed; see the separate gotcha on those. (4) `cargo check -p desktop --all-targets` = RED, but NOT from this ticket: the sole error is app/yah/desktop/src/agent.rs:6364:14 error[E0382] use of moved value `orig_job` (1 error in desktop lib, same 1 in desktop lib test). Zero occurrences of LegacyServiceConfig / LegacyMirrorConfig / legacy_services / compose / ServiceCommands anywhere in that build output. That run's input closure was unchanged for its whole duration (daemon skew verdict: no skew), so the result is trustworthy rather than a racing-tree artifact.")
279//! @yah:gotcha("DESKTOP IS RED FOR A REASON THAT IS NOT THIS TICKET — this retires R895-F1's standing assumption rather than confirming it. F1 was signed off assuming desktop was red ONLY because of T2's in-flight LegacyServiceConfig deletion. Measured 2026-09-13 after T2's deletion was complete and `-p yah` was green: `cargo check -p desktop --all-targets` still fails, with exactly one distinct error, and it is app/yah/desktop/src/agent.rs:6364:14 error[E0382] `use of moved value: orig_job` — a borrow-check error in a peer's UNCOMMITTED in-flight edit, not a missing-type error. Evidence it is not ours and not mine to fix: (a) the whole build output contains zero mentions of LegacyServiceConfig, LegacyMirrorConfig, legacy_services, compose or ServiceCommands; (b) `git status --short` shows agent.rs as ' M' (dirty), and the anchor commit e0530813af8f7d86f5eb7ea9a6b5a57d386bf30b has entirely different code at those lines, so the failing hunk exists only in the working tree; (c) `orig_job` appears 2x at the anchor vs 9x in the worktree, i.e. a peer is mid-way through expanding its use; (d) the failing site carries an `R198-T5-B` comment about restating `job` on a resumed session's SessionStarted event. Per shared-tree doctrine this is a peer's live file and was deliberately NOT edited. No session on camp.roster currently lists R198-T5-B, so the owner could not be resolved to a name from the roster — routing it to the R895 leader rather than asserting an author.")
280//! @yah:verify("CHECK 1 RE-VERIFIED CLEAN OF SKEW, and the raft failures MEASURED rather than inferred (session:218a3724, 2026-09-13). `cargo check -p yah --all-targets` re-run to final: EXITF=0, zero errors, and this run's input closure was UNCHANGED for its whole duration (daemon verdict: no skew) — the earlier green had carried a 3-input skew warning (peers editing camp.rs, prelude.rs, codex_oauth.rs), so this re-run is what actually settles it. RAFT SUITE: re-ran the full yubaba workspace a second time under materially lighter load (camp.machine at run 2: load 9.78/12.58/15.02 on 15 cpus, 2 rustc/cargo processes, vs 16.82 and 9 processes during run 1). Result 60 passed / 34 failed vs run 1's 61 passed / 33 failed, and run 1's 33 failing names are a STRICT SUBSET of run 2's 34 (the extra is raft_leader_pin::an_off_anchor_leader_hands_leadership_to_the_anchor_region). So these are persistent failures with one timing-sensitive straggler, NOT contention artifacts — halving the load did not fix them. Every panic is a leader-election timeout (\"voters never agreed on a leader\" x6, \"nodes never agreed on a leader\" x4, 22 \"timed out\" strings). Run 2 also carried a no-skew verdict. The pin `services_is_compose_independent_post_r556_f7_t3` passed in BOTH runs. Note camp.machine reported quiet:false on both occasions and never reached a formally quiet window — 4 live slots in the foreign `noisetable` camp, which camp.roster cannot see by contract — so this is labelled a lighter-load measurement, not a quiet-machine one.")
281//! @yah:gotcha("THE 33-34 RAFT/ROLLOUT FAILURES IN THE YUBABA WORKSPACE ARE PRE-EXISTING AND NOT THIS TICKET'S — attribution established four ways, not assumed. (a) STRUCTURAL: the deletion removed a struct FIELD, which is a compile-time failure mode (E0560/E0609); these binaries compiled and ran 60-61 passing tests each, so a removed field cannot be producing a runtime election timeout. (b) BY SEARCH: `grep -rn --include='*.rs' -E 'legacy_services|LegacyServiceConfig' oss/yubaba/` returns ONLY annotation/doc-comment prose (this ticket's own block in cloud/src/lib.rs, a config.rs:1432 doc line, a paths.rs R222-era handoff) — zero code references anywhere in the yubaba workspace, so the deleted field was never reachable from its raft paths. The only `cloud::CloudConfig` use anywhere under crates/yubaba/tests/ is pond_reconciler_smoke.rs:39,65, which is NOT among the failures. (c) BY REPETITION: failures reproduce across two independent runs at very different machine loads, run 1's set a strict subset of run 2's. (d) BY SYMPTOM: all are raft leader-election timeouts, a shape unrelated to config deserialization. These belong to whoever owns the yubaba raft/rollout suite, not to R895. SEPARATE TRAP WORTH KNOWING, found the hard way in this session: `rg` IS NOT ON PATH in the camp daemon's build shell (W302 relocates Bash calls to `sh -c` under the daemon). A command shaped `rg ... || echo \"(none)\"` therefore prints the innocent-looking fallback because the BINARY IS MISSING, not because there were no matches — a false negative that reads exactly like a clean sweep. Use `grep -rn --include='*.rs'` in any daemon-relocated command, or check the exit code explicitly. This is the same class as R901-B1 and nearly laundered a fabricated proof into this ticket.")
282//! @yah:handoff("LEADER SIGN-OFF (Ashguard:blade, R895 relay lead, 2026-09-13). Implemented across two couriers — @Miravel:griffin (session:49bbf176) did the deletion and the salvage and wound down cleanly at its context limit; @Ashguard:coffee (session:218a3724) finished the downstream fixes and all verification. DELIVERED: cloud::compose (765 lines) and cloud::mesh_service (158 lines) deleted; LegacyServiceConfig and CloudConfig::legacy_services removed; the shipped `yah cloud service deploy` CLI verb, the POST /compose route and handler, ComposeDeployRequest/Response, ServerState::compose_dir, DEFAULT_COMPOSE_DIR, with_compose_dir and the cloud-client mirror all removed; four downstream E0560 sites fixed (app/yah/cli/src/cloud.rs x3, app/yah/cli/src/lan_tunnel.rs:1001) plus reconciler/mesofact_bundle.rs. THE SALVAGE — the irreplaceable half — LANDED: W206/R558-T2's tenant network-isolation policy is transplanted into .yah/docs/working/W343-per-workload-mesh-addressing.md (\"Open, deliberately\") and the oss/kamaji/crates/kamaji/src/container_net.rs module docs, both marked required-and-unimplemented and citing R895-F3. legacy_mirrors/LegacyMirrorConfig was investigated and correctly KEPT as live — two peers independently misread this diff as a rename toward it; it is a deletion and that misreading is pinned in the gotchas.")
283//! @yah:verify("ALL FIVE CHECKS CLOSED at sign-off (session:218a3724). (1) `cargo check -p yah --all-targets` EXIT 0, re-run to a final NO-SKEW verdict so the earlier skew-flagged green is settled. (2) `cargo check -p desktop --all-targets` RED, and confirmed unrelated: sole distinct error is E0382 use-of-moved-value `orig_job` at app/yah/desktop/src/agent.rs:6364, a peer's uncommitted work, filed as R902-B1. Borrowck runs only after type-check, so desktop type-resolving proves this ticket's deletion is clean from desktop's side. (3) `cargo test -p cloud-client` 52 passed / 0 failed, renamed test services_returns_empty_array_when_no_workloads_deployed passing. (4) yubaba `--workspace` 2204 passed / 0 failed across 10 of 11 binaries, with services_is_compose_independent_post_r556_f7_t3 — the pin that keeps `yah cloud service status` alive after `deploy` dies — PASSING IN BOTH RUNS. The 11th binary's 33-34 raft leader-election failures were measured at two load levels, proved load-independent, attributed away from this ticket four ways, and filed as R903. (5) Test-count baseline -33, reconciled BY NAME rather than by count via `comm` against `git show <anchor>:<path>` — necessary because peers' uncommitted tests made two files' raw counts go UP across a deletion-only ticket.")
284
285pub mod almanac_dispatch;
286pub mod app_manifest;
287pub mod asset_journal;
288pub mod asset_status;
289#[cfg(test)]
290mod asset_status_tests;
291mod atomic_write;
292pub mod capability;
293pub mod cloud_init;
294pub mod config;
295pub mod envoy;
296pub mod identities;
297pub mod inner_door;
298// R374-F3: `local_runtime` + `provider::s3_sign` moved to the `local-driver`
299// crate so yubaba can own MinIO lifecycle without a reverse yubaba→cloud dep.
300// `local_driver_glue` carries the cloud-config adapter that used to live as
301// `LocalContainerSpec::from_provider_config`.
302pub mod local_driver_glue;
303pub mod mesh;
304pub mod migrate;
305pub mod multi_root;
306pub mod paths;
307pub mod proc_control;
308pub mod provider;
309pub mod provision;
310pub mod reconciler;
311pub mod recovery_journal;
312pub mod release_manifest;
313// R898-F1 / W348 §2.2: `DomainConfig.routes` compiled into ONE ordered table —
314// path, mode, RESOLVED origin, headers, auth — rendered to both front doors.
315pub mod route_table;
316pub mod state;
317pub mod status;
318pub mod topology;
319pub mod validate;
320
321pub use almanac_dispatch::dispatch_on_change;
322pub use asset_journal::{AssetState, AssetStatusEvent, AssetStatusJournal};
323pub use capability::Capability;
324pub use config::{
325 BucketLogEntry, CampCloudDbs, CloudConfig, CloudDb, ConnectSpec, DbCatalog, DevDb, GitSource,
326 IngressDecl, IngressEdge, IngressProvider, LegacyMirrorConfig,
327 MachineConfig, MirrorAssignment,
328 MirrorConfig, MirrorProviderSlot, MirrorShape, PondDb, PondDbKind, Provider, ProviderConfig,
329 ServiceComponent, ServiceConfig, TopologyConfig, WorkloadConfig, WorkloadConfigError,
330};
331pub use local_driver::pond_warden::{warden_container_label, warden_container_name};
332pub use local_driver::{
333 canonical_label, canonical_name, ContainerRunSpec, ContainerState, CustomDockerHostProvider,
334 DetectedRuntime, LocalContainerSpec, LocalDockerRuntime, LocalRuntime, OwnedContainer,
335 RuntimePref, RuntimeProvider, SocketRuntimeProvider, LABEL_KEY, NAME_PREFIX,
336};
337pub use local_driver_glue::local_container_spec_from_provider;
338pub use provider::{
339 on_ingress_owner_changed, reconcile_assignment, BucketAcl, BucketRef, CfAccountInfo,
340 CloudflareClient, CloudflareEnvoy, CreateR2BucketResult, CreateTokenResult, CreateTunnelResult,
341 DigitalOceanEnvoy, DnsRecordDetail, FloatingIpAssignOutcome, FloatingIpProvider,
342 FloatingIpState, FloatingIpTarget, GrantScope, HetznerDriver, HetznerEnvoy, HetznerFloatingIp,
343 Location, MachineProvider, OvhFloatingIp, ProjectId, R2BucketInfo, R2CustomDomain, ServerId,
344 ServerSpec, ServerStatus, ServerSummary, TokenGrant, TunnelConnState, TunnelDnsRecord,
345 TunnelDriftRow, TunnelDriftState, VultrFloatingIp, WorkerDeployResult, DOOR_DNS_GRANTS,
346 MESOFACT_STATIC_GRANTS, TUNNEL_EDIT_GRANTS,
347};
348#[cfg(feature = "local-docker")]
349pub use provider::{LocalDockerEnvoy, LocalDockerProvider};
350pub use reconciler::{
351 collect_live_derive_hashes, compute_derive_cache_candidates, compute_live_set,
352 compute_prune_candidates, compute_service, derive_minio_key, execute_derive_cache_prune,
353 execute_prune, load_service_and_mirror, mesofact_static::WORKER_SCRIPT, new_sync_id,
354 pond::MINIFLARE_SIM_SCRIPT, publish_to_pond, summarize, CellStatus, CloudflareWorkerReconciler,
355 ContainerOptions, ContainerReconciler, DeriveCacheLiveHashes, DerivePruneCandidate, DriftEntry,
356 HeadscaleReconciler, HealthState,
357 MesofactStaticReconciler,
358 MirrorObservation, PondOptions, PondPublishReport, PondState, ProviderScope, PruneCandidate,
359 PruneOutcome, PruneReport, ReconcileCtx, Reconciler, RunningWorkload, RunningWorkloadSummary,
360 Runtime, ServiceStatus, StaticAssetReconciler, StatusSummary, SyncHistoryEntry, SyncOutcome,
361 SyncState, WireContainerStatus,
362};
363// R918-F5 — the local-process reconciler is unix-only; see the gate on
364// `reconciler::local_process` for why it is gated whole rather than half-ported.
365#[cfg(unix)]
366pub use reconciler::LocalProcessReconciler;
367pub use status::{collect_machine_report, AgentProbe, DriftFinding, MachineReport};