Skip to main content

Crate cloud

Crate cloud 

Source
Expand description

@yah:relay(R040, “yah-cloud Phase 1: mirror bootstrap (Hetzner driver, Podman runtime, yah-yubaba)”) @yah:status(review) @yah:parent(Q058) @yah:next(“yah-A track (8 tickets) — substrate for noisetable cloud + future managed camps; gates noisetable C6 (E2E) and D1–D5 (corpus migration)”) @yah:next(“Adjacent relays: R019 (SSH-RPC remote camps), R032 (yah-agentd), R034 (identity registry — hostkeys overlap)”) @yah:next(“Reparented under Q058 cloud quest 2026-05-07. Largely superseded by R091 (yubaba orchestration + integration testing) and R092 (managed-camps CLI + .yah/cloud/ schema) — net new work should land on those relays. Remaining unique R040 scope (Cloudflare Tunnel cookbook F15, pg-on-mesh recipe F16, mesh-promote phasing F18-F19) stays here as historical/operational reference; consider archiving once R091+R092 reach review.”) @yah:next(“Refinement open questions: subcommand framework alignment with yah arch/board, where hostkey fingerprints live, reverse proxy choice. Crate-layout question soft-answered by this stub being a sibling crate; A1 may still refactor into a sub-module of an existing crate if desired (move the //! block + delete this stub).”) @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md) @arch:see(architecture/yah-managed-camps-topology.md)

@yah:ticket(R040-F1, “A1: .yah/cloud/ schema + parser (MachineConfig, MirrorConfig, ServiceConfig)”) @yah:at(2026-05-05T00:29:11Z) @yah:assignee(agent:claude) @yah:status(review) @yah:phase(P1) @yah:parent(R040) @yah:next(“Per-file TOML layout: machines/.toml, mirrors/.toml, services/.toml — diffability + parallel editing”) @yah:next(“Sketch types in arch doc; refine to fit actual yah crate structure (this stub crate is one option; folding into an existing crate is the other)”) @yah:verify(“cargo test -p cloud config::tests round-trips all three config types”) @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md) @yah:handoff(“Schema + parser complete: MachineConfig, MirrorConfig, ServiceConfig, BucketSpec, PortMapping in crates/yah/cloud/src/config.rs. Per-file TOML layout via CloudConfig::load. MachineConfig::save for write-back (A4 hostkey). 6 tests pass (round-trips + tempdir integration + save/reload). tempfile dev-dep added. No warnings.”) @yah:next(“A2: yah cloud subcommand stub — wire into yah CLI using same clap/subcommand pattern as yah board/arch. Commands: machine {status,provision,destroy}, mirror {show,status}, service {deploy,status}, agent {ping,services,logs}. All no-op except status/show which dump parsed .yah/cloud/ config. Verify: yah cloud mirror show noisetable prints declared regions/services.”)

@yah:ticket(R040-F3, “A3: Hetzner driver — MachineProvider trait + hcloud impl (server, bucket, status, destroy)”) @yah:at(2026-05-05T00:29:11Z) @yah:assignee(agent:claude) @yah:status(review) @yah:phase(P1) @yah:parent(R040) @yah:next(“API token from HETZNER_API_TOKEN env; document in .yah/cloud/SECRETS.md”) @yah:next(“Gotcha: Hetzner has TWO APIs (Cloud and Robot) — Phase 1 is Cloud only; Object Storage is S3-compat, separate surface”) @yah:next(“Verify before A6: Hetzner Hillsboro + Ashburn Object Storage both GA; CPX-22 available in PDX/IAD/FSN”) @yah:verify(“cargo test -p cloud provider::hetzner – –ignored (gated; needs token)”) @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md) @yah:handoff(“MachineProvider trait + HetznerDriver impl complete in crates/yah/cloud/src/provider/. Cloud API (create_server, server_status, destroy_server) uses HETZNER_API_TOKEN via reqwest+bearer. create_bucket uses Hetzner Object Storage S3-compat with AWS Sig V4 signing (HETZNER_S3_ACCESS_KEY + HETZNER_S3_SECRET_KEY). Location enum maps pdx/iad/fsn to hetzner cloud IDs and S3 endpoints. 9 unit tests pass; 2 integration tests behind #[ignore] gate. .yah/cloud/SECRETS.md created documenting all three token types.”) @yah:next(“A2: wire yah cloud subcommand stub into CLI (yah board/arch pattern). All commands no-op except machine status + mirror show which dump parsed .yah/cloud/ config. Verify: yah cloud mirror show noisetable prints declared regions/services.”)

@yah:ticket(R040-F4, “A4: cloud-init YAML + yah cloud machine provision (Debian + Podman + Tailscale + yah-yubaba)”) @yah:at(2026-05-05T00:29:11Z) @yah:status(review) @yah:phase(P1) @yah:parent(R040) @yah:next(“Template at .yah/cloud/cloud-init/mirror.yml; renderer substitutes YAH_WARDEN_BASE64 + HEADSCALE_PREAUTH_KEY + TAGS per machine”) @yah:next(“Hostkey generated ON the machine (ssh-keygen in cloud-init); private half never touches dev machines”) @yah:next(“Provision command polls cloud-init done, reads back hostkey fingerprint via yah-yubaba /identity, writes back to machine.toml”) @yah:next(“Surface /var/log/cloud-init.log and /var/log/cloud-init-output.log on failure”) @yah:verify(“yah cloud machine provision noisetable-pdx-1 –dry-run prints rendered cloud-init YAML”) @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md) @yah:handoff(“Cloud-init template + renderer + provision command landed. Template at .yah/cloud/cloud-init/mirror.yml (canonical) with embedded copy at crates/yah/cloud/templates/mirror.yml (drift-test in cloud_init::tests::embedded_template_matches_workspace_canonical). Renderer at crates/yah/cloud/src/cloud_init.rs substitutes {{KEY}} (no-spaces) for MACHINE_NAME, YAH_WARDEN_BASE64, HEADSCALE_PREAUTH_KEY, TAGS; doc comments use {{ KEY }} (with spaces) so they survive rendering. find_unsubstituted() trips on any unsubstituted real placeholder. Provision orchestrator at crates/yah/cloud/src/provision.rs decoupled from concrete provider (takes &dyn MachineProvider). CLI wired in app/yah/cli/src/cloud.rs: yah cloud machine provision [–dry-run] [–yubaba ]; –dry-run with no flags uses placeholders for the yubaba binary + preauth key (with NOTE lines explaining); live path requires –yubaba + HEADSCALE_PREAUTH_KEY + HETZNER_API_TOKEN. 20 unit tests pass (8 new for cloud_init + provision). Vocab: per-node infra daemon is yah-yubaba (was yah-mirror-agent — renamed 2026-05-01 because mirror is a deployment, not a node).”) @yah:next(“A8 yah-yubaba /identity endpoint: poll cloud-init done after Hetzner accepts the create_server call, GET /identity for hostkey fingerprint, write back to .yah/cloud/machines/.toml via MachineConfig::save(). Wiring stub already in handle_provision after rt.block_on — replace the next: print with the real poll loop.”) @yah:next(“Optional polish: surface /var/log/cloud-init.log + /var/log/cloud-init-output.log on cloud-init failure (probably via the yah-yubaba /diagnostics endpoint when A8 lands, since direct SSH from the dev machine isn’t part of Phase 1).”)

@yah:ticket(R040-T5, “A5: yah cloud machine status — drift report (declared vs Hetzner vs agent health)”) @yah:at(2026-05-05T00:29:11Z) @yah:status(review) @yah:phase(P2) @yah:parent(R040) @yah:next(“Drift categories: missing-server, wrong-machine-type, missing-bucket, hostkey-fingerprint-mismatch, missing-co-tenant, agent-unreachable, service-unhealthy”) @yah:next(“Exit non-zero on drift unless –quiet; bare ‘yah cloud machine status’ dumps all declared machines”) @yah:next(“Can land any time after A4 (parallel-safe with A6/A7/A8)”) @yah:verify(“Pre-A6: reports ‘not provisioned’ for each declared machine”) @yah:verify(“Post-A6+A7+A8: all three report ‘in sync’”) @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md) @yah:handoff(“Drift module landed at crates/yah/cloud/src/status.rs (DriftFinding + MachineReport + collect_machine_report). MachineProvider trait gained find_server_by_name (GET /servers?name=) + bucket_exists (S3 HEAD) — Hetzner impls share a refactored sign_s3_empty_body helper for PUT/HEAD signing. CLI: yah cloud machine status [name] [--quiet] queries Hetzner, prints declared+live state with per-finding [drift]/[note ] tags, and bails non-zero on real drift unless –quiet. Soft findings (missing creds, A8-not-landed, bucket-uncheckable) are surfaced as notes only. Pre-A6 verify: smoke-tested against this workspace’s noisetable-pdx-1 declaration with no token — reports ‘unknown’ + provider-error note, exit 0. 7 new drift unit tests pass via in-memory FakeProvider; 27 total pass in crates/yah/cloud.”) @yah:next(“A8 will plug in real agent-ping (yah-yubaba /health + /identity). The HostkeyFingerprintMismatch + AgentUnreachable variants are wired through the report rendering — A8 just needs to flip the AgentUnreachable emit-site in collect_machine_report() from the always-emit pre-A8 stub to a live probe.”)

@yah:ticket(R040-T6, “A6: provision noisetable-{pdx,iad,fsn}-1 (operational; yah dogfoods on noisetable’s mirror, no yah-cloud SaaS)”) @yah:at(2026-05-05T00:29:11Z) @yah:status(review) @yah:phase(P2) @yah:parent(R040) @yah:verify(“yah cloud machine status –path /Users/user/ss/noisetable — all three ‘in sync’”) @yah:verify(“aws s3 ls –endpoint-url s3://noisetable-assets--1 — all three list”) @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md) @yah:handoff(“Code path complete across R040-F1..F8 + cloud-init runcmd alignment. Gated on operator runbook (release tag → live provision against Hetzner). Open repo-config decision (noisetable side, not yah): noisetable/.yah/.gitignore line 3 ignores ‘/cloud’ wholesale — blocks tracking machines/, mirrors/, SECRETS.md. 30 cloud-crate + 17 yah cloud:: tests pass.”) @yah:next(“Operator runbook lives in events.jsonl history; the strategic next is the live mirror-status verify (yah cloud machine status --path /Users/user/ss/noisetable → all three ‘in sync’) which is gated on the release-tag + provision steps.”)

@yah:ticket(R040-F7, “A7: yah cloud service deploy — Podman compose generation + Cloudflare front + tier isolation”) @yah:at(2026-05-05T00:29:11Z) @yah:assignee(agent:claude) @yah:status(review) @yah:phase(P3) @yah:parent(R040) @yah:next(“Resolves service configs + mirror declarations → per-machine compose.yml → push via yah-yubaba (A8) → systemctl restart yah-cloud-services”) @yah:next(“Tier isolation at SOFTWARE layer: separate PG roles per service set, distinct Headscale tags (tier:t0 yah meta vs tier:t2 noisetable assets), mesh_only flag controls Cloudflare exposure”) @yah:next(“Cloudflare DNS records (manual, in SECRETS.md): pdx.cloud.noisetable.example etc. → public IPs via orange-cloud”) @yah:next(“Gotcha: Cloudflare free tier proxies HTTP/HTTPS but NOT raw TCP/UDP — Headscale mesh uses direct WireGuard between machines”) @yah:next(“Assumes: Cloudflare account exists + noisetable.example DNS delegated to it (else spawns Cloudflare-bootstrap sub-ticket)”) @yah:next(“Open question (refinement): Caddy vs nginx vs direct-to-Cloudflare for per-machine reverse proxy”) @yah:verify(“curl https://pdx.cloud.noisetable.example/healthz (asset-registry) returns 200”) @yah:verify(“curl https://pdx.cloud.yah.example/healthz (yah meta-directory) returns 200”) @yah:verify(“yah cloud service status –all — all green on all 3 machines”) @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md) @yah:handoff(“Compose generation + yubaba service deploy + CLI wiring landed. New file: crates/yah/cloud/src/compose.rs (generate_compose_bundle: tier-named networks, Caddy service for mesh_only:false, Caddyfile with hostname when cloud_domain is set in mirrors/*.toml). MirrorConfig gained cloud_domain: Option. Yubaba: replaced 501 stubs with real /services (podman compose ps –format json; empty array when no compose.yml) and /compose (write compose.yml + Caddyfile + write yah-cloud-services.service unit + systemctl enable –now). ServerState gained compose_dir: PathBuf (default /etc/yah-cloud) + with_compose_dir() builder. cloud-client: added ComposeDeployRequest / ComposeDeployResponse wire types + deploy_compose() method. CLI: yah cloud service deploy resolves machines via mirror→service lookup, generates per-machine bundle (using location+cloud_domain as public hostname when available), and POSTs to each machine’s yubaba. yah cloud service status [name|–all] queries /services on each machine and prints [+]/[-] per container. All tests pass: 18 yubaba, 47 cloud, 8 cloud-client, 120 yah bin.”) @yah:next(“Caddy in compose needs a Cloudflare origin cert or Let’s Encrypt config once real domains land — Caddyfile currently uses :port placeholders which work for testing. Operator action: set cloud_domain in mirrors/.toml then re-run yah cloud service deploy to regenerate with real hostnames.”) @yah:next(“Cloudflare Tunnel integration (R040-F15): add a cloudflared service to the compose stack — currently services are exposed via Caddy on public ports 80/443 (orange-cloud model). F15 will replace with no-public-port cloudflared tunnel approach.”) @yah:next(“podman compose ps –format json output format varies by podman-compose version — current implementation passes raw JSON through. If operators hit parse issues, add a normalization layer in get_services() that extracts Name+State from both podman-compose and Podman 4.x native formats.”) @yah:next(“yah cloud service rolling not yet implemented — stub left in place. Compose rolling restarts are just podman compose up –no-deps ; wire it when rolling deploys are needed.”)

@yah:ticket(R040-F8, “A8: yah-yubaba (Rust binary on machine) + yah-cloud-client + desktop Cloud panel”) @yah:at(2026-05-05T00:29:11Z) @yah:status(review) @yah:phase(P3) @yah:parent(R040) @yah:verify(“cargo test -p yubaba -p cloud -p cloud-client + cargo test -p yah –bin yah cloud:: — 18 + 47 + 9 + 17 tests pass”) @yah:verify(“cargo run -p yubaba – serve –bind 127.0.0.1:7450 –state /tmp/i.json then curl POST /register-hostkey with a real ssh-ed25519 pubkey, then yah cloud machine attach <machine> --host 127.0.0.1:7450 --wait 5 --path <camp> prints attached: <machine> (fingerprint recorded: SHA256:…) and writes the line into .yah/cloud/machines/.toml”) @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md) @arch:see(architecture/yah-managed-camps-topology.md) @yah:handoff(“Session 3 landed yah cloud machine attach <name> (the hostkey write-back follow-on). app/yah/cli/src/cloud.rs::handle_attach loads MachineConfig, polls cloud-client GET /identity until success or –wait seconds elapse (default 180s, retry every 3s on NotRegistered or Transport errors), then calls MachineConfig::save() with the fingerprint. Decision logic in decide_attach_action covers all four states: Write (no fingerprint on file), Idempotent (match — no-op write), Mismatch (bail with –force hint), Overwrite (mismatch + –force). Subcommand args harmonized with agent ping: –host (default http://:7443), –wait , –force, –path. handle_provision’s stale stub print (next: hostkey fingerprint write-back lands with A8) replaced with next: once cloud-init finishes, run yah cloud machine attach <name>. End-to-end smoke verified all four AttachAction paths against a live yubaba on 127.0.0.1:7450 — pre-registration polling + bail, fresh-write, idempotent re-run, mismatch refusal, –force overwrite. 5 new unit tests (decide_attach_action_{writes_when_no_declared,idempotent_on_match,bails_on_mismatch_without_force,overwrites_with_force,force_is_idempotent_when_match}) — total 9 cloud:: tests in the yah bin. Cumulative R040-F8 deliverables across sessions 1–3: yubaba binary (12 tests), cloud-client crate (6 tests + doctest), AgentProbe trait + drift wire-up (4 new in cloud, 29 total), yah cloud agent {ping,services,logs} CLI, yah cloud machine attach, provision next-step hint correctly points at attach. Out of scope still: Tailscale-mesh discovery, desktop Cloud panel, mTLS hardening.”) @yah:next(“Tailscale-mesh discovery for –host: yah cloud agent ping/services/attach defaults to http://:7443 which only resolves inside the mesh. Smart discovery would query tailscale status --json for the machine’s mesh IP. Until then, operators pass –host explicitly. Add a TS_DEVICES env shortcut or yah cloud agent ip <machine> helper if the manual path stays painful.”) @yah:next(“Desktop Cloud panel (out of session, larger): packages/yah/ui/src/cloud/ — machines list, per-machine service list, log tail. Tauri command bridges in app/yah/desktop/src/. Pulls on the cloud-client crate (already in workspace). Defer until at least Hostkey write-back lands so the panel has something useful to display per-machine.”) @yah:next(“Hardening parking lot (post-Phase-1): mTLS (server cert from machine hostkey, client cert from desktop user identity); /metrics endpoint (Prometheus); streaming /logs from journald with heartbeat. None of these gate Phase 1 verify; cloud-client’s CloudClient::with_timeout exists so log streaming has a hook.”) @yah:gotcha(“yubaba binds 0.0.0.0:7443 by default; cloud-init template adds ufw rules to restrict to tailscale0. If a future deployment skips ufw, public IP exposure is a real risk. Better long-term: bind to tailscale0 IP directly (cloud-init systemd unit can compute it via ip -4 addr show tailscale0).”) @yah:gotcha(“Re-running register-hostkey with a different pubkey replaces the stored identity (idempotent for same key, swaps for different). Probably what you want, but worth flagging if someone designs an audit log around ‘first registration wins’.”) @yah:gotcha(“Default URL resolution (http://<machine.name>:7443) is shared between WardenHealthProbe (status drift), agent ping/services, and machine attach — outside Tailscale mesh every default-host call against a real Hetzner machine will fail. Status emits AgentUnreachable as a soft note; agent/attach surface the transport error directly. Once mesh discovery lands, all three call sites should resolve via Tailscale before falling back to plain hostname.”)

@yah:ticket(R040-F18, “Phase 1a — yah mesh start: camp bootstrap with stable URL from day 1”) @yah:at(2026-05-05T00:29:11Z) @yah:status(review) @yah:assignee(agent:claude) @yah:parent(R040) @yah:phase(P1) @yah:next(“Implements Phase 1a from .yah/docs/architecture/A041-yah-mesh-bootstrap.md — design committed 2026-05-04. Phase 1b lands as R040-F19, Phase 2 (HA) as R040-F20+F21.”) @yah:next(“Subcommand surface: yah mesh start brings up Headscale locally (yah-desktop or camp daemon); auto-detects direct vs. cloudflared reachability; configures DNS for mesh.; stores URL in vault as mesh-url so provision picks it up.”) @yah:next(“Cloudflared bundling: yah-desktop should manage the cloudflared binary so the operator doesn’t need a separate install. Same daemon can carry yah serve ingress + Headscale tunnel under one install — that’s the user-acceptable story.”) @yah:next(“Stable URL contract: mesh.<your-domain> set ONCE here and never changes across Phase 1b promotion or Phase 2 leader changes. Nodes embed this URL in their tailscale config; downstream phases re-point DNS, not nodes.”) @yah:next(“Headscale binary provisioning: download + manage like cloudflared (precedent: yah already shells out to cloudflared/tailscale/cargo). Config template per arch doc ‘Headscale config’ section. Initial ACL writes a permissive tag:* policy.”) @yah:next(“Provision integration: when mesh-url is set, yah cloud machine provision calls Headscale API for a single-use preauth key and embeds {MESH_URL, PREAUTH_KEY, TAGS} into cloud-init’s tailscale up line. Static headscale-preauth-key in vault becomes optional.”) @yah:next(“Also lands yah mesh status (coordinator location, connected nodes, preauth count). yah mesh backup/restore can be Phase-1a stubs; full litestream comes with R040-F21.”) @yah:verify(“yah mesh start on a fresh camp brings up Headscale + DNS, prints stable URL, stores in vault”) @yah:verify(“yah cloud machine provision with mesh-url set generates preauth key and joins the new machine to the camp’s Headscale”) @yah:verify(“tailscale status on the new machine shows it connected via the camp’s tunnel; node-to-node ping works”) @arch:see(.yah/docs/architecture/A041-yah-mesh-bootstrap.md) @yah:handoff(“Phase 1a landed. New files: crates/yah/cloud/src/mesh.rs (HeadscaleClient + generate_headscale_config + headscale_download_url + DEFAULT_ACL_POLICY), app/yah/cli/src/mesh.rs (yah mesh start/status/stop/backup-stub/restore-stub). cloud_init::RenderInput gained mesh_url: Option; renderer substitutes {{MESH_LOGIN_SERVER_ARG}} → ’ –login-server ’ or ‘’ depending on field. provision::build_request now takes mesh_url: Option. template mirror.yml updated with MESH_LOGIN_SERVER_ARG placeholder. cloud.rs::handle_provision now reads mesh-url from vault/HEADSCALE_URL env; when set, calls HeadscaleClient::from_vault_or_env() to auto-generate a single-use preauth key (falls back to static headscale-preauth-key when no API key is available). yah mesh start downloads headscale binary from GitHub (HEADSCALE_VERSION=0.23.0), writes config.yaml + acls.yaml, creates default user, spawns headscale serve in background, saves PID, stores mesh-url in vault, prints DNS/tunnel operator instructions. yah mesh status shows PID, mesh-url, node count + per-node online status. yah mesh stop does SIGTERM+SIGKILL with 5s grace. Cloud crate: 36 tests pass (+4 new: render_with_mesh_url_adds_login_server, render_without_mesh_url_no_login_server, provision::build_request_with_mesh_url_adds_login_server, mesh::config_contains_server_url, mesh::config_all_paths_in_data_dir, mesh::download_url_current_platform). yah lib: 116 unit tests pass (+4 new mesh:: tests). arch dogfood integration tests are pre-existing failures unrelated to this work.”) @yah:next(“Operator runbook for Phase 1a: (1) run yah mesh start –url https://mesh. on the camp; (2) set up DNS + cloudflared (or direct port-forward) for that URL; (3) headscale apikeys create to get an API key; (4) yah keys set headscale-api-key; (5) yah cloud machine provision –yubaba-url –yubaba –path /Users/user/ss/noisetable — provision now auto-generates a preauth key from Headscale.”) @yah:next(“Phase 1b (R040-F19): yah mesh promote — migrates Headscale from camp to first cluster machine. The stable-URL contract from this ticket means no node reconfiguration is needed; only DNS re-points.”) @yah:next(“Cloudflared bundling (R040-F15): cloud-init template needs cloudflared install + token. yah mesh start should detect reachability and auto-configure cloudflared when –url points through a Cloudflare Tunnel. The operator-instruction path this session is the MVP fallback.”) @yah:next(“Headscale API key auto-generation: currently the operator must manually run headscale apikeys create. A future improvement: yah mesh start can use the local UNIX socket (headscale.sock) to generate the API key and store it in the vault automatically without needing the binary in PATH.”)

@yah:ticket(R040-F19, “Phase 1b — yah mesh promote: migrate Headscale from camp to first cluster machine”) @yah:at(2026-05-05T00:29:11Z) @yah:assignee(agent:claude) @yah:status(review) @yah:phase(P2) @yah:parent(R040) @yah:next(“Implements Phase 1b from .yah/docs/architecture/A041-yah-mesh-bootstrap.md. Depends on R040-F18 (Phase 1a) shipping the stable-URL contract and camp Headscale.”) @yah:next(“Subcommand: yah mesh promote . Refuses if target machine isn’t healthy (yubaba /health) or hostkey not registered.”) @yah:next(“Migration sequence: (1) stop Headscale on camp, (2) copy SQLite DB + config + private keys to target via yubaba, (3) start Headscale on target as a systemd unit yubaba manages, (4) update Cloudflare DNS A/AAAA for mesh. to point at target, (5) tear down camp Headscale + (if used) cloudflared.”) @yah:next(“Single-member raft is the degenerate case at this point — yubaba runs on the target, no quorum partners yet. Acceptable; HA shows up in Phase 2. Mark cluster as ‘single-machine’ in mesh status output so the operator knows the SPOF is real.”) @yah:next(“Atomicity / failure recovery: keep a snapshot of camp Headscale state until target is verified serving traffic. yah mesh promote –abort restores the camp coordinator if the new target proves unhealthy mid-migration. Dry-run mode prints the plan without executing.”) @yah:next(“Brief control-plane outage (~10s) is expected during DNS cutover. Data plane (existing WireGuard tunnels) is unaffected — verify by leaving a node-to-node ping running across the migration.”) @yah:next(“Out of scope: litestream replication, openraft, multi-member promotion. All of that lands in R040-F20/F21.”) @yah:verify(“yah mesh promote noisetable-pdx-1 migrates the coordinator; yah mesh status shows the new location; yah cloud machine provision against the migrated coordinator joins successfully”) @yah:verify(“Existing nodes’ tailnet connectivity uninterrupted across the migration (continuous ping from one node to another stays green)”) @arch:see(.yah/docs/architecture/A041-yah-mesh-bootstrap.md) @yah:handoff(“Phase 1b landed. New command: yah mesh promote [–dry-run] [–abort] [–host ] [–path

]. Migration sequence: (1) verify yubaba health + hostkey registered, (2) stop camp headscale, (3) POST SQLite DB + WireGuard keys + ACL policy to yubaba POST /headscale/deploy (yubaba downloads headscale binary from GitHub, writes files to configurable headscale_dir, starts systemd unit), (4) poll GET /headscale/health until running (90s timeout), (5) update Cloudflare DNS A record via CF API (CLOUDFLARE_API_TOKEN + CLOUDFLARE_ZONE_ID) or print manual instructions, (6) persist coordinator type to vault (mesh-coordinator-type=cluster, mesh-coordinator-machine=). –abort restarts camp headscale from existing local state. Rollback on mid-migration failure auto-restarts camp headscale. yah mesh status updated: shows ‘cluster [] (single-machine — SPOF)’ after promote, camp PID status before. New yubaba endpoints: POST /headscale/deploy (accepts base64-encoded state files, writes to configurable headscale_dir, curl-downloads binary, systemctl enable –now), GET /headscale/health (systemctl is-active + localhost:8080 HTTP probe). ServerSummary gained public_ipv4: Option (parsed from Hetzner public_net.ipv4.ip). Cloudflare DNS helper: update_cloudflare_dns() + cloudflare_credentials() in cloud::mesh. cloud-client: deploy_headscale() + headscale_health() methods. Tests: +8 yubaba (15 total), +2 cloud-client (9 total), +4 yah mesh:: (120 total). All pass.”) @yah:next(“Operator runbook: (1) yah mesh start –url https://mesh. on camp, (2) provision a machine, (3) yah cloud machine attach , (4) yah mesh promote [–host :7443 pre-mesh]. DNS update needs CLOUDFLARE_API_TOKEN + CLOUDFLARE_ZONE_ID or manual update.”) @yah:next(“Mesh URL connectivity: before Tailscale mesh is up, –host must be the machine’s public IP (pre-mesh access). Once mesh is established, the default http://:7443 works. Consider yah cloud agent ip helper or mesh-discovery to auto-resolve (same gap as R040-F8 Tailscale discovery).”) @yah:next(“HEADSCALE_VERSION is pinned to 0.23.0. If the operator’s local headscale.db was created by a different version, the remote headscale may reject the database. Add a version-mismatch warning to the deploy step.”)

@yah:relay(R085, “Infra-info MCP — yubaba / machine / mirror status read surface”) @yah:assignee(agent:claude) @yah:status(review) @yah:parent(Q082) @yah:handoff(“Four read-only cloud.* MCP tools landed in crates/yah/agent-tools/src/cloud_tools.rs. Registration: KgToolRegistry::with_cloud() in tools.rs (same builder pattern as with_board/with_scryer/with_task). cloud-client dep added to agent-tools Cargo.toml. Local-config tools (cloud.machines, cloud.mirror_state) read .yah/cloud/{machines,mirrors,services}/.toml from ctx.camp_root with lightweight serde structs — no cloud crate dep, no network. Network tools (cloud.warden_status, cloud.service_ports) probe yubaba HTTP API via cloud_client::CloudClient and return reachable:false rather than error when unreachable. 14 unit tests, all pass. Total agent-tools tests: 140.”) @yah:next(“Wire with_cloud() into the agent session startup (camp.rs or agent.rs in app/yah/cli) alongside with_board() and with_scryer() so the tools appear in the agent’s tool list automatically.”) @yah:next(“Feed cloud.warden_status into the drift detection surface (status.rs AgentProbe) so yubaba reachability populates machine status reports rather than being a separate probe path.”) @yah:verify(“cargo test -p agent-tools – cloud 2>&1 | grep ‘test result’ shows 14 passed; 0 failed”) @yah:verify(“KgToolRegistry::standard_read_only(ctx).with_cloud().schemas() returns 4 tools all named cloud.”) @arch:see(.yah/quests.md)

@yah:relay(R168, “yah-cloud TOML config schema + validation”) @yah:assignee(agent:claude) @yah:at(2026-05-13T19:07:07Z) @yah:status(review) @yah:parent(Q156) @yah:handoff(“Schema doc written at .yah/docs/architecture/A031-yah-cloud-config-shape.md. MirrorConfig.camp renamed to .camp with serde alias ‘camp’ for migration compat (R137 rename). load_mirrors() added to handle both flat mirrors/.toml and folder mirrors//mirror.toml layouts — yah-com/mirror.toml now loads. cloud_tools.rs local stub updated to match. All 90 cloud-crate tests pass + 4 new layout tests. agent-tools + yah bin check clean.”) @yah:next(“Remove serde alias ‘camp’ from MirrorConfig once all known mirrors are migrated (tracked in R137)”) @yah:verify(“cargo test -p cloud — 90 tests pass including mirror_folder_layout_loads + mirror_malformed_fails_with_field_path”) @yah:verify(“Schema doc links back to yah-public-site.md deployment-shape section”) @arch:see(.yah/docs/architecture/A031-yah-cloud-config-shape.md) @arch:see(.yah/docs/architecture/A045-yah-public-site.md)

@yah:relay(R260, “OpenRouter model almanac — track top-weekly models for backend selection”) @yah:assignee(agent:claude) @yah:at(2026-05-20T22:23:52Z) @yah:kind(spike) @yah:status(in-progress) @arch:see(architecture/yah-cloud-config-shape.md) @yah:next(“Unique signal: popularity + freshness + capability tags. Pricing is already covered by crates/yah/probe/data/pricing.json (LiteLLM snapshot) — almanac does NOT duplicate $/token; it adds rank, :free flag, modality, capability tags (code / vision / tool-use / long-context), and a last_seen_at so consumers can tell when an entry has gone stale.”) @yah:next(“Primary consumer: the OpenRouter character bundle at crates/yah/kg/preroll/openrouter/{agents,subclasses}.json. Currently a single preroll agent ("Quick Hand") pinned to qwen/qwen3-coder:free — this WILL silently break when OpenRouter rotates the free-tier list. Almanac’s first job is to keep that subclass pointed at a currently-available free model (and let the bundle expand into a roster of currently-hot free models, not just one).”) @yah:next(“Free-model rotation is the load-bearing constraint. Refresh cadence has to match how often OpenRouter shuffles the :free set (empirically weekly-ish). Storage options ranked by how well they handle rotation: (1) scheduled daemon refresh writing back into crates/yah/kg/preroll/openrouter/ — auto-heals; (2) on-demand fetch+cache in the runner — never stale at session start but offline-fragile; (3) build-time embedded snapshot like probe/data/pricing.json — simplest, goes stale between releases (probably wrong shape for free-tier).”) @yah:next(“Two approaches to evaluate side-by-side: (A) straight call to https://openrouter.ai/api/v1/models (free, no key, structured JSON — already returns context_length, pricing.prompt/pricing.completion, architecture.modality; crates/yah/runner/src/resolver/openrouter.rs already talks to OpenRouter so a fetch_models() sibling to its existing GET /api/v1/credits is the natural home); (B) inference-based tool that distills https://openrouter.ai/models?order=top-weekly HTML into our schema. Confirm (A) first — if the JSON exposes the top-weekly ordering and :free flag, (B) is unnecessary; if not, (B) becomes the rotation-detector.”) @yah:next(“Query-resolvable subclass (the real primitive): AgentSubclass.model at crates/yah/kg/src/party.rs:520 is a hard provider:model string today. Promote it to a sum type — literal provider:model OR ModelQuery { tier, capability_tags, min_context_tokens, max_price_per_token, … } — resolved by the almanac at session start. Slots into the existing fallbacks: Vec<FallbackRule> machinery at party.rs:540: emit ConfigSwitch{ModelResolved} (sibling of FallbackTriggered) when the query lands, so the UI can show "Quick Hand → qwen3-coder:free (current pick, refreshed 2h ago)". The preroll bundle then declares intent ("Thief-shaped, free, code-tuned") instead of a brittle pin.”) @yah:next(“Constraint vocabulary — the hard numbers behind "Thief-shaped, free, code-tuned", layered by cost-to-implement so we ship something useful at T0 and can grow. T0 (already free from /models): context_length, pricing.prompt/pricing.completion, architecture.modality come back in the same fetch. MVP heuristic = Pareto frontier of (context_window, price_per_token); at :free tier price collapses to 0 so it cleanly reduces to "max context with matching capability tag". This is enough to pick a non-broken Quick Hand. T1 (external quality scores): join on canonical model id against artificialanalysis.ai’s quality index (general + code subscore), Aider’s code leaderboard, or LiveBench. Cheap to add once the T0 pipe exists; lets queries express min_quality or prefer_code_subscore. T2 (almanac-as-evaluator): spend a few credits running a small fixed corpus (lint-fix, summarize, ack-and-route) through each new free candidate, score against a reference, persist scores into the almanac. The dogfood option — most signal, most expensive, most interesting. Out of scope for the spike; file as a follow-up child once T0/T1 are in flight.”) @yah:next(“Recommendation surface: popularity-ranked picker entries in packages/yah/ui/src/components/agent/Picker/toPickerAgents.ts + AgentProvidersPanel.tsx so adding an OpenRouter backend surfaces the currently-hot models first; query-resolved subclasses display their resolved pick + a freshness chip; stale literal-pinned :free references show a warning.”) @yah:next(“Spike output: a design note at .yah/docs/working/W091-openrouter-almanac.md + a working (A) prototype printing the top 10 models in our T0 schema (id, context_length, price, :free flag, capability tags, popularity rank). From there, refine into child tickets for: (i) refresh mechanism, (ii) ModelSpec::Literal | ModelSpec::Query sum type + T0 resolver, (iii) preroll-bundle migration from pinned-model to ModelQuery, (iv) UI recommendation + freshness chips, (v) T1 external-benchmark join, (vi) T2 self-eval as a stretch/follow-up.”) @yah:handoff(“Spike output delivered: working prototype + design note + 6 child tickets. Prototype lives in crates/yah/runner/src/resolver/openrouter.rs (AlmanacEntry struct + fetch_models() against /api/frontend/models/find?order=top-weekly + derive_capability_tags) plus app/yah/cli/src/agent.rs (yah agent almanac –top N –free-only –tag code –json). Design note at .yah/docs/working/W091-openrouter-almanac.md. Six follow-ups filed: R260-F1 (refresh scheduler), R260-F2 (ModelSpec sum type + T0 resolver), R260-T3 (preroll bundle migration → ModelSpec::Query), R260-F4 (UI freshness chips + stale-pin warning), R260-F5 (T1 benchmark join), R260-S6 (T2 self-eval, stretch). T0 prototype confirms the rotation problem: yah agent almanac --top 10 --free-only --tag code shows qwen/qwen3-coder:free (Quick Hand’s current pin) is NO LONGER in the top-10 free code models — current leaders are openrouter/owl-alpha, nvidia/nemotron-3-super, poolside/laguna-m.1.”) @yah:verify(“cargo run -p yah – agent almanac –top 10 — prints OpenRouter top-10 weekly models in T0 schema”) @yah:verify(“cargo run -p yah – agent almanac –top 10 –free-only –tag code — shows the currently-hot free-tier code-tuned roster (the Quick Hand replacement set)”) @yah:verify(“cargo check -p runner -p yah — clean (note: cargo test -p runner –lib is blocked by pre-existing R258-F3 test breakage on AgentSession; tracked as gotcha)”) @yah:gotcha(“Pre-existing R258-F3 test breakage in crates/yah/runner/src/{anthropic.rs:833, sessions.rs} — cargo test -p runner --lib fails with E0063 (AgentSession initialiser missing last_text + slot_id fields). R258-F3 is currently in review. R260’s new resolver::openrouter::tests entries are byte-perfect but un-runnable until R258-F3’s test fixtures are updated.”) @yah:gotcha(“OpenRouter’s /api/frontend/models/find?order=top-weekly is undocumented. The almanac treats it as best-effort and the design (see openrouter-almanac.md) calls for graceful degradation to public /api/v1/models alone when the join fails. ?order= is silently ignored on the documented endpoint (it always returns newest-first); approach (B) HTML scraping is unnecessary and was ruled out.”)

@yah:relay(R414, “yah-cloud pill sync + publish + Pills panel”) @yah:assignee(bundle-anthropic-miravel) @yah:at(2026-06-03T07:07:03Z) @yah:status(review) @yah:phase(P2) @yah:parent(Q410) @arch:see(.yah/docs/working/W141-pill-rings.md) @arch:see(.yah/docs/architecture/A031-yah-cloud-config-shape.md) @yah:depends_on(R411) @yah:handoff(“P2 delivered: resolve_pill_catalog (chips crate) + party_resolve_pill_catalog Tauri command + env wiring + PillsPanel in FullCharacterEditor. Pills panel shows all categorized pills grouped by category with per-character tri-state toggle (always-on/discoverable/off); edits agent.persona.chips.pill_state map via patchAgent.”) @yah:next(“P3: local pills.toml loader — add ~/.yah/pills.toml and .yah/pills.toml as new resolution layers (separate from chips.toml) per W141 layer stack”) @yah:next(“P3: MCP URL trust display — when a content-pill with mcp[] is toggled always-on, show one-time approval dialog listing the URLs before activating”) @yah:next(“P4: cloud sync stub — ‘Share to my account’ button in PillsPanel (disabled until account system exists)”) @yah:verify(“cargo test -p chips — 84 passed (3 new resolve_pill_catalog tests)”) @yah:verify(“bun run typecheck — no new errors (6 pre-existing unchanged)”)

@arch:see(.yah/docs/working/W193-asset-dependency-status-surface.md)

@yah:ticket(R743-T7, “yah-cloud: 5 test binaries to 1 (or 2 — pond_smoke may stay isolated)”) @yah:at(2026-08-11T01:18:56Z) @yah:status(review) @yah:phase(P2) @yah:parent(R743) @yah:next(“tests/main.rs mod’ing the siblings + autotests = false and [test] name = "main" in oss/yubaba/crates/cloud/Cargo.toml.”) @yah:verify(“cargo test -p yah-cloud – –list count unchanged; three green runs. One commit — oss subtree.”) @yah:gotcha(“tests/pond_smoke.rs:198 is the only real fixed-port bind in the whole in-scope set — MinIO on http://127.0.0.1:9000, an external dependency it does not own. It is a legitimate deliberate-isolation candidate: if merging it makes the suite flaky or order-dependent, keep it as its own [test] target and say so on the ticket rather than force-merging.”) @yah:gotcha(“pond_smoke.rs and mesofact_static_e2e.rs both host live @yah: annotations.”) @yah:handoff(“LANDED: 5 test binaries → 2. New tests/main.rs mods live_workspace_smoke, mesofact_static_e2e, pg_driver_live and whisper_derive_e2e; Cargo.toml gains autotests = false plus explicit [test] main (tests/main.rs) and [test] pond_smoke (tests/pond_smoke.rs). Files were mod’d, not concatenated, so mesofact_static_e2e.rs’s live R441-B2 annotation block stays where the harvester expects it and CARGO_MANIFEST_DIR (which three of the four walk up from) is unchanged. autotests = false only gates [test] discovery — examples/load_probe.rs is still autodiscovered. Test names gained a module prefix (e.g. whisper_derive_e2e::derive_pipeline_upload_skip_prune_reproducibility); substring filters still match, but --test <file-stem> for the four merged files is now --test main -- <module>::.”) @yah:handoff(“ISOLATION DECISION: pond_smoke STAYS its own [test] target — 2 targets, not 1, and deliberately so. Three reasons, strongest first. (1) DECISIVE: .yah/qed/pond-smoke.toml’s pond-spinup-budget step shells out to cargo test --release --locked -p cloud --test pond_smoke -- --nocapture. Folding pond_smoke into main turns that committed pipeline step into a hard cargo error, and that file is outside this crate (peer-owned path, not mine to edit). (2) It is the only test in the crate that reaches a FIXED external address — MinIO at http://127.0.0.1:9000 (pond_smoke.rs:199) — and the only one that creates/docker rm -fs named containers in a Drop guard, so its blast radius on failure is outside the process and a standalone runnable handle is worth keeping. xtask/src/lib.rs’s @yah:assumes census independently found this to be the ONLY real port bind across all ten in-scope crates. (3) It is benchmark-shaped and libtest parallelises within one binary: both its tests assert wall-clock budgets (WARM_RESTART_BUDGET 3s, WARDEN_COLD_BUDGET 5s, COLD_START_BUDGET 15s). Merging would drop whisper_derive_e2e (two in-process axum servers + BLAKE3 + tempdir IO, ungated) and live_workspace_smoke (a full CloudConfig::load walk of the real workspace, ungated) onto sibling threads inside that measurement window, and the pipeline runs it –nocapture for the timing report, which merging would interleave with every other test’s output. The four that DID merge have no such conflict: read-only fs, 127.0.0.1:0 ephemeral ports, tempdirs, and no process-global state (no set_var / set_current_dir / top-level statics anywhere in the set).”) @yah:verify(“BASELINE (HEAD layout, cargo test -p yah-cloud – –list, run with the change temporarily reverted): 6 test binaries — lib 924 tests; live_workspace_smoke 1; mesofact_static_e2e 1; pg_driver_live 1; pond_smoke 2; whisper_derive_e2e 1; Doc-tests cloud 1. Integration total 6 across 5 binaries.”) @yah:verify(“AFTER (same command, exit 0): lib 924 tests; tests/main.rs 4 tests; tests/pond_smoke.rs 2 tests; Doc-tests cloud 1. Integration total 6 across 2 binaries — COUNT UNCHANGED, every one of the 6 names accounted for, now module-prefixed inside main.”) @yah:verify(“THREE GREEN RUNS: cargo test -p yah-cloud –test main –test pond_smoke, three consecutive times — main 3 passed / 0 failed / 1 ignored, pond_smoke 2 passed / 0 failed, all three runs identical. No order-dependence observed.”) @yah:verify(“HONEST SCOPE OF WHAT ACTUALLY EXECUTED (confirmed with – –nocapture, container has no docker / no live pg / no mesofact-dev binary): REALLY RAN — live_workspace_smoke::live_yah_workspace_loads_cleanly (real CloudConfig::load of the checked-in .yah/ tree, real assertions) and whisper_derive_e2e::derive_pipeline_upload_skip_prune_reproducibility (in-process axum fake-S3 + fake upstream on 127.0.0.1:0, all three runs). SELF-SKIPPED at runtime — mesofact_static_e2e (prints ‘skipping: set YAH_RECONCILER_E2E_BIN’), pond_smoke::pond_spinup_budget and pond_smoke::warden_container_spinup_budget (both print ‘SKIP: set YAH_LOCAL_SIM_E2E=1’, need orbstack/colima/docker). IGNORED — pg_driver_live (#[ignore], needs a built yah-pg-dev binary + network). So the live/e2e legs were compiled and linked but NOT exercised here; the pipeline steps that do exercise them are unchanged for pond_smoke and reachable as --test main -- mesofact_static_e2e:: for the merged one.”) @yah:gotcha(“PRE-EXISTING, NOT CAUSED BY THIS TICKET: .yah/qed/local-sim-smoke.toml’s only step runs cargo test -p cloud --test local_sim_smoke, but tests/local_sim_smoke.rs does not exist and git log --all finds it at neither oss/yubaba/crates/cloud/tests/ nor the pre-OSS crates/yah/cloud/tests/ path. That pipeline was already dangling before this change (it most likely wants pond_smoke). Left alone: .yah/qed/ is outside this crate. Flagging it because autotests = false makes a target-name typo fail identically to this, and the next person to hit it should not blame R743-T7.”) @yah:gotcha(“Historical @yah:verify strings on already-landed tickets still cite the retired per-file target names — --test whisper_derive_e2e (reconciler/static_asset.rs:46,118), --test mesofact_static_e2e (the R441-B2 block inside tests/mesofact_static_e2e.rs itself). Those are landed records, not live invocations, so they were deliberately NOT rewritten; the working form is now cargo test -p yah-cloud --test main -- <module>::.”) @yah:gotcha(“CONTAINER, not code: the first attempt at the post-change –list died with error: linking with cc failed … ld terminated with signal 7 [Bus error] on both the merged main target and the untouched lib test target. Root cause was the container’s root filesystem at 100% (8.0K free) — oss/yubaba/target alone was 23G with no cargo-orphan-gc installed here to reclaim it. Cleared by deleting the regenerable incremental caches (oss/yubaba/target/debug/incremental, target/debug/incremental) while no cargo was running; the listing then exited 0. Worth knowing because the failure mode reads exactly like the R770 orphan-gc symptom described in CLAUDE.md but is plain ENOSPC.”) @yah:tier(Warrior)

Re-exports§

pub use almanac_dispatch::dispatch_on_change;
pub use asset_journal::AssetState;
pub use asset_journal::AssetStatusEvent;
pub use asset_journal::AssetStatusJournal;
pub use capability::Capability;
pub use compose::generate_compose_bundle;
pub use compose::ComposeBundle;
pub use config::BucketLogEntry;
pub use config::CampCloudDbs;
pub use config::CloudConfig;
pub use config::CloudDb;
pub use config::ConnectSpec;
pub use config::DbCatalog;
pub use config::DevDb;
pub use config::GitSource;
pub use config::IngressDecl;
pub use config::IngressEdge;
pub use config::IngressProvider;
pub use config::LegacyMirrorConfig;
pub use config::LegacyServiceConfig;
pub use config::MachineConfig;
pub use config::MirrorAssignment;
pub use config::MirrorConfig;
pub use config::MirrorProviderSlot;
pub use config::MirrorShape;
pub use config::PondDb;
pub use config::PondDbKind;
pub use config::Provider;
pub use config::ProviderConfig;
pub use config::ServiceComponent;
pub use config::ServiceConfig;
pub use config::TopologyConfig;
pub use config::WorkloadConfig;
pub use config::WorkloadConfigError;
pub use local_driver_glue::local_container_spec_from_provider;
pub use provider::on_ingress_owner_changed;
pub use provider::reconcile_assignment;
pub use provider::BucketAcl;
pub use provider::BucketRef;
pub use provider::CfAccountInfo;
pub use provider::CloudflareClient;
pub use provider::CloudflareEnvoy;
pub use provider::CreateR2BucketResult;
pub use provider::CreateTokenResult;
pub use provider::CreateTunnelResult;
pub use provider::DigitalOceanEnvoy;
pub use provider::FloatingIpAssignOutcome;
pub use provider::FloatingIpProvider;
pub use provider::FloatingIpState;
pub use provider::FloatingIpTarget;
pub use provider::GrantScope;
pub use provider::HetznerDriver;
pub use provider::HetznerEnvoy;
pub use provider::HetznerFloatingIp;
pub use provider::Location;
pub use provider::MachineProvider;
pub use provider::OvhFloatingIp;
pub use provider::ProjectId;
pub use provider::R2BucketInfo;
pub use provider::R2CustomDomain;
pub use provider::ServerId;
pub use provider::ServerSpec;
pub use provider::ServerStatus;
pub use provider::ServerSummary;
pub use provider::TokenGrant;
pub use provider::TunnelConnState;
pub use provider::TunnelDnsRecord;
pub use provider::TunnelDriftRow;
pub use provider::TunnelDriftState;
pub use provider::VultrFloatingIp;
pub use provider::WorkerDeployResult;
pub use provider::MESOFACT_STATIC_GRANTS;
pub use reconciler::collect_live_derive_hashes;
pub use reconciler::compute_derive_cache_candidates;
pub use reconciler::compute_live_set;
pub use reconciler::compute_prune_candidates;
pub use reconciler::compute_service;
pub use reconciler::derive_minio_key;
pub use reconciler::execute_derive_cache_prune;
pub use reconciler::execute_prune;
pub use reconciler::load_service_and_mirror;
pub use reconciler::mesofact_static::WORKER_SCRIPT;
pub use reconciler::new_sync_id;
pub use reconciler::pond::MINIFLARE_SIM_SCRIPT;
pub use reconciler::publish_to_pond;
pub use reconciler::summarize;
pub use reconciler::CellStatus;
pub use reconciler::CloudflareWorkerReconciler;
pub use reconciler::ContainerOptions;
pub use reconciler::ContainerReconciler;
pub use reconciler::DeriveCacheLiveHashes;
pub use reconciler::DerivePruneCandidate;
pub use reconciler::DriftEntry;
pub use reconciler::HeadscaleReconciler;
pub use reconciler::HealthState;
pub use reconciler::LocalProcessReconciler;
pub use reconciler::LocalStaticOptions;
pub use reconciler::MesofactStaticReconciler;
pub use reconciler::MirrorObservation;
pub use reconciler::PondOptions;
pub use reconciler::PondPublishReport;
pub use reconciler::PondState;
pub use reconciler::ProviderScope;
pub use reconciler::PruneCandidate;
pub use reconciler::PruneOutcome;
pub use reconciler::PruneReport;
pub use reconciler::ReconcileCtx;
pub use reconciler::Reconciler;
pub use reconciler::RunningWorkload;
pub use reconciler::RunningWorkloadSummary;
pub use reconciler::Runtime;
pub use reconciler::ServiceStatus;
pub use reconciler::StaticAssetReconciler;
pub use reconciler::StatusSummary;
pub use reconciler::SyncHistoryEntry;
pub use reconciler::SyncOutcome;
pub use reconciler::SyncState;
pub use reconciler::WireContainerStatus;
pub use status::collect_machine_report;
pub use status::AgentProbe;
pub use status::DriftFinding;
pub use status::MachineReport;

Modules§

almanac_dispatch
Dispatch [almanac::OnChangeConfig] actions after a feed run (R330-F4).
app_manifest
yah-app.toml schema, parser, closed-set required_when evaluator, .yah/apps.toml registry, and workspace discovery (R470-T7).
asset_journal
Append-only JSONL journal for static-asset reconciler decisions (R470-T1).
asset_status
Build the W193 workspace asset-status report as a serde_json::Value suitable for cloud.status RPC responses (R470-T4).
capability
Service capabilities and the mirror-side driver bindings that satisfy them (W265).
cloud_init
Cloud-init template renderer for mirror-machine bootstrap (R040-F4).
compose
Podman Compose file generation for yah-cloud service deployments (R040-F7).
config
@yah:ticket(R040-F16, “pg-on-mesh service recipe: bind tailscale0 + pg_hba.conf snippet + ufw rules”) @yah:at(2026-05-05T00:32:34Z) @yah:assignee(agent:claude) @yah:status(review) @yah:parent(R040) @yah:handoff(“Companion to R040-F15. Inter-node TCP (Postgres primary↔replica, NATS clusters, anything raw-protocol) lives on the Headscale mesh, not on Hetzner public IPs. Each node has a stable 100.64.x.x mesh IP that survives replacement of the underlying box, so DNS / config / pg_hba never churn when a CPX-11 is rebuilt. WireGuard already encrypts the wire — TLS becomes defense-in-depth, not load-bearing. This ticket carries the concrete pg-shaped recipe so the first stateful service deploy doesn’t have to re-derive the pattern; subsequent services (redis, NATS, etc.) cargo-cult from it.”) @yah:next(“ServiceConfig gains a bind_interface: Option<String> field (e.g. Some(\"tailscale0\") for mesh-only services). The cloud-init/podman compose renderer translates this into either --network host + pg listen_addresses = '<mesh-ip>' OR a podman macvlan/host-binding pattern that achieves the same.”) @yah:next(“Generated pg_hba.conf snippet: allow the mesh subnet (100.64.0.0/10) for replication + app users. Postgres binds to the node’s tailscale0 mesh IP only — listen_addresses is templated from the node’s tailscale ip --4 at first boot.”) @yah:next(“Generated ufw rules: ufw allow in on tailscale0 to any port 5432; ufw deny 5432 — mirrors the existing yah-yubaba 7443 pattern in mirror.yml. Same shape works for any mesh-only port.”) @yah:next(“Replica connection string uses primary’s mesh IP, NOT its public IP. Stable across box replacement.”) @yah:next(“Out of scope: pg_basebackup orchestration, failover, WAL archiving — those belong in noisetable’s domain; this ticket only standardizes the binding/firewall/auth shape so noisetable’s pg deployment doesn’t reinvent it.”)
envoy
Envoy framework types — Tier, AdapterFlavor, VerbCategory, InternalVerb, VerbDescriptor.
identities
Stub local identity registry: .yah/cloud/identities/<machine>.json.
local_driver_glue
Adapter between cloud’s TOML-driven ProviderConfig shape and local_driver::LocalContainerSpec.
mesh
Headscale API client and configuration helpers for yah mesh operations.
mesh_service
Recipe helpers for stateful services bound exclusively to the Headscale mesh (R040-F16).
migrate
W305/R742-F3 — planning a workload’s move between sovereign groups.
multi_root
Multi-root cloud config — union sibling .X/ config trees (W206 layout (b)).
paths
On-disk path resolver for camp deployment manifests.
proc_control
The process-control channel — how a yah-supervised process describes itself to its supervisor and to an agent, instead of being guessed at from the outside.
provider
@yah:relay(R409, “Envoy — providers and internal verb catalog (W144)”) @yah:at(2026-06-02T20:58:35Z) @yah:status(open) @arch:see(.yah/docs/working/W144-envoy-providers-and-tiering.md)
provision
Provisioning orchestrator: config → cloud-init render → MachineProvider call.
reconciler
Reconciler abstraction — bring a workload up against a mirror’s provider slots.
release_manifest
Yubaba release-manifest fetch + per-triple resolution (R330-F21).
state
Reconciler-written sidecar state for declared machines.
status
Drift detection between declared .yah/cloud/ config and live cloud state.
topology
R850 — what the declared topology does when a node dies, and whether it fits on the boxes it is declared against.
validate
Workspace-wide lint checks for the .yah/ declaration tree (R470-T3).

Structs§

ContainerRunSpec
Run-time spec for a single container managed by LocalRuntime::run.
CustomDockerHostProvider
RuntimeProvider for a raw DOCKER_HOST string supplied by the operator (e.g. "tcp://localhost:2375" or "unix:///path/to/custom.sock"). Always reported as available — the operator opted in explicitly; failures surface as docker CLI errors rather than probe misses.
LocalContainerSpec
Probe spec lifted from a kind = "local-container" provider TOML.
LocalDockerRuntime
WorkloadRuntime implementation backed by the docker CLI.
LocalRuntime
Reachable local container daemon — the winning provider from the cascade.
OwnedContainer
A container managed by this module, as observed via LocalRuntime::list_owned.
SocketRuntimeProvider
RuntimeProvider backed by a Unix socket path. Available iff the tilde-expanded path exists on disk.

Enums§

ContainerState
Container State.Status from docker inspect. Anything the docs don’t enumerate lands in [Unknown].
DetectedRuntime
Which runtime answered the probe cascade.
RuntimePref
Operator-declared runtime preference from the provider TOML’s top-level runtime field. auto walks the full cascade; any pinned value probes only that runtime.

Constants§

LABEL_KEY
Docker label key applied to every container this module owns. The value is <service>:<env>:<slot> so docker ps --filter label=yah.pond=<v> scopes orphan cleanup to a specific mirror.
NAME_PREFIX
Canonical name prefix for every container managed by this module.

Traits§

RuntimeProvider
Probe contract for a single Docker-compatible runtime back-end. Each implementation knows its socket path (or raw DOCKER_HOST) and can report whether it is currently reachable.

Functions§

canonical_label
Build the canonical label value for the same triple.
canonical_name
Build the canonical container name for a (service, env, slot) triple.
warden_container_label
Canonical label of the camp’s pond yubaba container: camp:<env>:yubaba.
warden_container_name
Canonical name of the camp’s pond yubaba container for env: yah-pond-camp-<env>-yubaba. The camp (which starts it) and the desktop probes (which inspect it) MUST agree on this — always go through here.