Skip to main content

Module config

Module config 

Source
Expand description

@yah:ticket(R040-F16, “pg-on-mesh service recipe: bind tailscale0 + pg_hba.conf snippet + ufw rules”) @yah:at(2026-05-05T00:32:34Z) @yah:assignee(agent:claude) @yah:status(review) @yah:parent(R040) @yah:handoff(“Companion to R040-F15. Inter-node TCP (Postgres primary↔replica, NATS clusters, anything raw-protocol) lives on the Headscale mesh, not on Hetzner public IPs. Each node has a stable 100.64.x.x mesh IP that survives replacement of the underlying box, so DNS / config / pg_hba never churn when a CPX-11 is rebuilt. WireGuard already encrypts the wire — TLS becomes defense-in-depth, not load-bearing. This ticket carries the concrete pg-shaped recipe so the first stateful service deploy doesn’t have to re-derive the pattern; subsequent services (redis, NATS, etc.) cargo-cult from it.”) @yah:next(“ServiceConfig gains a bind_interface: Option<String> field (e.g. Some(\"tailscale0\") for mesh-only services). The cloud-init/podman compose renderer translates this into either --network host + pg listen_addresses = '<mesh-ip>' OR a podman macvlan/host-binding pattern that achieves the same.”) @yah:next(“Generated pg_hba.conf snippet: allow the mesh subnet (100.64.0.0/10) for replication + app users. Postgres binds to the node’s tailscale0 mesh IP only — listen_addresses is templated from the node’s tailscale ip --4 at first boot.”) @yah:next(“Generated ufw rules: ufw allow in on tailscale0 to any port 5432; ufw deny 5432 — mirrors the existing yah-yubaba 7443 pattern in mirror.yml. Same shape works for any mesh-only port.”) @yah:next(“Replica connection string uses primary’s mesh IP, NOT its public IP. Stable across box replacement.”) @yah:next(“Out of scope: pg_basebackup orchestration, failover, WAL archiving — those belong in noisetable’s domain; this ticket only standardizes the binding/firewall/auth shape so noisetable’s pg deployment doesn’t reinvent it.”)

@yah:ticket(R323-F9, “Add sync-wave ordering to ServiceComponent (deploy-panel wave order)”) @yah:assignee(agent:claude) @yah:at(2026-05-26T15:20:25Z) @yah:status(review) @yah:phase(P2) @yah:parent(R323) @yah:next(“ServiceComponent gains a wave/order field (or depends_on between components) so the deploy panel (R323-F4) can group workload rollout rows into sync waves (wave 0 parallel, wait healthy, wave 1, …). Today all components are implicitly wave 0.”) @yah:next(“compute_service/compute_cell in reconciler/sync_status.rs surface the wave per workload so F4 doesn’t re-derive it.”) @yah:gotcha(“Until this lands, F4 should render every workload as wave 0 (no ordering).”) @yah:handoff(“Added wave: u32 (serde default=0, skip_serializing_if zero) to ServiceComponent in config.rs. Added is_zero_u32 helper. Fixed the three struct literal call-sites that now need wave: 0 (config.rs test, local_sim.rs x2, mesofact_static.rs). Added wave?: number to the TS ServiceComponent interface with a doc comment. Deploy panel now reads c.wave ?? 0 for each WorkloadRow instead of hardcoded 0. SyncFooter computes maxWave from the components array and renders ‘wave 0’ (all-zero case) or ‘waves 0–N’ (multi-wave). All 218 cloud lib tests pass; bun run typecheck clean.”) @yah:verify(“cargo test -p cloud –lib # 218 passed”) @yah:verify(“cd packages/yah/ui && bun run typecheck # no new errors”) @yah:verify(“In service.toml: add wave = 1 to a component, rebuild, open the deploy panel — that workload row shows ‘w1’ badge; SyncFooter shows ‘waves 0–1’”) @yah:verify(“Component with no wave field in TOML deserializes as wave=0 (default). Saving a wave=0 component omits the field from the output TOML (skip_serializing_if).”)

@arch:see(.yah/docs/working/W142-pond.md)

@yah:relay(R615, “Linked infra sources: sources.toml overlay so a camp can borrow another camp’s substrate”) @yah:at(2026-07-20T18:18:05Z) @yah:status(open) @arch:see(.yah/docs/working/W274-linked-infra-sources.md)

@yah:ticket(R615-F1, “InfraSource types + SourcesConfig::load(infra_dir) parsing .yah/infra/sources.toml”) @yah:status(review) @yah:assignee(agent:bundle-anthropic-miravel) @yah:at(2026-08-08T19:55:57Z) @yah:phase(P1) @yah:parent(R615) @yah:next(“Add InfraSourceKind { Path { path }, Git(GitSource) } + InfraSource { owner, kind, mode, select } to cloud/src/config.rs. Reuse the existing GitSource (config.rs:1205, { repo, ref, subdir }) verbatim — do not invent a second git-source shape.”) @yah:next(“SourcesConfig::load(infra_dir) reads .yah/infra/sources.toml (schema_version = 1, ordered [[source]] array). Absent file = empty list, never an error — every existing camp has no sources.toml.”) @yah:next(“mode is the write-gate: read-only (borrower cannot mutate) vs owner-manages. Model it as an enum, not a bool, so a future read-write-with-approval tier is additive.”) @yah:verify(“cargo check -p cloud && cargo test -p cloud”) @arch:see(.yah/docs/working/W274-linked-infra-sources.md) @yah:tier(Cleric) @yah:handoff(“InfraSourceKind{Path{path},Git(GitSource)} + SourceMode{ReadOnly,Manage} + InfraSource{owner,kind,mode,select} + SourcesConfig{schema_version,source} all landed in oss/yubaba/crates/cloud/src/config.rs (after default_git_ref, ~line 1550). GitSource reused verbatim – Git(GitSource) wraps the existing R561 type unchanged, no second git-source shape. InfraSourceKind is internally tagged (#[serde(tag="kind", rename_all="kebab-case")]) and flattened into InfraSource so a [[source]] table reads exactly like W274’s example: owner/kind/path-or-repo+ref+subdir/mode/select all at one table level. mode: SourceMode defaults ReadOnly via #[serde(default)] on the field (enum, not bool, per the ticket’s own instruction – Manage is the explicit escape hatch). SourcesConfig::load(infra_dir) returns Ok(default()) – schema_version=1, empty source list – when sources.toml is absent; only parses+errors when the file exists and is malformed.”) @yah:handoff(“Tree anchor 85801e7f. Pathspec: oss/yubaba/crates/cloud/src/config.rs (only file touched). Tests: cargo test -p yah-cloud –lib (from oss/yubaba) 710 passed / 0 failed / 4 ignored, +6 new over the 704 baseline your R707-T6 verification recorded (sources_load_is_empty_when_the_file_is_absent, sources_parses_a_path_kind_exactly_like_w274s_example, sources_parses_a_git_kind_reusing_gitsource_verbatim, sources_mode_defaults_to_read_only_and_manage_is_explicit, sources_preserves_declaration_order, sources_round_trips_through_serialize). cargo check -p cloud also green (implied by the test build).”) @yah:handoff(“Tree anchor at handoff: 85801e7f6b76b369c0c8ecd2e5c7874990cd9286 — the shared tree as I left it. Diff against it (git diff 85801e7f6b76b369c0c8ecd2e5c7874990cd9286..HEAD) to see what landed under you, and quote this SHA rather than ‘HEAD’ in any revert/restore instruction.”) @yah:next(“R615-F2 picks this straight up: overlay these sources into CloudConfig::load, tagging origin{owner,source} and merging camp-local-wins-on-collision.”) @yah:handoff(“Verified pre-existing work: InfraSourceKind{Path,Git(GitSource)} + SourceMode + InfraSource + SourcesConfig all present in oss/yubaba/crates/cloud/src/config.rs at tree anchor 871fde1c, matching the inline @yah:handoff notes already on this ticket. GitSource reused verbatim, no second git-source shape. This session added no new code – only ran verification and closed the board state, which a prior session left stuck in open despite the work being done (code + handoff notes landed, but board.review/handoff was never called).”) @yah:verify(“cargo check -p yah-cloud – clean (2 pre-existing unrelated warnings)”) @yah:verify(“cargo test -p yah-cloud –lib – 723 passed; 0 failed; 4 ignored (from oss/yubaba)”)

@yah:ticket(R615-F2, “Overlay loader: resolve sources in CloudConfig::load, tag origin, camp-local wins on collision”) @yah:status(review) @yah:assignee(agent:bundle-anthropic-miravel) @yah:at(2026-08-08T19:56:05Z) @yah:phase(P1) @yah:parent(R615) @yah:next(“In CloudConfig::load, after loading camp-local machines/providers/rules, resolve each source to an infra root (git sources read from the .yah/cache/infra/ sync cache — load stays offline), load that root’s machines/providers/rules, tag each entry with origin { owner, source }, and overlay UNDER camp-local. Camp-local wins on name collision.”) @yah:next(“The machine load site is config.rs:533 (load_dir::(paths::machines_dir(…))). Note config.rs:575 load_from_config_dir is a SECOND machine load site that deliberately skips the inherit_machines redirect for multi-root/sibling trees (W206) — decide explicitly whether sources overlay applies there too, and document the answer either way.”) @yah:verify(“cargo check -p cloud && cargo test -p cloud”) @yah:verify(“A camp with sources.toml [[source]] kind=path to a sibling camp sees that camp’s machines in CloudConfig::load, each tagged with the source owner”) @yah:gotcha(“Cross-camp MachineConfig schema skew is real: noisetable ships an older machine schema (location/server_type/hosts_mirrors) while yah’s use region/arch/[connect]. A borrowed source can carry fields the borrower’s binary predates. Overlay load MUST tolerate/skip unparseable foreign entries per-file and warn — never fail the whole load.”) @arch:see(.yah/docs/working/W274-linked-infra-sources.md) @yah:depends_on(R615-F1) @yah:tier(Warrior) @yah:handoff(“Overlay landed in CloudConfig::load (oss/yubaba/crates/cloud/src/config.rs). After camp-local machines/providers/legacy-merge finish, SourcesConfig::load(paths::infra_dir(workspace_root)) resolves + overlay_infra_sources() merges each source’s machines/providers UNDER what’s already there – camp-local wins any name collision, and among sources themselves the earlier-declared one wins (both proven by dedicated tests). Provenance is NOT a field on MachineConfig/ProviderConfig: added CloudConfig.machine_origins/provider_origins: BTreeMap<String, InfraOrigin> instead, keyed by name/id. Reason recorded in a doc comment on InfraOrigin – MachineConfig/ProviderConfig are constructed by struct literal in test helpers across several crates (including crates/yah/agent-tools/src/cloud_tools.rs, which is fenced/live-owned this session), so widening either shape would have forced an edit there for zero semantic gain; origin is a property of the LOAD, not the machine.”) @yah:handoff(“GOTCHA closed: added load_dir_tolerant() – a per-file-tolerant sibling of the existing (strict) load_dir – so one unparseable foreign machine/provider (schema skew) skips-with-a-tracing::warn! and never sinks the rest of that source’s directory or this camp’s own load. Proven by one_unparseable_foreign_machine_does_not_sink_the_rest_of_the_directory_or_the_load. load_dir itself is untouched – camp-local files still hard-fail on a bad TOML, which is correct, only borrowed roots get the tolerant path.”) @yah:handoff(“Git sources: InfraSource::infra_root() resolves kind=path to <workspace_root>//.yah/infra (live tree, no I/O beyond building the path) and kind=git to paths::infra_source_cache_dir(workspace_root, owner)/infra – a NEW path helper in paths.rs, also what R615-T3’s yah infra sync target directory must be so the two line up. An unsynced git source (cache dir absent) overlays nothing and is explicitly NOT an error (test: an_unsynced_git_source_overlays_nothing_and_is_not_an_error) – load() stays fully offline as W274 §3 requires.”) @yah:handoff(“select filtering implemented for machines only (name exact-match or literal mesh_tags membership – not a glob engine, matches W274’s own example verbatim) via machine_matches_select(); does NOT apply to providers – documented as a deliberate choice, nothing in W274 or the ticket describes a provider-scoped filter.”) @yah:handoff(“EXPLICIT DECISION on the config.rs:575-equivalent gotcha (now load_from_config_dir): sources overlay does NOT apply there. Multi-root sibling config dirs (W206 layout (b)) are a second config root INSIDE the same camp, not a second camp – .yah/infra/sources.toml is tied to paths::infra_dir(workspace_root) specifically, which has no well-defined meaning for an arbitrary config_dir. Documented in the function’s doc comment and proven by load_from_config_dir_never_applies_sources_overlay (a sources.toml at the real workspace root does NOT leak into a load_from_config_dir call against a sibling .noisetable/ dir under that same root).”) @yah:handoff(“Tree anchor 85801e7f. Pathspec: oss/yubaba/crates/cloud/src/config.rs, oss/yubaba/crates/cloud/src/paths.rs (added infra_source_cache_dir + 1 test), oss/yubaba/crates/cloud/src/reconciler/mesofact_bundle.rs (CloudConfig test-literal fixed for the 2 new fields), app/yah/cli/src/cloud.rs (3 CloudConfig test-literal sites fixed, same reason). Tests: cargo test -p yah-cloud –lib (from oss/yubaba) 720 passed / 0 failed / 4 ignored, +10 over R615-F1’s 710 baseline (9 overlay tests in config.rs + 1 in paths.rs). cargo build -p yah –lib (repo root) green – confirms nothing downstream (agent-tools, cloud.rs, hub) broke from CloudConfig’s two new fields.”) @yah:handoff(“Tree anchor at handoff: 85801e7f6b76b369c0c8ecd2e5c7874990cd9286 — the shared tree as I left it. Diff against it (git diff 85801e7f6b76b369c0c8ecd2e5c7874990cd9286..HEAD) to see what landed under you, and quote this SHA rather than ‘HEAD’ in any revert/restore instruction.”) @yah:next(“R615-T3 (yah infra sync) is unblocked and has everything it needs: paths::infra_source_cache_dir(workspace_root, owner) is the exact target directory to clone/pull git sources into, already matching what F2’s overlay reads from.”) @yah:next(“R615-F4 (Infra tab origin badge, not in my assigned lane) can read CloudConfig.machine_origins/provider_origins directly – no further backend plumbing needed for the badge itself.”) @yah:handoff(“Verified pre-existing work: overlay landed in CloudConfig::load (oss/yubaba/crates/cloud/src/config.rs) at tree anchor 871fde1c – SourcesConfig::load resolves sources, overlay_infra_sources() merges under camp-local with camp-local-wins and earlier-source-wins collision rules, machine_origins/provider_origins BTreeMaps added to CloudConfig, load_dir_tolerant() added for per-file-tolerant foreign schema skew, InfraSource::infra_root() resolves path/git kinds, load_from_config_dir explicitly does NOT get the overlay (documented). Matches this ticket’s own inline @yah:handoff notes. This session added no new code – only ran verification and closed board state that a prior session left stuck in open despite the work being done.”) @yah:verify(“cargo check -p yah-cloud – clean (2 pre-existing unrelated warnings)”) @yah:verify(“cargo test -p yah-cloud –lib – 723 passed; 0 failed; 4 ignored (from oss/yubaba), includes overlay tests + load_dir_tolerant test + infra_source_cache_dir test in paths.rs”)

@yah:ticket(R605-F12, “Sovereign groups have no voting axis, so non-voting membership is inexpressible and the raft guard is enforced by an absent field”) @yah:status(review) @yah:at(2026-08-20T05:15:30Z) @yah:assignee(agent:bundle-anthropic-ashguard) @yah:parent(R605) @arch:see(.yah/docs/working/W325-isolated-x86-build-capacity.md) @yah:next(“OPERATOR INTENT (2026-08-19) that the model cannot currently record: us-west-003 is a NON-VOTING member of the us-west-001-based (prod) sovereign group, and us-west-011 is a DIFFERENT sovereign (dev) from 001/003. The dev/prod split is already declared correctly. The non-voting membership is not — us-west-003.toml declares no sovereign_group at all.”) @yah:next(“THE GAP: MachineConfig::sovereign_group is a single Option, so membership is binary, and judge_join (oss/yubaba/crates/cloud/src/config.rs:459) permits a join IFF both sides declare the same non-None group. There is no way to say ‘in this blast radius, but not quorum-eligible’.”) @yah:next(“WHY THAT IS ACTIVELY BAD, not just missing: today the ONLY thing refusing us-west-003 into the prod raft at the join gate is its ABSENT stamp. Its own file is emphatic it must never hold a raft node id (‘a home-internet partition should never be able to stall the raft’), and that guarantee currently rests on a field nobody wrote. Stamping it prod to record the operator’s real intent would REMOVE the guard. This is precisely the W305 failure mode that produced R742-T4: no-voter sat inert on three nodes asserting something nothing enforced.”) @yah:next(“PROPOSED SHAPE (recommended): a second axis, e.g. sovereign_role = voter | non-voter (default voter for back-compat, or make it required), with judge_join permitting a same-group join only for voters. Then us-west-003 stamps prod + non-voter, the intent is machine-readable, and the raft guard stops depending on omission. us-west-004 (R605-T7) would take the same shape.”) @yah:next(“TOUCHES TWO COPIES OF THE PREDICATE, do not fix only one: cloud::judge_join renders the camp-side refusal, but the predicate itself lives in workload_spec::sovereign::join_permitted because yubaba’s POST /raft/add-learner gate asks the same question and there is deliberately no yubaba -> cloud edge. Also re-read yubaba serve --sovereign-group, whose node-side gate is narrower on purpose (an unset flag means ‘declared nothing’, not ‘declared standalone’).”) @yah:gotcha(“THE CODE AND THE OPERATOR CURRENTLY DISAGREE ABOUT 003, and a reader should know which is which before editing. judge_join’s own doc comment asserts ‘prod and dev are both stamped, and us-west-002/003/015 are deliberately not raft members’ — i.e. R742-F1 modelled 003 as STANDALONE. The operator’s model is that it is a NON-VOTING MEMBER of prod. Those are different claims, not a wording difference: standalone means no blast-radius relationship to 001 at all. Do not silently ‘correct’ either side; this ticket is the reconciliation.”) @yah:gotcha(“FLEET STATE AS DECLARED (2026-08-19): prod = us-west-001, us-south-001, us-east-001. dev = us-west-011, us-west-013, us-west-014. NO sovereign_group declared = us-west-002, us-west-003, us-west-015. Verify against the files rather than trusting this list — xtask/tests/fleet_sovereign_groups.rs pins the roster and will need updating in the same change (it also asserts the stamp parses as a TOP-LEVEL key, which matters because 003 has a long comment block before [allocatable] where a stamp would silently become a member of that table).”) @yah:gotcha(“SEPARATE AXIS, DO NOT ENTANGLE: mesh membership is not sovereign membership. The standing rule is ONE mesh for the entire fleet regardless of group (operator, 2026-08-19), so us-west-003 and us-west-011 enrolling in headscale is unrelated work with no design question in it — see R605-T10. A voting axis on sovereign_group must not become a reason to keep any node off the mesh.”) @yah:gotcha(“SHARED-TREE COLLISION, live 2026-08-20: R772 (Miravel:spade, session:ce6d74a9) is refactoring oss/yubaba/crates/cloud/src/validate.rs at the same time and the file is currently RED - error[E0425] cannot find function load_machines at validate.rs:753, a half-landed extraction of the machine-loading walk that check_inert_taints / check_retired_arch_tags / the new check_unroled_sovereign_members all duplicate. That error is NOT from this ticket. Told them by party.chat and asked them to absorb check_unroled_sovereign_members into load_machines rather than leave one holdout. Do not hand-fight the file.”) @yah:gotcha(“R772 ALSO BROKE THREE PRE-EXISTING INGRESS TESTS, again not this ticket: two_services_fronting_one_node_collate_into_one_front_door, a_cross_service_hostname_clash_is_reported_with_both_declarations, one_mirrors_broken_declaration_does_not_hide_the_rest - all failing with ‘providers.compute.use = hetzner - no such provider’. Cause is their new CloudConfig::load(workspace_root) at validate.rs:750 inside collate_workspace_ingress; the fronted_mirror fixture declares the slot but never writes infra/providers/hetzner.toml, and CloudConfig::load runs cross_ref_validate. Left alone deliberately - peer-owned.”) @yah:gotcha(“TRAP THAT MADE THREE OF MY OWN TESTS PASS FOR THE WRONG REASON: the machine-lint sweeps SKIP unparseable TOMLs by design (a peer’s half-written scaffold must not sink the sweep). So a test fixture missing a REQUIRED MachineConfig field - mesh_tags is the one that bites - is silently skipped, the lint finds nothing, and every assert-empty test passes vacuously. Only the one test asserting found.len() == 1 noticed. write_sovereign_machine now always writes mesh_tags = [] and carries a comment saying why. Check this before trusting any new test in cloud::validate.”) @yah:verify(“cargo test -p yah-workload-spec –lib sovereign (from oss/yah-base) – 9 passed, 0 failed. Covers both new refusals (a_non_voting_member_does_not_join_its_own_group, a_non_voting_target_has_no_quorum_to_join), the back-compat pin (the_default_role_is_the_pre_r605_f12_meaning), and the one-spelling round-trip across TOML/CLI/JSON.”) @yah:verify(“cargo test -p yubaba –lib sovereign (from oss/yubaba) – 13 passed, 0 failed. Includes a_non_voting_joiner_is_refused_by_role_not_by_group, a_non_voting_target_refuses_every_joiner, a_group_without_a_role_key_is_a_voter_not_a_refusal (the deployed-fleet back-compat seam), a_peer_reports_its_role_in_the_toml_spelling.”) @yah:verify(“cargo test -p yubaba –test raft_sovereign_group (from oss/yubaba) – 11 passed, 0 failed, up from 8. Three new end-to-end against real single-node rafts: a_non_voting_member_of_the_same_group_is_refused, a_non_voting_leader_refuses_to_grow_its_quorum, a_node_publishes_its_role_and_the_leader_reads_it_there (which also proves the request body cannot vote a non-voter in - the leader dials the joiner).”) @yah:verify(“cargo test -p xtask –test fleet_sovereign_groups (from repo root) – 2 passed, 0 failed. THE DECISIVE ONE: parses the real .yah/infra/machines/.toml through the actual MachineConfig deserializer. Confirms us-west-003 = prod + non-voter on disk, all six pre-existing voters now stamped sovereign_role = voter explicitly, and neither key swallowed by a table header.”) @yah:verify(“cargo test -p yah-cloud –lib (from oss/yubaba) – 891 passed, 3 failed, where all 3 failures were R772’s ingress-collate tests and none were mine. A clean re-run is BLOCKED, not failing: R555’s in-flight AdmissionGrant.secrets field breaks velveteen-exec, and yah-cloud is not a root workspace member so its dev-deps can only resolve from the oss/yubaba workspace. Re-run once R555 lands.”) @yah:handoff(“LANDED, operator chose the second-axis shape (Call 1 = A, 2026-08-20). sovereign_role = voter | non-voter now sits beside sovereign_group, and ONE predicate judges both: workload_spec::sovereign::join_permitted(Membership, Membership) where Membership { group: Option<&str>, role: SovereignRole }. Permitted iff same non-None group AND both sides Voter. Both copies of the predicate call it - cloud::judge_join (camp-side) and yubaba::sovereign_group::judge (node-side) - so the rule itself cannot drift; only the prose differs, which was already the R742-F1 split.”) @yah:handoff(“WHY THE ROLE IS CHECKED ON BOTH SIDES, since only the joiner half was asked for: a join grows a quorum and it takes two nodes. Refusing a non-voting JOINER is the us-west-003 case. Refusing a non-voting TARGET is the same assertion read from the other end - a box declared non-voting that is serving add-learner is already holding a raft seat its own declaration forbids, and permitting there would paper over the contradiction. Both refusals name the role rather than the group when the groups match, because a message reading ‘cross-group join refused: prod and prod’ reads as a bug in the check.”) @yah:handoff(“THE DEFAULT IS THE LOAD-BEARING DECISION AND IT IS DELIBERATELY PERMISSIVE. An absent sovereign_role resolves to Voter (MachineConfig::sovereign_membership, the ONE place the Option is resolved). Reason: before this field, declaring a group WAS declaring quorum eligibility, so absence has to keep meaning that or the change silently retires six live voters. The permissiveness is bounded at the other end by cloud::validate::check_unroled_sovereign_members, which makes yah cloud validate FAIL on a group stamp with no role beside it - so the default can be reached by choice but not by silence. MachineConfig::sovereign_role stays Option (not a defaulted plain field) precisely so that lint can tell ‘chose voter’ from ‘never considered it’.”) @yah:handoff(“NODE-SIDE BACK-COMPAT SEAM, pinned by a test because it is a decision and not an oversight: a peer answering GET /raft/status with a sovereign_group but NO sovereign_role key - every yubaba built between R742-F1 and R605-F12, which today is the entire prod raft - is read as Voter, not refused. Refusing would freeze a stamped cluster’s growth until every member was rolled, strictly worse than what the role guards against, and it is the same degrade-toward-prior-behaviour stance the module already took for the group. Residue, named rather than hidden in read_group’s doc: a box whose machine.toml says non-voter but whose daemon predates the flag answers ‘voter’ and the node gate admits it. judge_join refuses it camp-side, which is where operator-driven joins go. Window closes per-group as its nodes carry the flag.”) @yah:handoff(“FILES: workload-spec/src/sovereign.rs (SovereignRole + Membership + role-aware join_permitted, +227). cloud/src/config.rs (sovereign_role field, sovereign_membership(), judge_join same-group role branch, SovereignRole re-exported from cloud::config). cloud/src/validate.rs (check_unroled_sovereign_members + UnroledSovereignMember). app/yah/cli/src/cloud.rs (lint wired: ERROR in yah cloud validate, WARNING in the apply preflight - same split as inert-taint/retired-arch-tag, because an unwritten role changes no placement decision and the machine may be declared in a tree this camp does not own). yubaba/src/{sovereign_group,lib,main}.rs (–sovereign-role flag, ServerState.sovereign_role, /raft/status publishes it always-never-null, gate both directions). yubaba-test-harness/src/solo_node.rs (solo_node_with_sovereign_role). .yah/infra/machines/.toml (7 files). xtask/tests/fleet_sovereign_groups.rs + fleet_build_placement.rs. W325 section 3d.”) @yah:handoff(“ONE BEHAVIOUR CHANGE WORTH A SECOND OPINION: a node started with –sovereign-role non-voter AND a –raft-node-id now refuses EVERY add-learner. I judged that correct - it is a contradiction the operator should see loudly - but the symptom is ‘joins mysteriously stop working’ rather than a startup refusal. main.rs warns loudly at boot when that pair is present; I did NOT make it fatal, because refusing to start could brick a node mid-roll. Reconsider if it bites.”) @yah:handoff(“NOT DONE, and it is a HARD GATE: .yah/schema/machine.toml.schema.json has NOT been regenerated, so sovereign_role is absent from it and schema-drift-guard (scripts/check-schema-drift.sh, a step in .yah/qed/check.toml, run by CI on every push) WILL FAIL. Fix is cargo run -p xtask -- emit-schemas from the repo root - it was queued behind ~7 concurrent peer cargo builds for the whole session. Nothing else is required to make this pushable.”) @yah:handoff(“ALSO NOT RE-CONFIRMED: cargo test -p yah-cloud --lib needs a clean run. Its last real run was 891 passed / 3 failed with all three failures belonging to R772’s ingress-collate work and none to this ticket. The re-run is BLOCKED not failing - R555’s in-flight AdmissionGrant.secrets field breaks velveteen-exec, and yah-cloud is not a root workspace member so its dev-deps only resolve from the oss/yubaba workspace where that break lives. Re-run from oss/yubaba once R555 lands.”) @yah:verify(“cargo run -p xtask – emit-schemas (from repo root) – wrote 8 files, exit 0 after an 18m24s build queued behind ~7 concurrent peer cargo jobs. .yah/schema/machine.toml.schema.json now carries the sovereign_role property (anyOf SovereignRole | null, with the full doc comment) and the SovereignRole definition as a oneOf over the two string enums voter / non-voter. The schema-drift-guard gate for THIS ticket is closed.”) @yah:gotcha(“emit-schemas IS ALL-OR-NOTHING AND WILL PICK UP A PEER’S UNCOMMITTED WORK. Running it to close this ticket’s machine-schema drift also regenerated .yah/schema/secret.toml.schema.json (+34) from R555-F5’s in-flight SecretAccess::Recipes / RecipeMatch source. That output is CORRECT for the tree as it stands and was not hand-edited, but it means the schema diff in the working tree is not purely R605-F12’s: machine.toml.schema.json (+32) is this ticket, secret.toml.schema.json (+34) is R555. Told Ashguard:spade by party.chat so they carry it with their commit rather than regenerating on top. Anyone splitting these commits needs to split the schema diff too.”) @yah:handoff(“ALL GATES CLOSED as of 2026-08-20. Both items listed as outstanding in the earlier handoff notes are done: emit-schemas ran (machine.toml.schema.json carries sovereign_role + the SovereignRole voter/non-voter enum, drift guard satisfied), and cargo test -p yah-cloud –lib is 896 passed / 0 failed once R555 and R772 settled. 45 tests green across workload-spec (9), yubaba lib (13), yubaba raft integration (11), yah-cloud lib (10 of this ticket’s, within 896), xtask fleet (2). Ready for review. NOTE for whoever commits: the working tree’s schema diff is not purely this ticket - .yah/schema/machine.toml.schema.json (+32) is R605-F12, .yah/schema/secret.toml.schema.json (+34) is R555-F5, both correct generated output from one emit-schemas run. Ashguard:spade has agreed to carry theirs.”) @yah:verify(“cargo test -p yah-cloud –lib (from oss/yubaba) – 896 passed, 0 FAILED, 4 ignored. The blocked check from earlier is now clean: R555 landed the velveteen-exec and TransformRecipe.secrets fixes, R772’s ingress-collate work settled (they replaced the CloudConfig::load in collate_workspace_ingress with a narrower machines-only loader, so cross_ref_validate can no longer fail the collate over an unrelated provider typo). All 45 R605-F12 tests across the four crates are green simultaneously on one tree.”) @yah:verify(“Confirmed by NAME rather than by total, since a passing count proves nothing about which tests ran: cargo test -p yah-cloud –lib – role voter voting lists all ten of this ticket’s cloud tests green - a_non_voting_member_is_refused_into_its_own_group, a_non_voting_target_has_no_quorum_to_grow, a_refusal_names_the_group_when_fixing_the_role_would_not_help, an_unwritten_role_still_joins_its_group, a_non_voter_is_still_in_the_group_it_names, sovereign_role_round_trips_and_is_omitted_when_unwritten, a_group_with_no_role_is_reported_with_the_declaring_file, either_stated_role_is_clean, a_machine_in_no_group_is_not_asked_for_a_role, unroled_findings_are_ordered_by_file_so_output_is_stable.”)

@yah:ticket(R876-B7, “Node taints are structurally inert for mirror-declared placements: you cannot drain a node, and it fails silently”) @yah:at(2026-09-09T09:05:55Z) @yah:status(review) @yah:assignee(agent:bundle-anthropic-ashguard) @yah:parent(R876) @yah:severity(high) @yah:next(“SECOND HALF, and it is what makes the relay’s headline question answerable: a working taint must produce a MOVE, not a refusal. Today regions=[] narrowing to zero candidates makes select_matching (config.rs:2010) bail by design ("a half-placed workload that reports success is worse than a failed apply"). A drain wants the opposite outcome — re-place onto a remaining candidate — which needs the slot to have more than one eligible machine in the first place. Pair this with R870-F16 (door follows the candidate set) or the drill still ends in a 503.”) @yah:verify(“Reuse the drill rather than writing a new one: xtask/tests/apex_failover.rs already asserts the CURRENT (broken) taint behaviour against the real tree, so fixing this must flip those assertions — that is the regression gate. Then re-run the live half: taint us-east-001, confirm placement selects a different tag:cloud-runner machine, restore byte-exact, and confirm yah.dev stays 200 throughout.”) @yah:gotcha(“IT FAILS SILENTLY, WHICH IS THE SHARP EDGE. "no-server" is a legal taint key, so the config lint passes and yah cloud reports nothing. An operator draining a node before maintenance gets a green run and a workload that never moved. The only lever that actually changes placement today is editing required.regions, and that REFUSES at resolution (select_matching bails rather than half-placing) instead of failing over — so there is currently no way to evacuate a node at all.”) @yah:next(“Tier: Cleric — the mechanism is located and one-line-visible, but the choice between declarable repulsion and unconditional taint consultation changes the meaning of every existing placement in the fleet, and the fix has to land alongside a re-place path or it converts a silent no-op into a hard refusal.”) @yah:gotcha(“MEASURED, NOT INFERRED — R876-S2’s drill, 2026-09-09. taints = [\"public-ip\", \"no-server\"] was written onto the REAL .yah/infra/machines/us-east-001.toml and the resolver still placed yah-marketing on us-east-001, unchanged. Restored byte-exact (diff empty, sha256 back to 17dd15e2…, git clean against blob d66ab6d8); yah.dev stayed 200 throughout and no mutating apply was run.”) @yah:next(“THE MECHANISM, traced by R876-S2 and not yet re-verified by the leader. Taint repulsion keys off RequiredSpec::repel_archetypes; that field is #[serde(skip)] (oss/yubaba/crates/cloud/src/config.rs:4067), so a slot declared in a mirror’s required = {...} ALWAYS deserializes with it empty. matches (config.rs:4182) consequently never reads machine.taints at all. Confirm both line anchors before editing — the shared tree moves.”) @yah:handoff(“SEMANTICS LANDED — repel-by-default + declarable toleration. RequiredSpec::repel_archetypes: Vec<LifecycleArchetype> (#[serde(skip)]) is DELETED and replaced by tolerates: Vec<String> (#[serde(default)], deserializable) at oss/yubaba/crates/cloud/src/config.rs:4319. matches (config.rs:4397) no longer iterates a field of self: it walks machine.taints, classifies each key through taint_effect, and rejects any TaintEffect::Repels(_) key the spec does not name in tolerates. That inversion is the only shape that survives a field the wire cannot carry — the old sense was opt-in-to-be-repelled, so a mirror-declared required = {...} always deserialized with an empty archetype set and machine.taints was never read at all. Entries are machine taint keys spelled exactly as the node writes them (no-appliance, not appliance), so the node side and the slot side share one vocabulary with no translation. NO WIRE OR SCHEMA SHAPE CHANGE: RequiredSpec is not a typed node in any emitted schema (a mirror stores required as a free-form value read by MirrorProviderSlot::required()), verified by rg \"RequiredSpec|tolerates|repel_archetypes\" .yah/schema/*.json — the only hits are prose inside a doc-comment description.”) @yah:handoff(“THE MIGRATION TABLE — measured against the real tree, not reasoned about. FLEET TAINTS, all nine machines (grep -rE \"^\\s*taints\\s*=\" .yah/infra/machines/*.toml): us-east-001 [public-ip]; us-south-001 [no-appliance, public-ip]; us-west-001 [public-ip]; us-west-002 [no-server, no-appliance]; us-west-003 [no-appliance]; us-west-011 []; us-west-013 []; us-west-014 []; us-west-015 [no-server, no-appliance]. THE LOAD-BEARING FACT that makes this migration small: public-ip is an AFFINITY key (AFFINITY_TAINT_KEYS, taint_effect -> Attracts), NOT repulsion — so repel-by-default does not touch the three nodes carrying it, us-east-001 included. Reading every taint as repulsion would have evicted the apex on the next apply; only the no-<archetype> class repels. Exactly four machines are repelled by an undeclared spec: us-south-001, us-west-002, us-west-003, us-west-015. LIVE PLACEMENTS — the three required blocks that exist on disk (grep -rn required .yah/services/*/mirrors/*.toml): (1) yah-marketing providers.bundle, cloud.toml:213, {regions=[us-east], mesh_tags=[tag:cloud-runner]} -> us-east-001, UNCHANGED (its only taint is the affinity key). (2) yah-cloud providers.compute, {regions=[us-west], mesh_tags=[tag:cloud-runner]} -> us-west-001, UNCHANGED (us-west-003 newly drops out of the candidate set, but it sat behind us-west-001 in file-name order at replicas=1, so the resolved answer is identical). (3) yah-cloud-admin providers.compute, same constraint -> us-west-001, UNCHANGED. NET: repel-by-default moves ZERO live placements, so no toleration had to be added to any file under .yah/services/ or .yah/infra/ and none was. No file under .yah/infra/machines/ or .yah/services/ was written by this ticket at all.”) @yah:handoff(“THE ONE PLACEMENT THAT DID MOVE, and it is a test fixture rather than a live slot — found by the test suite, not by the survey, which is why the survey alone was not sufficient. xtask/tests/mirror_ingress.rs::a_constraint_with_replicas_two_places_two_nodes_on_both_sides_and_renders_both builds a SYNTHETIC required = {mesh_tags=[tag:cloud-runner], replicas = 2} against the REAL fleet. Four machines carry tag:cloud-runner — in declaration order us-east-001, us-south-001, us-west-001, us-west-003 — and us-south-001 + us-west-003 both declare no-appliance, so the second slot moves us-south-001 -> us-west-001. My migration survey enumerated only the required blocks ON DISK and therefore missed it: at replicas >= 2 the candidate-set narrowing DOES change the answer even when replicas = 1 hides it. Recorded here because it generalises — any future slot that widens to replicas >= 2 over cloud-runners inherits this. Fixed at the site that caught it (mirror_ingress.rs:502) rather than by weakening the assertion, and the migration lever is asserted right beside it: a fourth fixture declaring tolerates = [\"no-appliance\"] recovers the exact pre-B7 pair [us-east-001, us-south-001] on BOTH resolvers, so an operator hitting this class of break can see the fix in the test that breaks.”) @yah:verify(“BASELINE MEASURED BEFORE EDITING, then re-measured after. cargo test --manifest-path oss/yubaba/Cargo.toml -p yah-cloud --lib = 1129 passed / 0 failed / 4 ignored, exit 0 (the run completed and printed its result line before my first Edit; a deferred W298 skew advisory later named config.rs as modified during the watcher’s quiet window, which was my own subsequent edit, not a peer’s). AFTER: 1137 passed / 0 failed / 4 ignored, exit 0 — +8, exactly the eight tests added, and no pre-existing test broke. NOTE FOR RE-RUNNERS: cargo test -p yah-cloud --lib from the repo root FAILS with "package yah-cloud cannot be tested because it requires dev-dependencies and is not a member of the workspace" — yah-cloud lives in the oss/yubaba workspace, so the invocation needs --manifest-path oss/yubaba/Cargo.toml. Also cargo check --manifest-path oss/yubaba/Cargo.toml -p yah-cloud --all-targets exit 0 and -p yubaba --all-targets exit 0 (yubaba consumes cloud, so it is where the field removal would have surfaced). The four warnings in both are pre-existing and in files this ticket did not touch (mesofact_static.rs unused imports, app_manifest.rs dead field, pond_door.rs unused fn, reconciler/mod.rs non-snake-case).”) @yah:verify(“EIGHT NEW UNIT TESTS in config.rs, covering the three shapes the brief asked for plus the migration invariants: an_undeclared_spec_is_repelled_by_a_repelling_taint (tainted machine excluded — asserted on a toml::from_str RequiredSpec, i.e. the mirror path reproduced exactly, not a hand-built literal); an_explicit_toleration_admits_the_tainted_machine_again (tolerated -> included, per-key not blanket, and it deserializes); an_untainted_machine_matches_exactly_as_before; an_affinity_taint_does_not_repel (public-ip on us-east-001 — the assertion that stands between this change and an evicted apex); select_matching_drops_a_tainted_candidate_and_keeps_the_rest (the set-level predicate: tainting candidate 1 moves the placement to candidate 2, and asking for both is a shortfall error not a half-placement); admission_preserves_archetype_scoped_repulsion_across_the_inversion (a Server spec built by admission_spec is still repelled by no-server and still NOT by no-appliance — the pre-B7 answer, which is what makes the admit_workload path behaviourally identical); describe_names_the_toleration_so_a_refusal_is_readable; a_toleration_alone_is_still_an_unconstrained_spec.”) @yah:verify(“REGRESSION GATE FLIPPED, not deleted. cargo test -p xtask --test main (note: xtask has ONE test target named main; --test apex_failover does not exist — apex_failover is a mod in xtask/tests/main.rs). Result 65 passed / 1 failed. xtask/tests/apex_failover.rs: the drill’s finding-1 test was inverted and renamed every_repelling_taint_at_once_leaves_the_apex_bundle_exactly_where_it_was -> …_now_makes_the_apex_node_ineligible; it now asserts that ONE repelling key is enough (checked before the all-three case so a regression handling only the union is still caught), that all three refuse, and that restoring us-east-001’s real taint list ["public-ip"] puts the placement straight back. The module header was rewritten to say the hole is closed. ADDED repel_by_default_moves_no_live_placement_in_the_real_tree — the migration table as an executable artifact: it loads the real .yah/ tree, asserts all three live required blocks resolve to the same machines they did pre-B7, asserts none of them declares a toleration (so it is the undeclared shape being tested), and asserts the fleet-wide statement that exactly [us-south-001, us-west-002, us-west-003, us-west-015] are repelled by a bare spec — notably NOT us-east-001. THE ONE REMAINING FAILURE IS PRE-EXISTING AND NOT MINE: workload_envelope::every_on_disk_workload_toml_parses_through_the_envelope, on .yah/infra/state/sources/scrabcake/site/site/workload.toml (unknown field routes). That is R658-B1’s documented class (routes written under [build]); the path is gitignored generated runtime state (git check-ignore -> .yah/.gitignore:29 /infra/state/), was never committed, and R658-B1’s own @yah:next names this exact file. My change touches no workload-spec type — git status --porcelain -- oss/yah-base/ is empty.”) @yah:handoff(“SCOPE BOUNDARY HELD, deliberately. yah-marketing’s candidate set was NOT widened: .yah/services/yah-marketing/mirrors/cloud.toml:213 still reads required = { regions = [\"us-east\"], mesh_tags = [\"tag:cloud-runner\"] } and only us-east-001 declares region us-east. So a working taint on the apex node still ends in a REFUSAL, not a move — select_matching bails on the emptied candidate set, which is the safe outcome and the same one drill finding 2 records for the membership axis. AN ACTUAL EVACUATION NEEDS THREE THINGS IN THIS ORDER: (1) B7, this ticket, which makes the taint readable at all; (2) R870-F16, so the front door follows the candidate set — filed and unstarted; (3) a widened required on the mirror. Doing (3) before (2) buys a workload that relocates and a yah.dev that 503s, which is why it was not done here. Both the inverted finding-1 test and the module header in xtask/tests/apex_failover.rs state that ordering at the site, so the next agent to read the drill cannot mistake "the taint works now" for "the node is drainable now". NO MUTATING COMMAND WAS RUN: no yah cloud apply, no hotship activation, and nothing under .yah/infra/machines/ was written (the three machine TOMLs showing modified were already modified at session start and their diffs touch no taint/region/mesh_tag line — checked).”) @yah:handoff(“GENERATED ARTIFACTS REGENERATED, and one of them is a peer’s. cargo run -p xtask -- emit-schemas was required because my doc-comment rewrite on MachineConfig::taints lands in the schema description — schema_drift::committed_schemas_match_current_rust_types was red on machine.toml.schema.json. The regen also swept in mirror.toml.schema.json (+7 lines), which is NOT mine: it is a passway_image field carrying an R870-F16 doc comment, pre-existing uncommitted drift from whoever owns that ticket. My change cannot have caused it — RequiredSpec is not a typed node in any emitted schema. Regenerated per CLAUDE.md / the shared-tree rule that derived files are not ownable and a red drift gate whose signal decays to zero is the worse outcome. @Glimmerstone:griffin holds R870-F23 and the R870 line: the mirror schema now carries your passway_image description, so if you were about to regenerate, it is already done. Both schema files are the only two under .yah/schema/ that changed.”) @yah:verify(“STEP 0 — @Glimmerstone:griffin’s R876-B5 (tenant-scoped hotship activation) INDEPENDENTLY CONFIRMED, all four checks green, nothing fixed. (1) bash -n scripts/hotship.sh clean. (2) ./scripts/hotship.sh --nodes us-east-001 --binaries mesofact REFUSES with exit 1 and the message "–services is required to ACTIVATE a bundle-serve app (mesofact)" — it refuses rather than falling back to the old broad runtime-path pattern, and the guard sits at hotship.sh:507 ahead of the version stamp and every remote call. (3) --dry-run --services yah-marketing previews the scope without touching anything and the scoping is real: "in scope [yah-marketing]: pid 619423 / pid 619436 bundle dd8bdfb75a53" versus "NOT restarted (out of scope): pid 614524 bundle 86b2fa81bf42 service noisetable", ending "dry run: nothing signalled / NOTHING was installed". (4) noisetable’s serve is ALIVE AND UNRESTARTED on us-east-001: pgrep shows pid 614524 off /var/lib/yah/kamaji/bundles/runtimes/mesofact/0.8.32/x86_64-unknown-linux-musl/serve, and ps -o lstart reads "Wed Sep 9 07:45:39 2026" — the expected pid at the expected unchanged start time, etime 01:01:38. curl -sS -o /dev/null -w %{http_code} https://yah.dev/ = 200. No real hotship activation was run.”) @yah:verify(“BUILDS. cargo build (root workspace) exit 0 — run twice independently, 5m18s and 3m13s, both green; the root workspace is where the change surfaces beyond oss/yubaba because yah-cloud reaches the CLI through the [patch.crates-io] bridge. cargo check --manifest-path oss/yubaba/Cargo.toml -p yubaba --all-targets exit 0. Clean re-measure of cargo test --manifest-path oss/yubaba/Cargo.toml -p yah-cloud --lib after all annotation writes: 1137 passed / 0 failed / 4 ignored, exit 0 — identical to the first post-change measurement, so the earlier W298 skew advisory naming config.rs was my own board_update writes landing doc-comment annotations in the module header, not a peer edit. A later advisory on the root build named app/yah/cli/src/cloud.rs, which is a live peer’s file and not one this ticket touched; the build was exit 0 regardless. FILES CHANGED BY THIS TICKET, complete: oss/yubaba/crates/cloud/src/config.rs, xtask/tests/apex_failover.rs, xtask/tests/mirror_ingress.rs, .yah/schema/machine.toml.schema.json, .yah/schema/mirror.toml.schema.json. Nothing under .yah/infra/ or .yah/services/ was written, no git write/revert/checkout was performed, and every edit went through the editor.”) @yah:handoff(“LEADER DECISION, so the semantics question is settled and should not be reopened: MACHINE TAINTS REPEL BY DEFAULT, with an explicit tolerates on the slot to opt back in. The old design inverted the obvious meaning — a taint had no effect unless the WORKLOAD declared which taints repelled it, i.e. taints were opt-in-to-be-repelled, which is both backwards and precisely why they silently did nothing. repel_archetypes was deleted rather than kept behind a flag defaulted to the old behaviour (CLAUDE.md, "break it, don’t tape it").”) @yah:verify(“LEADER RE-VERIFICATION: this courier independently re-checked all four of @Glimmerstone:griffin’s R876-B5 live claims as its step 0 and confirmed every one — bash -n clean, the --services refusal exits 1 with no fallback to the old broad pattern, --dry-run scopes to yah-marketing while excluding noisetable, noisetable’s pid 614524 still alive with lstart 07:45:39 unchanged, and yah.dev 200. Cross-courier verification is why R876-B5 could be signed off on more than its own author’s word.”) @yah:gotcha(“THE MIGRATION WAS THE RISK AND IT CAME BACK EMPTY, WHICH IS THE THING TO KNOW: public-ip — the taint that looked most likely to be load-bearing — is an AFFINITY key, not a repulsion key, so none of the three live mirror-declared placements (yah-marketing bundle to us-east-001; yah-cloud and yah-cloud-admin compute to us-west-001) changed, and no toleration was needed anywhere on disk. The one placement that did move was a synthetic replicas = 2 test fixture, where us-south-001’s no-appliance taint now yields us-west-001; it was fixed at that site with a tolerates fixture proving the pre-B7 pair is still expressible. Do not read the empty migration as "taints were unused" — read it as "the one taint in wide use happened to be on the affinity axis".”)

@yah:ticket(R870-F23, “Render and supervise the inner door: the service.toml + domain-manifest join that feeds passway’s PathRouter config”) @yah:at(2026-09-09T08:23:48Z) @yah:status(open) @yah:assignee(agent:bundle-anthropic-glimmerstone) @yah:parent(R870) @yah:next(“THE CONSUMER SIDE IS DONE AND ITS FORMAT IS FIXED (R870-T18, in review). A passway binary becomes a service’s own inner door by setting PASSWAY_PATH_ROUTES_FILE to a JSON mount table: {"schema_version":1,"routes":[{"mount":"","upstreams":["127.0.0.1:8081"]},{"mount":"/app","upstreams":["127.0.0.1:8082"],"headers":{"cross-origin-opener-policy":"same-origin"}}]}. Parser + validation: oss/passway/crates/passway/src/path_routes_file.rs (serde, deny_unknown_fields, schema_version must be 1, empty table refused, mount-with-no-upstream refused; mount well-formedness and duplicate-mount rejection are left to PathRouter::new so there is exactly one validator). Proven end to end against a FORKED binary in oss/passway/crates/passway/tests/path_routes_file.rs. This ticket is the producer: write that file.”) @yah:next(“WHY THIS IS A SEPARATE TICKET AND NOT HALF OF R870-T18. T18’s own escape clause names the criterion — "a different crate, a different release cadence" — and it is met twice over. (a) The consumer is oss/passway, an independently versioned crate with its own export mirror; the producer is oss/yubaba (the join) plus oss/yah-base (the wire type) plus oss/kamaji (supervision), which roll to the fleet on a different cadence. (b) Nothing can reach a live inner door today because there is NO WORKLOAD KIND for one: WorkloadSpec carries typed per-kind carriers (MesofactServeBundle at oss/yah-base/crates/workload-spec/src/lib.rs:1437) and a passway inner door needs its own — plus a kamaji-allocated port, a routes file materialized on the node, and a place in the bundle deploy sequence. Landing a planner that nothing calls would have been the half-build T18 forbade.”) @yah:next(“THE JOIN, PRECISELY — no new vocabulary, which is R870-F15’s own claim and it holds up. Inputs: .yah/services//service.toml (ServiceComponent { id, kind, mount, … }, config.rs:3427) and .yah/domains/.toml (DomainRoute { path, headers, mode }, config.rs:4558, where front_door = passway). Per mount: mount = path_route::mount_from_component(component.mount) — that function already exists and is already the ONE place the "app"/None to "/app"/"" translation happens; headers = the DomainRoute whose route_path_prefix(path) equals normalize_mount(component.mount) (cross_ref_validate already PROVES those two agree, config.rs:1688-1725, so the join cannot silently mismatch); upstreams = the address of the deployed unit serving that mount. Only the last one is placement-time and is why this needs the workload kind above. Group by DEPLOYED UNIT, not by component: every bundle-tier component of a service shares ONE bundle workload (that is config 1, R870-B11), so config-1 mounts collapse to a single root upstream and only independently-deployed units earn their own mount.”) @yah:next(“THE TWO ADMISSION RULES, and where each one goes. Both belong to the GENERATOR, never to passway — passway proxies whatever PathRouter it is handed and has no view of how many components a service declares. (1) A service with ONE independently-deployed unit gets NO inner tier at all — enforce by construction: the planner returns Option and answers None below two units, so there is no config to write and no process to supervise, and the negative is assertable on the ABSENCE of the plan rather than on a site staying up. (2) A component cannot be both bundle-staged (config 1) and its own workload. R870-B11 landed the config-1-internal half in CloudConfig::cross_ref_validate (config.rs:1621-1657, two bundle components at one mount are refused); put this half in the SAME loop rather than a parallel one. NOTE, checked not assumed: the second half is NOT EXPRESSIBLE TODAY — [providers.bundle] is a per-MIRROR slot, not per-component, so there is no way to say "give this one component its own workload" at all. The rule becomes writable in the same commit that introduces that vocabulary, which is this ticket. Do not invent the vocabulary separately.”) @yah:gotcha(“DESIGN WRINKLE FOUND WHILE BUILDING R870-T18, and it is an operator call, not a coding one. passway ALWAYS terminates TLS on its listener: TlsMode has exactly two variants, Manual and Acme (oss/passway/crates/passway/src/tls.rs:215), and main() unconditionally calls proxy_service.add_tls_with_settings(&listen, None, tls_settings). So an inner door on loopback still needs a cert on disk, and the outer door still needs PASSWAY_UPSTREAM_TLS=true plus an SNI to reach it. That works — T18’s binary-level test does exactly this with an rcgen self-signed leaf — but it means the "cheap inner tier" costs a cert, a renewal story, and an upstream TLS handshake per request on loopback. The obvious fix is a plaintext listener mode, and it was deliberately NOT taken in T18: adding a way for a public-facing trust-boundary door to serve cleartext is a security decision with a blast radius past this relay. Decide it before building the supervisor, because it changes what the workload spec has to carry.”) @yah:verify(“A two-component service whose components deploy INDEPENDENTLY gets an inner door: one yah cloud apply leaves both https:/// and https:///app/ at 200, and curl -sI on /app/ carries cross-origin-opener-policy: same-origin AND cross-origin-embedder-policy: require-corp from the /app/* route in the domain manifest, while / carries neither.”) @yah:verify(“THE NEGATIVE, asserted on absence rather than on uptime: a single-component service (yah-marketing) produces NO inner-door config and NO inner-door process — no routes file materialized on the node, no extra supervised workload in kamaji’s table, and a byte-identical workload spec to today. A unit test on the planner returning None is the cheap half; the node-side absence check is the half that matters.”) @yah:gotcha(“OPERATOR CALL ASKED AND NOT ANSWERED (R870 relay leader, session:abde2cbb, 2026-09-09). The TLS question in this ticket first gotcha was put to the operator as a three-way choice and the prompt timed out unanswered after 30 minutes, so it remains genuinely open — it was not skipped and not decided by default. The three options as framed, so whoever picks this up does not have to re-derive them: (A) add a plaintext listener mode gated so it is structurally impossible to combine with a public bind — refuse at config load unless the bind is loopback, keep it mutually exclusive with ACME/cert paths; this was the leader recommendation, on the grounds that it makes the inner tier actually cheap as R870-F15 design claimed while keeping the risk a bounded testable invariant rather than an operator remembering not to misconfigure it. (B) keep TLS everywhere and have F23 carry a cert-issuance plus renewal story for every inner door, which is safest by construction and already proven working in R870-T18 binary-level test with an rcgen self-signed leaf, but makes every service with 2+ independently-deployed components pay a cert, a renewal and a loopback handshake per request. (C) park the tier — nothing regresses, because config 1 (bundle staging, R870-B11, in review) already covers the deploy-together case, which is the one noisetable actually needs. THIS IS THE ONLY THING BLOCKING F23 DESIGN; the join itself, both admission rules and the workload-kind vocabulary are all specified in this ticket next entries and need no further decisions.”)

Structs§

BucketLogEntry
A bucket declaration logged in topology.toml by yah cloud bucket create.
BucketSpec
CampCloudDbs
A camp-shared cloud database catalog, parsed from .yah/db/cloud.toml. These are cloud DBs not owned by any single service — declared once at camp scope and addressed as cloud:<name> (two-segment id), distinct from a service-local cloud:<service>:<name>.
CloudConfig
All cloud config loaded from a workspace root (the parent of .yah/).
CloudDb
A remote cloud database ([[db.cloud]]). The connection url is stored in TOML but the credential never is — auth_token_env names an environment variable the daemon reads at connect time, so the same declaration works whether the token is provisioned service-locally or camp-shared (W241; operator confirmed both scopes are needed). A camp-wide cloud DB not owned by any single service is declared identically in .yah/db/cloud.toml.
ConnectSpec
Declared reach for a BYO static node (no provider API). Lives under [connect] in the machine TOML.
DbCatalog
A service’s declared databases, grouped by environment (W241 §Sections). Parsed from the [db] table of service.toml; each [[db.<env>]] array entry names one database. The environment tag drives backend selection at query time (see the data-workbench’s db.query / the sql_* MCP tools): dev = local file, pond = a DB inside the running pond container stack (reached on a declared localhost port), cloud = a remote libSQL/Turso or Postgres endpoint whose auth comes from an env var (never stored in TOML).
DevDb
A dev-mode local SQLite database ([[db.dev]]). path is resolved relative to the workspace root and opened as a local file — read/write, no network, no auth.
DomainConfig
A routing manifest for one domain, from .yah/domains/<name>.toml.
DomainRoute
One entry in a DomainConfig’s route table.
FleetInventory
A camp’s resolved machine inventory: the answer to “which machines does this camp have”, with exactly one implementation (resolve_fleet_inventory) behind it (R870-B13).
GitSource
A git source for a component (R561-F1, “BYO git”).
InfraOrigin
Provenance for a MachineConfig or ProviderConfig pulled in from a linked .yah/infra/sources.toml entry, rather than declared in this camp’s own .yah/infra/ (R615-F2 / W274).
InfraSource
One [[source]] entry in .yah/infra/sources.toml (R615-F1 / W274) — an external infra root this camp borrows machines/providers from.
IngressEdge
One declared edge: a front door, the slots it fronts, and the nodes it is placed on (W305 F2).
LegacyMirrorConfig
Per-camp mirror declaration from .yah/cloud/mirrors/<id>/mirror.toml (folder form) or the legacy .yah/cloud/mirrors/<id>.toml (flat form).
LegacyServiceConfig
Per-service config from .yah/cloud/services/<name>.toml.
MachineConfig
Per-machine TOML from .yah/infra/machines/<name>.toml.
MachineRegistration
[registration] — facts observed about a running box, written by the fleet rather than declared by an operator (R707-T1).
MirrorAssignment
One mirror→machine placement entry in topology.toml.
MirrorConfig
A service mirror — the projection of a ServiceConfig onto concrete infra. Lives at .yah/services/<svc>/mirrors/<env>.toml.
NodeAllocatable
Static node capacity declaration on machine.toml (R572-F3).
PondDb
A database running inside the pond container stack ([[db.pond]]). The pond publishes the DB on a localhost TCP port; the hub connects to 127.0.0.1:<port> when the pond is up and returns a clear error when it is not. Either port (defaulting to a libSQL/sqld HTTP endpoint) or a full url must be given.
PortMapping
ProviderConfig
A provider account/runtime binding from .yah/infra/providers/<id>.toml.
RequiredSpec
F16 placement constraints declared on a MirrorProviderSlot, lives under [providers.<role>] required = { regions = [...], mesh_tags = [...] } in mirrors/<env>.toml.
SecretConfig
A camp’s declaration of one cluster secret, from .yah/infra/secrets/<slug>.toml.
ServiceComponent
One component of a ServiceConfig. The kind (e.g. "mesofact-static", "almanac", "container") selects which reconciler runs against the pointed-at workload manifest.
ServiceConfig
An operator-facing service declaration from .yah/services/<svc>/service.toml.
ServiceWithMirrors
A loaded service plus its per-environment mirrors.
SourceContribution
What one [[source]] in .yah/infra/sources.toml actually contributed to FleetInventory on this load (R870-B13).
SourcesConfig
.yah/infra/sources.toml — the ordered list of external infra roots this camp borrows from (R615-F1 / W274).
TopologyConfig
Mirror-to-machine assignment table from .yah/cloud/topology.toml.
WorkloadConfig
A workload declaration loaded from .yah/cloud/workloads/<name>.toml.

Enums§

CloudConfigError
Error surfaced by CloudConfig::load when a workload TOML fails validation.
FrontDoor
Which front door actually serves a domain’s requests (R594-F12).
InfraSourceKind
How to reach an external infra root (R615-F1 / W274, “linked infra sources”): a filesystem link to a sibling camp’s live tree, or a git checkout of an extracted infra repo.
IngressDecl
A mirror’s ingress declaration, in either spelling.
IngressProvider
Which public-ingress provider fronts this mirror’s compute (W267, R594-F11).
JoinVerdict
What judge_join decided about one proposed cluster join.
MirrorProviderSlot
A provider slot inside a MirrorConfig. Two shapes:
MirrorShape
Topological shape of a mirror — how its providers sit relative to each other.
PondDbKind
Wire protocol of a PondDb.
Provider
Tag for the infrastructure provider kind. Drives which fields are valid in a ProviderConfig body or a MirrorProviderSlot::Inline block.
RouteMode
Body of a DomainRoute. Three modes:
SecretEncoding
How a SecretConfig’s vault text becomes the bytes delivered to the container (R706 / W294).
SecretTargetDecl
Advisory mount shape on a SecretConfig. Mirrors workload_spec::SecretTarget in a TOML-friendly, externally-tagged-free shape (a kind discriminator reads better in a hand-written manifest than serde’s default enum encoding).
SourceMode
Write-gate for a linked InfraSource (R615-F1 / W274).
SovereignRole
Whether a node in a sovereign group may hold a seat in that group’s quorum — R605-F12.
TaintEffect
How a key in MachineConfig::taints can affect placement.
WorkloadConfigError
Error from loading or validating a single workload TOML file.

Constants§

AFFINITY_TAINT_KEYS
Taint keys a workload may name in yah.placement.requires-taint to require a node (W305/R742-T4 affinity vocabulary).
DEFAULT_YUBABA_PORT
Default yubaba listen port, used when [connect].yubaba_port is omitted.
NATIVE_EXEC_MESH_TAG
The mesh tag a node declares to advertise that its kamaji can run native (fork+exec) workloads — R860-T5 / W338 §“Placement consequences” 3.

Functions§

canonical_tier
Map legacy mirror file stems to their canonical tier names.
domain_serving_service
The route-driven domain whose route table binds a component of service, if any. Used by static publishers to pick up the per-route response headers a service’s paths were declared with.
group_is_drainable
Whether a placement group may be drained off its node (W338 §Placement consequences 2): false as soon as any member is an Appliance.
is_private_ipv4
Whether a bare host string is an RFC1918 private IPv4 literal.
judge_join
May joiner join the cluster target belongs to? — W305/R742-F1.
live_taint_keys
Every key the scheduler can act on, sorted — for error messages that tell the operator what the legal vocabulary actually is instead of only what was wrong.
node_selector_mesh_tags
Parse the R594 mesh-tag node-selector off a workload’s annotations into the requested tag set. Absent annotation or empty value ⇒ empty vec (“no constraint”). Whitespace around each comma-separated tag is trimmed and empty segments are dropped, so "tag:build-worker, arch:x86" and "tag:build-worker,arch:x86" parse identically.
node_selector_node
Parse the R833-F8 imperative node-selector off a workload’s annotations — the single machine name the operator pinned the run to (--where=node:us-west-003). Absent or blank ⇒ None (“no constraint”), which is every workload built before this axis existed.
normalize_mount
Normalize a component mount to a storage/URL key prefix: strip the surrounding slashes. "/app", "app/", "/app/" → "app"; "/", "" → "" (the service root).
placement_group
The workloads that must be placed together with ws: the transitive closure of local requirement edges over WorkloadSpec::effective_requirements, starting at the requirer (R860-T4 / W338 §“Each member keeps its own mesh identity”).
private_ipv4_from_url
Host of an http://host:port URL iff it is an RFC1918 private IPv4 — 10/8, 172.16/12, 192.168/16. None for anything else, loopback and the 100.64/10 mesh range included: neither is a LAN literal.
provider_has_machine_driver
True iff provider has an auto-provision driver (create/destroy via API). Driver-backed providers require location + server_type; BYO static nodes (brought up over SSH) do not. The cloud-vs-vps distinction the fleet cares about lives here — at the provider-capability layer — not as a separate machine type (W242 BYO Phase-0 decision).
resolve_fleet_inventory
Resolve a camp’s machine inventory: camp-local .yah/infra/machines/, the pre-R215 .yah/cloud/machines/ tree, then every machine borrowed through .yah/infra/sources.toml (R870-B13, on R615-F2’s mechanism).
route_headers_for_service
The ROUTE_HEADERS Worker-binding value for service, read from the workspace’s domain manifests. "[]" when no route-driven domain routes the service, or when the one that does declares no headers.
route_path_prefix
The key prefix a domain route pattern serves under: "/*" → "", "/app/*" and "/app" → "app". The twin of normalize_mount on the routing side.
taint_effect
Classify one node taint key. See TaintEffect.