cloud/reconciler/mesofact_bundle.rs
1//! W272 bundle tier for `mesofact-static` / `mesofact-spa` components — the
2//! config half of the services-tab sync arm.
3//!
4//! Part of R599-F8 — the canonical ticket annotation lives in
5//! `app/yah/cli/src/cloud.rs`, which owns the orchestration half. This module
6//! only owns *what the mirror declares*: parsing the `[providers.bundle]` slot
7//! and resolving which machines the built bundle gets deployed to.
8//!
9//! **The component kind does not change.** A mesofact site is a mesofact site;
10//! the mirror decides how its bytes are distributed. A mirror with a
11//! `[providers.static]` slot rides the historical build-and-publish-to-CDN path
12//! ([`super::mesofact_static`]); a mirror that declares `[providers.bundle]`
13//! rides the W272 chain instead:
14//!
15//! ```text
16//! build → bundle assembly (per-file blake3) → R2 publish → workload deploy
17//! → node materializes → kamaji forks the serve binary
18//! ```
19//!
20//! The deploy leg lives at the apply layer (`app/yah/cli/src/cloud.rs`) rather
21//! than in a `Reconciler::up`, for the same reason
22//! [`super::mesofact_runner`] does: machine resolution needs [`CloudConfig`],
23//! which [`ReconcileCtx`] deliberately does not carry. What runs here is the
24//! validation a desktop-side bring-up can still do offline —
25//! [`MesofactBundleReconciler`] checks the slot parses and the placement
26//! resolves, then bails with a pointer at the CLI.
27//!
28//! @yah:ticket(R703-T7, "Stamp a publish beacon into the W272 bundle so a passway apex can be serving-verified too")
29//! @yah:status(review)
30//! @yah:at(2026-08-08T23:55:25Z)
31//! @yah:assignee(agent:bundle-anthropic-ashguard)
32//! @yah:parent(R703)
33//! @yah:next("R703-B4 added a publish beacon (prefix/.well-known/yah-publish.json, oss/yubaba/crates/cloud/src/reconciler/publish_beacon.rs) written by the R2 publish path, and the reconciler fetches it back through the declared front door to fail an apply whose bytes nobody serves. A passway apex serves a W272 bundle, NOT the R2 prefix, so it has nothing to answer that probe with -- the front-door check can only ever pass there once the bundle carries an equivalent stamp.")
34//! @yah:next("SCOPE: stamp a PublishBeacon into the bundle at build time using the same digest shape (PublishBeacon::new + digest_of are already public and take a BTreeMap of key -> sha256; reuse them rather than inventing a second digest). It must be reachable at /.well-known/yah-publish.json through mesofact serve, which serves bundle paths directly, so it needs to be a bundle entry at exactly that path.")
35//! @yah:next("THEN: the mesofact_bundle sync path gains the same call mesofact_static::verify_serving makes. That is where the check becomes symmetric -- today only the R2 arm can prove it is being read.")
36//! @yah:verify("A bundle built for yah-marketing contains /.well-known/yah-publish.json, and curl https://yah.dev/.well-known/yah-publish.json through a passway apex returns a beacon whose digest matches the bundle that was synced.")
37//! @yah:gotcha("GATED ON R546 REGARDLESS. The bundle tier cannot sync at all until the musl serve_bins at target/x86_64-unknown-linux-musl/release/mesofact exists; slot_ready is false and .yah/services/yah-marketing/mirrors/cloud.toml falls back to the static chain. There is nothing to verify until that lands, which is why R703-B4 filed this rather than doing it in-pass.")
38//! @yah:tier(Cleric) — the digest and probe shapes are already built and public; this is threading a known artifact through the bundle builder, not a design.
39//! @yah:handoff("SHIPPED. A W272 bundle now carries a publish beacon and the bundle sync arm verifies it through the apex, so both serving tiers can prove they are being read rather than merely written.")
40//! @yah:handoff("publish_beacon.rs: BUNDLE_BEACON_PATH = app/dist/html/.well-known/yah-publish.json (the one bundle entry mesofact serve answers /.well-known/yah-publish.json from), PublishBeacon::for_bundle, bundle_beacon(), stamp_bundle(). Reuses digest_of over a BTreeMap of path -> hash as the ticket asked; no second digest was invented.")
41//! @yah:handoff("DESIGN CALL worth reviewing: the bundle stamp is clock-free. published_at became Option<String> (serde default, so beacons already in R2 still parse) and for_bundle sets None. A wall clock inside an entry of a content-addressed unit would flip the bundle digest on every assembly, breaking W272 immutability, the blob dedupe that makes a re-publish a no-op, and the assembly_is_deterministic test. The digest already names the exact immutable unit, so nothing diagnostic is lost; messages render via published_label().")
42//! @yah:handoff("Symmetry, the third next bullet: mesofact_static::verify_serving and the new cloud.rs verify_bundle_serving both call one shared publish_beacon::check_serving -> ServingVerdict, rather than the bundle arm growing a second copy that drifts. probe_urls() is the pure half (three collapses: static probes prefix + apex separately, a bundle collapses onto the apex, bucket-direct collapses onto the origin) so it is testable with no network. Probe budgets split: EDGE_PROBE (4 x 5s, CDN propagation) vs NODE_PROBE (20 x 6s) because a bundle deploy has to fetch blobs, materialize, and restart the serve process.")
43//! @yah:handoff("Stamped in assemble_component_bundle_with_sidecars (app/yah/cli/src/cloud.rs), not in the sync arm, so yah cloud bundle build and a sync still emit byte-identical trees. BundleSlot gained verify_serving (default true, non-bool rejected rather than defaulted) and an optional zone (defaults to the service domain) + serving_zone(). Documented both in the [providers.bundle] block of .yah/services/yah-marketing/mirrors/cloud.toml.")
44//! @yah:handoff("DISCOVERED WORK, outside the ticket title, done in-pass. Two claims this ticket rests on were inference, not verification, so I pinned them. (1) oss/mesofact/crates/mesofact/src/server.rs:1450 — serves_the_publish_beacon_from_a_dot_well_known_path proves GET /.well-known/yah-publish.json really returns 200 application/json through Server::from_bundle (a leading-dot directory is exactly the shape a static server tends to reject or rewrite), plus an_unstamped_bundle_does_not_answer_the_beacon_url_with_200 so an unstamped bundle 404s instead of 200-ing HTML. (2) oss/yah-base/crates/mesofact-bundle/src/store.rs:439 — a_dot_directory_entry_publishes_and_materializes proves publish_bundle + materialize_bundle round-trip the first dot-directory entry a bundle has ever carried; if checked_rel were ever tightened to a naive no-dot-segment rule, a node would refuse to materialize a bundle it had already accepted.")
45//! @yah:verify("cargo test -p yah-cloud --lib (in oss/yubaba): 741 passed, 0 failed. 23 in reconciler::publish_beacon (13 new), 34 in reconciler::mesofact_bundle (4 new).")
46//! @yah:verify("cargo test -p yah --lib: 1035 passed, 0 failed. Includes the new cloud::bundle_assembly_tests::an_assembled_bundle_carries_its_publish_beacon, and assembly_is_deterministic still passes with the stamp in place, which is the evidence the clock-free design holds W272 immutability.")
47//! @yah:verify("cargo test -p mesofact --lib server:: (in oss/mesofact): 34 passed, 0 failed. cargo test -p yah-mesofact-bundle --features store (in oss/yah-base): 31 passed, 0 failed.")
48//! @yah:verify("cargo check --workspace --exclude desktop: clean. cargo check --workspace in oss/yubaba: clean. cargo test -p xtask --test schema_drift: 3 passed, so no generated-artifact drift. yah cloud validate --path .: ok, no alias or port collisions, re-run after the mirror comment edit.")
49//! @yah:gotcha("THE SECOND HALF OF THE VERIFY LINE IS NOT DONE AND COULD NOT BE. curl https://yah.dev/.well-known/yah-publish.json still 404s, because the bundle tier cannot sync at all until R546 produces target/x86_64-unknown-linux-musl/release/{mesofact,almanac-feed}. slot_ready is false, yah-marketing still falls back to the static chain, and no bundle has been assembled by this code against the live apex. Everything is proven by test, nothing by a live apply. R703 now carries a notify_on(R546) that spells out the live run.")
50//! @yah:gotcha("When the bundle tier first turns on, expect the apply to FAIL the serving check for a while, and read that as the check working. us-east-001 is serving a hand-placed bundle from before this code existed, which carries no stamp, so the apex will answer the probe 404 (Missing) until a bundle assembled by THIS code is deployed there. Do not reach for verify_serving = false; deploy the stamped bundle.")
51//! @yah:next("LIVE VERIFY, gated on R546 and the only thing left: build the two musl binaries, yah cloud apply --service yah-marketing --env cloud, confirm it takes the bundle arm, then curl https://yah.dev/.well-known/yah-publish.json and check the digest equals bundle_beacon() over the synced manifest.")
52//!
53//! @yah:ticket(R752-B7, "revalidate routes allowlist is parsed, shipped, then dropped - the receiver accepts pokes for every route")
54//! @yah:status(review)
55//! @yah:at(2026-08-13T00:22:14Z)
56//! @yah:assignee(agent:bundle-anthropic-ashguard)
57//! @yah:parent(R752)
58//! @yah:severity(medium)
59//! @yah:gotcha("Found 2026-08-12 while wiring R330-F13's sidecar to the live receiver. `[providers.bundle.revalidate] routes` is documented as an allowlist ('empty = all routes', mesofact_bundle.rs:218), is parsed into RevalidateSlot.routes, is copied into MesofactRevalidateReceiver.routes (mesofact_bundle.rs:264), and is shipped over the wire to kamaji. kamaji then never reads it: bundle_workload_spec_revalidate (oss/kamaji/crates/kamaji-bin/src/server.rs:2117) builds the receiver's argv from publish_config + listen and its env from receiver.env, and `routes` appears nowhere. grep confirms server.rs touches receiver.feeds / feed_interval_secs / feed_project_prefix / publish_config / env and never receiver.routes.")
60//! @yah:gotcha("MEASURED, not inferred: with .yah/services/yah-marketing/mirrors/cloud.toml declaring routes = [\"/releases\"], POST http://100.64.0.3:8081/revalidate {\"routes\":[\"/issues\"]} returned 202 on us-east-001 and went on to re-render and republish /issues. An undeclared route was accepted and acted on.")
61//! @yah:gotcha("Severity is medium not high because the receiver is not publicly reachable (mesh IP, and it is the tenant's own render path) — but it IS an unauthenticated write-shaped endpoint today: its process env carries no MESOFACT_MIRROR_KEY, so mirror_key_env is unresolved too. The declared scoping control and the declared bearer are BOTH inert, which is worth knowing before anyone treats either as a boundary.")
62//! @yah:next("Decide whether the allowlist is real. If yes, pass it to the receiver (argv or env) in bundle_workload_spec_revalidate and enforce it there; if no, delete the field rather than leaving a documented control that does nothing.")
63//! @yah:next("If it becomes enforced, .yah/services/yah-marketing/mirrors/cloud.toml already lists both \"/releases\" and \"/issues\" — R330-F13 added /issues precisely so enforcement does not silently break the now-working issue-filing path.")
64//! @yah:next("Same question for mirror_key_env: it resolves to nothing today, so the receiver runs open. Whatever change starts resolving it must set the matching ALMANAC_MIRROR_KEY on the issue-tracker unit on us-east-001 in the SAME change, or the sidecar's poke starts 401ing and /issues silently stops updating.")
65//! @yah:handoff("OPERATOR CALL 2026-08-12: the allowlist is real - enforce it IF present. Auth is a separate, pluggable axis (cheers auth, preshared key, or unauthenticated are all legitimate for an almanac route); the allowlist is scoping, not authentication, and the two are now independent controls end to end.")
66//! @yah:handoff("Node leg (oss/kamaji/crates/kamaji-bin/src/server.rs, bundle_workload_spec_revalidate): each declared route is rendered as one `--allow-route <route>` on the receiver's argv. An empty list emits no flag at all, which keeps the documented 'empty = all routes' meaning - `--allow-route \"\"` would have scoped the receiver to a route that cannot exist and silently killed every revalidation.")
67//! @yah:handoff("Receiver leg (oss/mesofact/crates/mesofact/src/revalidate.rs): RevalidateConfig gained `routes`, fed by a new repeatable `--allow-route` flag on `mesofact serve`. Enforced in BOTH shapes a poke can take - an explicit `{\"route\": ...}` outside the list gets a synchronous 403 and never enqueues, and a whole-site poke (no route named) is NARROWED to the list at render time. The narrowing is the half that matters: the escape actually measured on us-east-001 sent {\"routes\":[\"/issues\"]}, which the receiver's body type does not have a field for, so it deserialized to route=None and ran as a whole-site render. A handler-only check would still have let that through.")
68//! @yah:handoff("Route selection was split out of render_routes into a pure `render_targets(workload, route, allow)` so the scoping rule is testable without booting V8 - a security-shaped control whose only evidence was 'it compiles' is how this got shipped inert in the first place. It also errors on a disallowed explicit route rather than rendering nothing, so an in-process caller cannot get a silent success.")
69//! @yah:handoff("The allowlist is intersected with the manifest, not unioned: a listed route the manifest cannot render (ssr, deferred, or a typo) is skipped instead of turning every whole-site poke into an error.")
70//! @yah:handoff("Config docs corrected where they now lie: RevalidateSlot.routes in oss/yubaba/crates/cloud/src/reconciler/mesofact_bundle.rs and the block in .yah/services/yah-marketing/mirrors/cloud.toml both said the field was inert. The cloud.toml note now says the list is LOAD-BEARING - a route absent from it stops being republished after the next deploy of that mirror.")
71//! @yah:handoff("tenants.rs (multi-tenant receiver) passes an empty allowlist with a comment naming the shape to copy - tenants/<id>.toml has no routes key yet, so per-tenant scoping is unmodelled rather than silently unenforced.")
72//! @yah:verify("cargo test -p mesofact --all-features (oss/mesofact) - 104 passed, 0 failed, including 8 new: out-of-list route 403s and does not enqueue, in-list route accepted, a correct mirror_key does NOT widen the allowlist, empty allowlist accepts anything, whole-site poke accepted then narrowed, whole-site targets = manifest INTERSECT allowlist, an allowlisted route absent from the manifest is not rendered, explicit disallowed route errors at render time.")
73//! @yah:verify("cargo test -p kamaji-bin --all-features (oss/kamaji) - 239 passed, 0 failed. The pre-existing revalidate_spec_argv_matches_mesofact_serve_clap_shape test is the one that should have caught this: it declared routes = [\"/releases\"] and pinned an argv that never mentioned it, green the whole time. It now asserts the --allow-route pair, plus two new tests for the empty-list and two-route cases.")
74//! @yah:verify("cargo test -p yah-cloud --lib mesofact_bundle (oss/yubaba) - 44 passed, 0 failed.")
75//! @yah:verify("cargo test -p xtask --test schema_drift - 3 passed; the doc-comment edits touch no schemars-derived type, so no generated artifact moved.")
76//! @yah:verify("cargo clippy --all-features --all-targets on both changed crates - no new warnings from the changed files (mesofact-core/mesofact-build/server.rs warnings are pre-existing).")
77//! @yah:verify("Checked the roll is safe BEFORE it happens: the only mirror in the tree declaring [providers.bundle.revalidate] is yah-marketing/cloud.toml, and it lists both /releases and /issues. The only live pokers name exactly those - issue-tracker sends Poke::route(\"/issues\") (crates/yah/issue-tracker/src/main.rs:86) and the almanac on_change arms in .yah/almanac/{releases,yah-desktop}.toml both name /releases. fleet.toml uses kind=\"reload\", which pokes almanac's own receiver, not this one. So nothing that works today starts 403ing.")
78//! @yah:gotcha("NOT DEPLOYED - code only. Enforcement starts at the next `yah cloud` sync of yah-marketing, which re-forks the receiver with the new argv. Deliberately not rolled from this session: it is an outward-facing change to a live node, and deployment belongs to R330-F13/R523. Before that roll, the live receiver still accepts a poke for any route.")
79//! @yah:next("mirror_key_env is still inert and the receiver still runs OPEN - untouched here, because the operator's call put auth on its own axis. Whatever change starts resolving it must set the matching ALMANAC_MIRROR_KEY on the issue-tracker unit on us-east-001 in the SAME change, or the sidecar's poke starts 403ing and /issues silently stops updating.")
80//! @yah:next("Public front door: R752-F9 filed for the low-security platform key that ships with the browser bundle for POST /api/issues (the operator's second decision). Different endpoint, different key namespace - do not collapse it with MESOFACT_MIRROR_KEY.")
81//! @yah:gotcha("BEHAVIOUR CHANGE worth knowing: after the roll, a whole-site poke at yah-marketing re-renders ONLY /releases and /issues, not / and /404. That is the intended reading of the declared list, but it means the landing page can no longer be refreshed by poking the receiver - it is republished by a full deploy. If someone wants / kept fresh from a feed, add it to routes in cloud.toml.")
82//!
83//! @yah:relay(R876, "Node-side mesofact: prove the hot-ship loop on real hardware, then prove the load can move off us-east-001")
84//! @yah:at(2026-09-09T04:10:09Z)
85//! @yah:assignee(agent:bundle-anthropic-ashguard)
86//! @arch:see(.yah/docs/working/W267-sovereign-public-ingress.md)
87//! @yah:next("SESSION CONTEXT (2026-09-08, chat session, ashguard/spade). This relay exists because the node-side mesofact iteration loop was measured end-to-end and found not to close, a hot-ship arm was built to close it, and the first live exercise of that arm turned up a separate availability problem worth its own drill. Both children are LIVE-FLEET work that a chat session deliberately did not run.")
88//! @yah:next("WHAT ALREADY LANDED IN THE TREE (uncommitted; camp git-policy is `defer`). (1) scripts/hotship.sh: registry gained `source` and `dest` columns + a `mesofact` entry `mesofact|mesofact||oss/mesofact|bundle-serve|artifact:mesofact|runtime:mesofact`; new `bundle-serve` activation; new `unpack_artifact` resolving the newest qed-produced tarball out of .yah/cache/artifacts/named/; runtime-asset install arm that overwrites only RUNNING versions and writes a `serve.hotship` stamp. (2) .yah/qed/hotship.toml: `binaries` param description updated. (3) oss/qed/crates/qed/src/runner.rs: execute_step_local_container now publishes/injects/discards `source_context` — that is the arm-leg fix, with two new tests.")
89//! @yah:verify("Measured, not assumed — `mesofact-musl` x86_64 leg took 9m56s / 9m57s / 11m10s across its three successful runs (qed run records 23ef24bb, 936d0be1, d62e9bdd). Its aarch64 leg failed in ~1.4s on all 13 recorded runs, so the pipeline never reached `[[pipeline.on_success]]` and its publish has NEVER fired.")
90//! @yah:gotcha("THE CACHE-HIT SHORT-CIRCUIT IS THE LOAD-BEARING FACT FOR BOTH CHILDREN. `ensure_runtime_asset` (oss/yah-base/crates/mesofact-bundle/src/runtime.rs:488) returns on a bare `dest.is_file()` — no re-hash, no manifest GET, deliberately and documented. Consequences, both real: (a) a node that has ever resolved `mesofact/<ver>` will NEVER re-fetch it, so the pre-hotship iteration loop required a new version + a bundle republish + an apply for every single change; (b) dropping bytes at that path IS a working hot ship needing no R2 write — which is what makes the arm fit hotship's never-writes-the-CDN charter — but it is also invisible afterwards unless something records it. That is why the arm writes a `serve.hotship` stamp beside each binary. R746-T3 (us-east-001 reporting `kamaji 0.8.22` while carrying none of it) is the same failure this prevents.")
91//! @yah:gotcha("RELAY-LEVEL BLOCKER FOUND 2026-09-11 AND IT GATES BOTH REMAINING CHILDREN: A PROTOCOL SKEW BETWEEN THE TREE AND THE FLEET. Three versions, measured: the TREE speaks ProtocolVersion::V11 (at 8aa2dc6a), the last RELEASE v0.8.37 speaks V9, and us-east-001 is a matched **V10** pair. The node's version is the one nothing reports directly — it was dated by `grep -c -a` on /usr/local/bin/{kamaji,yubaba} finding `ephemeral_storage_mb` PRESENT and `scratch-floor-mb` ABSENT, i.e. both binaries predate R885-T6 (the change V11 exists for) and postdate V9. `strings` is not installed on that node; use `grep -c -a`. Consequence: kamaji-bin/src/server.rs:1390 refuses any handshake whose version != CURRENT, so a tree-built yubaba or CLI does not degrade — every call fails at connect while the unit still reports active and yubaba silently falls back to its in-process containerd runtime, wiring no network namespace while service records still advertise addresses. That is what makes `--allow-proto-skew` look survivable when on the node serving noisetable.com it is a worse outage than the bugs being fixed. Note hotship.sh's own guard compares tree-vs-LAST-RELEASE and so assumes the node runs the release, which here it does not — the refusal is right, its stated reason is not the operative one.")
92//! @yah:next("THE RELAY'S REMAINING WORK IS ONE CONSOLIDATED PRODUCTION DRILL, SHARED BY R876-B11 AND R876-B12, AND IT IS DELIBERATELY ONE RESTART RATHER THAN TWO. Both children's code is landed and green in the tree; both carry `depends_on(R885)` as the machine-checkable form of the precondition. Sequence when R885's kamaji work quiesces: (1) ONE paired `scripts/hotship.sh --nodes us-east-001 --binaries yubaba,kamaji` — the pair is required because a V11 binary cannot go to that V10 node alone, and it must wait for R885 because `--binaries kamaji` builds kamaji from the working tree and would otherwise push unfinished native-exec work onto a production node whose kamaji restart kills and replays every supervised workload. (2) Execute R876-B11's RUNBOOK STEP 0/6-6/6 — the kamaji restart plus the service-record survival check IS B11's verification. (3) Execute R876-B12's runbook against the `noisetable` slot — the same restart supplies B12's \"reproduce B11, then recover with the new verb\" test. (4) Check the apex per-origin with `--resolve` (east 51.81.85.145, south 45.32.194.254, west 15.204.89.240), never through round-robin DNS. If it 503s, the recovery is `yah cloud mirror up noisetable-marketing --env cloud`, which rebuilds wasm — that is an operator call, not the drill runner's, and closing that gap is precisely what B12 exists for.")
93//! @yah:notify_on(R885, "R885's kamaji work has quiesced — R876's consolidated production drill is now unblocked. Both R876-B11 (yubaba serving-declaration fix) and R876-B12 (`yah cloud service rolling` no-build redeploy) are landed, green, and waiting on ONE paired `scripts/hotship.sh --nodes us-east-001 --binaries yubaba,kamaji` plus ONE kamaji restart. The reason it waited: the tree is ProtocolVersion V11 while us-east-001 is a matched V10 pair, so a tree-built binary cannot go alone, and `--binaries kamaji` builds kamaji FROM THE TREE — shipping it mid-R885 would have pushed unfinished native-exec work onto the node serving noisetable.com. Run R876-B11's RUNBOOK STEP 0/6-6/6, then R876-B12's runbook, then check the apex per-origin with --resolve. See R876's @yah:next for the full sequence.")
94//! @yah:gotcha("PROTOCOL-VERSION PROVENANCE — DO NOT CITE THE V10 IN THIS RELAY'S SKEW GOTCHA AS A MEASUREMENT. It is an INFERENCE: `grep -c -a` on us-east-001's binaries found `ephemeral_storage_mb` present and `scratch-floor-mb` absent, therefore predating R885-T6 and postdating the v0.8.37 release. @Glimmerstone:eclipse subsequently read the node DIRECTLY — `/health` reports `kamaji_version: 0.8.38-h5` — which is a hot-shipped build ahead of the release, and that direct read is the citable number. The two agree and the operational conclusion is unchanged (the tree is a bump ahead of the node, so the ship must be the pair), but they are not independent confirmations of each other and should not be quoted as such. SEPARATE AXIS, DO NOT CROSS THEM: `/health`'s `cluster_protocol: 7` is the raft/cluster protocol, NOT the kamaji wire ProtocolVersion this skew is about.")
95//! @yah:gotcha("Cross-relay note from @Ashguard:coffee (leading R584), 2026-09-12 — a red gate that traces to this ticket's files. `cargo test -p xtask` fails `cluster_epochs::tests::the_declaration_records_current_per_input_digests_for_every_axis` with \"recorded digest for `rust-lines oss/yubaba/crates/yubaba/src/lib.rs containing /raft/` is stale\" (xtask --lib 51 pass / 1 fail; xtask --test main 68 / 1). R876's session held ~182 uncommitted insertions / 10 writes in that file when this was observed, which is why it is recorded here. Two R584 couriers hit this gate and both deliberately declined to re-record the digest, on reasoning worth preserving: a cluster-protocol epoch declaration is NOT a derived artifact. Schemas and TS bindings are pure functions of the tree, so anyone may regenerate them (R584-F2/F3/F4 all did) — but an epoch digest is a semantic assertion that the raft protocol did not change, and only the author of those raft lines can truthfully make it. Re-recording it from outside would launder an unreviewed protocol change through a green gate. If R876's changes to the /raft/ lines are protocol-neutral, re-record and the gate clears; if they are not, the declaration itself needs updating. Either way it should not ship red.")
96//!
97//! @yah:ticket(R876-S2, "Failover drill: move the yah.dev mesofact load off us-east-001 and back, and find out what actually blocks it")
98//! @yah:status(review)
99//! @yah:at(2026-09-09T07:33:42Z)
100//! @yah:kind(spike)
101//! @yah:assignee(agent:bundle-anthropic-ashguard)
102//! @yah:parent(R876)
103//! @arch:see(.yah/docs/working/W267-sovereign-public-ingress.md)
104//! @yah:next("Tier: Cleric — the answer is a design call about the production apex's availability, not a mechanical edit. OPERATOR-REQUESTED 2026-09-08, verbatim intent: \"run a drill to move a mesofact load from us-east-001 to another and back (by tainting or something)\", framed by the standing observation \"I've been pretty clear nodes go down in this system\".")
105//! @yah:gotcha("THE MISSING RUNTIME ASSET ON THE OTHER NODES IS A SYMPTOM, NOT THE CONSTRAINT — do not start by seeding caches. The runtime asset is a lazy cache: `ensure_runtime_asset` fetches on miss from KAMAJI_BUNDLE_ORIGIN (https://cdn.yah.dev) and blake3-verifies, and mesofact/0.8.32's manifest is published and serving 200. A cold node can pull it. us-south-001 has none simply because it has never been asked to run one.")
106//! @yah:gotcha("WHAT ACTUALLY PINS IT — TWO INDEPENDENT PINS, both in .yah/services/yah-marketing/mirrors/cloud.toml, and a drill that only clears one will still fail. (1) The bundle slot declares `required = { regions = [\"us-east\"], mesh_tags = [\"tag:cloud-runner\"] }`. Measured: us-east-001, us-south-001 and us-west-001 ALL carry tag:cloud-runner + arch:x86 + os:linux + tag:voter-candidate, so `regions = [\"us-east\"]` is doing 100% of the narrowing and exactly one machine declares that region. Candidate set of one. (2) `upstream_host` is hardcoded to 100.64.0.3, so even if placement moved, the front door keeps sending traffic to the old node. `shape = \"single-machine\"`.")
107//! @yah:gotcha("READ R772's ANNOTATIONS AT THE TOP OF THAT MIRROR BEFORE EDITING IT — the obvious fix was already tried and reverted, with the reason recorded at the pin. Dropping `regions` ALONE makes things worse, not better: `plan_ingress` is deliberately pure (no network, no credentials, no CloudConfig) so it cannot resolve a constraint, and `IngressPlan::workload_machine()` goes Some(us-east-001) -> None, which is what upstream discovery is aimed at (app/yah/cli/src/cloud.rs ~4935). Today that damage is INVISIBLE because upstream_host is pinned. The recorded order is: teach the ingress planner to resolve a placement (or hand it a pre-resolved one) FIRST, then convert the slot. Preferred shape, also recorded: the CALLER resolves and passes the machine in, rather than plan_ingress growing a CloudConfig parameter.")
108//! @yah:next("SHAPE OF THE DRILL. Off: taint or otherwise make us-east-001 ineligible, and observe what the system actually does rather than what it should do — does anything attempt a re-place at all, does the bundle materialize on the new node, does the runtime asset cold-fetch from cdn.yah.dev, does the front door follow. Back: reverse it and confirm the load returns and yah.dev stays 200 throughout. The deliverable is the WRITTEN ANSWER to \"what breaks first, and in what order\" — a spike, not a fix. Expect the honest outcome to be that nothing moves, because of the two pins above; that result is worth having recorded and measured rather than inferred.")
109//! @yah:next("GUARD THE DRILL ITSELF. yah.dev is live and 200 on / and /releases; us-east-001 is also PROD raft voter 3 (100.64.0.3), so a taint that reaches raft membership is a quorum event, not just a placement one. Establish the rollback and the blast radius before tainting anything, and do not run this in the same window as R876-T1's activation — one unproven change at a time against the only node serving the apex.")
110//! @yah:next("SIBLING WORK, do not duplicate: R869 covers raft state having no off-fleet copy, and R870 covers the second-tenant/sovereign-front-door tiers — both live near this. This spike is narrower: can ONE declared-singleton workload move between nodes at all. If the answer needs a design change, file it under R870 or as its own relay rather than growing this spike.")
111//! @yah:gotcha("PIN #2 AS FILED IS STALE — `upstream_host` IS ALREADY GONE. Removed by R844-T10 on 2026-09-04; .yah/services/yah-marketing/mirrors/cloud.toml:262 records the removal, `fronted = true` (line 254) replaced `port = 8080`, and xtask/tests/mirror_ingress.rs::the_apex_derives_the_backend_host_it_no_longer_pins asserts on the real file that NO pin is declared. So the filed claim `upstream_host is hardcoded to 100.64.0.3` was five days out of date at filing. Pin #1 is still exactly right: `required = { regions = [\"us-east\"], mesh_tags = [\"tag:cloud-runner\"] }` at line 213, and only us-east-001 declares region = \"us-east\".")
112//! @yah:gotcha("THE FRONT-DOOR PIN DID NOT GO AWAY, IT MOVED OUT OF THE TREE — and that is worse for this drill, because no test, no `ingress collate` and no mirror diff can see it any more. Measured live 2026-09-09: BOTH yah.dev doors carry `PASSWAY_YUBABA_URL=http://100.64.0.3:7443` — us-east-001 /etc/passway-test.env and us-south-001 /etc/passway.env. SOUTH'S POINTS AT EAST, NOT AT ITSELF. And /service-records is strictly node-local, which I proved by query rather than reading the invariant: 100.64.0.2 (south) answers only `headscale`; 100.64.0.3 (east) answers `noisetable` + `yah-marketing`. So if yah-marketing moved off east, both doors keep polling 100.64.0.3, find no yah-marketing record, and yah.dev 503s no matter where the workload actually landed. app/yah/cli/src/mesh.rs:119 already records the south-points-at-east half; the 503-on-move consequence for yah.dev is the part to carry into the drill.")
113//! @yah:handoff("WHAT BREAKS FIRST, AND IN WHAT ORDER — the ticket's deliverable, answered from read-only measurement on 2026-09-09 with NOTHING tainted and NOTHING mutated. (1) PLACEMENT NEVER MOVES. `required.regions = [\"us-east\"]` matches exactly one machine, so tainting us-east-001 empties the candidate set rather than selecting a new node; the reconciler has nowhere to put the workload and the drill stops here. (2) IF you widen `regions`, THE FRONT DOOR DOES NOT FOLLOW. `passway_discovery_env` (app/yah/cli/src/cloud.rs:4429) renders the correct `PASSWAY_YUBABA_URL` from the placement node's mesh_ipv4 — but the apply PRINTS that env (cloud.rs:566 handoff) rather than writing /etc/passway.env on the node, so both doors keep polling 100.64.0.3 until a human edits two files on two boxes and reloads. Nothing reports an error; yah.dev just 503s. (3) ONLY THEN does the runtime-asset cold-fetch matter, and that half is genuinely fine — the filed gotcha is right that it is a lazy cache and https://cdn.yah.dev/runtimes/mesofact/0.8.32/x86_64-unknown-linux-musl.toml is live and 200. The honest answer the ticket predicted is confirmed, and the reason is one step earlier than filed.")
114//! @yah:handoff("GROUNDING FOR STEP (2) ABOVE, read from code rather than from an annotation: the ONLY production caller of `passway_discovery_env` is the `cloud::IngressProvider::Passway` arm at app/yah/cli/src/cloud.rs:7527, and it pushes the rendered lines into `out` — the vec the apply PRINTS. Nothing in that path writes /etc/passway.env or reloads a door. Independently, the `CoordinatorPin` doc at cloud.rs:4460-4475 already states the structural half in as many words (\"a **remote** door can never discover a workload placed on another node\"), and cites the same R844-B11 per-node invariant. So the design limit was known and written down; what this spike adds is the live measurement that BOTH yah.dev doors are currently the remote case with respect to any move — us-south-001's /etc/passway.env points at 100.64.0.3, not at itself — and therefore that the apex 503s on a move regardless of destination, not merely for the N-1 doors the doc anticipates.")
115//! @yah:handoff("THE DRILL IS SAFE TO RUN AND WILL NOT MOVE ANYTHING — both halves grounded in code, so the taint does not have to be spent to learn it. `select_matching` (oss/yubaba/crates/cloud/src/config.rs:2010) BAILS with \"no candidates matching {req} — {pool}: {names}\" whenever fewer machines match than `want`, and its doc says why in as many words: it \"never [returns] a one-element vec: a half-placed workload that reports success is worse than a failed apply\". So tainting us-east-001 out makes `yah cloud apply` fail at RESOLUTION, before any deploy or teardown — the running yah-marketing workload is untouched and yah.dev keeps serving. That is the good news and the disappointing news at once: nothing attempts a re-place, so the taint measures the resolver's refusal and nothing downstream of it. Note also that a taint in .yah/infra/machines/*.toml is inert until someone runs an apply; it is a tree edit, not a live fleet action.")
116//! @yah:gotcha("CORRECTION TO MY OWN EARLIER ENTRY, AND IT NARROWS THE GAP CONSIDERABLY: passway DOES know how to follow a moving backend. R844-F23 (shipped 2026-09-05, oss/passway/crates/passway/src/discovery.rs) makes `PASSWAY_YUBABA_URL` a LIST — the door polls N yubabas, unions the matching records, and holds last-known-good PER SOURCE so one node's yubaba restarting cannot drain the other's backends. It is tested end to end through the real LoadBalancer + TcpHealthCheck. So the mechanism for following a move exists and works. THE ACTUAL GAP IS ONE STEP UPSTREAM: `passway_discovery_env` renders one URL per PLACEMENT node — where the workload IS — not per CANDIDATE node, where it MAY GO. A single-machine placement therefore renders a single-yubaba door by construction, which is why all three apex doors list only 100.64.0.3. The failover-capable rail was handed a candidate set of one.")
117//! @yah:next("THE FIX THIS SPIKE ARGUES FOR, small and with an obvious home: render the door's yubaba list from the slot's `required = {...}` CANDIDATE SET rather than from the resolved placement. `CloudConfig::resolve_machines` (oss/yubaba/crates/cloud/src/config.rs:1788) already returns a Vec of matching machines, so the shape exists. Polling a node that does not hold the workload is harmless by design — discovery.rs's \"answered with none retires only its own share\" rule — and the budget fits: `base_urls.len() * PASSWAY_YUBABA_TIMEOUT_SECS` must stay under PASSWAY_UPDATE_INTERVAL_SECS, which at the 5s/30s defaults allows six nodes, and only four machines in the fleet carry tag:cloud-runner (us-east-001, us-south-001, us-west-001, us-west-003). With that change a `regions` edit, a taint, or a node dying moves the workload and every door follows within one 30s tick with no human edit. SECOND HALF, still needed: the apply must WRITE the door env rather than print it, or the re-render never reaches the node.")
118//! @yah:verify("THE SINGLE-DISCOVERY-SOURCE RISK IS NOT THEORETICAL — IT WAS MEASURED ACCIDENTALLY BY R876-T1 THIRTY MINUTES AGO. During T1's activation on us-east-001, a ~2-second gap in east's mesofact serve process 502'd the yah.dev apex on EVERY probe in that window (04:39:16 and 04:39:17, one GET/s), not one probe in three. yah.dev is round-robin across three origins — us-east-001, us-south-001, us-west-001 (west promoted to origin 3 on 2026-09-08, oss/yubaba/crates/yubaba/src/cert_store.rs:90) — so a third of requests should have survived if the origins were independent. They are not: all three doors carry PASSWAY_YUBABA_URL=http://100.64.0.3:7443, so east's serve process is a single point of failure for the whole apex regardless of how many origins front it. Three doors, one backend, one blast radius.")
119//! @yah:gotcha("GROUNDED FROM CODE, replacing the annotation-derived version of this claim: `passway_discovery_env` (app/yah/cli/src/cloud.rs:4370) opens with `let placed = plan.workload_machines()` and builds one poll URL per entry, under a comment that states the design assumption in as many words — \"One poll URL per placement node. The record store is node-local, so this list IS the set of places the workload can be seen from.\" That comment is TRUE IN THE PRESENT TENSE and is exactly why the door cannot follow a move: the set of places a workload can be seen from RIGHT NOW is not the set it could be seen from after it moves, and the door is configured with the former. The refusals in the same function (unplaced, machine absent, no mesh_ipv4, no hostname rules) confirm there is no candidate-set path — every branch resolves against `placed`. So the change R876-S2 argues for is one line of intent: feed this loop the slot's resolved CANDIDATE set instead of `plan.workload_machines()`.")
120//! @yah:handoff("DRILL RUN FOR REAL 2026-09-09 (session:d9a29d70), off and back, apex 200 throughout — and IT DID NOT MATCH THE PREDICTION. The prediction was \"taint us-east-001 and resolution refuses\". The refusal half is confirmed; the TAINT half is wrong, and that is the most valuable thing this drill produced. THE TAINT LEVER DOES NOT EXIST FOR THIS WORKLOAD CLASS. Node taints repel by ARCHETYPE: `RequiredSpec::matches` (oss/yubaba/crates/cloud/src/config.rs:4182) reads `machine.taints` only inside `for arch in &self.repel_archetypes`, and that field is `#[serde(skip)]` (config.rs:4067). A mirror's `required = { regions, mesh_tags }` is deserialized straight from TOML, so on the `resolve_bundle_machines` -> `resolve_machines` path the set is ALWAYS empty, the loop body never runs, and the taint list is never read at all. Only `admit_workload`, which builds the spec from a WorkloadSpec, populates the axis. MEASURED AGAINST THE REAL FILE, not just in a fixture: with `taints = [\"public-ip\", \"no-server\"]` written into .yah/infra/machines/us-east-001.toml, `mirror_ingress::the_apex_bundle_places_on_the_node_set_its_upstreams_are_pinned_to` still passed — us-east-001, unchanged. All three repelling keys at once (`no-server`, `no-appliance`, `no-job`) are equally inert. CONSEQUENCE WORTH CARRYING: every placement declared by a mirror's `required` is DRAIN-PROOF, while every placement that arrives through `admit_workload` is not. An operator told \"drain a node by tainting it\" would edit the file, see the lint pass (`no-server` is a legal key, so `check_inert_taints` does not fire), run an apply, and get a successful deploy onto the node they meant to evacuate.")
121//! @yah:handoff("THE ONE LEVER THAT WORKS, AND THE ACTUAL ERROR TEXT VERBATIM. With no taint lever, the only way to make us-east-001 ineligible is the membership axis — `region`. Pulled it for real (`region = \"us-east\"` -> `\"us-east-DRAINED\"` in .yah/infra/machines/us-east-001.toml) and the deploy-side resolver refused. NOTE THE TWO LAYERS, because they say different things and `{}` shows only the first: OUTER (anyhow `to_string()`) = `F16 placement: cannot place providers.bundle.required (required.regions=[us-east] + required.mesh_tags=[tag:cloud-runner]) onto 1 machine(s) - check .yah/services/yah-marketing/mirrors/cloud.toml against .yah/infra/machines/*.toml`. FULL CHAIN (`{:#}`) appends `: no candidates matching required.regions=[us-east] + required.mesh_tags=[tag:cloud-runner] - declared machines: us-east-001, us-south-001, us-west-001, us-west-002, us-west-003, us-west-011, us-west-013, us-west-014, us-west-015`. The inner layer is the one naming the pool searched, so an operator who sees only the outer line is told a placement failed but not what was considered and rejected. The ingress planner refuses too, and more thinly: `resolving placement for [providers.bundle] required = { ... }` with the same cause beneath it. AGAINST THE REAL TREE this took SEVEN of the eleven mirror_ingress tests down at once, including `the_whole_camp_collates_onto_its_nodes_without_conflict` — so the camp has a real standing guard against this drift, which is the reassuring half.")
122//! @yah:handoff("\"AND BACK\", AND THE DRILL LEFT NOTHING BEHIND. Restored `region = \"us-east\"` by editor write (never `git checkout`/`restore` — shared tree) and proved it byte-exact three independent ways: `diff` against a pre-edit `cp` at /tmp/us-east-001.toml.R876S2-orig was empty, sha256 back to 17dd15e29a0a8cf185879b5b0f49b7a54a134f5837a9ef80e01e3d065efe42d2 (identical to pre-drill), and `git status --porcelain` on the path empty, i.e. matching HEAD blob d66ab6d84d63a2c15b5eeec931e76139069a789e. Post-restore: mirror_ingress 11 passed / 0 failed, apex_failover 4 passed / 0 failed, https://yah.dev/ 200 and /releases 200 — same as the pre-drill baseline. THE PROD-SAFETY ARGUMENT, now measured rather than reasoned: no mutating `yah cloud apply` was run, and none was needed. The drill window against the real machine file was seconds, because the test binary was already compiled and was invoked directly (./target/debug/deps/main-<hash>) instead of through cargo. Even had a peer run an apply inside that window, `select_matching` (config.rs:2010) bails on a shortfall rather than half-placing, so the failure mode is a refused apply and an untouched running workload — the resolver fails CLOSED. Nothing on any node was touched: no headscale, no raft membership, no /etc/passway*.env.")
123//! @yah:handoff("THE DRILL IS NOW A STANDING TEST, not a story about an afternoon — NEW FILE xtask/tests/apex_failover.rs (4 tests, registered in xtask/tests/main.rs). It loads the REAL .yah/services/yah-marketing/mirrors/cloud.toml and the REAL .yah/infra/machines/, then makes us-east-001 ineligible in the loaded CloudConfig — which is the same experiment as a tree edit, since `CloudConfig::load` is the only thing between those files and the resolver, and it is repeatable by anyone with no window of wrong bytes on disk. The four: (1) `every_repelling_taint_at_once_leaves_the_apex_bundle_exactly_where_it_was` — pins the headline finding so that if someone ever wires `repel_archetypes` through the mirror path, this test fails and tells them the drain lever just started working; (2) `making_the_apex_node_ineligible_refuses_to_resolve_rather_than_failing_over` — asserts BOTH error layers, and asserts the chain names the pool, so the diagnostic quality itself is now guarded; (3) `restoring_the_region_puts_the_apex_bundle_back_on_the_same_node` — the \"and back\" half, proving the refusal is not sticky; (4) `the_apex_candidate_set_has_exactly_one_member_and_regions_is_why` — measures that exactly one machine declares region us-east while FOUR carry tag:cloud-runner, so it records which half of the constraint is doing the pinning and will fail the day someone widens it.")
124//! @yah:handoff("FIX FILED AS R870-F16 (child of R870, which is live/active), NOT implemented here — the spike says the design change belongs elsewhere and it does. Title: \"Render the door's yubaba poll list from the slot's CANDIDATE set, and WRITE the door env onto the node instead of printing it\". Annotation anchored in app/yah/cli/src/cloud.rs. Both halves are in the ticket body as separate @yah:next entries with the argument for why NEITHER works alone: (a) alone renders a better list nobody installs, (b) alone installs the same one-node list. It carries the measured single-point-of-failure fact (three yah.dev doors, all `PASSWAY_YUBABA_URL=http://100.64.0.3:7443`, R876-T1's ~2s outage 502ing every probe), the budget constraint VERIFIED from source rather than quoted (`env_secs(\"PASSWAY_YUBABA_TIMEOUT_SECS\", 5)` at oss/passway/crates/passway/src/main.rs:976 and `env_secs(\"PASSWAY_UPDATE_INTERVAL_SECS\", 30)` at main.rs:1139 — 5s/30s, six nodes fit, four machines carry the tag), and a SCOPE BOUNDARY gotcha stating that F16 makes the door FOLLOW a move but does not make a move POSSIBLE: the `regions` candidate-set-of-one and the missing drain lever both remain, and F16 should land BEFORE the slot is widened (the recorded R772 order).")
125//! @yah:verify("HOW EVERY CLAIM ABOVE WAS CHECKED, with exit-visible results rather than inference. BASELINE before touching anything: `curl -sS -o /dev/null -w '%{http_code}' https://yah.dev/` = 200, /releases = 200. FILE BASELINE: `git status --porcelain .yah/infra/machines/us-east-001.toml` empty (clean), HEAD blob `git rev-parse HEAD:.yah/infra/machines/us-east-001.toml` = d66ab6d84d63a2c15b5eeec931e76139069a789e, worktree sha256 = 17dd15e29a0a8cf185879b5b0f49b7a54a134f5837a9ef80e01e3d065efe42d2, copy saved to /tmp/us-east-001.toml.R876S2-orig. TAINT ARM: added `\"no-server\"` to the real file, ran the prebuilt binary directly — `mirror_ingress::the_apex_bundle_places_on_the_node_set_its_upstreams_are_pinned_to` = 1 passed / 0 failed (taint inert, confirmed). REGION ARM: `region = \"us-east-DRAINED\"` in the real file, `mirror_ingress::` = 4 passed / 7 FAILED, error text captured verbatim (recorded in the handoff above). RESTORE: editor write, then `diff` vs the /tmp copy = empty, sha256 = 17dd15e2... (unchanged), `git status --porcelain` on the path = empty. AFTER: `mirror_ingress::` 11 passed / 0 failed; `apex_failover::` 4 passed / 0 failed; yah.dev / = 200 and /releases = 200, matching the opening baseline exactly.")
126//! @yah:verify("CAVEATS ON THE ABOVE, stated rather than buried. (1) The first `cargo test` invocation of apex_failover.rs came back with a PostToolUse advisory that two build inputs (crates/yah/cloud-client/src/lib.rs, oss/yubaba/crates/yubaba/src/domain_issuer.rs) were edited by a peer mid-run. Neither is on this drill's path, and the one failure in that run was my own assertion targeting the wrong anyhow formatting (`{}` shows only the outer context; the pool-naming layer needs `{:#}`) — fixed in the test and the re-run was clean, so the advisory is noted but did not affect the result. (2) The real-file arms deliberately ran the ALREADY-COMPILED test binary (./target/debug/deps/main-<hash>) rather than `cargo test`, to keep the window of wrong bytes on the prod machine file to seconds instead of a build. That means those two arms exercised the source as of the immediately preceding compile, which is the same source the clean 11/0 and 4/0 runs used. (3) NOT DONE, and deliberately: no mutating `yah cloud apply`, so the drill measures the RESOLVER's refusal and nothing downstream of it. Whether a re-place would actually materialize a bundle on a cold node, cold-fetch the runtime asset, and come up serving is STILL UNMEASURED — it is unreachable without either widening `regions` for real or landing R870-F16 first.")
127//! @yah:next("RECOMMENDATION FOR THE LEADER — a SECOND ticket this drill argues for, deliberately not filed (the brief scoped this session to one). The missing drain lever is separable from R870-F16 and is arguably the more dangerous of the two, because it fails SILENTLY in the operator's favour: `taints = [\"no-server\"]` on a node is accepted by `check_inert_taints` (it is a legal repelling key), passes lint, and then places the workload onto the node anyway, because a mirror-declared `required` never populates `RequiredSpec::repel_archetypes` (`#[serde(skip)]`, config.rs:4067). Shape of the fix, if wanted: either populate `repel_archetypes` on the mirror path from the slot's known archetype, or give `RequiredSpec` an explicit deserialized drain axis — and either way extend xtask/tests/fleet_taints.rs, which today only checks that no node declares an INERT taint and would not have caught this. The pinning test `apex_failover::every_repelling_taint_at_once_leaves_the_apex_bundle_exactly_where_it_was` will fail the moment that lands, which is the intended signal, not a regression.")
128//! @yah:handoff("THE DESIGN FIX THIS SPIKE ARGUED FOR IS FILED, NOT GROWN INTO THE SPIKE, per the ticket's own routing instruction: R870-F16 under the live R870 relay, carrying both halves — (a) render the door's yubaba poll list from the slot's `required` CANDIDATE set instead of `plan.workload_machines()`, and (b) make the apply WRITE the door env onto the node rather than print it. The 5s/30s PASSWAY_YUBABA_TIMEOUT_SECS / PASSWAY_UPDATE_INTERVAL_SECS budget constants were re-read from oss/passway/crates/passway/src/discovery.rs rather than copied from the annotation.")
129//! @yah:verify("LEADER RE-VERIFICATION, by a second independent courier session (session:8e21f595) rather than self-report: `apex_failover` 4 passed / 0 failed, and `git status --porcelain .yah/infra/machines/us-east-001.toml` empty. Both as claimed.")
130//! @yah:handoff("DRILL RUN FOR REAL 2026-09-09 (the operator-requested empirical step a prior read-only session declined to spend), AND IT REFUTED HALF THE PREDICTION. The refusal half held; the TAINT half did not — there is no working drain lever for this workload class at all. `taints = [\"public-ip\", \"no-server\"]` written onto the real .yah/infra/machines/us-east-001.toml left the apex placing on us-east-001, unchanged, because taint repulsion keys off `RequiredSpec::repel_archetypes`, which is `#[serde(skip)]`, so a mirror-declared `required = {...}` always has it empty and `matches` never consults `machine.taints`. It fails silently — `no-server` is a legal key and the lint passes. Filed as R876-B7. The only lever that moves anything is `region`, and pulling it REFUSES at resolution rather than failing over.")
131//!
132//! @yah:ticket(R870-B11, "Bundle tier is one workload per SERVICE, so a multi-component service loses every component but the last")
133//! @yah:status(review)
134//! @yah:assignee(agent:bundle-anthropic-miravel)
135//! @yah:at(2026-09-09T08:02:46Z)
136//! @yah:parent(R870)
137//! @yah:severity(high)
138//! @yah:verify("The isolation headers survive the merge: 'curl -sI https://noisetable.com/app/' carries cross-origin-opener-policy: same-origin AND cross-origin-embedder-policy: require-corp, from the /app/* route in .yah/domains/noisetable-com.toml. Without both, SharedArrayBuffer is undefined and the wasm demo throws — the R749-F3 failure mode, one route over.")
139//! @yah:verify("Single-component regression: yah-marketing (one component, no mount) deploys byte-identically — same workload name, same digest for an unchanged tree.")
140//! @arch:see(.yah/docs/working/W267-sovereign-public-ingress.md)
141//! @yah:tier(Wizard)
142//! @yah:next("THE STANDARD, settled by the operator 2026-09-09. THREE supported configurations for composing paths on one hostname, each with an owner, and they are NOT three implementations — (2) and (3) are one passway feature at two scopes, filed as R870-F15, and (1) is assembly plus what mesofact's server already does. (1) MESOFACT SPLITS: one bundle, one serve process, components staged at their mounts inside dist/. One digest, one restart, no extra hop. Use when the components deploy together. (2) OUTER PASSWAY SPLITS: the public door path-routes to N bundle workloads. RESERVED for surfaces the door itself owns and that must answer while the upstream is down — /.well-known/*, the holding page, status. (3) INNER PASSWAY: the service runs its own non-TLS door. The DEFAULT for path-splitting an application, because a service's routes are a build-time fact about its own site and do not belong in shared public ingress. The decision rule is one question: do these components deploy together? Together -> 1. Independently -> 3. (1) and (3) compose; they are not a ladder.")
143//! @yah:next("AN EARLIER DRAFT OF THIS TICKET PROPOSED CONTRACT V2 — a manifest carrying [[components]] with per-component kind and a mount dispatcher in serve. That is NOT the plan and should not be revived for the static case: the flat content map already expresses it, and mesofact's static handler already serves it. A contract bump only becomes necessary if a single bundle must carry TWO components with DIFFERENT SERVE-TIME KINDS — e.g. a second SSR project at its own mount, which needs a second isolate. Nothing declares that today. When something does, that is a new ticket, not a widening of this one.")
144//! @yah:next("DO NOT LET CONFIG 1 AND CONFIG 3 BOTH CLAIM A MOUNT. mesofact's route table dispatches WITHIN a bundle; passway's dispatches BETWEEN bundles. A mount is owned by exactly one of them, and yah should refuse a config where a component is both staged into another component's bundle and given its own workload.")
145//! @yah:gotcha("REPORTED BY THE NOISETABLE CAMP while standing up noisetable.com, immediately after R870-B6 made a second tenant's bundle materializable at all. noisetable-marketing declares TWO static components — 'site' (mesofact-spa, web/landing, no mount) and 'app' (mesofact-static, app/browser, mount = \"/app\", the Trunk-built wasm demo). One apply reconciles both. Both assemble their OWN bundle and both deploy it under the SAME workload name — [providers.bundle].name = \"noisetable\" in the mirror — so the second reconcile replaces the first and the last component in service.toml becomes the whole of the hostname.")
146//! @yah:gotcha("MEASURED 2026-09-09 on the live apply, both orderings. With 'site' first: 'component site -> bundle \"noisetable\": 26 file(s) digest 21dabdf8', then 'component app -> bundle \"noisetable\": 7 file(s) digest c130e77e', and the door served the 7-file app bundle — https://noisetable.com/ = 404 AND https://noisetable.com/app/ = 404, because Trunk output is rooted at '/' so that bundle holds neither the landing index nor anything under /app/. With 'app' first the marketing page comes back 200 and /app/ stays 404. Two components, one upstream, no ordering that serves both.")
147//! @yah:gotcha("ROOT CAUSE, read not guessed: 'mount' has ZERO occurrences in mesofact_bundle.rs. It is honoured only by the STATIC tier, where publish_prefix(service, env, mount) extends the R2 key (mesofact_static.rs:903, :1435) — that is what kept these two components from colliding under the retired Cloudflare Worker door, which resolved a request by path out of R2. passway has no path resolver in front: it proxies a hostname to ONE upstream, and the upstream is the single mesofact-serve workload. So the separation that exists in the object store does not exist at the door, and the bundle tier never learned about mounts.")
148//! @yah:gotcha("THIS IS THE SAME DEFECT CLASS AS R870-B6 AND MesofactServeBundle::port BEFORE R844-F2 — a per-service or per-node singleton where a per-workload value belongs. B6 was the store axis (one KAMAJI_BUNDLE_ORIGIN per node), R844-F2 was the port axis (one KAMAJI_BUNDLE_PORT per node), this is the identity axis (one bundle workload per service). Each was invisible while exactly one thing existed and became a silent overwrite the moment there were two.")
149//! @yah:handoff("CONFIG1-INTERNAL GUARD LANDED AND TESTED, in oss/yubaba/crates/cloud/src/config.rs's `cross_ref_validate` (NOT cloud.rs — this file had no live-peer WIP). For every service, if two or more `mesofact-static`/`mesofact-spa` components declare the same normalized mount (via `normalize_mount`, None treated as the service root), CloudConfig::load now bails naming both component ids and the mount, before anything stages to disk. This is the config1-internal half of the ticket's overlap guard: a mount is owned by exactly one bundle-tier component. Three new tests added directly below `mount_and_route_prefix_normalization_agree` in config.rs's test module: two_bundle_components_at_the_same_mount_are_rejected, two_bundle_components_with_no_mount_are_rejected (the literal noisetable-shape footgun if `app`'s mount had been omitted instead of declared), bundle_components_at_distinct_mounts_still_load (regression guard using the existing write_two_component_service fixture, the noisetable.com shape). cargo test -p yah-cloud --lib: 1120 passed/0 failed/4 ignored before these 3 tests existed (verified by inspection — the new validation is a wholly new, early-return-free loop no pre-existing test path could have hit; git stash to get a literal pre-edit number was refused by this camp's git policy, defer mode), 1123 passed/0 failed/4 ignored after. Zero regressions.")
150//! @yah:handoff("COLLISION DISCOVERED AND RESPECTED, exactly the shape the dispatch note pre-authorized reporting rather than resolving. app/yah/cli/src/cloud.rs, oss/yah-base/crates/mesofact-bundle/src/assemble.rs, and oss/yah-base/crates/mesofact-bundle/src/lib.rs are ALL currently mid-turn-edited (party.agent_status on session:6c6bce91 read in_progress:true at the time of this session) by @Ashguard:blade under leader session:241139dd/R877 — NOT the B12-determinism work the dispatch note attributed to that session's cloud.rs dirtiness. The in-flight code is titled 'R870-B11' in its own comments and already implements the OPERATOR-MANDATED design, not the earlier contract-v2 draft: oss/yah-base/crates/mesofact-bundle/src/assemble.rs gained `pub fn collect_component_files(project_root, out_dir, mount: Option<&str>, include_config: bool)`, which stages a component's dist tree at `app/dist/<mount>/` (root when mount is None) exactly as the ticket's `next` specifies, and re-exports it from lib.rs. cloud.rs's `deploy_mesofact_bundle` and `reconcile_component` already detect multi-component bundle-tier services (`bundle_component_ids.len() > 1`, filtered on kind mesofact-static/mesofact-spa) and pick a shared `.yah/infra/state/bundles/<service>/__service` staging dir instead of one per component, and non-primary components now return `RunningWorkload::adopted(...)` as a no-op instead of deploying separately.")
151//! @yah:handoff("THAT WIRING IS INCOMPLETE AS OF THIS SESSION, not a design disagreement — grepped cloud.rs for every call site of `collect_component_files`: zero. `deploy_mesofact_bundle` picks the shared staging dir and the primary component's build info, but nothing in the current diff actually calls the new primitive to merge each component's dist/ tree into that shared staging dir before the manifest is assembled, so as committed today the merge would still produce a bundle containing only the primary component (an improvement over the current silent-overwrite bug — the second component would be a clean no-op instead of clobbering the workload — but not yet the fix: /app/ would still 404, not 200).")
152//! @yah:handoff("Tree anchor at handoff: 5f4c7b8b956196ae292ffba3c6b50aec52e81760 — the shared tree as I left it. Diff against it (`git diff 5f4c7b8b956196ae292ffba3c6b50aec52e81760..HEAD`) to see what landed under you, and quote this SHA rather than 'HEAD' in any revert/restore instruction.")
153//! @yah:verify("cargo test -p yah-cloud --lib (from oss/yubaba) — 1123 passed, 0 failed, 4 ignored, run twice (once confirming the 3 new tests individually, once the full suite) after the config.rs guard landed.")
154//! @yah:notify_on(R877-F2, "R877-F2's courier edited one character inside deploy_mesofact_bundle while it was your in-flight, uncommitted work — re-check it. The call to component_workload_dir passed `cfg.workspace_root` (a PathBuf) where the fn takes `&Path`, which red-lined the whole `yah` crate for every session in the camp; it was fixed forward with a `&` rather than reverted, because the fn is new and no last-good SHA existed. Your logic is untouched — the site is now `component_workload_dir(&cfg.workspace_root, ctx.service, primary)` (app/yah/cli/src/cloud.rs, in the is_multi_component arm). Overwrite freely if your own version differs; just don't drop the borrow.")
155//! @yah:handoff("AUTHORSHIP RESOLVED PER THE LEADER'S CORRECTION: no camp.who_wrote tool exists in this build (ToolSearch for that name and for \"who wrote\"/\"authorship\" returned nothing), so I used party.btw on the R870 leader (session:abde2cbb) instead. Its transcript recall: R870-B11 was first dispatched to @Kriek:polaris (bundle-kimi-krieg), flagged immediately as the wrong tier for a Rust refactor and superseded — that session is no longer in camp.roster/camp.sessions (confirmed absent from both just now), so nothing live owns it. @Ashguard:blade (session:6c6bce91) was independently confirmed NOT the author — its parentSessionId is the R877 leader, its activeToolCall was a musl cargo check (read-only), and party.agent_status showed in_progress:true but on that unrelated build, not a Write/Edit. So the mount-staging code sitting in the tree was Kriek's abandoned WIP: unowned, matching the operator-mandated design, safe to finish.")
156//! @yah:handoff("WIRED END TO END, in the three files the abandoned WIP had already started (I finished, did not restart): oss/yah-base/crates/mesofact-bundle/src/assemble.rs gained `assemble_bundle_from_files(dest, name, runtime_version, files, serve_bins, sidecar_bins, built_against)` — the multi-component counterpart to `assemble_self_bundle_with`/`assemble_vanilla_bundle`, picking vanilla vs self-contained the same way, but taking an already-merged `Vec<BundleFile>` instead of collecting from one `(project_root, out_dir)`. Re-exported from lib.rs. app/yah/cli/src/cloud.rs gained `assemble_multi_component_bundle`, called from `deploy_mesofact_bundle` in a new `is_multi_component` branch: for each bundle-tier component it resolves the project dir (`component_workload_dir`, already present), runs its build, and calls `yah_mesofact_bundle::collect_component_files(project, out_dir, mount, include_config = idx==0)` — the mount-aware primitive the abandoned WIP had already written — merging every component's files into ONE list before calling `assemble_bundle_from_files`. `reconcile_component`'s pre-existing guard (unchanged, verified consistent: same filter predicate and iteration order) already only calls `deploy_mesofact_bundle` for the first bundle-tier component, so `ctx.component` is provably the primary throughout.")
157//! @yah:handoff("TESTED AT EVERY LEVEL AVAILABLE WITHOUT A LIVE APPLY. New assembler-level tests in cloud.rs's `bundle_assembly_tests` module (the ticket's own required tests): `two_mounted_components_merge_into_one_bundle` — a site (no mount) + app (mount /app) component assemble into ONE manifest whose content map carries `app/dist/index.html` AND `app/dist/app/index.html`, both files verified present on disk with their distinct content. `single_component_via_multi_path_matches_single_component_path` — the multi-component assembler, given exactly one component, produces a byte-identical manifest (same digest, same content keys) to the pre-existing single-component path for the same fixture — the strongest available form of \"yah-marketing deploys byte-identically\", since the single-component code path in `deploy_mesofact_bundle` is untouched by this change (still calls `assemble_component_bundle_with_sidecars` exactly as before) and is now also proven equivalent at the primitive level.")
158//! @yah:handoff("ISOLATION HEADERS SURVIVE THE MERGE BY CONSTRUCTION, not by a change I made: `add_declared_route_headers(&mut serve_env, ...)` in `deploy_mesofact_bundle` already reads ALL of a service's declared route headers from the domain config (service-scoped, not component-scoped) and runs once regardless of is_multi_component — so the /app/* COOP/COEP rule was already flowing into the one shared `serve_env` before this change and still does now that there's only one bundle to carry it. Not independently re-verified by a new test (would need full domain-config + `mesofact::Server` integration, out of assembler-test scope) — the live `curl -sI` check from the ticket's own verify list is the real proof and is part of the unrun live gate below.")
159//! @yah:verify("cargo test -p yah-mesofact-bundle --lib (from oss/yah-base): 34 passed, 0 failed — no regressions from assemble_bundle_from_files.")
160//! @yah:verify("cargo test -p yah --lib -- cloud:: (whole cloud module, from workspace root): 165 passed, 0 failed, 1 ignored — includes both new R870-B11 assembler tests plus every pre-existing bundle_assembly_tests test (assembly_is_deterministic, sidecars_without_serve_bins_are_rejected, a_self_contained_bundle_carries_the_feed_sidecar, etc.), all still green.")
161//! @yah:verify("cargo test -p yah-cloud --lib (from oss/yubaba): 1128 passed, 0 failed, 4 ignored — includes the three R870-B11 mount-ownership-guard tests from the prior handoff. Re-run twice across both edit sessions; camp build-skew advisories fired on shared, unrelated files each time (demux_routes.rs, validate.rs, qed/publish.rs) — none overlapped my four touched files, confirmed by content, not just by filename absence from the warning.")
162//! @yah:verify("All four touched files (assemble.rs, lib.rs, cloud.rs, config.rs) are captured in the camp's automatic sync commits c6cd94fd and 718dfacb (verified by `git show <sha>:<path> | grep` for my actual function/test names, not just diffstat line counts) — nothing was lost when the working tree went clean mid-session.")
163//! @yah:verify("LEADER RE-VERIFICATION (session:abde2cbb, 2026-09-09), independent of the courier: `cargo test --manifest-path oss/yah-base/Cargo.toml -p yah-mesofact-bundle` = 34 passed / 0 failed, and `cargo test --manifest-path oss/yubaba/Cargo.toml -p yah-cloud --lib` = 1128 passed / 0 failed / 4 ignored. Both match the courier's reported counts exactly. (Note for anyone re-running these: neither crate is a member of the ROOT workspace, so a bare `cargo test -p yah-cloud` fails with \"requires dev-dependencies and is not a member of the workspace\" — the manifest-path form above is the one that works.)")
164//! @yah:gotcha("PROCESS NOTE WORTH MORE THAN THE FIX, because it nearly cost this ticket twice. The mount-staging work was started by a FIRST courier that was superseded mid-flight (a Kimi-tier carrier misrouted by the R879-B1 slot-allocator bug), which left half-wired code in the tree with no live owner. The SECOND courier found it, inferred from the dispatch's shared-tree warning that a live peer (@Ashguard:blade) owned it, and stopped — returning `ok` on a ticket whose core was unbuilt. The premise was wrong: camp.roster showed blade's parentSessionId was session:241139dd, the R877 relay leader, and its activeToolCall was a musl cross-check. ABANDONED WIP IS NOT A PEER'S IN-FLIGHT WORK, and the two are indistinguishable from `git status` alone — the discriminator is the roster's parent/ticket/activeToolCall, not the dirty-file list. Also recorded because the doctrine points at a tool that does not exist on this surface: the second courier reported `camp.who_wrote` is unavailable and had to establish authorship via party.btw instead.")
165//! @yah:gotcha("THE LIVE GATE HAS NOW BEEN RUN AND IT PASSES — 2026-09-10, from the noisetable camp, on a yah CLI built from this fix (0.8.37, `yah qed run yah-cli-install`). `yah cloud apply --env cloud --service noisetable-marketing` -> 2 components reconciled, one merged bundle. Every route on the hostname: / 200, /market 200, /account 200, /credits 200, /app 200, /app/ 200. `curl -sI https://noisetable.com/app/` carries BOTH cross-origin-opener-policy: same-origin and cross-origin-embedder-policy: require-corp. The mount's real assets serve, not just its index: /app/noise_table_browser-4fa9b4e37cb997d3.js 200 (87928 B), the same-hash _bg.wasm 200 (22822173 B, content-type application/wasm), /app/audio_worklet_processor.js 200. /.well-known/yah-publish.json 200, which also proves the site stayed the config-owning primary. This ticket's own verify list is met end to end.")
166//! @yah:gotcha("THE FIX WAS ONE LINE PLUS ITS TEST, and it is NOT the design choice the earlier gotcha left open. `collect_component_files` (oss/yah-base/crates/mesofact-bundle/src/assemble.rs) now builds `dist_prefix = format!(\"app/dist/html/{m}\")` for a mounted component instead of `format!(\"app/dist/{m}\")`. The server was left alone deliberately: `app/dist/html/` is a SERVER-SIDE CONSTANT, not a build-output coincidence — `Server::from_bundle` hands `from_workload` the bundle's `app/` dir and `from_workload` does `DistPointer::new(workload.join(\"dist\").join(\"html\"))` (oss/mesofact/crates/mesofact/src/server.rs:268). Corroborated independently by BUNDLE_BEACON_PATH in yubaba's publish_beacon.rs, which is already `app/dist/html/.well-known/yah-publish.json`. So `dist/html` is where a bundle's servable tree lives, full stop, and teaching the server a SECOND root at `dist/<mount>/` would have added a resolution rule to serve files that simply belong one directory over. The unmounted primary still stages at `app/dist` and needs no `html/` of its own: a mesofact-spa/static build already emits one inside its out_dir. Grepped for a second staging site before editing — `format!(\"app/dist` has exactly one occurrence in the tree, so there is no other path to keep in step.")
167//! @yah:verify("cargo test --manifest-path oss/yah-base/Cargo.toml -p yah-mesofact-bundle: 36 passed, 0 failed (34 before, +2 for the new served-root tests). No regressions from the prefix change — the unmounted path is untouched.")
168//! @yah:verify("cargo test -p yah --lib -- cloud:: (workspace root): 175 passed, 0 failed, 1 ignored — includes the repointed two_mounted_components_merge_into_one_bundle and single_component_via_multi_path_matches_single_component_path plus every pre-existing bundle_assembly_tests case. The single-component byte-identical test still passes, so yah-marketing's shape is unchanged by the mount fix.")
169//! @yah:verify("LIVE, 2026-09-10 from ~/ss/noisetable: `yah cloud apply --env cloud --service noisetable-marketing` then / 200, /market 200, /account 200, /credits 200, /app 200, /app/ 200; `curl -sI /app/` carries both isolation headers; /app/*.js, /app/*_bg.wasm (application/wasm, 22.8 MB) and /app/audio_worklet_processor.js all 200. The stopgap ordering question is settled separately — see the note below.")
170//! @yah:gotcha("WHY THE UNIT TEST WENT GREEN ON A BROKEN PATH, AND WHAT THE FIXTURE NOW DOES INSTEAD — the durable half of this ticket, because it is the reason a wrong path reached a live deploy at all. `two_mounted_components_merge_into_one_bundle` used to assert the content map carried `app/dist/index.html` AND `app/dist/app/index.html`. Both keys were exactly what the code produced and NEITHER was what mesofact serves: the shared `component()` fixture wrote `dist/index.html`, a convenience shape no real mesofact build emits (a mesofact-spa build emits `dist/html/index.html`). The fixture encoded the same wrong layout the implementation had, so both halves could stage outside the served root and still agree — the test's oracle WAS the implementation's own path construction, which proves only self-consistency. FIXED, not just noted: a new `spa_component()` fixture in app/yah/cli/src/cloud.rs writes `dist/html/index.html` (kind = mesofact-spa, the served layout) and the merge test uses it instead of `component()` — deliberately a NEW fixture, since `component()` is shared by many unrelated tests and its dist-root shape is fine for them. `mounted_component()` keeps `dist/index.html`, which IS correct: that is Trunk's real output. The assertions now derive from the server: a `SERVED_ROOT = \"app/dist/html/\"` constant, a loop asserting every `app/dist/**` key starts with it, then the two specific keys. Two more tests in assemble.rs cover the primitive directly — `a_mounted_component_stages_inside_the_served_root` and `an_unmounted_component_stages_its_own_html_tree_at_the_dist_root`, both asserting against hand-written literals rather than a recomputed prefix.")
171//! @yah:verify("Ordering: both routes serve 200 with 'site' declared first. NOT verified in the reverse order, and deliberately so — see the gotcha on why order remains load-bearing for config staging. Serving is order-independent; the staged config is not.")
172//! @yah:gotcha("THE ORDERING NOTE IN noisetable's service.toml DOES NOT GET DELETED, and this ticket's original verify item asking for that is WRONG — repointed rather than quietly dropped. The premise was 'order stops mattering once the overwrite race is gone'. The overwrite race IS gone, but declaration order stayed load-bearing for an unrelated reason the merge introduced: `assemble_multi_component_bundle` calls `collect_component_files(..., include_config = idx == 0)`, so `mesofact.routes.ts`, `mesofact.config.toml` and the built manifest's `data_inputs` are staged from the FIRST bundle-tier component only, and `deploy_mesofact_bundle` likewise takes the workload envelope's build provenance from `bundle_components[0]`. On noisetable.com only web/landing owns those files — app/browser is a Trunk project with none — so 'app' first stages NO config and silently drops the site's. Proven on a real apply, not argued: with 'site' first, `app/mesofact.routes.ts` appears in the staged manifest and /.well-known/yah-publish.json serves 200; the old order dropped it. So the note was REWRITTEN to name the new reason (and the `app/dist/html/<mount>/` staging path), not removed. If yah wants order to genuinely stop mattering, that is a separate ticket: pick the config-owning primary declaratively — the component that HAS a `routes` field, or an explicit `primary = true` — instead of by index.")
173//! @yah:next("SCOPE AS SHIPPED — CONFIG 1, ASSEMBLY-ONLY. No contract bump, no serving change: `BundleManifest.content` is a flat path map and contract v1 clause 2 only requires built assets under `app/dist/**`, so `app/dist/html/<mount>/**` is expressible on contract 1 today. mesofact::Server already serves any file under its dist root by path, with a clean-URL .html fallback, and already applies the per-route header table as its outermost layer. The fix was the assembler staging each mounted component's build output INSIDE the served root at `app/dist/html/<normalized mount>/` — an earlier draft of this line said `app/dist/<mount>/`, which is the bug that reached production. The serving half needed nothing.")
174//!
175//! @yah:ticket(R870-B12, "A mesofact-spa component rebuilds to a different bundle digest every apply, defeating W272 blob dedupe")
176//! @yah:status(review)
177//! @yah:at(2026-09-09T07:53:11Z)
178//! @yah:assignee(agent:bundle-anthropic-ashguard)
179//! @yah:parent(R870)
180//! @yah:severity(medium)
181//! @yah:gotcha("MEASURED, not inferred (R870-T9, 2026-09-08). Three consecutive `yah cloud apply --service noisetable-marketing --env cloud` runs from an UNCHANGED bundle source tree produced three different bundle stamps for the `site` component, read straight off the live door at https://noisetable.com/.well-known/yah-publish.json: 44c3268ca235 -> eb800ef6a775 -> 78d4396f8d51, all reporting `files: 25`. Same file COUNT, different content digest, so some entry's bytes change on every build. The only edits between runs were to the mirror TOML (.yah/services/noisetable-marketing/mirrors/cloud.toml), which is not bundle content.")
182//! @yah:gotcha("DO NOT GO LOOKING IN THE ASSEMBLER — it is not the assembler. `assembly_is_deterministic` (app/yah/cli/src/cloud.rs) passes and is genuinely testing what it claims: identical inputs assemble to an identical digest, and R703-T7 deliberately made the publish beacon clock-free to keep that true. The inputs are what differ. `site` is `kind = \"mesofact-spa\"` (web/landing), and a mesofact-spa build emits a per-build id that lands in the prerendered HTML and the hydrate paths (`/{build_id}/hydrate/...`, oss/mesofact/crates/mesofact/src/server.rs), so the built tree is different bytes each time. That last sentence is INFERENCE from the shape of the evidence plus the hydrate-path convention — the exact generation site was not located, and locating it is step one.")
183//! @yah:next("STEP 1, LOCALIZE: find where a mesofact-spa build derives its build id and confirm it is the (only) source of per-build drift. Cheapest proof: assemble the same component twice with `yah cloud bundle build`, diff the two staging trees entry-by-entry, and name the files whose hashes moved. If it is only the build-id-bearing files, the fix is to derive the id from content (a hash over the build inputs) instead of from a clock/counter — which is the same move R703-T7 already made for the publish beacon and for the same reason.")
184//! @yah:next("WHY IT MATTERS, so nobody files this as cosmetic. W272 §1 immutability is what makes a re-publish a no-op: matching blobs dedupe, the node skips materializing, and an unchanged site does not restart its serve process. A digest that moves on every apply defeats all three — every apply re-uploads, re-materializes and re-forks, which costs a real ~2s of 502 on whatever that serve process fronts (measured on us-east-001 2026-09-09, R876-T1, recorded in scripts/hotship.sh's header). So an apply that changes nothing still takes the site down for two seconds.")
185//! @yah:verify("Two `yah cloud bundle build` runs of .yah/services/noisetable-marketing component `site`, from an unchanged tree, produce the same manifest digest — and the door's beacon at https://noisetable.com/.well-known/yah-publish.json stops moving across repeated no-op applies.")
186//! @yah:handoff("ROOT CAUSE CONFIRMED, and the ticket's INFERENCE was right: mesofact_build::pipeline::default_build_id() (oss/mesofact/crates/mesofact-build/src/pipeline.rs:59, pre-change) was a SystemTime::now() UTC stamp at one-second resolution. LOCALIZED BY MEASUREMENT, not by reading: built the `spa` fixture twice with the PRE-CHANGE binary and diffed entry-by-entry. Exactly three files drifted and no others -- dist/html/app.html (the hydration weave bakes `/{build_id}/hydrate/app.CZuUyrnp.js`), dist/manifest.json (`build_id` field), dist/tag-index.json (`build_id` field). Hydrate bundle names are content-hashed by rolldown and did NOT move, which is why the live door reported the same `files: 25` on all three drifting stamps. Second half of step 1 (is the build id the ONLY drift source): built four richer fixtures -- head-sitemap, static-assets, spa-parametric, static-islands -- twice each with an EXPLICIT fixed --build-id; every pair was byte-identical, so with the clock pinned nothing else in the pipeline is non-deterministic.")
187//! @yah:verify("GATE MET on the real component. Two `yah cloud bundle build web/landing --out <tmp> --run-build` runs from ~/ss/noisetable (component `site` of .yah/services/noisetable-marketing), each doing a full `bun run build:cloud`, both printed `files: 34, digest: 896b4c39fcba7492a28c337df2565ace94de7c118bfe07ead2ff59c9934cdb6a`. Also ran the raw binary twice against /Users/leif/ss/noisetable/web/landing into two /tmp out-dirs: same 23-file tree byte-for-byte, same build_id 441813d4564564aae8524b8500e3c57a. TESTS: cargo test -p mesofact-build -> lib 111 passed / 0 failed / 3 ignored; tests/pipeline.rs 12 passed (9 pre-existing + 3 new); check_cli 5, conformance 4, render 7; all green, exit 0. Baseline caveat stated plainly: the pre-change baseline I measured was BEHAVIOURAL (the three-file drift above, with the pre-change binary); I did not run the cargo suite before editing, so the no-regression claim rests on all 9 pre-existing pipeline tests plus all 111 lib tests passing after.")
188//! @yah:handoff("FIX LANDED in oss/mesofact/crates/mesofact-build/src/pipeline.rs (+ tests/pipeline.rs). default_build_id() and its civil_from_days() date helper are DELETED -- no clock left in the build. The id is now derived from the built tree, following R703-T7's shape (compute after the artifacts exist; the self-referential part is excluded because a digest covering itself has no fixed point). Mechanism, three new items in pipeline.rs: (1) BUILD_ID_PLACEHOLDER = \"__mesofact_build_id__\" is woven by prerender in place of the id, breaking the cycle where the id names a tree that cannot be finished without it; (2) scan_staged_tree() walks out_dir after prerender, hashing every file to BLAKE3 (yah_mesofact_bundle::BundleHash, the same hash the W272 bundle uses) and noting which files carry the placeholder -- it skips .mesofact-build/ (scratch, deleted before return) and top-level manifest.json / tag-index.json / sitemap.xml, which are written later and whose stale copies from a previous build must not feed the id; (3) derive_build_id() hashes that path->hash map plus the serialized manifest and keeps 32 hex digits (128 bits -- collision-proof, short enough to read in the `/{build_id}/hydrate/...` and `<build_id>/html/...` paths where it is actually seen), then substitute_build_id() rewrites the placeholder in place. An explicit BuildOptions.build_id still wins and skips derivation entirely, so every existing caller and test is unaffected. Hashing the OUTPUT rather than the input sources is deliberate and stronger than the ticket's suggested \"hash over the build inputs\": noisetable's `build:cloud` differs from `build` only by the NOISETABLE_API_ORIGIN env var, and a data_inputs change moves prerendered HTML without moving any source file -- an input hash would miss both, an output hash cannot. Three tests pin it in tests/pipeline.rs: two_builds_of_an_unchanged_tree_are_byte_identical (fails on exactly the three files named above without the fix, and its assert names the drifting paths), a_different_project_derives_a_different_build_id (guards the opposite failure -- a stable-but-constant id would pass the first test while serving stale bytes from an immutable prefix forever), and no_placeholder_survives_into_the_built_tree.")
189//! @yah:gotcha("THE LIVE HALF OF THE VERIFY IS UNRUN. Repeated no-op applies against noisetable.com are an outward-facing deploy to live infra, so this session stopped short of them. Exact command, from ~/ss/noisetable: run `yah cloud apply --service noisetable-marketing --env cloud` two or three times with NO edits in between, reading `curl -s https://noisetable.com/.well-known/yah-publish.json` after each -- the `digest` must be identical across all runs (it moved 44c3268ca235 -> eb800ef6a775 -> 78d4396f8d51 on 2026-09-08, which is what filed this ticket). Note the bundle the apply assembles is NOT digest-comparable to the local `yah cloud bundle build` runs recorded in verify: the reconciler names it per the service component and stamps the publish beacon, so its digest differs by construction. The invariant to check is that it stops MOVING, not that it matches any local value. Note also that the site component's `dist/` is now content-addressed while the `app` component (mesofact-static, Trunk) was never checked for its own per-build drift -- if the beacon still moves after this, `app` is the next place to diff.")
190//! @yah:cleanup("The TS pipeline still has the identical clock bug: defaultBuildId() at oss/mesofact/packages/mesofact-build/src/index.ts:422 is `new Date().toISOString()`, consumed at index.ts:135. Deliberately NOT fixed here. Nothing in this camp builds through it -- app/yah/web/{marketing,dashboard,analytics} and noisetable's web/landing all shell to the Rust binary (scripts/mesofact-build.sh / `cargo run -p mesofact-build`), and tests/pipeline.rs:160 already records that the Rust pipeline is the sole production build path. Porting the derivation to TS means reimplementing the placeholder weave in prerender.ts plus a tree walk and BLAKE3 in TS, with no camp build that would catch a divergence. Worth doing if the TS pipeline ever regains a production consumer.")
191//! @yah:verify("LEADER RE-VERIFICATION (session:abde2cbb, 2026-09-09), run independently of the courier. Read the change by content first: `oss/mesofact/crates/mesofact-build/src/pipeline.rs` deletes `default_build_id()`'s `std::time::SystemTime::now()` stamp outright (not shimmed beside it), weaves a `BUILD_ID_PLACEHOLDER` through prerender, then derives the real id in `derive_build_id(&StagedTree, manifest_json)` and substitutes it via `substitute_build_id` — R703-T7's content-hash shape, applied to the same defect one layer down. An explicit `opts.build_id` still wins, so the caller-supplied path is unchanged. Then ran the suite myself: `cargo test -p mesofact-build` is fully green — 12 passed / 0 failed in tests/pipeline.rs (including the two that pin this ticket, `two_builds_of_an_unchanged_tree_are_byte_identical` and `a_different_project_derives_a_different_build_id`), 7 passed / 0 failed in tests/render.rs, 0 failures anywhere. The determinism is pinned by test rather than by a one-off manual run, which is what this ticket needed — its failure mode was \"works today, drifts again in a month\".")
192//!
193//! @yah:ticket(R870-B13, "A borrowing camp cannot render sovereign apex A records: domain phase demands .yah/infra/machines/ the camp does not have")
194//! @yah:status(review)
195//! @yah:at(2026-09-09T08:07:39Z)
196//! @yah:assignee(agent:bundle-anthropic-glimmerstone)
197//! @yah:parent(R870)
198//! @yah:severity(medium)
199//! @yah:gotcha("SURFACED 2026-09-08 (R870-T9) AND IT WAS PREVIOUSLY MASKED — read that before assuming it is a regression. `yah cloud apply --service noisetable-marketing --env cloud` from ~/ss/noisetable now reaches its domain phase for the first time (the marketing service used to hard-fail at the serving check before domains ran) and reports: `domain api-noisetable-com (api.noisetable.com): rendering sovereign apex A records from the ingress collation (DNS-only)` -> `FAILED: front door collated onto machine \"us-east-001\", which has no .yah/infra/machines/*.toml — cannot resolve its public address`. The SERVICE half is green in the same run (`noisetable-marketing ok 2 component(s) reconciled`), so noisetable.com is unaffected and serving; what fails is the api.noisetable.com DNS render.")
200//! @yah:gotcha("THE CAUSE IS ALREADY WRITTEN DOWN, in ~/ss/noisetable/.yah/services/noisetable-marketing/mirrors/cloud.toml's placement note: this camp BORROWS yah's fleet without importing its inventory — `.yah/infra/machines/` under ~/ss/noisetable is an EMPTY DIRECTORY, while the machine files live in ~/ss/yah and are read-only from there. That is the same constraint that forces `[providers.bundle]` to pin `machines = [\"us-east-001\"]` instead of declaring `required = { regions, mesh_tags }`, and the same reason noisetable-api's `[providers.compute]` writes `kind = \"static\"` + `machine = \"us-west-001\"` rather than `use = \"hetzner\"`. A pin is expressible because it is just a name; resolving that name to a PUBLIC ADDRESS is not, and the apex A-record render needs the address.")
201//! @yah:next("THE SHAPE OF THE FIX IS AN OPERATOR CALL, not a code call, which is why this is filed rather than fixed. Two expressible answers and they are not equivalent: (a) the borrowing camp declares its own `.yah/infra/machines/` entries — cheap, immediate, and a second copy of the fleet inventory that will drift from ~/ss/yah's silently; (b) yah exposes its inventory to borrowing camps over some read path, so there is one copy — the right shape, more work, and it decides how a camp names another camp's fleet. The mirror's own placement note already anticipates exactly this fork (\"TO CONVERT TO CONSTRAINTS LATER, one of two things has to happen first\"), so whichever is chosen also unblocks constraint-based placement in that camp, not just this DNS render.")
202//! @yah:next("OPERATOR ANSWERED 2026-09-09 (asked by the R870 relay leader, session:abde2cbb): option (b) — yah exposes its inventory to borrowing camps over a read path, so there is ONE copy. Option (a) (the borrowing camp declaring its own .yah/infra/machines/ entries) is REJECTED: a second copy of the fleet inventory drifts from ~/ss/yah silently and nothing detects the drift until a render is already wrong. So this ticket is no longer blocked on a decision — it is a design-plus-implementation task. Its scope now includes deciding how a camp NAMES another camp's fleet, because option (b) cannot be built without that, and per the mirror's own placement note the same answer also unblocks constraint-based placement (`required = { regions, mesh_tags }`) in the borrowing camp rather than only this DNS render.")
203//! @yah:handoff("DESIGN, AND WHAT IT REPLACES. The defect was not in the apex renderer — it was that a camp's machine inventory had TWO readers that disagreed about what the inventory is. CloudConfig::load applied the .yah/infra/sources.toml overlay inline (R615-F2, landed long ago); validate::load_machine_tomls did not, and the two callers that resolve a machine NAME to a machine — collate_workspace_ingress and reconciler::domain::plan_passway_apex — both read the latter. So in a borrowing camp (empty local machines/, one [[source]] link) the apex render failed on a machine that was declared all along, one directory over. The fix is one function: config::resolve_fleet_inventory(workspace_root) -> FleetInventory, extracted OUT of CloudConfig::load, which is now a caller of it rather than a second implementation. Three callers, one answer.")
204//! @yah:handoff("THE NAMING DECISION (the design call this ticket assigned, made and justified rather than escalated): a camp names another camp's fleet through the [[source]] entry R615-F1 already defines — `owner` is the logical name an operator sees, `kind = \"path\"` resolves against the borrowing camp's own root, `kind = \"git\"` against `yah infra sync`'s cache. NO second naming scheme was invented, deliberately. A camp that could name a foreign fleet two ways is a camp whose inventory can drift from itself, which is precisely what option (a) was rejected for. It also honours the ticket's camp-boundary constraint by construction: kind=path resolves to <path>/.yah/infra, i.e. inside the same .yah/ whose camp.toml defines the boundary — noisetable's existing sources.toml already spells it and needed no change.")
205//! @yah:handoff("ONE COPY, argued rather than asserted. kind=path reads the owner's live tree at <path>/.yah/infra/ on EVERY load — the borrowing camp persists nothing, so the two cannot disagree. kind=git reads a synced checkout, which IS a copy, but an explicit one with a named refresh verb (yah infra sync) and a pinned ref; that is the cache-with-an-invalidation-story the brief allows, as against a hand-maintained second inventory. noisetable uses kind=path, so for the camp in the defect there is literally one copy of the fleet, in ~/ss/yah.")
206//! @yah:handoff("BREAK-DON'T-TAPE, no fallback added. validate::MachineLoadMode is DELETED (its Strict variant existed only for the two resolution callers, which now read the inventory). load_machine_tomls is RENAMED to load_camp_local_machine_tomls, unconditionally tolerant, and its doc now states it is the LINT loader and not the fleet inventory — camp-local is its whole contract, because a lint exists to name a file the operator can edit and a borrowed machine lives in a tree they cannot. There is no read-local-then-fall-back-to-borrowed path anywhere: resolve_fleet_inventory is always camp-local-then-overlay, with camp-local winning any name collision and earlier sources beating later ones (R615-F2's rules, unchanged, now in one place).")
207//! @yah:handoff("FILES (all in ~/ss/yah — nothing in ~/ss/noisetable was touched). oss/yubaba/crates/cloud/src/config.rs: new pub FleetInventory { machines, origins, sources, contributions } + pub SourceContribution + pub resolve_fleet_inventory(); overlay_infra_sources() split into overlay_source_machines() (returns per-source contributions) and overlay_source_providers(); CloudConfig::load now calls resolve_fleet_inventory for its machine half (the legacy .yah/cloud/machines/ merge moved in with it, precedence preserved) and overlay_source_providers for providers. oss/yubaba/crates/cloud/src/validate.rs: MachineLoadMode deleted, loader renamed, collate_workspace_ingress reads the inventory. oss/yubaba/crates/cloud/src/reconciler/domain.rs: plan_passway_apex reads the inventory; public_origins' error text no longer says \"has no .yah/infra/machines/*.toml\" (it was wrong even in spirit) but names both surfaces.")
208//! @yah:handoff("THE DIAGNOSTIC SEAM, added because the failure class is indistinguishable from the name alone. \"no such machine\" reads identically whether a camp declared no link, aimed one at a directory that is not a camp, or filtered the machine out with `select`. So SourceContribution records per-source { owner, source, root, root_exists, machines-added } and FleetInventory::describe_sources() renders it; plan_passway_apex attaches it to the error ONLY on failure (map_err, and only when non-empty, so a camp with no sources gets no dangling header). A source contributing zero machines is deliberately NOT an error at load time — an unsynced kind=git source is legitimately empty and CloudConfig::load must stay offline (R615-F2) — so the fact is carried to whoever actually fails for want of a machine.")
209//! @yah:handoff("CONSTRAINT-BASED PLACEMENT IS UNBLOCKED IN THE BORROWING CAMP — the mirror's own \"TO CONVERT TO CONSTRAINTS LATER\" fork is answered, and this reaches further than the DNS render. The quoted failure in that placement note (`ingress declaration does not plan — ... no candidates matching required.regions=[us-east] + required.mesh_tags=[tag:cloud-runner] — declared machines: (no machines declared under .yah/infra/machines/)`) is IngressProblem::Declaration raised from collate_workspace_ingress, which is exactly the caller fixed here: resolve_ingress_placements now gets the borrowed fleet as its candidate set. Pinned by a_borrowing_camp_can_place_by_constraint_rather_than_by_pin (validate.rs), which plans a mirror carrying only `required = { regions, mesh_tags }` and no pin. The apply path was never affected — it resolves from cfg.machines off CloudConfig::load, which already had the overlay.")
210//! @yah:handoff("CONVERTING NOISETABLE'S PINS IS A SEPARATE TICKET AND NOT MINE TO FILE OR MAKE. ~/ss/noisetable is a different camp; the edits would be to .yah/services/noisetable-marketing/mirrors/cloud.toml ([providers.bundle] machines = [\"us-east-001\"] -> required = {...}) and .yah/services/noisetable-api/mirrors/cloud.toml ([providers.compute] kind=\"static\" + machine=\"us-west-001\" -> use=\"hetzner\" + required). Two reasons to keep them separate rather than fold them in: (1) that camp's own mirror documents why us-east-001 is NOT an arbitrary pick — the three live doors poll http://100.64.0.3:7443 because service records are node-local, so a constraint that resolved elsewhere would need PASSWAY_YUBABA_URL repointed in /etc/passway-noisetable.env on all three nodes; converting is a fleet change, not a config tidy. (2) The inline `required = {...}` form is mandatory there — the [providers.bundle.required] header form silently reparents fronted/zone/origin and drops the service out of the ingress plan (R844/R772). Recommend the noisetable camp file it against its own board.")
211//! @yah:verify("BASELINE MEASURED FIRST, before any edit: cargo test -p yah-cloud --lib (from oss/yubaba) = 1123 passed / 0 failed / 4 ignored. AFTER: 1128 passed / 0 failed / 4 ignored — +5, exactly the five tests added, zero regressions. Full crate incl. integration targets: 1128 + 3 (2 ignored, live/network) + 2, all green. cargo check --workspace --all-targets in oss/yubaba: 0 errors. cargo build -p yah and cargo check -p yah -p xtask --all-targets in the root workspace: 0 errors.")
212//! @yah:verify("REAL-TREE REGRESSION, the suite that plans yah's actual .yah/ through the changed collation: cargo test -p xtask --test main -- mirror_ingress apex = 15 passed / 0 failed, including the_yah_dev_apex_plans_one_front_door_per_declared_origin, the_apex_collates_onto_both_live_origins and all four apex_failover cases. yah's own camp (camp-local machines, no sources.toml) is byte-unchanged in behaviour: `yah cloud validate -p .` ok, `yah cloud ingress collate -p .` still renders us-east-001 + us-south-001 fronting yah.dev.")
213//! @yah:verify("THE FIVE NEW TESTS, and why none is vacuous. validate.rs: a_borrowing_camp_collates_a_front_door_on_a_machine_it_declares_nowhere (asserts the camp-local loader returns EMPTY in the same test that the inventory returns the machine — that control is the non-vacuity proof); a_borrowing_camp_can_place_by_constraint_rather_than_by_pin; a_broken_link_is_reported_as_absent_rather_than_as_an_empty_fleet; the_fleet_inventory_fails_the_whole_load_on_one_unparseable_camp_local_toml (the Strict property, re-homed off the deleted MachineLoadMode); the_lint_loader_skips_... (renamed). domain.rs: a_borrowing_camp_renders_its_apex_from_the_owners_machine_declaration (the ticket's defect end to end, borrower + owner camps on disk, through plan_passway_apex) and an_unresolvable_front_door_names_the_links_that_were_consulted.")
214//! @yah:verify("LIVE PROOF AGAINST THE REAL BORROWING CAMP, read-only, with the rebuilt binary (target/debug/yah, NOT ~/.local/bin/yah): `yah cloud ingress collate -p .` from ~/ss/noisetable now renders `us-east-001 — passway ... api.noisetable.com → 100.64.0.3:4332` and the same for us-south-001, and `yah cloud validate -p .` reports `ok ... 2 node front door(s) collate cleanly`. THE 100.64.0.3 IS THE PROOF, not decoration: the api mirror pins no upstream_host, so that address is resolved by machine_mesh_addrs from us-east-001's [registration].mesh_ipv4 — and 100.64.0.3 appears NOWHERE in ~/ss/noisetable/.yah except inside one prose comment. It can only have come from ~/ss/yah/.yah/infra/machines/us-east-001.toml through the [[source]] link. Both front-door nodes carry taints=[\"public-ip\"] and a public [connect].address (51.81.85.145 / 45.32.194.254), so public_origins resolves both.")
215//! @yah:verify("SCHEMA/ARTIFACT GATES: ./scripts/check-schema-drift.sh exits 0, and git status on .yah/schema/ + packages/yah/workload-spec/ is clean — the new types (FleetInventory, SourceContribution) carry no schemars derive and feed no generated artifact, as expected.")
216//! @yah:gotcha("I DID NOT RUN THE MUTATING CROSS-CAMP APPLY, deliberately, and this is the one item left for the operator. `yah cloud apply` has NO --dry-run (checked the clap definition in app/yah/cli/src/cloud.rs: Apply takes env/path/service/continue-on-error/format/config-root/namespace and nothing else), so running it would write live A records at api.noisetable.com AND reconcile two components of a different camp's production service. That is outward-facing and irreversible, so I proved the resolution path read-only instead (see the verify entries) rather than performing it. EXACT COMMAND when the operator wants it: cd ~/ss/noisetable && /Users/leif/ss/yah/target/debug/yah cloud apply --service noisetable-marketing --env cloud . Expect the domain phase to reach `domain api-noisetable-com (api.noisetable.com): rendering sovereign apex A records from the ingress collation (DNS-only)` WITHOUT the `no .yah/infra/machines/*.toml` failure, and `noisetable-marketing ok 2 component(s) reconciled` in the same run.")
217//! @yah:gotcha("THE FIX IS NOT INSTALLED. It is in the source and in target/debug/yah only. ~/.local/bin/yah (the operator's shell, every QED step's argv=[\"yah\"]) and /Applications/yah.app/Contents/MacOS/yah (every agent's MCP surface) both still serve the old binary, so the noisetable apply will keep failing the same way until `cargo xtask install` and, for agent sessions, `cargo xtask install --dest /Applications/yah.app/Contents/MacOS/yah`. Not done here: installing over the live app is an operator action per app/yah/cli/CLAUDE.md.")
218//! @yah:gotcha("ADJACENT GAP, FOUND NOT FIXED, and named so it is not re-derived: crates/yah/agent-tools/src/cloud_tools.rs has its OWN camp-local walk of .yah/infra/machines/ for the cloud.machines / cloud.mirror_state agent tools, so those tools still report an empty fleet in a borrowing camp. It is a DISPLAY surface, not a resolution one — nothing renders wrong from it, it just shows less than the camp has — and it is a different crate outside this ticket's blast radius, so I did not widen into it. The one-line fix once someone owns that crate: swap the walk for cloud::config::resolve_fleet_inventory and badge borrowed rows off FleetInventory::origins, the same provenance the Infra tab (R615-F4) already uses.")
219//! @yah:handoff("Option (b) built: `config::resolve_fleet_inventory` is now the single reader of a camp's machine inventory (camp-local + legacy + every fleet borrowed through `.yah/infra/sources.toml`), extracted out of `CloudConfig::load` and adopted by the two resolution callers that were stuck on a camp-local-only loader — `validate::collate_workspace_ingress` and `reconciler::domain::plan_passway_apex`. `MachineLoadMode` deleted, `load_machine_tomls` renamed to `load_camp_local_machine_tomls` and doc'd as lint-only; no fallback shim anywhere. Naming reuses R615-F1's `[[source]]` owner/kind rather than inventing a second scheme, so there stays exactly one copy of the fleet. Full design rationale, file-level account and the noisetable-side follow-on are in the appended handoff entries above.")
220//! @yah:verify("Baseline measured first: `cargo test -p yah-cloud --lib` 1123/0/4 -> after 1128/0/4 (+5, the five tests added, zero regressions). `cargo test -p xtask --test main -- mirror_ingress apex` 15/0 against the real tree. Root workspace and oss/yubaba `--all-targets` clean. Live read-only proof from ~/ss/noisetable with the rebuilt binary; the mutating apply was NOT run (no `--dry-run`, writes live DNS on another camp) — exact command recorded in a gotcha.")
221//! @yah:verify("LEADER RE-VERIFICATION (session:abde2cbb, 2026-09-09), independent of the courier. Tests: `cargo test --manifest-path oss/yubaba/Cargo.toml -p yah-cloud --lib` = 1128 passed / 0 failed / 4 ignored, matching the courier's count against the 1123 baseline it measured first (+5). Then checked the two structural properties the ticket actually turned on, rather than trusting the count. (1) ONE READER: `resolve_fleet_inventory` exists once, at config.rs:3283, and `load_camp_local_machine_tomls` at validate.rs:345 is now explicitly the lint loader. (2) NO SHIM: a repo-wide grep for `MachineLoadMode` across oss/yubaba/crates/cloud/src/ and app/yah/ returns hits ONLY inside annotation prose describing its deletion — zero live references. That matters more than the tests here, because the rejected option (a) would have reappeared as a read-local-then-fall-back-to-borrowed path, which is exactly the shim the root CLAUDE.md forbids and exactly the drift the operator rejected. The courier also declined to run the mutating cross-camp apply (it writes live DNS in another camp and has no --dry-run) and recorded the exact command instead — correct call, and the read-only proof from ~/ss/noisetable was run.")
222
223use std::collections::BTreeMap;
224use std::path::PathBuf;
225
226use anyhow::{bail, Context, Result};
227use async_trait::async_trait;
228
229use workload_spec::{
230 BlakeHash, BundleLifecycle, MesofactRevalidateReceiver, MesofactServeBundle, Millis,
231};
232
233use super::{ReconcileCtx, Reconciler, RunningWorkload};
234use crate::config::{CloudConfig, RequiredSpec, BUNDLE_SERVING_MESH_TAG};
235use crate::MirrorConfig;
236
237/// Mirror provider role that opts a mesofact component into the bundle tier.
238pub const SLOT_ROLE: &str = "bundle";
239
240/// Sub-key under `[providers.bundle]` that declares the revalidate receiver
241/// (R330-F12 almanac push endpoint).
242pub const REVALIDATE_KEY: &str = "revalidate";
243
244/// Default idle TTL for an `on-demand` (JIT) bundle when the slot doesn't name
245/// one: five minutes with zero connections before kamaji reaps the process.
246pub const DEFAULT_IDLE_TTL_MS: u64 = 300_000;
247
248/// Every key `[providers.bundle]` is allowed to carry (R556-B14).
249///
250/// The mirror schema's `MirrorProviderSlot` is `additionalProperties: true` by
251/// construction — it is one flattened `BTreeMap<String, toml::Value>` shared by
252/// every provider role, so it cannot know what any single role reads. That
253/// leniency is fine at the schema layer and is the wrong default here: a slot
254/// whose key nobody reads is not "extra metadata", it is an operator's
255/// instruction being ignored. Both instances that motivated this were
256/// **parses clean, deploys, wrong at request time** — a typo'd `prot = 8081`
257/// falls back to kamaji's node default, which post-R599-F12 is whatever OTHER
258/// bundle already holds 8080 on that node; and a `[providers.bundle.env]` block
259/// was, before R556-T12, read by nothing at all while looking exactly like it
260/// worked.
261///
262/// The set is the UNION of what every consumer of this slot reads, not just
263/// what [`BundleSlot::parse`] reads — `plan_ingress` reads four of its own off
264/// the same table (`reconciler::ingress`), and `MirrorProviderSlot::required`
265/// reads `required`. Scoping it to one consumer would reject live mirrors.
266///
267/// `use` / `kind` are absent deliberately: they are captured by the
268/// `MirrorProviderSlot` enum variant itself and never appear in `fields()`.
269const ALLOWED_SLOT_KEYS: &[&str] = &[
270 // BundleSlot::parse
271 "account",
272 "bucket",
273 "env",
274 "idle_ttl_ms",
275 "lifecycle",
276 "machines",
277 "name",
278 "origin",
279 "port",
280 REVALIDATE_KEY,
281 "runtime_version",
282 "serve_bins",
283 "serve_build",
284 "verify_serving",
285 "zone",
286 // MirrorProviderSlot::required — F16 placement, read via the slot, not here
287 "required",
288 // reconciler::ingress::plan_ingress — the front-door planner reads the same
289 // table. `machines`, `port` and `zone` are shared with the list above.
290 "machine",
291 "upstream_host",
292 // R844-F5 split participation from the port value, but only taught the
293 // planner about it — so a bundle slot spelling the portless shape it
294 // introduced (`fronted = true`, no `port`) was rejected here as an unknown
295 // key, and the deletion that ticket exists to enable would have failed the
296 // apply. Found and fixed from R844-F8.
297 "fronted",
298];
299
300/// True when this mirror opts its mesofact components into the W272 bundle
301/// tier — i.e. declares a `[providers.bundle]` slot.
302///
303/// Checked at the dispatch layer before the static reconciler runs, so the
304/// two tiers are mutually exclusive per mirror rather than per component.
305pub fn slot_declared(mirror: &MirrorConfig) -> bool {
306 mirror.providers.contains_key(SLOT_ROLE)
307}
308
309/// Resolve a slot-declared binary path against the workspace root.
310///
311/// Slot paths are workspace-relative unless absolute — the operator writes them
312/// in a mirror file, not from a shell cwd.
313pub fn resolve_slot_path(workspace_root: &std::path::Path, path: &std::path::Path) -> PathBuf {
314 if path.is_absolute() {
315 path.to_path_buf()
316 } else {
317 workspace_root.join(path)
318 }
319}
320
321/// Declared-but-absent binaries, as `(label, resolved path)`.
322///
323/// The label is the config coordinate (`providers.bundle.serve_bins.<triple>`)
324/// so a caller can name the exact line an operator has to fix.
325pub fn missing_bins(slot: &BundleSlot, workspace_root: &std::path::Path) -> Vec<(String, PathBuf)> {
326 let serve = slot
327 .serve_bins
328 .iter()
329 .map(|(triple, path)| (format!("providers.{SLOT_ROLE}.serve_bins.{triple}"), path));
330 let feed = slot.revalidate.iter().flat_map(|rv| {
331 rv.feed_bins.iter().map(|(triple, path)| {
332 (
333 format!("providers.{SLOT_ROLE}.{REVALIDATE_KEY}.feed_bins.{triple}"),
334 path,
335 )
336 })
337 });
338 serve
339 .chain(feed)
340 .filter_map(|(label, path)| {
341 let resolved = resolve_slot_path(workspace_root, path);
342 (!resolved.is_file()).then_some((label, resolved))
343 })
344 .collect()
345}
346
347/// True when the bundle tier can actually **serve**, not merely when it has
348/// been declared.
349///
350/// R330-B43 — THIS DISTINCTION IS THE WHOLE POINT, and getting it wrong froze
351/// yah.dev for 19 days. Dispatch used to switch tiers on [`slot_declared`]
352/// alone, so writing a `[providers.bundle]` block instantly disabled the
353/// working `[providers.static]` publish chain — while the bundle tier itself
354/// could not come up, because its `serve_bins` binaries had never been built.
355/// The old path was off, the new path could not turn on, and the site quietly
356/// served stale bytes at HTTP 200 with no error anywhere.
357///
358/// A cut-over must never be able to disable a serving path before its
359/// successor can serve. So the switch keys on the binaries EXISTING, and a
360/// declared-but-unready slot falls back to the static chain (loudly) instead
361/// of taking over and stranding the site.
362///
363/// A slot with no `serve_bins` at all is "ready" here on purpose: that is the
364/// vanilla-runtime shape, which fails later for a different, well-reported
365/// reason rather than being a half-built self-contained bundle.
366///
367/// R746-F2: a `serve_build` slot is likewise ready, and for a stronger reason —
368/// the sync can *produce* the binary it needs by dispatching the declared QED
369/// recipe, so there is no such thing as a path an operator forgot to build.
370/// That is the whole point of the declaration: B43's failure was "declared but
371/// nobody can build it here", and a recipe is exactly the thing that removes
372/// the "here".
373pub fn slot_ready(slot: &BundleSlot, workspace_root: &std::path::Path) -> bool {
374 missing_bins(slot, workspace_root).is_empty()
375}
376
377/// Parsed `[providers.bundle]` slot — everything the sync arm needs that is
378/// *declared* rather than *derived*.
379///
380/// ```toml
381/// [providers.bundle]
382/// use = "cloudflare" # R2 credentials resolve via this provider
383/// bucket = "yah-dev-bundles" # the append-only bundle store
384/// origin = "https://cdn.yah.dev" # public origin serving that bucket;
385/// # omit on a fleet whose nodes already
386/// # point at it (R870-B6)
387/// machines = ["us-east-001"] # explicit placement (or `required = {…}`)
388/// name = "yah-marketing" # stable workload handle; defaults to the service name
389/// lifecycle = "keep-alive" # or "on-demand"
390/// idle_ttl_ms = 300000 # on-demand only
391/// runtime_version = "0.8.20" # vanilla bundles only (no serve binary)
392/// serve_bins = { x86_64-unknown-linux-musl = "target/…/mesofact-serve" }
393/// # …or, instead of naming pre-built paths, name the recipe that builds them:
394/// # [providers.bundle.serve_build]
395/// # pipeline = "mesofact-musl"
396/// # binary = "mesofact"
397/// # triples = ["x86_64-unknown-linux-musl"]
398/// zone = "yah.dev" # front door to serving-verify; defaults
399/// # to the service's own domain
400/// verify_serving = true # default; see the field docs
401///
402/// # Environment for the serve process, as source URIs resolved at deploy
403/// # (R556-T12). An SSR route reading a private source needs this or it gets
404/// # a credential-less server on the node.
405/// [providers.bundle.env]
406/// ANALYTICS_R2_ACCESS_KEY = "vault:cloudflare-r2-access-key-id"
407/// ANALYTICS_R2_BUCKET = "yah-analytics" # bare literal: not a secret
408/// ```
409#[derive(Debug, Clone, PartialEq, Eq)]
410pub struct BundleSlot {
411 /// R2 bucket holding the bundle store. Append-only, blob-deduped.
412 pub bucket: String,
413 /// Public HTTPS origin serving `bucket` — the R2 custom domain bound to it,
414 /// e.g. `"https://cdn.noisetable.com"` (R870-B6).
415 ///
416 /// `None` → the node fetches from its own `KAMAJI_BUNDLE_ORIGIN`, which is
417 /// correct exactly while the store the mirror publishes to is the store the
418 /// fleet was configured for. A second tenant publishing to its own bucket
419 /// must declare this or the node materializes from the wrong store and
420 /// fails with `missing blob manifests/<digest>` after a clean admission.
421 ///
422 /// Declared rather than derived from `bucket` or `zone`: the bucket→origin
423 /// mapping is a Cloudflare R2 custom-domain binding that exists or doesn't,
424 /// and guessing `https://cdn.<zone>` would put an unreachable URL on the
425 /// wire for every mirror that has not bound one. Same call
426 /// `providers.static.asset_origin` makes, one tier over.
427 pub origin: Option<String>,
428 /// Cloudflare account id override. `None` → resolve from the workspace's
429 /// cloudflare provider config / `CF_ACCOUNT_ID`.
430 pub account: Option<String>,
431 /// Stable operator-facing workload name. yubaba requires one for a bundle
432 /// deploy: the digest is the *content* and changes on every rebuild, so it
433 /// is not a usable handle for `list` / `stop`.
434 pub name: Option<String>,
435 /// Explicitly named target machines, in deploy order. Empty → fall back to
436 /// the slot's `required = {…}` placement spec.
437 pub machines: Vec<String>,
438 /// Stock runtime version recorded as `runtime = "mesofact/<version>"` for a
439 /// vanilla bundle. Ignored when `serve_bins` is non-empty. `None` → the
440 /// caller's own version.
441 pub runtime_version: Option<String>,
442 /// `<triple> → <path to serve binary>`. Any entry makes this a
443 /// `runtime = "self"` bundle that carries its own serve binaries.
444 pub serve_bins: BTreeMap<String, PathBuf>,
445 /// Build the serve binaries on demand instead of naming pre-built paths
446 /// (R746-F2). Mutually exclusive with `serve_bins`; either one makes this a
447 /// `runtime = "self"` bundle.
448 pub serve_build: Option<BinBuild>,
449 /// How kamaji supervises the served bundle.
450 pub lifecycle: BundleLifecycle,
451 /// Port the served bundle listens on (R599-F12). `None` → kamaji's
452 /// node-wide default (8080), which is only correct while the node hosts a
453 /// single bundle; declare one per workload to put several on a node.
454 pub port: Option<u16>,
455 /// `[providers.bundle.env]` — environment for the **serve** process, as
456 /// `NAME → source URI` (R556-T12).
457 ///
458 /// Values are the source *declaration*, kept verbatim and resolved
459 /// deploy-side by `yah cloud apply` — `vault:<slot>`, `env:<VAR>`, a
460 /// pipe-joined fallback chain of either, or a bare literal for a
461 /// known-non-secret value. Same grammar `~/.yah/qed/secrets.toml` uses, so
462 /// there is one source-URI vocabulary in the camp rather than two.
463 ///
464 /// Parsing stays here and resolution does not: this crate is offline by
465 /// construction (a misconfigured mirror must fail before a build runs), and
466 /// only the syncing machine has the vault. The `RevalidateSlot::mirror_key_env`
467 /// → [`RevalidateSlot::to_workload_payload`] split is the same shape one
468 /// level down.
469 pub env: BTreeMap<String, String>,
470 /// Optional revalidate receiver config (R330-F12). `Some` → the deploy
471 /// also stands up a `mesofact serve --revalidate` process.
472 pub revalidate: Option<RevalidateSlot>,
473 /// Public zone whose front door is checked after a deploy (R703-T7).
474 /// `None` → the service's own `domain`, which is the shape every mirror in
475 /// tree uses; declare one only when the bundle serves a zone that isn't it.
476 ///
477 /// Unlike `[providers.static]`, this is optional: the static slot's `zone`
478 /// is load-bearing for the Worker route and cache purge, whereas here it
479 /// only names what to probe.
480 pub zone: Option<String>,
481 /// Whether a deploy is checked against the live front door (R703-T7).
482 ///
483 /// Defaults to **true**, and the only reason to turn it off is a
484 /// deliberately in-flight front-door migration — with a comment naming the
485 /// ticket. It is declared in config rather than passed as a CLI flag for
486 /// the same reason the static slot's is: switching it off should be a
487 /// reviewable diff, not an invocation habit that quietly becomes permanent.
488 pub verify_serving: bool,
489}
490
491/// A binary the bundle needs, declared as **the recipe that builds it** rather
492/// than as a path someone is expected to have already produced (R746-F2).
493///
494/// ```toml
495/// [providers.bundle.serve_build]
496/// pipeline = "mesofact-musl" # .yah/qed/<name>.toml
497/// binary = "mesofact" # matches a step's `produces.binary`
498/// triples = ["x86_64-unknown-linux-musl"] # what the placed nodes run
499/// ```
500///
501/// # Why this is a declaration and not a fallback
502///
503/// The alternative shape — "use `serve_bins` if the path exists, otherwise
504/// build" — makes the deployed artifact a function of what happens to be on the
505/// operator's disk. Two machines syncing the same mirror would then ship
506/// different binaries, and the one with a stale path would ship the stale one
507/// silently. The mirror says which shape it is; the sync obeys.
508///
509/// Declaring both this and `serve_bins` is refused for the same reason.
510#[derive(Debug, Clone, PartialEq, Eq)]
511pub struct BinBuild {
512 /// QED pipeline name, resolved under `.yah/qed/<pipeline>.toml`.
513 pub pipeline: String,
514 /// Logical binary name, matched against a step's `[[steps.produces]]
515 /// binary`.
516 pub binary: String,
517 /// Target triples to resolve, in declaration order. Non-empty: a build
518 /// declaration that names no target builds nothing.
519 pub triples: Vec<String>,
520}
521
522/// Parsed `[providers.bundle.revalidate]` sub-slot — declares the almanac
523/// revalidate receiver to fork alongside the static bundle server (R330-F12).
524///
525/// ```toml
526/// [providers.bundle.revalidate]
527/// routes = ["/releases"] # allowlist (empty = all routes)
528/// mirror_key_env = "YAH_MARKETING_MIRROR_KEY" # env var holding the bearer
529/// publish_config = "mesofact.config.toml" # default
530/// feeds = ["releases"] # .yah/almanac/<name>.toml to keep fresh
531/// feed_interval_secs = 300 # default
532/// feed_runtime = "almanac-feed/0.8.22" # vanilla: node resolves the fetcher
533/// # …or, for a self-contained bundle, stage it in and name the built paths:
534/// # feed_bins = { x86_64-unknown-linux-musl = "target/…/almanac-feed" }
535/// ```
536#[derive(Debug, Clone, PartialEq, Eq)]
537pub struct RevalidateSlot {
538 /// Routes the receiver accepts pokes for (allowlist).
539 /// Empty → all routes in the workload manifest.
540 ///
541 /// Enforced on the node since R752-B7: kamaji renders this list as one
542 /// `--allow-route` per entry on the receiver's argv, `mesofact serve`
543 /// refuses an explicit poke outside it (403) and narrows a whole-site poke
544 /// to it. Before that it was parsed here, shipped over the wire, and read
545 /// by nobody — declaring it bought exactly nothing. Scoping only: `who may
546 /// poke` is `mirror_key_env`, and the two are independent.
547 pub routes: Vec<String>,
548 /// Env var name holding the tenant bearer secret. Deploy resolves it
549 /// and sets `MESOFACT_MIRROR_KEY` on the receiver process.
550 /// `None` → open receiver (no bearer check).
551 pub mirror_key_env: Option<String>,
552 /// Path to `mesofact.config.toml` with the `[publish]` block, relative
553 /// to the workload directory. `None` → default `"mesofact.config.toml"`.
554 pub publish_config: Option<PathBuf>,
555 /// Almanac feed names (`.yah/almanac/<name>.toml`) the on-node fetch tier
556 /// keeps fresh (R330-F31). Empty → no fetcher, and the receiver re-renders
557 /// whatever data the bundle was built with.
558 pub feeds: Vec<String>,
559 /// Seconds between feed-fetch ticks. `None` → the spec default.
560 pub feed_interval_secs: Option<u64>,
561 /// Per-triple path to the `almanac-feed` binary staged into the bundle as a
562 /// sidecar. The self-contained shape's answer to "how does the fetcher
563 /// reach the node".
564 ///
565 /// Mutually exclusive with [`feed_runtime`](Self::feed_runtime), for the
566 /// same reason `serve_bins` and `serve_build` are: the mirror declares
567 /// which shape it is, and a use-whichever-exists fallback would make the
568 /// deployed binary a function of the syncing machine's disk.
569 pub feed_bins: BTreeMap<String, PathBuf>,
570 /// Runtime ref the fetcher resolves from the node's shared runtime-asset
571 /// cache — `feed_runtime = "almanac-feed/0.8.22"` (R746-T3).
572 ///
573 /// This is the **vanilla** shape's answer, and it is what makes a vanilla
574 /// bundle with a feed tier possible at all: `feed_bins` is a path someone
575 /// must have cross-built, so a bundle that carries no serve binary but
576 /// still needs a sidecar path has only moved the toolchain requirement,
577 /// not removed it.
578 pub feed_runtime: Option<String>,
579}
580
581impl RevalidateSlot {
582 /// Build the [`MesofactRevalidateReceiver`] payload for the workload spec,
583 /// given the env vars resolved at deploy time and the feed definitions read
584 /// from the camp's `.yah/almanac/` tree.
585 ///
586 /// Feed definitions travel by value: reading them is the deploy side's job
587 /// (it is the only participant that has the camp checkout), and the node
588 /// gets a self-contained payload.
589 pub fn to_workload_payload(
590 &self,
591 env: BTreeMap<String, String>,
592 feeds: Vec<workload_spec::AlmanacFeed>,
593 feed_project_prefix: Option<String>,
594 secrets: Vec<workload_spec::SecretMount>,
595 ) -> MesofactRevalidateReceiver {
596 MesofactRevalidateReceiver {
597 routes: self.routes.clone(),
598 publish_config: self
599 .publish_config
600 .as_ref()
601 .map(|p| p.to_string_lossy().into_owned())
602 .unwrap_or_else(|| "mesofact.config.toml".to_string()),
603 mirror_key_env: self.mirror_key_env.clone(),
604 env,
605 feeds,
606 feed_interval_secs: self
607 .feed_interval_secs
608 .unwrap_or(DEFAULT_FEED_INTERVAL_SECS),
609 feed_project_prefix,
610 feed_runtime: self.feed_runtime.clone(),
611 secrets,
612 }
613 }
614}
615
616/// Mirrors `workload_spec`'s own default. Duplicated rather than exported
617/// because the spec keeps its serde defaults private; the parse tests below
618/// pin the two together.
619pub const DEFAULT_FEED_INTERVAL_SECS: u64 = 300;
620
621/// Bundle path segment the fetch tier's sidecar binary is staged under —
622/// `bins/<triple>/almanac-feed`, next to `bins/<triple>/serve` — and the
623/// filename it lands under in the node runtime-asset cache when a *vanilla*
624/// bundle resolves it by name instead (R746-T3).
625///
626/// Re-exported from `yah_mesofact_bundle` rather than re-typed: this crate and
627/// kamaji both used to declare their own copy, pinned together only by an
628/// argv-shape test. One `const` in the crate they both already depend on
629/// removes the drift instead of detecting it.
630pub use yah_mesofact_bundle::FEED_BIN as FEED_BIN_NAME;
631
632impl BundleSlot {
633 /// Parse the mirror's `[providers.bundle]` slot.
634 ///
635 /// Every failure names the offending field plus the service and env, so the
636 /// operator gets a file to open rather than a type error. Validation is
637 /// total and offline — nothing here touches the network, so a misconfigured
638 /// mirror fails before a build runs (R330-B5 fail-fast discipline).
639 pub fn parse(mirror: &MirrorConfig, service: &str, env: &str) -> Result<Self> {
640 let slot = mirror.providers.get(SLOT_ROLE).with_context(|| {
641 format!(
642 "mirror has no `providers.{SLOT_ROLE}` slot — required for the W272 bundle tier \
643 (service={service}, env={env})"
644 )
645 })?;
646 let fields = slot.fields();
647
648 // R556-B14. Unknown keys are rejected BEFORE anything is read, so the
649 // operator gets the typo rather than a downstream complaint about the
650 // field the typo was supposed to be. Nearest-match is offered because
651 // the realistic failure is one transposed character, and an error that
652 // only says "unknown" makes the reader diff the docs by eye.
653 for key in fields.keys() {
654 if ALLOWED_SLOT_KEYS.contains(&key.as_str()) {
655 continue;
656 }
657 let hint = nearest_slot_key(key)
658 .map(|k| format!(" — did you mean `{k}`?"))
659 .unwrap_or_default();
660 bail!(
661 "providers.{SLOT_ROLE} has an unknown key `{key}`{hint} (service={service}, \
662 env={env}). Every key this slot reads is one of: {}. An unrecognized key is \
663 refused rather than ignored because the failure it hides is silent: a typo'd \
664 `port` deploys onto whatever bundle already holds the node default, and a \
665 mistyped credential block deploys a serve process with no credentials at all.",
666 ALLOWED_SLOT_KEYS.join(", "),
667 );
668 }
669
670 let bucket = fields
671 .get("bucket")
672 .and_then(|v| v.as_str())
673 .filter(|s| !s.is_empty())
674 .with_context(|| {
675 format!(
676 "providers.{SLOT_ROLE} has no `bucket` — name the R2 bundle store in \
677 .yah/services/{service}/mirrors/{env}.toml"
678 )
679 })?
680 .to_string();
681
682 // R870-B6. Parsed strictly, and with the scheme required: the value's
683 // only consumer is `HttpReadOnlyObjectStore`, which joins keys onto it
684 // as path segments, so a bare hostname (`cdn.noisetable.com`) produces
685 // a relative URL that fails on the node — after a clean apply, at
686 // materialize time, which is the far side of the feedback loop this
687 // slot's validation exists to stay on.
688 let origin = match fields.get("origin") {
689 None => None,
690 Some(v) => {
691 let raw = v
692 .as_str()
693 .map(str::trim)
694 .filter(|s| !s.is_empty())
695 .with_context(|| {
696 format!(
697 "providers.{SLOT_ROLE}.origin must be a non-empty public HTTPS origin \
698 serving the bundle bucket, e.g. \"https://cdn.{service}.example\" \
699 (service={service}, env={env})"
700 )
701 })?;
702 if !(raw.starts_with("https://") || raw.starts_with("http://")) {
703 bail!(
704 "providers.{SLOT_ROLE}.origin = {raw:?} has no scheme (service={service}, \
705 env={env}) — kamaji fetches bundle objects by joining keys onto this \
706 value, so it must be a full origin URL like \
707 \"https://cdn.example.com\", not a bucket or hostname"
708 );
709 }
710 Some(raw.trim_end_matches('/').to_string())
711 }
712 };
713
714 let account = fields
715 .get("account")
716 .and_then(|v| v.as_str())
717 .filter(|s| !s.is_empty())
718 .map(str::to_string);
719
720 let name = fields
721 .get("name")
722 .and_then(|v| v.as_str())
723 .filter(|s| !s.is_empty())
724 .map(str::to_string);
725
726 let machines = match fields.get("machines") {
727 None => Vec::new(),
728 Some(v) => {
729 let list = v.as_array().with_context(|| {
730 format!(
731 "providers.{SLOT_ROLE}.machines must be an array of machine names \
732 (service={service}, env={env})"
733 )
734 })?;
735 list.iter()
736 .map(|entry| {
737 entry
738 .as_str()
739 .filter(|s| !s.is_empty())
740 .map(str::to_string)
741 .with_context(|| {
742 format!(
743 "providers.{SLOT_ROLE}.machines holds a non-string (or empty) \
744 entry (service={service}, env={env})"
745 )
746 })
747 })
748 .collect::<Result<Vec<_>>>()?
749 }
750 };
751
752 let runtime_version = fields
753 .get("runtime_version")
754 .and_then(|v| v.as_str())
755 .filter(|s| !s.is_empty())
756 .map(str::to_string);
757
758 let serve_bins = match fields.get("serve_bins") {
759 None => BTreeMap::new(),
760 Some(v) => {
761 let table = v.as_table().with_context(|| {
762 format!(
763 "providers.{SLOT_ROLE}.serve_bins must be a table of \
764 <target-triple> = <path> (service={service}, env={env})"
765 )
766 })?;
767 table
768 .iter()
769 .map(|(triple, path)| {
770 let path = path.as_str().filter(|s| !s.is_empty()).with_context(|| {
771 format!(
772 "providers.{SLOT_ROLE}.serve_bins.{triple} must be a non-empty \
773 path (service={service}, env={env})"
774 )
775 })?;
776 Ok((triple.clone(), PathBuf::from(path)))
777 })
778 .collect::<Result<BTreeMap<_, _>>>()?
779 }
780 };
781
782 let serve_build = parse_bin_build(
783 fields.get("serve_build"),
784 &format!("providers.{SLOT_ROLE}.serve_build"),
785 service,
786 env,
787 )?;
788
789 if serve_build.is_some() && !serve_bins.is_empty() {
790 bail!(
791 "providers.{SLOT_ROLE} declares BOTH `serve_bins` and `serve_build` — pick one \
792 (service={service}, env={env}). `serve_bins` names binaries you have already \
793 built; `serve_build` names the QED recipe that builds them. Accepting both \
794 would make the deployed binary depend on what happens to be on the syncing \
795 machine's disk, which is how one operator ships a stale binary while another \
796 ships a fresh one from the same mirror."
797 );
798 }
799
800 // R599-F12. Parsed strictly: a port is either absent or a real one, and
801 // a typo that silently fell back to 8080 would collide with whatever
802 // bundle already holds that port on the node — a failure that surfaces
803 // as the wrong site being served, not as an error.
804 let port = match fields.get("port") {
805 None => None,
806 Some(v) => {
807 let n = v.as_integer().with_context(|| {
808 format!(
809 "providers.{SLOT_ROLE}.port must be an integer TCP port \
810 (service={service}, env={env})"
811 )
812 })?;
813 Some(u16::try_from(n).ok().filter(|p| *p != 0).with_context(|| {
814 format!(
815 "providers.{SLOT_ROLE}.port = {n} is not a usable TCP port \
816 (1..=65535) (service={service}, env={env})"
817 )
818 })?)
819 }
820 };
821
822 // R556-T12. Env for the serve process. Declared as source URIs and
823 // stored verbatim — resolution is the deploy side's job (see the field
824 // docs). Every value is required to be a non-empty string: an empty
825 // source is a var that would silently reach the node unset, which is
826 // the exact failure mode this slot exists to remove.
827 let serve_env = match fields.get("env") {
828 None => BTreeMap::new(),
829 Some(v) => {
830 let table = v.as_table().with_context(|| {
831 format!(
832 "providers.{SLOT_ROLE}.env must be a table of <ENV_NAME> = \
833 \"<source-uri>\" (service={service}, env={env})"
834 )
835 })?;
836 table
837 .iter()
838 .map(|(name, source)| {
839 let source = source
840 .as_str()
841 .filter(|s| !s.trim().is_empty())
842 .with_context(|| {
843 format!(
844 "providers.{SLOT_ROLE}.env.{name} must be a non-empty source \
845 string — \"vault:<slot>\", \"env:<VAR>\", a pipe-joined \
846 chain of either, or a bare literal for a non-secret \
847 (service={service}, env={env})"
848 )
849 })?;
850 Ok((name.clone(), source.to_string()))
851 })
852 .collect::<Result<BTreeMap<_, _>>>()?
853 }
854 };
855
856 let lifecycle = parse_lifecycle(
857 fields.get("lifecycle").and_then(|v| v.as_str()),
858 fields.get("idle_ttl_ms").and_then(|v| v.as_integer()),
859 service,
860 env,
861 )?;
862
863 let revalidate = parse_revalidate_slot(fields, service, env)?;
864
865 let zone = fields
866 .get("zone")
867 .and_then(|v| v.as_str())
868 .filter(|s| !s.is_empty())
869 .map(str::to_string);
870
871 // R703-T7. Parsed strictly rather than `unwrap_or(true)` on a bad type:
872 // `verify_serving = "false"` silently reading as *enabled* is the
873 // friendlier-looking failure, but an operator who typed it believes the
874 // check is off and will be surprised by an apply that fails on a
875 // migration they thought they had silenced.
876 let verify_serving = match fields.get("verify_serving") {
877 None => true,
878 Some(v) => v.as_bool().with_context(|| {
879 format!(
880 "providers.{SLOT_ROLE}.verify_serving must be a boolean \
881 (service={service}, env={env})"
882 )
883 })?,
884 };
885
886 // R746-T3: a vanilla bundle carries no `bins/` by construction, so a
887 // sidecar declared as a PATH has nowhere to be staged into. Caught here
888 // rather than at assembly so the operator gets the mirror file and the
889 // remedy, offline, before a build runs.
890 let slot = Self {
891 bucket,
892 origin,
893 account,
894 name,
895 machines,
896 runtime_version,
897 serve_bins,
898 serve_build,
899 lifecycle,
900 port,
901 env: serve_env,
902 revalidate,
903 zone,
904 verify_serving,
905 };
906 if !slot.is_self_contained() {
907 if let Some(rv) = slot.revalidate.as_ref() {
908 if !rv.feed_bins.is_empty() {
909 anyhow::bail!(
910 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.feed_bins is declared but this is \
911 a VANILLA bundle (no serve_bins / serve_build), which carries no bins/ \
912 at all — replace it with feed_runtime = \"{FEED_BIN_NAME}/<version>\" \
913 and publish that asset once per triple with `yah cloud bundle \
914 publish-runtime` (service={service}, env={env})"
915 );
916 }
917 }
918 }
919 Ok(slot)
920 }
921
922 /// The zone whose front door a deploy of this bundle is checked against:
923 /// the slot's `zone`, else the service's own domain.
924 pub fn serving_zone<'a>(&'a self, service_domain: &'a str) -> &'a str {
925 self.zone.as_deref().unwrap_or(service_domain)
926 }
927
928 /// Stable workload handle: the slot's `name`, else the service name.
929 pub fn workload_name<'a>(&'a self, service: &'a str) -> &'a str {
930 self.name.as_deref().unwrap_or(service)
931 }
932
933 /// True when the assembled bundle carries its own serve binaries
934 /// (`runtime = "self"`) rather than resolving a stock node runtime asset.
935 ///
936 /// Keyed on the *declaration*, not on what is on disk: a `serve_build` slot
937 /// is self-contained before its binary has ever been built, because the
938 /// mirror said so. Deriving the shape from disk state instead is the bug
939 /// this relay exists to remove — it makes a bundle's shape depend on which
940 /// machine ran the sync.
941 pub fn is_self_contained(&self) -> bool {
942 !self.serve_bins.is_empty() || self.serve_build.is_some()
943 }
944
945 /// Build the `{digest, runtime, lifecycle}` triple a `mesofact-static`
946 /// workload carries once its bundle is published.
947 ///
948 /// `runtime` wire-mirrors `yah_mesofact_bundle::BundleRuntime`, so it is
949 /// taken from the manifest the assembler actually wrote rather than
950 /// re-derived here — the manifest is what the node will verify against.
951 ///
952 /// `env` is the **resolved** serve environment, passed in rather than read
953 /// off `self.env`: this crate holds source URIs, and only the syncing
954 /// machine can turn a `vault:<slot>` into a value. Same by-value handoff
955 /// [`RevalidateSlot::to_workload_payload`] takes, for the same reason —
956 /// the node must never see a keystore slot name (R556-T12).
957 pub fn serve_bundle(
958 &self,
959 digest: &str,
960 runtime: &str,
961 env: BTreeMap<String, String>,
962 ) -> MesofactServeBundle {
963 MesofactServeBundle {
964 digest: BlakeHash(digest.to_string()),
965 runtime: runtime.to_string(),
966 lifecycle: self.lifecycle.clone(),
967 // R599-F12: the slot's declared `port`, or `None` for kamaji's
968 // node-wide default. (@Ashguard:blade parked a `None` here to
969 // unblock the camp's build while this ticket was mid-flight; this
970 // is the real threading it named.)
971 port: self.port,
972 env,
973 // R870-B6: the store this bundle was published to, so the node
974 // fetches from it rather than from whichever store the *node* was
975 // pointed at. `None` keeps the node-wide `KAMAJI_BUNDLE_ORIGIN`,
976 // which is why no yah-owned mirror needs an edit.
977 origin: self.origin.clone(),
978 }
979 }
980}
981
982/// Closest [`ALLOWED_SLOT_KEYS`] entry to `key`, or `None` when nothing is
983/// close enough to be worth suggesting (R556-B14).
984///
985/// The threshold scales with the key's length — one edit for a short key like
986/// `port`, two for a longer one — so `prot` suggests `port` while an entirely
987/// invented key suggests nothing. A confidently wrong suggestion is worse than
988/// none: it sends the operator to fix a line that was never the problem.
989fn nearest_slot_key(key: &str) -> Option<&'static str> {
990 let budget = if key.len() <= 5 { 1 } else { 2 };
991 ALLOWED_SLOT_KEYS
992 .iter()
993 .map(|candidate| (edit_distance(key, candidate), *candidate))
994 .filter(|(d, _)| *d <= budget)
995 .min()
996 .map(|(_, candidate)| candidate)
997}
998
999/// Optimal string alignment (Damerau-Levenshtein restricted to adjacent
1000/// transpositions), three-row DP. Byte-wise: every key in this grammar is
1001/// ASCII, and a multi-byte typo is not a case worth carrying a char-vec for.
1002///
1003/// Transposition counts as ONE edit, not two, and that is the whole reason to
1004/// carry the extra row: `prot` for `port` is the motivating typo of R556-B14,
1005/// and plain Levenshtein scores it 2 — far enough away that a threshold tight
1006/// enough to avoid nonsense suggestions would refuse to suggest the one that
1007/// matters.
1008fn edit_distance(a: &str, b: &str) -> usize {
1009 let (a, b) = (a.as_bytes(), b.as_bytes());
1010 let mut prev2 = vec![0usize; b.len() + 1];
1011 let mut prev: Vec<usize> = (0..=b.len()).collect();
1012 let mut cur = vec![0usize; b.len() + 1];
1013 for (i, &ac) in a.iter().enumerate() {
1014 cur[0] = i + 1;
1015 for (j, &bc) in b.iter().enumerate() {
1016 let mut d = (prev[j] + usize::from(ac != bc))
1017 .min(prev[j + 1] + 1)
1018 .min(cur[j] + 1);
1019 if i > 0 && j > 0 && ac == b[j - 1] && a[i - 1] == bc {
1020 d = d.min(prev2[j - 1] + 1);
1021 }
1022 cur[j + 1] = d;
1023 }
1024 std::mem::swap(&mut prev2, &mut prev);
1025 std::mem::swap(&mut prev, &mut cur);
1026 }
1027 prev[b.len()]
1028}
1029
1030/// Parse a `[…serve_build]`-shaped table into a [`BinBuild`] (R746-F2).
1031///
1032/// Taken as a helper rather than inlined because the revalidate tier's
1033/// `feed_bins` has the identical "a path someone must have built" problem and
1034/// will want the identical declaration once a recipe produces `almanac-feed`.
1035/// Every message names the full config coordinate so the operator gets a line
1036/// to open.
1037fn parse_bin_build(
1038 value: Option<&toml::Value>,
1039 label: &str,
1040 service: &str,
1041 env: &str,
1042) -> Result<Option<BinBuild>> {
1043 let Some(value) = value else {
1044 return Ok(None);
1045 };
1046 let table = value.as_table().with_context(|| {
1047 format!("{label} must be a table of pipeline/binary/triples (service={service}, env={env})")
1048 })?;
1049
1050 let pipeline = table
1051 .get("pipeline")
1052 .and_then(|v| v.as_str())
1053 .filter(|s| !s.is_empty())
1054 .with_context(|| {
1055 format!(
1056 "{label}.pipeline must name a QED pipeline (.yah/qed/<name>.toml) \
1057 (service={service}, env={env})"
1058 )
1059 })?
1060 .to_string();
1061
1062 let binary = table
1063 .get("binary")
1064 .and_then(|v| v.as_str())
1065 .filter(|s| !s.is_empty())
1066 .with_context(|| {
1067 format!(
1068 "{label}.binary must name the produced binary — it is matched against the \
1069 pipeline's `[[steps.produces]] binary` (service={service}, env={env})"
1070 )
1071 })?
1072 .to_string();
1073
1074 let triples = table
1075 .get("triples")
1076 .and_then(|v| v.as_array())
1077 .with_context(|| {
1078 format!("{label}.triples must be an array of target triples (service={service}, env={env})")
1079 })?
1080 .iter()
1081 .map(|entry| {
1082 entry
1083 .as_str()
1084 .filter(|s| !s.is_empty())
1085 .map(str::to_string)
1086 .with_context(|| {
1087 format!("{label}.triples holds a non-string (or empty) entry (service={service}, env={env})")
1088 })
1089 })
1090 .collect::<Result<Vec<_>>>()?;
1091
1092 if triples.is_empty() {
1093 bail!(
1094 "{label}.triples is empty — a build declaration that names no target builds \
1095 nothing, and the bundle would assemble with no serve binary at all \
1096 (service={service}, env={env})"
1097 );
1098 }
1099
1100 Ok(Some(BinBuild {
1101 pipeline,
1102 binary,
1103 triples,
1104 }))
1105}
1106
1107/// `lifecycle = "keep-alive" | "on-demand"` (+ `idle_ttl_ms` for the latter).
1108fn parse_lifecycle(
1109 raw: Option<&str>,
1110 idle_ttl_ms: Option<i64>,
1111 service: &str,
1112 env: &str,
1113) -> Result<BundleLifecycle> {
1114 match raw.unwrap_or("keep-alive") {
1115 "keep-alive" | "keepalive" => {
1116 if idle_ttl_ms.is_some() {
1117 bail!(
1118 "providers.{SLOT_ROLE}.idle_ttl_ms only applies to `lifecycle = \"on-demand\"` \
1119 — a keep-alive bundle is never reaped (service={service}, env={env})"
1120 );
1121 }
1122 Ok(BundleLifecycle::KeepAlive)
1123 }
1124 "on-demand" | "ondemand" | "jit" => {
1125 let ttl = idle_ttl_ms.unwrap_or(DEFAULT_IDLE_TTL_MS as i64);
1126 if ttl <= 0 {
1127 bail!(
1128 "providers.{SLOT_ROLE}.idle_ttl_ms must be positive, got {ttl} \
1129 (service={service}, env={env})"
1130 );
1131 }
1132 Ok(BundleLifecycle::OnDemand {
1133 idle_ttl: Millis::from_ms(ttl as u64),
1134 })
1135 }
1136 other => bail!(
1137 "providers.{SLOT_ROLE}.lifecycle must be \"keep-alive\" or \"on-demand\", got \
1138 {other:?} (service={service}, env={env})"
1139 ),
1140 }
1141}
1142
1143/// Parse the optional `[providers.bundle.revalidate]` sub-table.
1144///
1145/// `None` → no revalidate receiver declared (the common case). `Some` → the
1146/// deploy also stands up a `mesofact serve --revalidate` process.
1147fn parse_revalidate_slot(
1148 fields: &BTreeMap<String, toml::Value>,
1149 service: &str,
1150 env: &str,
1151) -> Result<Option<RevalidateSlot>> {
1152 let sub = match fields.get(REVALIDATE_KEY) {
1153 None => return Ok(None),
1154 Some(v) => v.as_table().with_context(|| {
1155 format!(
1156 "providers.{SLOT_ROLE}.{REVALIDATE_KEY} must be a TOML table \
1157 (service={service}, env={env})"
1158 )
1159 })?,
1160 };
1161
1162 let routes = match sub.get("routes") {
1163 None => Vec::new(),
1164 Some(v) => {
1165 let list = v.as_array().with_context(|| {
1166 format!(
1167 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.routes must be an array of route \
1168 patterns (service={service}, env={env})"
1169 )
1170 })?;
1171 list.iter()
1172 .map(|entry| {
1173 entry
1174 .as_str()
1175 .filter(|s| !s.is_empty())
1176 .map(str::to_string)
1177 .with_context(|| {
1178 format!(
1179 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.routes holds a non-string \
1180 (or empty) entry (service={service}, env={env})"
1181 )
1182 })
1183 })
1184 .collect::<Result<Vec<_>>>()?
1185 }
1186 };
1187
1188 let mirror_key_env = sub
1189 .get("mirror_key_env")
1190 .and_then(|v| v.as_str())
1191 .filter(|s| !s.is_empty())
1192 .map(str::to_string);
1193
1194 let publish_config = sub
1195 .get("publish_config")
1196 .and_then(|v| v.as_str())
1197 .filter(|s| !s.is_empty())
1198 .map(PathBuf::from);
1199
1200 // ── Feed-fetch tier (R330-F31) ──────────────────────────────────────────
1201 let feeds = match sub.get("feeds") {
1202 None => Vec::new(),
1203 Some(v) => {
1204 let list = v.as_array().with_context(|| {
1205 format!(
1206 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.feeds must be an array of almanac \
1207 feed names (service={service}, env={env})"
1208 )
1209 })?;
1210 list.iter()
1211 .map(|entry| {
1212 entry
1213 .as_str()
1214 .filter(|s| !s.is_empty())
1215 .map(str::to_string)
1216 .with_context(|| {
1217 format!(
1218 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.feeds holds a non-string \
1219 (or empty) entry (service={service}, env={env})"
1220 )
1221 })
1222 })
1223 .collect::<Result<Vec<_>>>()?
1224 }
1225 };
1226
1227 let feed_interval_secs = match sub.get("feed_interval_secs") {
1228 None => None,
1229 Some(v) => {
1230 let secs = v.as_integer().filter(|n| *n > 0).with_context(|| {
1231 format!(
1232 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.feed_interval_secs must be a positive \
1233 integer number of seconds (service={service}, env={env})"
1234 )
1235 })?;
1236 Some(secs as u64)
1237 }
1238 };
1239
1240 let feed_bins = match sub.get("feed_bins") {
1241 None => BTreeMap::new(),
1242 Some(v) => {
1243 let table = v.as_table().with_context(|| {
1244 format!(
1245 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.feed_bins must be a table of \
1246 <target-triple> = <path> (service={service}, env={env})"
1247 )
1248 })?;
1249 table
1250 .iter()
1251 .map(|(triple, path)| {
1252 let p = path.as_str().filter(|s| !s.is_empty()).with_context(|| {
1253 format!(
1254 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.feed_bins.{triple} must be \
1255 a non-empty path string (service={service}, env={env})"
1256 )
1257 })?;
1258 Ok((triple.clone(), PathBuf::from(p)))
1259 })
1260 .collect::<Result<BTreeMap<_, _>>>()?
1261 }
1262 };
1263
1264 // R746-T3: the vanilla shape's fetcher. A ref, not a path — the node
1265 // resolves it from the shared runtime-asset cache the same way it resolves
1266 // `serve`, so no cross-built binary has to exist on the syncing machine.
1267 let feed_runtime = match sub.get("feed_runtime") {
1268 None => None,
1269 Some(v) => {
1270 let s = v.as_str().filter(|s| !s.is_empty()).with_context(|| {
1271 format!(
1272 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.feed_runtime must be a non-empty \
1273 runtime reference like \"{FEED_BIN_NAME}/0.8.22\" (service={service}, \
1274 env={env})"
1275 )
1276 })?;
1277 // Parse offline so a typo fails the apply with a file to open,
1278 // rather than a node failing to resolve it twenty minutes later.
1279 yah_mesofact_bundle::RuntimeRef::parse(s).with_context(|| {
1280 format!(
1281 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.feed_runtime (service={service}, \
1282 env={env})"
1283 )
1284 })?;
1285 Some(s.to_string())
1286 }
1287 };
1288
1289 // Declared, never inferred — the same rule serve_bins/serve_build follow.
1290 // "Use the path if it happens to exist, else the ref" would make the
1291 // deployed fetcher a function of the syncing machine's disk.
1292 if !feed_bins.is_empty() && feed_runtime.is_some() {
1293 anyhow::bail!(
1294 "providers.{SLOT_ROLE}.{REVALIDATE_KEY} declares BOTH feed_bins and feed_runtime — \
1295 pick one: feed_bins stages the `{FEED_BIN_NAME}` fetcher into the bundle (the \
1296 self-contained shape), feed_runtime resolves it from the node's runtime-asset \
1297 cache (the vanilla shape) (service={service}, env={env})"
1298 );
1299 }
1300
1301 // Declaring feeds without shipping the fetcher is the failure that looks
1302 // like success: the deploy goes green, the receiver serves, and the data
1303 // never moves again. Catch it here, offline, with the file to edit.
1304 if !feeds.is_empty() && feed_bins.is_empty() && feed_runtime.is_none() {
1305 anyhow::bail!(
1306 "providers.{SLOT_ROLE}.{REVALIDATE_KEY}.feeds declares {} feed(s) but neither \
1307 feed_bins nor feed_runtime — the node has no way to get the `{FEED_BIN_NAME}` \
1308 fetcher, so nothing would ever refresh them (service={service}, env={env})",
1309 feeds.len()
1310 );
1311 }
1312
1313 Ok(Some(RevalidateSlot {
1314 routes,
1315 mirror_key_env,
1316 publish_config,
1317 feeds,
1318 feed_interval_secs,
1319 feed_bins,
1320 feed_runtime,
1321 }))
1322}
1323
1324/// Resolve the machines a published bundle deploys to, in deploy order.
1325///
1326/// Two declaration forms, checked in that order:
1327/// 1. `machines = ["us-east-001", …]` — explicit, ordered, and the shape to
1328/// prefer while a bundle binds loopback (F10: one bundle per node, passway
1329/// co-located), because *which* nodes serve is then an operator decision
1330/// rather than a scheduler outcome.
1331/// 2. `required = { regions = […], mesh_tags = […], replicas = N }` — F16
1332/// placement. Resolves to the first `N` machines the constraint matches
1333/// (`replicas` absent = one, the only shape on disk before R844-F8).
1334///
1335/// An undeclared / unresolvable placement is an error, not an empty deploy —
1336/// silently publishing a bundle nobody serves is the failure mode this avoids.
1337/// So is a *short* one: `replicas = 2` matching a single machine fails here
1338/// rather than deploying one copy, because the front door would then publish a
1339/// hostname whose backend set is quietly half of what the mirror declared.
1340///
1341/// **R844-F8: the constraint arm shares its selector with the ingress
1342/// planner's.** [`CloudConfig::resolve_machines`] and
1343/// [`super::ingress::resolve_ingress_placements`] both bottom out in the same
1344/// N-selecting `select_matching` over the same `cfg.machines` slice, so the
1345/// deployer and the discovery fanout cannot pick different subsets of a
1346/// scale-N placement. The `machines = [...]` arm above needs no such
1347/// guarantee — the planner reads that literal list off the slot directly.
1348///
1349/// **R885-T14: both arms require the node to declare
1350/// [`BUNDLE_SERVING_MESH_TAG`].** Serving a bundle needs a kamaji built with
1351/// the `bundle-serving` feature and started with `--bundle-cache-dir` —
1352/// node-local startup facts a mirror author cannot see and should not have to
1353/// restate. The constraint arm gets it through
1354/// [`with_bundle_capability`], the literal arm through
1355/// [`ensure_bundle_capable`]; the ingress planner applies the same tag for the
1356/// same role (`reconciler::ingress::resolve_ingress_placements`), so the
1357/// set-for-set agreement promised above survives the extra axis.
1358pub fn resolve_bundle_machines<'a>(
1359 cfg: &'a CloudConfig,
1360 mirror: &MirrorConfig,
1361 slot: &BundleSlot,
1362 service: &str,
1363 env: &str,
1364) -> Result<Vec<&'a crate::MachineConfig>> {
1365 if !slot.machines.is_empty() {
1366 return slot
1367 .machines
1368 .iter()
1369 .map(|name| {
1370 let machine = cfg.machine(name).with_context(|| {
1371 format!(
1372 "providers.{SLOT_ROLE}.machines names {name:?}, which is not declared in \
1373 .yah/infra/machines/ (service={service}, env={env})"
1374 )
1375 })?;
1376 ensure_bundle_capable(machine, service, env)?;
1377 Ok(machine)
1378 })
1379 .collect();
1380 }
1381
1382 let required = mirror
1383 .providers
1384 .get(SLOT_ROLE)
1385 .and_then(|s| s.required())
1386 .filter(|r| !r.is_unconstrained())
1387 .with_context(|| {
1388 format!(
1389 "providers.{SLOT_ROLE} declares neither `machines = [...]` nor a constrained \
1390 `required = {{ … }}` placement — a bundle must name the nodes that serve it \
1391 (service={service}, env={env})"
1392 )
1393 })?;
1394
1395 let required = with_bundle_capability(&required);
1396
1397 cfg.resolve_machines(&required).with_context(|| {
1398 format!(
1399 "F16 placement: cannot place providers.{SLOT_ROLE}.required ({}) onto {} machine(s) \
1400 — check .yah/services/{service}/mirrors/{env}.toml against .yah/infra/machines/*.toml",
1401 required.describe(),
1402 required.replica_count(),
1403 )
1404 })
1405}
1406
1407/// The slot's declared constraint plus the capability the backend actually
1408/// needs — R885-T14.
1409///
1410/// The operator writes `required = { regions, mesh_tags, replicas }` in terms
1411/// of *where they want the bundle*; whether a matching node's kamaji has a
1412/// bundle backend at all is a fact about the node, not a preference, so it is
1413/// added here rather than asked of every mirror author. Idempotent: a mirror
1414/// that already spells [`BUNDLE_SERVING_MESH_TAG`] out by hand resolves
1415/// identically.
1416///
1417/// `mesh_tags` is the right axis because it is already an AND-ed superset check
1418/// against `MachineConfig::mesh_tags` and [`RequiredSpec::describe`] already
1419/// renders it — so a refusal reads `mesh_tags=[…,cap:bundle-serving]` with no
1420/// new field, no wire change and no schema drift. See
1421/// [`BUNDLE_SERVING_MESH_TAG`] for why a positive capability and not a taint.
1422pub fn with_bundle_capability(required: &RequiredSpec) -> RequiredSpec {
1423 let mut out = required.clone();
1424 if !out.mesh_tags.iter().any(|t| t == BUNDLE_SERVING_MESH_TAG) {
1425 out.mesh_tags.push(BUNDLE_SERVING_MESH_TAG.to_string());
1426 }
1427 out
1428}
1429
1430/// Refuse a hand-pinned `providers.bundle.machines` entry whose node does not
1431/// declare [`BUNDLE_SERVING_MESH_TAG`] — R885-T14.
1432///
1433/// The constraint arm gets this for free (an incapable node simply stops
1434/// matching); the literal arm has no matcher to fall out of, so the check is
1435/// explicit. Refusing it is deliberate rather than warning: a pin is the shape
1436/// in which the operator makes the strongest claim about a node, and the
1437/// alternative is a green `yah cloud mirror up` whose deploy leg comes back
1438/// `BackendRefused` after the build and the upload have already run.
1439fn ensure_bundle_capable(machine: &crate::MachineConfig, service: &str, env: &str) -> Result<()> {
1440 if machine
1441 .mesh_tags
1442 .iter()
1443 .any(|t| t == BUNDLE_SERVING_MESH_TAG)
1444 {
1445 return Ok(());
1446 }
1447 let name = &machine.name;
1448 bail!(
1449 "providers.{SLOT_ROLE}.machines pins {name:?}, whose declaration does not carry \
1450 `{BUNDLE_SERVING_MESH_TAG}` — that node is not declared able to serve a bundle, so \
1451 kamaji would answer the deploy with BackendRefused after the build and upload had \
1452 already run. Confirm its kamaji was built with the `bundle-serving` feature and is \
1453 started with --bundle-cache-dir plus --bundle-origin, then add \
1454 \"{BUNDLE_SERVING_MESH_TAG}\" to mesh_tags in .yah/infra/machines/{name}.toml \
1455 (service={service}, env={env})"
1456 )
1457}
1458
1459/// Desktop-side (offline) half of the bundle tier: validate the mirror's
1460/// declaration and bail with a pointer at the CLI.
1461///
1462/// The real chain — build, assemble, publish, deploy — runs at the apply layer
1463/// where [`CloudConfig`] is in hand. This exists so a desktop bring-up of a
1464/// bundle-tier mirror reports a *configuration* verdict instead of "no
1465/// reconciler wired".
1466pub struct MesofactBundleReconciler;
1467
1468impl MesofactBundleReconciler {
1469 pub fn new() -> Self {
1470 Self
1471 }
1472}
1473
1474impl Default for MesofactBundleReconciler {
1475 fn default() -> Self {
1476 Self::new()
1477 }
1478}
1479
1480#[async_trait]
1481impl Reconciler for MesofactBundleReconciler {
1482 fn kind(&self) -> &'static str {
1483 super::mesofact_static::WORKLOAD_KIND
1484 }
1485
1486 async fn up(&self, ctx: ReconcileCtx<'_>) -> Result<RunningWorkload> {
1487 let slot = BundleSlot::parse(ctx.mirror, &ctx.service.name, ctx.env)?;
1488 // Name the DECLARED shape, not a count. R746-F2 added a third shape, and
1489 // a bare `0 serve binaries` reads identically for "vanilla, resolves the
1490 // node's stock runtime" and "self-contained, builds its binary on
1491 // demand" — two different deploys.
1492 let shape = match (&slot.serve_build, slot.serve_bins.len()) {
1493 (Some(build), _) => format!(
1494 "self-contained, serve binary built by QED recipe `{}` for [{}]",
1495 build.pipeline,
1496 build.triples.join(", "),
1497 ),
1498 (None, 0) => format!(
1499 "vanilla, node resolves runtime mesofact/{}",
1500 slot.runtime_version.as_deref().unwrap_or("<caller version>"),
1501 ),
1502 (None, n) => format!("self-contained, {n} declared serve binary path(s)"),
1503 };
1504 bail!(
1505 "bundle tier validated (bucket={}, workload={}, {shape}) for service={}, \
1506 env={}, but the sync arm runs at the apply layer — deploy with \
1507 `yah cloud mirror up {} --env {}` (machine placement needs the workspace's \
1508 machine set, which a desktop bring-up does not load)",
1509 slot.bucket,
1510 slot.workload_name(&ctx.service.name),
1511 ctx.service.name,
1512 ctx.env,
1513 ctx.service.name,
1514 ctx.env,
1515 )
1516 }
1517}
1518
1519#[cfg(test)]
1520mod tests {
1521 use super::*;
1522 use crate::config::{MachineConfig, MirrorProviderSlot, MirrorShape, TopologyConfig};
1523 use std::path::PathBuf;
1524
1525 fn mirror_from(slots: BTreeMap<String, MirrorProviderSlot>) -> MirrorConfig {
1526 MirrorConfig {
1527 schema_version: 1,
1528 shape: MirrorShape::SingleMachine,
1529 providers: slots,
1530 ingress: Default::default(),
1531 ingress_machines: Vec::new(),
1532 drivers: Default::default(),
1533 asset_aliases: BTreeMap::new(),
1534 build: Default::default(),
1535 }
1536 }
1537
1538 /// Build a mirror whose `[providers.bundle]` slot is exactly `slot_toml`.
1539 fn mirror_with(slot_toml: &str) -> MirrorConfig {
1540 let slot: MirrorProviderSlot = toml::from_str(slot_toml).unwrap();
1541 let mut providers = BTreeMap::new();
1542 providers.insert(SLOT_ROLE.to_string(), slot);
1543 mirror_from(providers)
1544 }
1545
1546 /// A node that CAN serve a bundle — i.e. one that carries
1547 /// [`BUNDLE_SERVING_MESH_TAG`] (R885-T14).
1548 ///
1549 /// The tag is in the default fixture rather than opted into per test
1550 /// because every placement test here is asking "where does this land",
1551 /// which presupposes an eligible pool. Use
1552 /// [`machine_without_bundle_backend`] when the *absence* is the subject.
1553 fn machine(name: &str, region: &str) -> MachineConfig {
1554 MachineConfig {
1555 mesh_tags: vec![BUNDLE_SERVING_MESH_TAG.to_string()],
1556 ..machine_without_bundle_backend(name, region)
1557 }
1558 }
1559
1560 /// A node whose kamaji has no bundle backend: same in every other respect,
1561 /// so a test using it isolates exactly the capability axis.
1562 fn machine_without_bundle_backend(name: &str, region: &str) -> MachineConfig {
1563 MachineConfig {
1564 name: name.into(),
1565 provider: "static".into(),
1566 location: None,
1567 server_type: None,
1568 hosts_mirrors: vec![],
1569 mesh_tags: vec![],
1570 region: Some(region.into()),
1571 zone: None,
1572 arch: Some("x86_64".into()),
1573 bucket: None,
1574 vendor: None,
1575 nickname: None,
1576 legacy_hostkey_fingerprint: None,
1577 registration: Default::default(),
1578 ssh_keys: vec![],
1579 cloudflared: None,
1580 hosts_operator_bridge: false,
1581 connect: None,
1582 allocatable: None,
1583 taints: vec![],
1584 sovereign_group: None,
1585 sovereign_role: None,
1586 ingress_floating_ip: None,
1587 }
1588 }
1589
1590 fn cfg_with(machines: Vec<MachineConfig>) -> CloudConfig {
1591 CloudConfig {
1592 workspace_root: PathBuf::new(),
1593 machines,
1594 providers: vec![],
1595 machine_origins: BTreeMap::new(),
1596 provider_origins: BTreeMap::new(),
1597 recovery_measurements: BTreeMap::new(),
1598 services: BTreeMap::new(),
1599 domains: BTreeMap::new(),
1600 legacy_mirrors: vec![],
1601 workloads: vec![],
1602 topology: TopologyConfig::default(),
1603 }
1604 }
1605
1606 #[test]
1607 fn slot_declared_keys_off_the_bundle_role() {
1608 assert!(slot_declared(&mirror_with(
1609 r#"use = "cloudflare"
1610bucket = "b""#
1611 )));
1612 assert!(!slot_declared(&mirror_from(BTreeMap::new())));
1613 }
1614
1615 /// R330-B43 regression pin. This is the exact shape that froze yah.dev:
1616 /// a fully-valid `[providers.bundle]` slot whose serve binary was never
1617 /// built. `slot_declared` says yes (it only reads config), so dispatching
1618 /// on it alone handed the component to a tier that could not come up while
1619 /// taking the working static chain out of the picture. `slot_ready` is what
1620 /// the dispatch gate must ask instead.
1621 #[test]
1622 fn a_declared_slot_whose_serve_bin_is_absent_is_not_ready() {
1623 let root = tempfile::tempdir().unwrap();
1624 let mirror = mirror_with(
1625 r#"
1626use = "cloudflare"
1627bucket = "yah-dev"
1628
1629[serve_bins]
1630x86_64-unknown-linux-musl = "target/x86_64-unknown-linux-musl/release/mesofact"
1631"#,
1632 );
1633 let slot = BundleSlot::parse(&mirror, "yah-marketing", "prod").unwrap();
1634
1635 assert!(slot_declared(&mirror), "config declares the slot");
1636 assert!(
1637 !slot_ready(&slot, root.path()),
1638 "but it cannot serve — the binary does not exist"
1639 );
1640
1641 let missing = missing_bins(&slot, root.path());
1642 assert_eq!(missing.len(), 1);
1643 assert_eq!(
1644 missing[0].0, "providers.bundle.serve_bins.x86_64-unknown-linux-musl",
1645 "the label must name the exact config line to fix"
1646 );
1647 }
1648
1649 #[test]
1650 fn a_slot_becomes_ready_once_its_bins_exist() {
1651 let root = tempfile::tempdir().unwrap();
1652 let bin = root.path().join("target/x86_64-unknown-linux-musl/release");
1653 std::fs::create_dir_all(&bin).unwrap();
1654 std::fs::write(bin.join("mesofact"), b"#!/bin/sh\n").unwrap();
1655
1656 let mirror = mirror_with(
1657 r#"
1658use = "cloudflare"
1659bucket = "yah-dev"
1660
1661[serve_bins]
1662x86_64-unknown-linux-musl = "target/x86_64-unknown-linux-musl/release/mesofact"
1663"#,
1664 );
1665 let slot = BundleSlot::parse(&mirror, "yah-marketing", "prod").unwrap();
1666 assert!(slot_ready(&slot, root.path()));
1667 assert!(missing_bins(&slot, root.path()).is_empty());
1668 }
1669
1670 /// R746-F2. The shape B43 could not express: self-contained, declared, and
1671 /// buildable *from any machine* — so it is ready without anyone having a
1672 /// binary on disk, and there is no `missing:` line to print because nothing
1673 /// was ever promised to be there.
1674 #[test]
1675 fn a_serve_build_slot_is_self_contained_and_ready_with_no_binary_on_disk() {
1676 let root = tempfile::tempdir().unwrap();
1677 let mirror = mirror_with(
1678 r#"
1679use = "cloudflare"
1680bucket = "yah-dev"
1681
1682[serve_build]
1683pipeline = "mesofact-musl"
1684binary = "mesofact"
1685triples = ["x86_64-unknown-linux-musl"]
1686"#,
1687 );
1688 let slot = BundleSlot::parse(&mirror, "yah-marketing", "prod").unwrap();
1689
1690 let build = slot.serve_build.as_ref().expect("serve_build parsed");
1691 assert_eq!(build.pipeline, "mesofact-musl");
1692 assert_eq!(build.binary, "mesofact");
1693 assert_eq!(build.triples, vec!["x86_64-unknown-linux-musl".to_string()]);
1694
1695 assert!(slot.is_self_contained(), "declared shape, not disk state");
1696 assert!(slot_ready(&slot, root.path()));
1697 assert!(missing_bins(&slot, root.path()).is_empty());
1698 }
1699
1700 /// R746-F2 verify #1, at the only layer that can pin it offline: a vanilla
1701 /// slot carries no build declaration at all, so the sync has nothing to
1702 /// dispatch. The cheapness of the vanilla path is structural, not a
1703 /// heuristic someone has to keep true.
1704 #[test]
1705 fn a_vanilla_slot_declares_no_build_so_a_sync_has_nothing_to_dispatch() {
1706 let mirror = mirror_with(
1707 r#"
1708use = "cloudflare"
1709bucket = "yah-dev"
1710runtime_version = "0.8.22"
1711"#,
1712 );
1713 let slot = BundleSlot::parse(&mirror, "yah-marketing", "prod").unwrap();
1714 assert!(slot.serve_build.is_none());
1715 assert!(slot.serve_bins.is_empty());
1716 assert!(!slot.is_self_contained());
1717 assert_eq!(slot.runtime_version.as_deref(), Some("0.8.22"));
1718 }
1719
1720 /// The shape must stay DECLARED, never derived — so the two ways of naming
1721 /// a serve binary are mutually exclusive rather than one falling back to
1722 /// the other. A fallback would make the deployed binary a function of the
1723 /// syncing machine's disk.
1724 #[test]
1725 fn serve_bins_and_serve_build_together_are_refused() {
1726 let mirror = mirror_with(
1727 r#"
1728use = "cloudflare"
1729bucket = "yah-dev"
1730
1731[serve_bins]
1732x86_64-unknown-linux-musl = "some/path/mesofact"
1733
1734[serve_build]
1735pipeline = "mesofact-musl"
1736binary = "mesofact"
1737triples = ["x86_64-unknown-linux-musl"]
1738"#,
1739 );
1740 let err = BundleSlot::parse(&mirror, "yah-marketing", "prod").unwrap_err();
1741 let msg = err.to_string();
1742 assert!(msg.contains("BOTH `serve_bins` and `serve_build`"), "{msg}");
1743 }
1744
1745 /// Each field is load-bearing, so each absence is refused by name rather
1746 /// than defaulted into a build that produces nothing.
1747 #[test]
1748 fn a_serve_build_missing_a_field_is_refused_naming_the_coordinate() {
1749 let cases = [
1750 (
1751 r#"[serve_build]
1752binary = "mesofact"
1753triples = ["x86_64-unknown-linux-musl"]"#,
1754 "serve_build.pipeline",
1755 ),
1756 (
1757 r#"[serve_build]
1758pipeline = "mesofact-musl"
1759triples = ["x86_64-unknown-linux-musl"]"#,
1760 "serve_build.binary",
1761 ),
1762 (
1763 r#"[serve_build]
1764pipeline = "mesofact-musl"
1765binary = "mesofact""#,
1766 "serve_build.triples",
1767 ),
1768 (
1769 r#"[serve_build]
1770pipeline = "mesofact-musl"
1771binary = "mesofact"
1772triples = []"#,
1773 "serve_build.triples is empty",
1774 ),
1775 ];
1776 for (fragment, expected) in cases {
1777 let mirror = mirror_with(&format!(
1778 "use = \"cloudflare\"\nbucket = \"yah-dev\"\n\n{fragment}\n"
1779 ));
1780 let err = BundleSlot::parse(&mirror, "yah-marketing", "prod").unwrap_err();
1781 let msg = format!("{err:#}");
1782 assert!(
1783 msg.contains(expected),
1784 "expected {expected:?} in error, got: {msg}"
1785 );
1786 }
1787 }
1788
1789 /// A declared feed tier is part of "can it serve" — R330-F31 stages the
1790 /// fetcher as a sidecar, so a missing feed_bin strands the feed tier the
1791 /// same way a missing serve_bin strands the server.
1792 #[test]
1793 fn a_missing_feed_bin_also_blocks_readiness() {
1794 let root = tempfile::tempdir().unwrap();
1795 let bin = root.path().join("target/musl");
1796 std::fs::create_dir_all(&bin).unwrap();
1797 std::fs::write(bin.join("mesofact"), b"x").unwrap();
1798
1799 let mirror = mirror_with(
1800 r#"
1801use = "cloudflare"
1802bucket = "yah-dev"
1803
1804[serve_bins]
1805x86_64-unknown-linux-musl = "target/musl/mesofact"
1806
1807[revalidate]
1808routes = ["/releases"]
1809feeds = ["releases"]
1810
1811[revalidate.feed_bins]
1812x86_64-unknown-linux-musl = "target/musl/almanac-feed"
1813"#,
1814 );
1815 let slot = BundleSlot::parse(&mirror, "yah-marketing", "prod").unwrap();
1816 let missing = missing_bins(&slot, root.path());
1817 assert_eq!(missing.len(), 1, "only the feed binary is absent");
1818 assert!(missing[0].0.contains("revalidate.feed_bins"));
1819 assert!(!slot_ready(&slot, root.path()));
1820 }
1821
1822 /// A vanilla-runtime slot declares no binaries at all. That is a different
1823 /// shape, not a half-built one, so it stays "ready" here and fails later
1824 /// with its own specific message rather than being silently downgraded.
1825 #[test]
1826 fn a_vanilla_slot_declaring_no_bins_is_ready() {
1827 let root = tempfile::tempdir().unwrap();
1828 let mirror = mirror_with(
1829 r#"use = "cloudflare"
1830bucket = "b""#,
1831 );
1832 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
1833 assert!(!slot.is_self_contained());
1834 assert!(slot_ready(&slot, root.path()));
1835 }
1836
1837 #[test]
1838 fn parses_a_self_contained_keep_alive_slot() {
1839 let mirror = mirror_with(
1840 r#"
1841use = "cloudflare"
1842bucket = "yah-dev-bundles"
1843machines = ["us-east-001"]
1844name = "yah-marketing"
1845
1846[serve_bins]
1847x86_64-unknown-linux-musl = "target/x86_64-unknown-linux-musl/release/mesofact-serve"
1848"#,
1849 );
1850 let slot = BundleSlot::parse(&mirror, "yah-marketing", "ha").unwrap();
1851 assert_eq!(slot.bucket, "yah-dev-bundles");
1852 assert_eq!(slot.machines, vec!["us-east-001".to_string()]);
1853 assert_eq!(slot.workload_name("yah-marketing"), "yah-marketing");
1854 assert!(slot.is_self_contained());
1855 assert_eq!(slot.lifecycle, BundleLifecycle::KeepAlive);
1856 }
1857
1858 #[test]
1859 fn workload_name_falls_back_to_the_service_name() {
1860 let mirror = mirror_with(
1861 r#"use = "cloudflare"
1862bucket = "b""#,
1863 );
1864 let slot = BundleSlot::parse(&mirror, "scrabcake", "ha").unwrap();
1865 assert_eq!(slot.workload_name("scrabcake"), "scrabcake");
1866 assert!(!slot.is_self_contained());
1867 }
1868
1869 #[test]
1870 fn on_demand_takes_the_default_idle_ttl() {
1871 let mirror = mirror_with(
1872 r#"use = "cloudflare"
1873bucket = "b"
1874lifecycle = "on-demand""#,
1875 );
1876 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
1877 assert_eq!(
1878 slot.lifecycle,
1879 BundleLifecycle::OnDemand {
1880 idle_ttl: Millis::from_ms(DEFAULT_IDLE_TTL_MS)
1881 }
1882 );
1883 }
1884
1885 #[test]
1886 fn on_demand_honors_an_explicit_idle_ttl() {
1887 let mirror = mirror_with(
1888 r#"use = "cloudflare"
1889bucket = "b"
1890lifecycle = "on-demand"
1891idle_ttl_ms = 15000"#,
1892 );
1893 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
1894 assert_eq!(
1895 slot.lifecycle,
1896 BundleLifecycle::OnDemand {
1897 idle_ttl: Millis::from_ms(15_000)
1898 }
1899 );
1900 }
1901
1902 /// An idle TTL on a keep-alive bundle is a config mistake that would
1903 /// otherwise be silently ignored — the process is never reaped.
1904 #[test]
1905 fn idle_ttl_on_a_keep_alive_slot_is_rejected() {
1906 let mirror = mirror_with(
1907 r#"use = "cloudflare"
1908bucket = "b"
1909idle_ttl_ms = 15000"#,
1910 );
1911 let err = BundleSlot::parse(&mirror, "s", "e")
1912 .unwrap_err()
1913 .to_string();
1914 assert!(err.contains("idle_ttl_ms"), "{err}");
1915 assert!(err.contains("on-demand"), "{err}");
1916 }
1917
1918 #[test]
1919 fn unknown_lifecycle_names_the_legal_values() {
1920 let mirror = mirror_with(
1921 r#"use = "cloudflare"
1922bucket = "b"
1923lifecycle = "serverless""#,
1924 );
1925 let err = BundleSlot::parse(&mirror, "s", "e")
1926 .unwrap_err()
1927 .to_string();
1928 assert!(err.contains("keep-alive"), "{err}");
1929 assert!(err.contains("on-demand"), "{err}");
1930 }
1931
1932 /// R599-F12: the declared serving port reaches the workload spec. Without
1933 /// it every bundle rides kamaji's node-wide default, so a node can host
1934 /// exactly one.
1935 #[test]
1936 fn a_declared_port_reaches_the_serve_bundle() {
1937 let slot = BundleSlot::parse(
1938 &mirror_with(
1939 r#"use = "cloudflare"
1940bucket = "b"
1941port = 8081"#,
1942 ),
1943 "s",
1944 "e",
1945 )
1946 .unwrap();
1947 assert_eq!(slot.port, Some(8081));
1948 assert_eq!(
1949 slot.serve_bundle("a".repeat(64).as_str(), "self", BTreeMap::new())
1950 .port,
1951 Some(8081)
1952 );
1953
1954 // Absent → kamaji's node default, the pre-R599-F12 behaviour.
1955 let bare = BundleSlot::parse(
1956 &mirror_with(
1957 r#"use = "cloudflare"
1958bucket = "b""#,
1959 ),
1960 "s",
1961 "e",
1962 )
1963 .unwrap();
1964 assert_eq!(bare.port, None);
1965 assert_eq!(
1966 bare.serve_bundle("a".repeat(64).as_str(), "self", BTreeMap::new())
1967 .port,
1968 None
1969 );
1970 }
1971
1972 /// R870-B6: the store the mirror publishes to reaches the workload, so the
1973 /// node fetches from it rather than from whatever the *node* was pointed at.
1974 /// A second tenant on the fleet publishes to its own bucket; before this,
1975 /// its deploy passed admission and then failed to materialize.
1976 #[test]
1977 fn a_declared_origin_reaches_the_serve_bundle() {
1978 let slot = BundleSlot::parse(
1979 &mirror_with(
1980 r#"use = "cloudflare"
1981bucket = "noisetable-marketing"
1982origin = "https://cdn.noisetable.com""#,
1983 ),
1984 "noisetable-marketing",
1985 "prod",
1986 )
1987 .unwrap();
1988 assert_eq!(slot.origin.as_deref(), Some("https://cdn.noisetable.com"));
1989 assert_eq!(
1990 slot.serve_bundle("a".repeat(64).as_str(), "self", BTreeMap::new())
1991 .origin
1992 .as_deref(),
1993 Some("https://cdn.noisetable.com")
1994 );
1995 }
1996
1997 /// The single-tenant regression, asserted at the wire type rather than by
1998 /// watching yah.dev stay up: a mirror that declares no `origin` — which is
1999 /// every yah-owned mirror — produces the same spec it did before R870-B6,
2000 /// so the node keeps using `KAMAJI_BUNDLE_ORIGIN` and needs no edit.
2001 #[test]
2002 fn no_declared_origin_leaves_the_node_wide_one_in_charge() {
2003 let slot = BundleSlot::parse(
2004 &mirror_with(
2005 r#"use = "cloudflare"
2006bucket = "yah-dev""#,
2007 ),
2008 "yah-marketing",
2009 "prod",
2010 )
2011 .unwrap();
2012 assert_eq!(slot.origin, None);
2013 assert_eq!(
2014 slot.serve_bundle("a".repeat(64).as_str(), "self", BTreeMap::new())
2015 .origin,
2016 None
2017 );
2018 }
2019
2020 /// A trailing slash is trimmed at parse rather than at three consumers:
2021 /// `HttpReadOnlyObjectStore` joins keys onto this value, and
2022 /// `https://cdn.x//blobs/…` is a different object to an S3-shaped origin.
2023 #[test]
2024 fn a_trailing_slash_on_the_origin_is_trimmed() {
2025 let slot = BundleSlot::parse(
2026 &mirror_with(
2027 r#"use = "cloudflare"
2028bucket = "b"
2029origin = "https://cdn.example.com/""#,
2030 ),
2031 "s",
2032 "e",
2033 )
2034 .unwrap();
2035 assert_eq!(slot.origin.as_deref(), Some("https://cdn.example.com"));
2036 }
2037
2038 /// A bucket name (or bare hostname) where an origin belongs is refused
2039 /// offline. Accepted, it would deploy clean and fail on the node at
2040 /// materialize time — the far side of the feedback loop.
2041 #[test]
2042 fn an_origin_without_a_scheme_is_refused_naming_the_shape() {
2043 for bad in ["cdn.noisetable.com", "noisetable-marketing"] {
2044 let err = BundleSlot::parse(
2045 &mirror_with(&format!(
2046 "use = \"cloudflare\"\nbucket = \"b\"\norigin = {bad:?}"
2047 )),
2048 "s",
2049 "e",
2050 )
2051 .unwrap_err()
2052 .to_string();
2053 assert!(err.contains("scheme"), "{err}");
2054 assert!(err.contains(bad), "{err}");
2055 }
2056 }
2057
2058 /// R556-B14: a misspelled key fails the parse naming itself, rather than
2059 /// deploying a wrong-but-plausible workload.
2060 ///
2061 /// `prot = 8081` is the motivating instance: it parses clean today, the
2062 /// port falls back to kamaji's node-wide default, and post-R599-F12 that
2063 /// default is whatever OTHER bundle already holds 8080 on the node. The
2064 /// operator sees the wrong site served, with nothing in any log naming the
2065 /// typo.
2066 #[test]
2067 fn an_unknown_slot_key_is_rejected_naming_the_key() {
2068 let err = BundleSlot::parse(
2069 &mirror_with(
2070 r#"use = "cloudflare"
2071bucket = "b"
2072prot = 8081"#,
2073 ),
2074 "yah-marketing",
2075 "prod",
2076 )
2077 .unwrap_err()
2078 .to_string();
2079 assert!(err.contains("prot"), "the error must name the typo: {err}");
2080 assert!(
2081 err.contains("did you mean `port`"),
2082 "one transposed character is the realistic failure — suggest the \
2083 fix rather than making the operator diff the docs: {err}"
2084 );
2085 }
2086
2087 /// The suggester's distance metric counts a transposition as ONE edit.
2088 /// Plain Levenshtein scores `prot`→`port` at 2, which is far enough away
2089 /// that any threshold tight enough to suppress nonsense suggestions would
2090 /// also suppress the single typo this ticket was filed about.
2091 #[test]
2092 fn the_key_suggester_treats_a_transposition_as_one_edit() {
2093 assert_eq!(edit_distance("prot", "port"), 1);
2094 assert_eq!(edit_distance("bukcet", "bucket"), 1);
2095 assert_eq!(nearest_slot_key("prot"), Some("port"));
2096 assert_eq!(nearest_slot_key("bucket"), Some("bucket"));
2097 assert_eq!(nearest_slot_key("zzzzzzzzzzzzzz"), None);
2098 }
2099
2100 /// No suggestion when nothing is close. A confidently wrong hint sends the
2101 /// operator to edit a line that was never the problem.
2102 #[test]
2103 fn an_unrecognizable_slot_key_is_rejected_without_a_bogus_suggestion() {
2104 let err = BundleSlot::parse(
2105 &mirror_with(
2106 r#"use = "cloudflare"
2107bucket = "b"
2108ingress_tunnel_hostname = "analytics.yah.dev""#,
2109 ),
2110 "yah-analytics",
2111 "prod",
2112 )
2113 .unwrap_err()
2114 .to_string();
2115 assert!(err.contains("ingress_tunnel_hostname"), "{err}");
2116 assert!(!err.contains("did you mean"), "{err}");
2117 }
2118
2119 /// R556-B14's regression criterion: the allowed set is the UNION of every
2120 /// consumer's reads, not just `BundleSlot::parse`'s. `plan_ingress` reads
2121 /// `machine` / `machines` / `port` / `upstream_host` off this same table
2122 /// and `MirrorProviderSlot::required` reads `required` — scoping the set to
2123 /// one consumer would reject the live yah-marketing mirror, which carries
2124 /// `upstream_host`.
2125 #[test]
2126 fn keys_read_by_other_consumers_of_this_slot_are_allowed() {
2127 // Every non-comment key of .yah/services/yah-marketing/mirrors/cloud.toml's
2128 // [providers.bundle] block, as of R556-B14.
2129 let slot = BundleSlot::parse(
2130 &mirror_with(
2131 r#"use = "cloudflare"
2132verify_serving = false
2133bucket = "yah-dev"
2134name = "yah-marketing"
2135machines = ["us-east-001"]
2136port = 8080
2137zone = "yah.dev"
2138upstream_host = "100.64.0.3"
2139lifecycle = "keep-alive"
2140runtime_version = "0.8.23"
2141
2142[revalidate]
2143routes = ["/releases", "/issues"]
2144mirror_key_env = "YAH_MARKETING_MIRROR_KEY"
2145feeds = ["releases", "yah-desktop"]
2146feed_interval_secs = 5
2147feed_runtime = "almanac-feed/0.8.22""#,
2148 ),
2149 "yah-marketing",
2150 "prod",
2151 )
2152 .unwrap();
2153 assert_eq!(slot.bucket, "yah-dev");
2154 assert_eq!(slot.port, Some(8080));
2155 assert!(!slot.verify_serving);
2156
2157 // …and the F16 placement form, whose `required` is read through the
2158 // slot rather than by `parse`.
2159 BundleSlot::parse(
2160 &mirror_with(
2161 r#"use = "cloudflare"
2162bucket = "yah-dev"
2163
2164[required]
2165regions = ["us-east"]"#,
2166 ),
2167 "s",
2168 "e",
2169 )
2170 .unwrap();
2171
2172 // …and R844-F5's portless shape: `fronted = true` with no `port`. This
2173 // is the same union rule one ticket later — the key is read only by
2174 // `plan_ingress`, but it is declared on THIS table, so rejecting it here
2175 // would have made the pin deletion R844-F5 exists to enable fail the
2176 // apply rather than land as a no-op.
2177 let portless = BundleSlot::parse(
2178 &mirror_with(
2179 r#"use = "cloudflare"
2180bucket = "yah-dev"
2181zone = "yah.dev"
2182fronted = true"#,
2183 ),
2184 "s",
2185 "e",
2186 )
2187 .unwrap();
2188 assert_eq!(portless.port, None);
2189 }
2190
2191 /// R556-T12: `[providers.bundle.env]` parses into source URIs, kept
2192 /// verbatim. Resolution is deliberately NOT done here — this crate is
2193 /// offline by construction and only the syncing machine holds the vault.
2194 #[test]
2195 fn env_sources_are_parsed_verbatim_and_not_resolved() {
2196 let slot = BundleSlot::parse(
2197 &mirror_with(
2198 r#"use = "cloudflare"
2199bucket = "b"
2200
2201[env]
2202ANALYTICS_R2_ACCESS_KEY = "vault:cloudflare-r2-access-key-id"
2203ANALYTICS_R2_SECRET_KEY = "vault:cloudflare-r2-secret-key|env:R2_SECRET"
2204ANALYTICS_R2_BUCKET = "yah-analytics""#,
2205 ),
2206 "s",
2207 "e",
2208 )
2209 .unwrap();
2210 assert_eq!(slot.env.len(), 3);
2211 assert_eq!(
2212 slot.env.get("ANALYTICS_R2_ACCESS_KEY").map(String::as_str),
2213 Some("vault:cloudflare-r2-access-key-id"),
2214 "the SOURCE is stored, never a resolved secret — this struct is \
2215 parsed on any machine and printed by diagnostics",
2216 );
2217 assert_eq!(
2218 slot.env.get("ANALYTICS_R2_SECRET_KEY").map(String::as_str),
2219 Some("vault:cloudflare-r2-secret-key|env:R2_SECRET"),
2220 "a pipe-joined fallback chain survives parsing intact",
2221 );
2222 assert_eq!(
2223 slot.env.get("ANALYTICS_R2_BUCKET").map(String::as_str),
2224 Some("yah-analytics"),
2225 "a bare literal is a legitimate non-secret source",
2226 );
2227
2228 // Absent block → empty, and the serve bundle carries whatever the
2229 // deploy resolved (nothing, here).
2230 let bare = BundleSlot::parse(
2231 &mirror_with("use = \"cloudflare\"\nbucket = \"b\""),
2232 "s",
2233 "e",
2234 )
2235 .unwrap();
2236 assert!(bare.env.is_empty());
2237 }
2238
2239 /// The resolved env reaches the workload payload — the leg that was missing
2240 /// entirely (R556-T12). Before it, `MesofactServeBundle` had nowhere to put
2241 /// credentials, so kamaji forked the serve process with an empty
2242 /// environment and an SSR route reading a private source 500'd per request.
2243 #[test]
2244 fn resolved_env_reaches_the_serve_bundle() {
2245 let slot = BundleSlot::parse(
2246 &mirror_with(
2247 r#"use = "cloudflare"
2248bucket = "b"
2249
2250[env]
2251ANALYTICS_R2_ACCESS_KEY = "vault:cloudflare-r2-access-key-id""#,
2252 ),
2253 "s",
2254 "e",
2255 )
2256 .unwrap();
2257
2258 let mut resolved = BTreeMap::new();
2259 resolved.insert("ANALYTICS_R2_ACCESS_KEY".to_string(), "AKIA".to_string());
2260 let sb = slot.serve_bundle(&"a".repeat(64), "self", resolved);
2261
2262 assert_eq!(
2263 sb.env.get("ANALYTICS_R2_ACCESS_KEY").map(String::as_str),
2264 Some("AKIA"),
2265 "the node receives the VALUE; a keystore slot name must never \
2266 cross the wire",
2267 );
2268 }
2269
2270 /// An env entry that is not a usable source string must fail the parse.
2271 /// The whole point of the slot is that a credential problem surfaces at
2272 /// sync, in milliseconds, rather than as a per-request 500 on a node.
2273 #[test]
2274 fn an_unusable_env_source_is_rejected() {
2275 for bad in [
2276 "[env]\nFOO = \"\"",
2277 "[env]\nFOO = \" \"",
2278 "[env]\nFOO = 8081",
2279 "env = \"vault:x\"",
2280 ] {
2281 let toml = format!("use = \"cloudflare\"\nbucket = \"b\"\n{bad}");
2282 let err = BundleSlot::parse(&mirror_with(&toml), "yah-marketing", "ha")
2283 .unwrap_err()
2284 .to_string();
2285 assert!(err.contains("env"), "{bad}: {err}");
2286 }
2287 }
2288
2289 /// A port typo must fail the parse, not silently fall back to 8080 — that
2290 /// fallback would land the workload on whatever bundle already holds the
2291 /// default port, and surface as the wrong site being served.
2292 #[test]
2293 fn an_unusable_port_is_rejected_rather_than_defaulted() {
2294 for bad in ["port = 0", "port = 70000", r#"port = "8081""#] {
2295 let toml = format!("use = \"cloudflare\"\nbucket = \"b\"\n{bad}");
2296 let err = BundleSlot::parse(&mirror_with(&toml), "yah-marketing", "ha")
2297 .unwrap_err()
2298 .to_string();
2299 assert!(err.contains("port"), "{bad}: {err}");
2300 }
2301 }
2302
2303 // ── serving verification (R703-T7) ──────────────────────────────────────
2304
2305 /// The check is on by default and probes the service's own domain, so a
2306 /// mirror that says nothing about it still gets verified.
2307 #[test]
2308 fn serving_verification_is_on_by_default_and_targets_the_service_domain() {
2309 let slot = BundleSlot::parse(
2310 &mirror_with(
2311 r#"use = "cloudflare"
2312bucket = "b""#,
2313 ),
2314 "yah-marketing",
2315 "prod",
2316 )
2317 .unwrap();
2318 assert!(slot.verify_serving);
2319 assert_eq!(slot.zone, None);
2320 assert_eq!(slot.serving_zone("yah.dev"), "yah.dev");
2321 }
2322
2323 #[test]
2324 fn an_explicit_zone_overrides_the_service_domain() {
2325 let slot = BundleSlot::parse(
2326 &mirror_with(
2327 r#"use = "cloudflare"
2328bucket = "b"
2329zone = "staging.yah.dev""#,
2330 ),
2331 "yah-marketing",
2332 "prod",
2333 )
2334 .unwrap();
2335 assert_eq!(slot.serving_zone("yah.dev"), "staging.yah.dev");
2336 }
2337
2338 #[test]
2339 fn verify_serving_can_be_switched_off_for_an_in_flight_migration() {
2340 let slot = BundleSlot::parse(
2341 &mirror_with(
2342 r#"use = "cloudflare"
2343bucket = "b"
2344verify_serving = false"#,
2345 ),
2346 "s",
2347 "e",
2348 )
2349 .unwrap();
2350 assert!(!slot.verify_serving);
2351 }
2352
2353 /// `verify_serving = "false"` reading as *enabled* would leave an operator
2354 /// certain they had silenced a check that then fails their apply.
2355 #[test]
2356 fn a_non_boolean_verify_serving_is_rejected_rather_than_defaulted() {
2357 let err = BundleSlot::parse(
2358 &mirror_with(
2359 r#"use = "cloudflare"
2360bucket = "b"
2361verify_serving = "false""#,
2362 ),
2363 "yah-marketing",
2364 "prod",
2365 )
2366 .unwrap_err()
2367 .to_string();
2368 assert!(err.contains("verify_serving"), "{err}");
2369 assert!(err.contains("boolean"), "{err}");
2370 }
2371
2372 #[test]
2373 fn a_slot_without_a_bucket_names_the_file_to_edit() {
2374 let mirror = mirror_with(r#"use = "cloudflare""#);
2375 let err = BundleSlot::parse(&mirror, "yah-marketing", "ha")
2376 .unwrap_err()
2377 .to_string();
2378 assert!(err.contains("bucket"), "{err}");
2379 assert!(
2380 err.contains(".yah/services/yah-marketing/mirrors/ha.toml"),
2381 "{err}"
2382 );
2383 }
2384
2385 #[test]
2386 fn explicit_machines_resolve_in_declaration_order() {
2387 let mirror = mirror_with(
2388 r#"use = "cloudflare"
2389bucket = "b"
2390machines = ["us-south-001", "us-east-001"]"#,
2391 );
2392 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2393 let cfg = cfg_with(vec![
2394 machine("us-east-001", "us-east"),
2395 machine("us-south-001", "us-south"),
2396 ]);
2397 let resolved = resolve_bundle_machines(&cfg, &mirror, &slot, "s", "e").unwrap();
2398 let names: Vec<_> = resolved.iter().map(|m| m.name.as_str()).collect();
2399 assert_eq!(names, vec!["us-south-001", "us-east-001"]);
2400 }
2401
2402 #[test]
2403 fn an_undeclared_machine_is_an_error_not_a_skip() {
2404 let mirror = mirror_with(
2405 r#"use = "cloudflare"
2406bucket = "b"
2407machines = ["us-west-999"]"#,
2408 );
2409 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2410 let cfg = cfg_with(vec![machine("us-east-001", "us-east")]);
2411 let err = resolve_bundle_machines(&cfg, &mirror, &slot, "s", "e")
2412 .unwrap_err()
2413 .to_string();
2414 assert!(err.contains("us-west-999"), "{err}");
2415 assert!(err.contains(".yah/infra/machines/"), "{err}");
2416 }
2417
2418 #[test]
2419 fn falls_back_to_f16_required_placement() {
2420 let mirror = mirror_with(
2421 r#"use = "cloudflare"
2422bucket = "b"
2423required = { regions = ["us-east"] }"#,
2424 );
2425 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2426 let cfg = cfg_with(vec![
2427 machine("us-east-001", "us-east"),
2428 machine("us-south-001", "us-south"),
2429 ]);
2430 let resolved = resolve_bundle_machines(&cfg, &mirror, &slot, "s", "e").unwrap();
2431 assert_eq!(resolved.len(), 1);
2432 assert_eq!(resolved[0].name, "us-east-001");
2433 }
2434
2435 /// R844-F8: a constraint with `replicas = N` places N machines, and the
2436 /// ingress planner's resolver picks the SAME N.
2437 ///
2438 /// The set-for-set half is the assertion that matters. Both sides returning
2439 /// two while disagreeing about *which* two aims the discovery fanout at a
2440 /// node the bundle was never deployed to, and the front door then renders a
2441 /// subset of the backends with every line in the mirror still reading
2442 /// correctly. They agree here because they are one selector over one
2443 /// candidate slice, not two implementations that happen to match.
2444 #[test]
2445 fn a_replica_count_places_n_machines_and_the_ingress_planner_picks_the_same_n() {
2446 let mirror = mirror_with(
2447 r#"use = "cloudflare"
2448bucket = "b"
2449zone = "scaled.yah.dev"
2450port = 8080
2451required = { regions = ["us-east"], replicas = 2 }"#,
2452 );
2453 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2454 let cfg = cfg_with(vec![
2455 machine("us-east-001", "us-east"),
2456 machine("us-east-002", "us-east"),
2457 machine("us-east-003", "us-east"),
2458 machine("us-south-001", "us-south"),
2459 ]);
2460
2461 let deployed: Vec<&str> = resolve_bundle_machines(&cfg, &mirror, &slot, "s", "e")
2462 .unwrap()
2463 .iter()
2464 .map(|m| m.name.as_str())
2465 .collect();
2466 assert_eq!(
2467 deployed,
2468 vec!["us-east-001", "us-east-002"],
2469 "two asked for, two placed — NOT the three the constraint matches, or \
2470 adding a box to the fleet would scale a production front door"
2471 );
2472
2473 let planned = super::super::ingress::resolve_ingress_placements(&cfg.machines, &mirror)
2474 .unwrap()
2475 .remove("bundle")
2476 .expect("the constraint slot resolves for the planner too");
2477 assert_eq!(planned, deployed, "set for set, not merely in count");
2478 }
2479
2480 /// Never a partial placement. One of two reported as success is the
2481 /// failure that looks like it worked.
2482 #[test]
2483 fn fewer_matches_than_replicas_fails_the_deploy_resolver() {
2484 let mirror = mirror_with(
2485 r#"use = "cloudflare"
2486bucket = "b"
2487required = { regions = ["us-east"], replicas = 2 }"#,
2488 );
2489 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2490 let cfg = cfg_with(vec![
2491 machine("us-east-001", "us-east"),
2492 machine("us-south-001", "us-south"),
2493 ]);
2494 // `{:#}` — the shortfall is the *source* of the placement failure, and
2495 // the outer context only names the constraint and the count wanted.
2496 let err = format!(
2497 "{:#}",
2498 resolve_bundle_machines(&cfg, &mirror, &slot, "s", "e").unwrap_err()
2499 );
2500 assert!(err.contains("onto 2 machine(s)"), "{err}");
2501 assert!(err.contains("only 1 of 2"), "{err}");
2502 assert!(err.contains("required.regions=[us-east]"), "{err}");
2503 assert!(
2504 err.contains("us-south-001"),
2505 "names the pool it searched: {err}"
2506 );
2507 }
2508
2509 /// Publishing a bundle no node serves is the silent failure this guards.
2510 #[test]
2511 fn no_placement_at_all_is_rejected() {
2512 let mirror = mirror_with(
2513 r#"use = "cloudflare"
2514bucket = "b""#,
2515 );
2516 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2517 let cfg = cfg_with(vec![machine("us-east-001", "us-east")]);
2518 let err = resolve_bundle_machines(&cfg, &mirror, &slot, "s", "e")
2519 .unwrap_err()
2520 .to_string();
2521 assert!(err.contains("machines"), "{err}");
2522 assert!(err.contains("required"), "{err}");
2523 }
2524
2525 #[test]
2526 fn serve_bundle_carries_the_manifest_runtime_verbatim() {
2527 let mirror = mirror_with(
2528 r#"use = "cloudflare"
2529bucket = "b""#,
2530 );
2531 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2532 let digest = "a".repeat(64);
2533 let sb = slot.serve_bundle(&digest, "mesofact/0.8.20", BTreeMap::new());
2534 assert_eq!(sb.digest.0, digest);
2535 assert_eq!(sb.runtime, "mesofact/0.8.20");
2536 assert_eq!(sb.lifecycle, BundleLifecycle::KeepAlive);
2537 }
2538
2539 // ── revalidate receiver parsing (R330-F12) ──────────────────────────────
2540
2541 #[test]
2542 fn parses_revalidate_slot_with_routes_and_mirror_key_env() {
2543 let mirror = mirror_with(
2544 r#"use = "cloudflare"
2545bucket = "b"
2546machines = ["us-east-001"]
2547
2548[revalidate]
2549routes = ["/releases"]
2550mirror_key_env = "YAH_MARKETING_MIRROR_KEY"
2551"#,
2552 );
2553 let slot = BundleSlot::parse(&mirror, "yah-marketing", "prod").unwrap();
2554 let rv = slot.revalidate.expect("revalidate slot should parse");
2555 assert_eq!(rv.routes, vec!["/releases"]);
2556 assert_eq!(
2557 rv.mirror_key_env.as_deref(),
2558 Some("YAH_MARKETING_MIRROR_KEY")
2559 );
2560 assert!(rv.publish_config.is_none());
2561 }
2562
2563 #[test]
2564 fn parses_revalidate_with_custom_publish_config() {
2565 let mirror = mirror_with(
2566 r#"use = "cloudflare"
2567bucket = "b"
2568machines = ["us-east-001"]
2569
2570[revalidate]
2571routes = ["/releases", "/downloads"]
2572publish_config = "custom-mesofact.config.toml"
2573"#,
2574 );
2575 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2576 let rv = slot.revalidate.unwrap();
2577 assert_eq!(rv.routes.len(), 2);
2578 assert_eq!(
2579 rv.publish_config.unwrap(),
2580 PathBuf::from("custom-mesofact.config.toml")
2581 );
2582 assert!(rv.mirror_key_env.is_none());
2583 }
2584
2585 #[test]
2586 fn no_revalidate_when_section_absent() {
2587 let mirror = mirror_with(
2588 r#"use = "cloudflare"
2589bucket = "b""#,
2590 );
2591 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2592 assert!(slot.revalidate.is_none());
2593 }
2594
2595 #[test]
2596 fn revalidate_with_empty_routes_is_open_allowlist() {
2597 let mirror = mirror_with(
2598 r#"use = "cloudflare"
2599bucket = "b"
2600
2601[revalidate]
2602mirror_key_env = "BEARER"
2603"#,
2604 );
2605 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2606 let rv = slot.revalidate.unwrap();
2607 assert!(rv.routes.is_empty());
2608 assert_eq!(rv.mirror_key_env.as_deref(), Some("BEARER"));
2609 }
2610
2611 #[test]
2612 fn to_workload_payload_maps_fields() {
2613 let slot = RevalidateSlot {
2614 routes: vec!["/releases".into()],
2615 mirror_key_env: Some("MY_KEY".into()),
2616 publish_config: Some(PathBuf::from("cfg.toml")),
2617 ..bare_revalidate_slot()
2618 };
2619 let mut env = BTreeMap::new();
2620 env.insert("MESOFACT_S3_ACCESS_KEY_ID".into(), "ak".into());
2621 env.insert("MESOFACT_MIRROR_KEY".into(), "bearer1".into());
2622 let payload = slot.to_workload_payload(env.clone(), vec![], None, vec![]);
2623 assert_eq!(payload.routes, vec!["/releases"]);
2624 assert_eq!(payload.publish_config, "cfg.toml");
2625 assert_eq!(payload.mirror_key_env.as_deref(), Some("MY_KEY"));
2626 assert_eq!(payload.env.get("MESOFACT_S3_ACCESS_KEY_ID").unwrap(), "ak");
2627 assert_eq!(payload.env.get("MESOFACT_MIRROR_KEY").unwrap(), "bearer1");
2628 }
2629
2630 #[test]
2631 fn to_workload_payload_defaults_publish_config() {
2632 let payload = bare_revalidate_slot().to_workload_payload(BTreeMap::new(), vec![], None, vec![]);
2633 assert_eq!(payload.publish_config, "mesofact.config.toml");
2634 assert!(payload.mirror_key_env.is_none());
2635 assert!(payload.routes.is_empty());
2636 }
2637
2638 // ── Feed-fetch tier (R330-F31) ──────────────────────────────────────────
2639
2640 fn bare_revalidate_slot() -> RevalidateSlot {
2641 RevalidateSlot {
2642 routes: vec![],
2643 mirror_key_env: None,
2644 publish_config: None,
2645 feeds: vec![],
2646 feed_interval_secs: None,
2647 feed_bins: BTreeMap::new(),
2648 feed_runtime: None,
2649 }
2650 }
2651
2652 #[test]
2653 fn parses_feed_tier_declaration() {
2654 // A staged sidecar belongs to a self-contained bundle, so this fixture
2655 // declares one — R746-T3 refuses feed_bins on a vanilla slot.
2656 let mirror = mirror_with(
2657 r#"use = "cloudflare"
2658bucket = "b"
2659
2660[serve_bins]
2661x86_64-unknown-linux-musl = "target/x86_64-unknown-linux-musl/release/mesofact"
2662
2663[revalidate]
2664routes = ["/releases"]
2665feeds = ["releases", "yah-desktop"]
2666feed_interval_secs = 60
2667
2668[revalidate.feed_bins]
2669x86_64-unknown-linux-musl = "target/x86_64-unknown-linux-musl/release/almanac-feed"
2670"#,
2671 );
2672 let rv = BundleSlot::parse(&mirror, "s", "e")
2673 .unwrap()
2674 .revalidate
2675 .unwrap();
2676 assert_eq!(rv.feeds, vec!["releases", "yah-desktop"]);
2677 assert_eq!(rv.feed_interval_secs, Some(60));
2678 assert_eq!(rv.feed_bins.len(), 1);
2679 assert!(rv.feed_bins["x86_64-unknown-linux-musl"].ends_with("almanac-feed"));
2680 assert!(rv.feed_runtime.is_none());
2681 }
2682
2683 /// R746-T3: the vanilla shape's feed tier. This is the declaration that
2684 /// makes yah-marketing deployable from a machine with no Rust toolchain —
2685 /// no path to a cross-built fetcher anywhere in it.
2686 #[test]
2687 fn a_vanilla_slot_declares_its_fetcher_as_a_runtime_ref() {
2688 let mirror = mirror_with(
2689 r#"use = "cloudflare"
2690bucket = "b"
2691runtime_version = "0.8.22"
2692
2693[revalidate]
2694routes = ["/releases"]
2695feeds = ["releases"]
2696feed_runtime = "almanac-feed/0.8.22"
2697"#,
2698 );
2699 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2700 assert!(!slot.is_self_contained());
2701 let rv = slot.revalidate.unwrap();
2702 assert_eq!(rv.feed_runtime.as_deref(), Some("almanac-feed/0.8.22"));
2703 assert!(rv.feed_bins.is_empty());
2704 }
2705
2706 /// The whole point: a vanilla slot with a feed tier is READY with nothing
2707 /// on disk. `feed_bins` would have kept the cross-built-binary requirement
2708 /// alive on the syncing machine while pretending the bundle was vanilla.
2709 #[test]
2710 fn a_vanilla_feed_tier_needs_no_binary_on_the_syncing_machine() {
2711 let mirror = mirror_with(
2712 r#"use = "cloudflare"
2713bucket = "b"
2714runtime_version = "0.8.22"
2715
2716[revalidate]
2717feeds = ["releases"]
2718feed_runtime = "almanac-feed/0.8.22"
2719"#,
2720 );
2721 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2722 let empty = std::path::Path::new("/nonexistent-workspace-root");
2723 assert!(missing_bins(&slot, empty).is_empty());
2724 assert!(slot_ready(&slot, empty));
2725 }
2726
2727 /// Declared, never inferred — the rule serve_bins/serve_build already
2728 /// follow. "Use the path if it exists, else the ref" would make the
2729 /// deployed fetcher a function of the syncing machine's disk.
2730 #[test]
2731 fn feed_bins_and_feed_runtime_together_are_refused() {
2732 let mirror = mirror_with(
2733 r#"use = "cloudflare"
2734bucket = "b"
2735
2736[serve_bins]
2737x86_64-unknown-linux-musl = "target/mesofact"
2738
2739[revalidate]
2740feeds = ["releases"]
2741feed_runtime = "almanac-feed/0.8.22"
2742
2743[revalidate.feed_bins]
2744x86_64-unknown-linux-musl = "target/almanac-feed"
2745"#,
2746 );
2747 let err = BundleSlot::parse(&mirror, "s", "e").unwrap_err().to_string();
2748 assert!(err.contains("feed_bins") && err.contains("feed_runtime"), "got {err}");
2749 }
2750
2751 /// A vanilla bundle carries no `bins/`, so a path-declared sidecar has
2752 /// nowhere to be staged. Caught at parse, with the remedy in the message.
2753 #[test]
2754 fn feed_bins_on_a_vanilla_slot_is_refused_naming_feed_runtime() {
2755 let mirror = mirror_with(
2756 r#"use = "cloudflare"
2757bucket = "b"
2758runtime_version = "0.8.22"
2759
2760[revalidate]
2761feeds = ["releases"]
2762
2763[revalidate.feed_bins]
2764x86_64-unknown-linux-musl = "target/almanac-feed"
2765"#,
2766 );
2767 let err = BundleSlot::parse(&mirror, "s", "e").unwrap_err().to_string();
2768 assert!(err.contains("VANILLA"), "got {err}");
2769 assert!(err.contains("feed_runtime"), "got {err}");
2770 }
2771
2772 /// A typo in the ref fails the apply offline, not on a node twenty minutes
2773 /// into a deploy.
2774 #[test]
2775 fn an_unparseable_feed_runtime_is_refused_at_parse() {
2776 let mirror = mirror_with(
2777 r#"use = "cloudflare"
2778bucket = "b"
2779runtime_version = "0.8.22"
2780
2781[revalidate]
2782feeds = ["releases"]
2783feed_runtime = "almanac-feed"
2784"#,
2785 );
2786 let err = BundleSlot::parse(&mirror, "s", "e").unwrap_err().to_string();
2787 assert!(err.contains("feed_runtime"), "got {err}");
2788 }
2789
2790 /// The payload the node acts on must carry the ref, or kamaji has nothing
2791 /// to resolve and the fetcher silently never forks.
2792 #[test]
2793 fn the_feed_runtime_ref_reaches_the_workload_payload() {
2794 let mut slot = bare_revalidate_slot();
2795 slot.feed_runtime = Some("almanac-feed/0.8.22".to_string());
2796 let payload = slot.to_workload_payload(BTreeMap::new(), vec![], None, vec![]);
2797 assert_eq!(payload.feed_runtime.as_deref(), Some("almanac-feed/0.8.22"));
2798 }
2799
2800 /// A receiver with no feed tier is the existing shape and must keep parsing
2801 /// — the fetcher is additive, not a new requirement on every mirror.
2802 #[test]
2803 fn revalidate_without_a_feed_tier_stays_empty() {
2804 let mirror = mirror_with(
2805 r#"use = "cloudflare"
2806bucket = "b"
2807
2808[revalidate]
2809routes = ["/releases"]
2810"#,
2811 );
2812 let rv = BundleSlot::parse(&mirror, "s", "e")
2813 .unwrap()
2814 .revalidate
2815 .unwrap();
2816 assert!(rv.feeds.is_empty());
2817 assert!(rv.feed_bins.is_empty());
2818 assert_eq!(rv.feed_interval_secs, None);
2819 }
2820
2821 /// Feeds declared with no fetcher binary is the silent-staleness trap: the
2822 /// deploy would go green and the data would never move. Fail at parse.
2823 #[test]
2824 fn feeds_without_feed_bins_is_rejected() {
2825 let mirror = mirror_with(
2826 r#"use = "cloudflare"
2827bucket = "b"
2828
2829[revalidate]
2830feeds = ["releases"]
2831"#,
2832 );
2833 let err = BundleSlot::parse(&mirror, "s", "e")
2834 .unwrap_err()
2835 .to_string();
2836 assert!(err.contains("feed_bins"), "got {err}");
2837 assert!(err.contains("feed_runtime"), "got {err}");
2838 assert!(err.contains(FEED_BIN_NAME), "got {err}");
2839 }
2840
2841 #[test]
2842 fn zero_feed_interval_is_rejected() {
2843 let mirror = mirror_with(
2844 r#"use = "cloudflare"
2845bucket = "b"
2846
2847[revalidate]
2848feed_interval_secs = 0
2849"#,
2850 );
2851 let err = BundleSlot::parse(&mirror, "s", "e")
2852 .unwrap_err()
2853 .to_string();
2854 assert!(err.contains("positive integer"), "got {err}");
2855 }
2856
2857 /// The reconciler's default and the workload-spec serde default are two
2858 /// copies of one number; this pins them together.
2859 #[test]
2860 fn feed_interval_default_matches_the_workload_spec_default() {
2861 let payload = bare_revalidate_slot().to_workload_payload(BTreeMap::new(), vec![], None, vec![]);
2862 assert_eq!(payload.feed_interval_secs, DEFAULT_FEED_INTERVAL_SECS);
2863
2864 let from_spec: workload_spec::MesofactRevalidateReceiver =
2865 serde_json::from_str("{}").expect("all receiver fields have serde defaults");
2866 assert_eq!(from_spec.feed_interval_secs, DEFAULT_FEED_INTERVAL_SECS);
2867 }
2868
2869 // ── R885-T14: cap:bundle-serving as a placement precondition ─────────────
2870 //
2871 // Every test below keeps a CAPABLE and an INCAPABLE node in the same
2872 // `CloudConfig`, so a pass provably comes from the capability axis rather
2873 // than from an empty pool or a region that happens to narrow to one.
2874
2875 /// The literal-pin arm. An operator naming a node by hand is making exactly
2876 /// the claim this tag exists to check, and the refusal has to name the tag
2877 /// and the file to edit — the whole value is that it arrives before the
2878 /// build and the upload, not after them as `BackendRefused`.
2879 #[test]
2880 fn a_node_without_the_bundle_backend_cannot_be_pinned_to_serve_a_bundle() {
2881 let cfg = cfg_with(vec![
2882 machine("us-east-001", "us-east"),
2883 machine_without_bundle_backend("us-south-001", "us-south"),
2884 ]);
2885
2886 let pinned_incapable = mirror_with(
2887 r#"use = "cloudflare"
2888bucket = "b"
2889machines = ["us-south-001"]"#,
2890 );
2891 let slot = BundleSlot::parse(&pinned_incapable, "s", "e").unwrap();
2892 let err = resolve_bundle_machines(&cfg, &pinned_incapable, &slot, "s", "e")
2893 .expect_err("a pin at a node with no bundle backend must be refused")
2894 .to_string();
2895 assert!(
2896 err.contains(BUNDLE_SERVING_MESH_TAG),
2897 "the refusal must name the tag the operator has to add; got {err}"
2898 );
2899 assert!(
2900 err.contains(".yah/infra/machines/us-south-001.toml"),
2901 "the refusal must name the file to edit; got {err}"
2902 );
2903
2904 // Same resolver, same fleet, same shape of declaration — only the
2905 // node's capability differs, so the refusal above cannot be coming from
2906 // anything else.
2907 let pinned_capable = mirror_with(
2908 r#"use = "cloudflare"
2909bucket = "b"
2910machines = ["us-east-001"]"#,
2911 );
2912 let slot = BundleSlot::parse(&pinned_capable, "s", "e").unwrap();
2913 let resolved = resolve_bundle_machines(&cfg, &pinned_capable, &slot, "s", "e")
2914 .expect("a pin at a capable node still resolves");
2915 assert_eq!(
2916 resolved.iter().map(|m| m.name.as_str()).collect::<Vec<_>>(),
2917 vec!["us-east-001"],
2918 );
2919 }
2920
2921 /// The constraint arm. Two nodes satisfy everything the mirror declared;
2922 /// only one can actually serve a bundle, and placement must pick that one
2923 /// rather than the file-order first match.
2924 ///
2925 /// `us-south-001` sorts after `us-east-001` on purpose being no help here —
2926 /// the incapable node is declared FIRST, so a resolver that ignored the
2927 /// capability would return it and this test would fail.
2928 #[test]
2929 fn a_constraint_places_past_a_matching_node_that_cannot_serve_bundles() {
2930 let mirror = mirror_with(
2931 r#"use = "cloudflare"
2932bucket = "b"
2933required = { regions = ["us-east"] }"#,
2934 );
2935 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2936
2937 let cfg = cfg_with(vec![
2938 machine_without_bundle_backend("us-east-000", "us-east"),
2939 machine("us-east-001", "us-east"),
2940 ]);
2941 let resolved = resolve_bundle_machines(&cfg, &mirror, &slot, "s", "e")
2942 .expect("the capable us-east node takes it");
2943 assert_eq!(
2944 resolved.iter().map(|m| m.name.as_str()).collect::<Vec<_>>(),
2945 vec!["us-east-001"],
2946 "placement must skip the region-matching node with no bundle backend",
2947 );
2948
2949 // Drop the capable node and the SAME declaration over the SAME region
2950 // now refuses — so the selection above was the capability talking.
2951 let cfg = cfg_with(vec![machine_without_bundle_backend(
2952 "us-east-000",
2953 "us-east",
2954 )]);
2955 let err = resolve_bundle_machines(&cfg, &mirror, &slot, "s", "e")
2956 .expect_err("a region with no bundle-capable node has nowhere to place")
2957 .to_string();
2958 assert!(
2959 err.contains(BUNDLE_SERVING_MESH_TAG),
2960 "the refusal must render the derived axis, not just the declared one; got {err}"
2961 );
2962 }
2963
2964 /// The deployer and the ingress planner must stay set-for-set across the
2965 /// new axis (`resolve_bundle_machines`'s own doc promises it). A capable
2966 /// and an incapable node both match the constraint; both sides must land on
2967 /// the capable one, and the door must not widen its poll set past it.
2968 #[test]
2969 fn the_ingress_planner_applies_the_same_capability_as_the_deployer() {
2970 let mirror = mirror_with(
2971 r#"use = "cloudflare"
2972bucket = "b"
2973zone = "yah.dev"
2974required = { regions = ["us-east"] }"#,
2975 );
2976 let slot = BundleSlot::parse(&mirror, "s", "e").unwrap();
2977 let machines = vec![
2978 machine_without_bundle_backend("us-east-000", "us-east"),
2979 machine("us-east-001", "us-east"),
2980 ];
2981 let cfg = cfg_with(machines.clone());
2982
2983 let deployed: Vec<String> = resolve_bundle_machines(&cfg, &mirror, &slot, "s", "e")
2984 .expect("deployer places")
2985 .iter()
2986 .map(|m| m.name.clone())
2987 .collect();
2988 let planned = super::super::ingress::resolve_ingress_placements(&machines, &mirror)
2989 .expect("ingress planner places");
2990 assert_eq!(
2991 planned.get(SLOT_ROLE),
2992 Some(&deployed),
2993 "the front door must be aimed at exactly the nodes the bundle is deployed to",
2994 );
2995
2996 let candidates = super::super::ingress::resolve_ingress_candidates(&machines, &mirror);
2997 assert_eq!(
2998 candidates.get(SLOT_ROLE),
2999 Some(&vec!["us-east-001".to_string()]),
3000 "widening the poll set must not widen past the capability — polling a node that \
3001 can never hold the bundle is a door aimed at nothing",
3002 );
3003 }
3004
3005 /// Idempotent: a mirror that already spells the tag out by hand resolves
3006 /// identically and does not end up declaring it twice.
3007 #[test]
3008 fn a_mirror_that_declares_the_capability_itself_is_unchanged() {
3009 let declared = RequiredSpec {
3010 regions: vec!["us-east".into()],
3011 mesh_tags: vec![BUNDLE_SERVING_MESH_TAG.to_string()],
3012 ..Default::default()
3013 };
3014 let widened = with_bundle_capability(&declared);
3015 assert_eq!(
3016 widened.mesh_tags, declared.mesh_tags,
3017 "the tag must not be pushed a second time",
3018 );
3019 assert_eq!(widened.regions, declared.regions, "nothing else moves");
3020
3021 let bare = RequiredSpec {
3022 regions: vec!["us-east".into()],
3023 ..Default::default()
3024 };
3025 assert_eq!(
3026 with_bundle_capability(&bare).mesh_tags,
3027 vec![BUNDLE_SERVING_MESH_TAG.to_string()],
3028 );
3029 }
3030
3031 /// Only `providers.bundle` implies a capability. A non-bundle slot's
3032 /// placement must be byte-identical to what it was before R885-T14 — this
3033 /// is the regression guard on the role mapping, and it is the reason that
3034 /// mapping is one function rather than an `if` at each call site.
3035 #[test]
3036 fn a_non_bundle_slot_requires_no_capability() {
3037 let machines = vec![machine_without_bundle_backend("us-east-000", "us-east")];
3038 let mirror = mirror_from(BTreeMap::from([(
3039 "static".to_string(),
3040 toml::from_str::<MirrorProviderSlot>(
3041 "use = \"cloudflare\"\nrequired = { regions = [\"us-east\"] }\n",
3042 )
3043 .expect("slot parses"),
3044 )]));
3045 let planned = super::super::ingress::resolve_ingress_placements(&machines, &mirror)
3046 .expect("a static slot places on a node with no bundle backend");
3047 assert_eq!(planned.get("static"), Some(&vec!["us-east-000".to_string()]));
3048 }
3049}