Skip to main content

Module cloud_init

Module cloud_init 

Source
Expand description

Cloud-init template renderer for mirror-machine bootstrap (R040-F4).

Loads the YAML template from .yah/cloud/cloud-init/mirror.yml (with a built-in fallback for portability) and substitutes per-machine values: machine name, yah-yubaba release-archive URL + sha256, Headscale pre-auth key, and mesh tags. The output is the user_data string passed to MachineProvider::create_server.

Hetzner caps user_data at 32KiB so the yubaba binary cannot be base64-embedded (R040-F11). The first-boot runcmd fetches the release tar.gz (matching what .github/workflows/release.yml publishes), verifies sha256 against the archive, extracts, and installs /usr/local/bin/yubaba before the systemd hand-off.

@yah:ticket(R040-F11, “yah-yubaba delivery: fetch from URL instead of base64 (Hetzner 32KiB user_data cap)”) @yah:at(2026-05-05T00:00:42Z) @yah:status(review) @yah:assignee(agent:claude) @yah:parent(R040) @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md) @yah:verify(“cargo test -p cloud”) @yah:verify(“cargo test -p yah –bin yah cloud::”) @yah:verify(“yah cloud machine provision noisetable-pdx-1 –path /Users/user/ss/noisetable –dry-run –yubaba-url https://example.com/yubaba –yubaba /tmp/yubaba — renders curl + sha256sum runcmd lines, computes sha from local file”) @yah:handoff(“Cloud-init now fetches yah-yubaba via URL + sha256 verify instead of base64 inline (Hetzner 32KiB user_data cap). RenderInput swapped warden_base64 for warden_url + warden_sha256; template runcmd does curl -o + chmod +x + ‘echo SHA bin | sha256sum -c -’. CLI provision flags: –yubaba-url required for live; –yubaba-sha256 and/or –yubaba (compute or assert sha); –dry-run falls back to placeholders. New helper resolve_warden_delivery() in app/yah/cli/src/cloud.rs covers all five flag combos with 7 unit tests. Cumulative: 30 cloud-crate tests pass (incl. new render_stays_under_hetzner_user_data_cap which fails the build if rendering ever blows past 32 KiB), 17 yah cloud:: tests pass. PHASE_1_MIRROR_BOOTSTRAP.md updated. Side fix: status.rs FakeProvider needed a one-line delete_bucket stub to compile (R040-F12 has the real impl in review).”) @yah:next(“Operator runbook (R040-T6) is now unblocked: publish yah-yubaba Linux musl binary at a stable URL (GitHub release artifact for the workspace.package version), then run yah cloud machine provision noisetable-$region-1 --yubaba-url <URL> --yubaba <local-copy> --path /Users/user/ss/noisetable for region in pdx iad fsn.”) @yah:next(“Optional polish: derive a default –yubaba-url from workspace.package.version + a hardcoded GitHub repo so operators don’t have to retype the URL per region. Skip until release-plz (R038-F3) is wired so the artifact actually exists at a predictable path.”)

@yah:ticket(R040-F15, “Cloudflare Tunnel in cloud-init: cloudflared install + token, no public ports needed”) @yah:at(2026-05-05T00:00:42Z) @yah:assignee(agent:claude) @yah:status(review) @yah:parent(R040) @yah:handoff(“Architecture decision (this session, recorded so future-self doesn’t re-litigate): public ingress for yah-cloud nodes is Cloudflare Tunnels — cloudflared runs on each origin, makes an OUTBOUND connection to CF edge, CF proxies inbound traffic through the tunnel. Origin needs zero inbound exposure (no port 80/443 open, no stable IPv4, no floating IP plumbing). DNS records point at <tunnel-id>.cfargotunnel.com and never churn when boxes are replaced. Inter-node TCP (Postgres replication, gRPC streams, NATS) goes over Headscale mesh — see R040-F16. Combined effect: Hetzner inline IPs are fine forever, IPv6-only is even an option (saves ~€0.50/mo per box; needs validating that tailscale install + apt mirrors work over v6).”) @yah:next(“cloud-init template: add cloudflared service install --token {{CLOUDFLARED_TOKEN}} + systemctl enable --now cloudflared to runcmd, parallel to the existing tailscale install. Token comes from RenderInput.cloudflared_token, sourced from keys::KeysStore::open().get(\"cloudflare-tunnel-token\") in cli/src/cloud.rs::handle_provision.”) @yah:next(“MachineConfig.cloudflared: Option — tunnel-id this machine joins. Empty/None means "no tunnel" (mesh-only nodes that don’t need public ingress). When set, the cloud-init renderer wires the right token in.”) @yah:next(“Optional: yah cloud tunnel {create,list,destroy} subcommand against the CF API. Lower-priority — operators can use cloudflared CLI directly until this matters at fleet scale.”) @yah:next(“Caveat to flag in handoff: Free + Pro tier CF Tunnels is HTTP/WebSocket/gRPC only. Raw TCP/UDP ingress (rare for yah, but possible — e.g. a public Postgres for some integration) needs CF Spectrum (paid) OR a primary IP for that one service. Don’t pre-build the second path; design records this as a known constraint.”) @arch:see(architecture/PHASE_1_MIRROR_BOOTSTRAP.md)

@yah:ticket(R441-B3, “cloud_init mirror.yml drift between cloud/templates and .yah/infra/cloud-init”) @yah:assignee(agent:claude) @yah:at(2026-06-04T22:56:14Z) @yah:status(review) @yah:parent(R441) @yah:next(“embedded_template_matches_workspace_canonical at cloud_init.rs:489 panics: the two mirror.yml files diverged. crates/yah/cloud/templates/mirror.yml (canonical, R406-T13/W154) has yubaba + kamaji + yubaba.slice systemd units plus the kamaji.service unit; .yah/infra/cloud-init/mirror.yml still has the older single yah-yubaba.service shape with no kamaji.”) @yah:next(“Pick the source of truth (the canonical-template path that the embedded test enforces was meant to be the cloud/templates copy) and sync the other file to match. Re-run the test to confirm.”) @yah:verify(“cargo test -p cloud –lib cloud_init::tests::embedded_template_matches_workspace_canonical passes”) @yah:handoff(“Overwrote crates/yah/cloud/templates/mirror.yml to match .yah/infra/cloud-init/mirror.yml (W154/R406-T13 yubaba+kamaji+yubaba.slice format). The workspace file was the canonical updated version; the embedded template hadn’t been synced. cargo test -p cloud –lib cloud_init::tests::embedded_template_matches_workspace_canonical passes.”)

@yah:ticket(R589-F1, “Rename service/binary emitters warden→yubaba in the provision path so new nodes provision as yubaba”) @yah:status(review) @yah:at(2026-07-06T06:13:22Z) @yah:assignee(agent:claude) @yah:parent(R589) @yah:handoff(“Hard-cut warden→yubaba across the provision-path emitters (no shims/aliases). cloud_init.rs: RenderInput fields warden_url/warden_sha256/warden_channel/warden_cosign_identity_regexp → yubaba_; placeholders {{YAH_WARDEN_URL}}/{{YAH_WARDEN_SHA256}}/{{WARDEN_CHANNEL}}/{{UFW_WARDEN_RULE}} → {{YAH_YUBABA_URL}}/{{YAH_YUBABA_SHA256}}/{{YUBABA_CHANNEL}}/{{UFW_YUBABA_RULE}}; PLACEHOLDER_WARDEN_URL/SHA256 → PLACEHOLDER_YUBABA_URL/SHA256; DEFAULT_WARDEN_CHANNEL → DEFAULT_YUBABA_CHANNEL; compute_warden_sha256() → compute_yubaba_sha256(). templates/mirror.yml + its workspace-canonical twin .yah/infra/cloud-init/mirror.yml (drift-tested against each other) got matching placeholder/env renames, incl. the systemd drop-in’s Environment=WARDEN_CHANNEL→YUBABA_CHANNEL. provision.rs’s build_request() params + release_manifest.rs’s WardenReleaseManifest→YubabaReleaseManifest / DEFAULT_WARDEN_COSIGN_IDENTITY→DEFAULT_YUBABA_COSIGN_IDENTITY followed (they feed straight into RenderInput). yubaba-test-harness/src/lib.rs: YAH_WARDEN_URL/SHA256 env vars → YAH_YUBABA_URL/SHA256, plus the local build_smoke_cloud_init()/wait_for_warden_health() helpers renamed to match. local_docker.rs’s cloud_init_boots_warden_service test renamed + its template placeholders updated. Cross-crate consumers outside oss/yubaba also updated in the same pass (real compile deps, not board-protocol scope creep): app/yah/cli/src/{cloud.rs,cli.rs,yubaba_fetch.rs} + tests/camp_yubaba_fetch.rs. Fenced OFF (belongs to R592-T4, wire-layer type/client renames): WardenHandle, YubabaRaft, YubabaRequest, WardenRoute wire types in yubaba/raft/, yubaba-test-harness’s WardenHandle, and camp.rs’s WardenContainerSpec/build_warden_run_spec/DEFAULT_WARDEN_IMAGE/DEFAULT_WARDEN_HTTP_PORT/read_warden_pond_port/WardenDeploy (pond continuous-deploy runtime, not the cloud-init provisioning path). Also left untouched: name-neutral state paths (/var/lib/yah-cloud/identity.json, /run/constable/constable.sock) per this ticket’s own scope fence, and stale warden_test_harness/warden_test_macros doc-comment crate names in yubaba-test-macros (pre-existing drift from an earlier, unrelated crate rename — not provision-path env/template residue). Verify: cargo check/test -p cloud -p yubaba-test-harness (oss/yubaba workspace) clean; cargo check -p yah + cargo test -p yah –lib yubaba_fetch:: + –test camp_yubaba_fetch clean from repo root. Zero remaining WARDEN_ env/placeholder refs in the provision path (grep-verified).”)

@yah:relay(R702, “Bounded disk everywhere — no process on a yah box grows without an explicit ceiling”) @yah:at(2026-08-03T01:29:40Z) @yah:status(open) @yah:assignee(agent:bundle-anthropic-glimmerstone) @yah:next(“DOCTRINE (operator, 2026-08-02, emphatic): a full disk brings the best software to its knees, and apps - Docker especially - love to eat all the space. Every single process that ever runs on a yah box must be explicitly forbidden from unbounded growth. A default that happens to be a fraction of the disk is NOT a bound. This relay makes that auditable rather than aspirational.”) @yah:next(“DONE ALREADY, not part of the remaining children: journald bounded on all three cloud nodes via /etc/systemd/journald.conf.d/10-yah-disk-bounds.conf (SystemMaxUse=500M, SystemKeepFree=1G, SystemMaxFileSize=50M, RuntimeMaxUse=64M) and the same block added to the TOP of runcmd in BOTH twin cloud-init templates (oss/yubaba/crates/cloud/templates/mirror.yml + .yah/infra/cloud-init/mirror.yml) so it lands before anything else on the box starts writing. us-south-001 went 755.5M -> 483M on restart. RuntimeMaxUse matters twice over: /run is tmpfs, so it bounds RAM too.”) @yah:next(“Children to file as work begins: (F) kamaji bundle-cache budget - THE live unbounded one, see gotcha; (T) containerd image/snapshot GC policy; (T) apt archive + old-kernel retention; (T) surface per-node disk headroom in yubaba node status so pressure is visible before it is fatal; (D) a W doc carrying the doctrine plus a review checklist item - every new on-disk writer declares its ceiling at the same time it declares the path.”) @yah:gotcha(“LIVE UNBOUNDED CACHE IN OUR OWN CODE, RIGHT NOW. kamaji’s mesofact bundle cache has working, unit-tested LRU eviction (BundleCache::evict_to_budget, oss/yah-base/crates/mesofact-bundle/src/store.rs:243) that is DISABLED in production: BundleBackend::cache_budget defaults to 0 (oss/kamaji/crates/kamaji-bin/src/server.rs:297) and 0 explicitly means ‘unbounded, eviction off’ (store.rs:207). There is no –bundle-cache-budget CLI flag in kamaji-bin/src/main.rs, and the live drop-in /etc/systemd/system/kamaji.service.d/20-bundle.conf on us-east-001 passes only –bundle-cache-dir and –bundle-origin. Net effect: every release wave materializes a new digest-keyed tree under /var/lib/yah/kamaji/bundles/ and NOTHING ever removes the old one. Only 6.3M on east today because there have been few waves - it grows monotonically with release count, forever.”) @yah:gotcha(“Fix shape for that child: add the flag, and make the DEFAULT non-zero rather than shipping another opt-in bound (an opt-in ceiling is how this happened). If 0-means-unbounded stays as an escape hatch it should require passing 0 explicitly, so the default path is always bounded.”) @yah:gotcha(“Measured baseline 2026-08-02 for whoever picks this up. us-south-001 (30G disk, 1 vCPU / 961MB): 7.2G used, /var 1.4G of which journal 755M, /var/lib/apt 295M, /var/cache 177M. us-east-001 (99G): 2.0G used, /usr 1.1G, /var 786M, containerd 154M, kamaji bundles 6.3M. us-west-001 (99G): 1.9G used. Nothing is near full today - this relay is about the derivative, not the current level.”) @yah:verify(“cargo test -p yah-cloud –lib cloud_init # 25/25 pass after the template edit, incl. embedded_template_matches_workspace_canonical (twin-drift guard) and rendered_runcmd_entries_are_all_strings (the colon-space YAML footgun guard) - ALREADY GREEN 2026-08-02”) @yah:verify(“Audit gate for closing this relay: on a freshly provisioned node, every path under /var that any yah-owned process writes to has a named ceiling, and each ceiling is asserted somewhere (unit test, systemd directive, or config) rather than assumed.”) @yah:verify(“journalctl –disk-usage on each cloud node stays under 500M across a week of normal operation.”) @yah:assumes(“The LAN/appliance nodes (us-west-011/013/014/015) need the same treatment and probably need it MORE (us-west-014 is the arm64 Pi appliance prototype - SD-card-class storage). They were not touched in the 2026-08-02 pass, which covered only the three cloud VMs.”) @yah:gotcha(“us-west-014 (Pi 5) VERIFIED 2026-08-02 and it is a two-filesystem box, which changes what ‘bounded’ means there. The NVMe design IS implemented and working - /dev/nvme0n1p1 is bind-mounted onto /var/lib/docker, /var/lib/yah-cloud and /srv/build, plus an 8 GB swapfile (enabled, 0 used), 234 G at 4%. But / is a 2.5 G SD-backed ext4 at 68% with ~745 M free, and /var itself is NOT on the NVMe - only those three subdirectories are. So /var/log (123 M) and /var/cache apt (170 M) sit on the tightest filesystem in the fleet. journald capped at 200M/400M-keepfree there (deliberately smaller than the cloud nodes’ 500M). The general rule for this box: anything new writing under /var lands on the SD unless it gets its own bind, so a per-node ceiling has to be sized to the filesystem it actually lands on, not copied from the cloud nodes.”) @yah:gotcha(“Docker on us-west-014 is the ONE docker install in the fleet that is already safe by placement (data-root bind-mounted to a 234 G NVMe at 4%). Do not let that make the containerd/docker GC child look optional - the cloud nodes have containerd on the root filesystem with no GC policy (154 M on us-east-001 today).”) @yah:handoff(“PI APPLIANCE IMAGE BOUNDED (2026-08-02), ahead of flashing us-west-011 + us-west-013. Two unbounded growers ship with stock docker and both hit build workers hardest: json-file logging defaults to no max-size/max-file, and the BuildKit cache grows forever without builder.gc. Neither had any config - /etc/docker/daemon.json did not exist on us-west-014 or in the image layer. Added to .yah/infra/pi-image/layer/yah-build-worker.yaml: a mkdir -p hook plus a baked /etc/docker/daemon.json (json-file, max-size 10m, max-file 3, builder.gc enabled with defaultKeepStorage 20GB) and /etc/systemd/journald.conf.d/10-yah-disk-bounds.conf at 200M/400M-keepfree/25M-maxfile/32M-runtime. Layer YAML re-parses (14 hooks) and the emitted JSON body validates.”) @yah:handoff(“LOCKSTEP DEBT PAID. mirror.yml’s new journald runcmd block obliged a matching change in its documented twin .yah/infra/cloud-init/stand-up-yubaba.sh (the SSH-deliverable transcription used for LAN nodes, W257 step 6). Written WRITE-IF-ABSENT on purpose: the standup default is the cloud-node 500M, but the Pi image bakes a tighter 200M sized to its 2.5 G SD root, and the standup runs AFTER first boot - an unconditional write would have silently regressed every Pi it touched. An existing ceiling always wins and the script echoes what it kept. bash -n clean.”) @yah:handoff(“W257 runbook step 5 gained a disk-ceiling verification block. The load-bearing assertion is systemctl is-active docker: dockerd refuses to start on a daemon.json it cannot parse, and a docker-less build worker is the entire purpose of the node gone. Runbook names the likely culprit after a docker bump (defaultKeepStorage deprecated in favour of builder.gc.policy in 27+; trixie ships 26.1.5, which honours the old key).”) @yah:handoff(“X86 FLEET PROVISIONING LANDED (2026-08-02), and it closes the arch-specific-ceiling gap this relay opened. New .yah/infra/preseed/{yah-x86-worker.cfg, build-iso.sh, .gitignore}. Deliberately the SAME shape as the Pi path rather than a new idiom: a Debian container does the work, the operator key is injected at build time instead of committed, and one artifact provisions every x86 box with per-box identity applied after install. build-iso.sh caches the stock trixie amd64 netinst, renders the preseed with ~/.ssh/yah.pub substituted for @@SSH_PUBKEY@@, injects it into the installer initrd (so the install is hands-off with no boot prompt to type at), regenerates md5sum.txt, and repacks a UEFI hybrid ISO with xorriso. Neither path is a roll-your-own distro - one configures rpi-image-gen, the other configures debian-installer.”) @yah:handoff(“PARTITIONING IS THE STRUCTURAL HALF OF THIS RELAY, and the preseed is where it finally gets decided up front instead of discovered. Separate LVs for /, /var and /var/lib/docker on VG yah, ~40 GB left unallocated for online lvextend. A runaway BuildKit cache now fills /var/lib/docker and NOTHING else - root stays writable, sshd keeps accepting, journald keeps recording, the box stays reachable to clean up. That is the backstop for a MISSING ceiling; it does not replace the ceilings, which the preseed late_command also writes (journald 500M + docker daemon.json with log rotation and builder.gc at 100GB, sized to the 512 GB disk).”) @yah:handoff(“DOCKER CEILING IS NO LONGER A PROPERTY OF ONE IMAGE. stand-up-yubaba.sh now applies /etc/docker/daemon.json write-if-absent on any node where docker is present (it installs containerd, not docker, so this is a conditional), and restarts docker THERE so a parse failure surfaces during standup rather than at the next reboot - dockerd refuses to start on a daemon.json it cannot parse. Three provisioning paths now converge on the same ceilings: Pi image (baked), preseed late_command (baked), standup script (retrofit). bash -n clean on both scripts.”) @yah:handoff(“Scaffolded .yah/infra/machines/us-west-012.toml for the GEEKOM A5 (Ryzen 7 5825U 8C/16T, 16 GB, 512 GB NVMe): arch:x86 + os:linux + build-worker/qed, taints copied from us-west-002 for day one. Recorded WHY it is a better arch:x86 host than 002 - 002 is a WSL2 box that sleeps and reboots with Windows, this is dedicated always-on Debian - so dropping no-server/no-appliance later is a deliberate re-decision rather than drift. Stays no-voter regardless: a residential uplink must never be able to stall the raft. allocatable is from the vendor spec with an explicit instruction to re-verify via nproc + free -m on the box, since the OVH nodes shipped wrong for months. Parses: cargo test -p yah-cloud –lib machine 24/24, yah cloud validate ok.”) @yah:verify(“UNPROVEN, and the one thing to watch: the partman-auto/expert_recipe in yah-x86-worker.cfg has never been run. A malformed recipe fails mid-install with an error that does not always name the offending stanza. Watch the first install of any ISO revision; once it completes cleanly the same ISO is proven for every later box. It also assumes ONE disk (early_command picks list-devices disk | head -n1), so a two-disk box needs the target pinned.”) @yah:verify(“UNPROVEN: build-iso.sh has not been executed - the initrd inject + xorriso repack path is written but not run, and the ISO URL pins DEBIAN_VERSION=13.1.0 which should be bumped to whatever trixie point release is current at build time (override with the env var).”) @yah:verify(“Cheap de-risk available before touching hardware: boot the built ISO in QEMU against a scratch qcow2 and let the unattended install run to completion. That proves the recipe, the initrd inject and the late_command without burning a USB or a trip to the box.”) @yah:handoff(“RENUMBERED (operator correction, 2026-08-02): the GEEKOM mini PC is us-west-003, NOT us-west-012 — the 01x block is reserved for the small LAN/appliance class and another Pi is taking 012. The convention is by HARDWARE CLASS, not arrival order: 00x = PC/server (001 OVH VPS, 002 WSL2 gamer box, 003 GEEKOM), 01x = small LAN boxes (011-014 Pis, 015 arm64 Mac). Recorded at the top of us-west-003.toml and in W257, because it is not derivable from the existing files and it is what tells a reader which of the two provisioning paths a node took. The us-west-012.toml scaffold was untracked and never committed, so it was removed rather than renamed; no stale refs remain (grep clean, yah cloud validate ok).”) @yah:gotcha(“LAN IP for us-west-003 is ASSUMED, not confirmed. W257’s convention is us-west-0NN -> 192.168.10.NN and it holds for every 01x node, but the only existing 00x LAN box breaks it: us-west-002 sits at 192.168.10.30, not .2. The scaffold uses .3 with a CONFIRM-BEFORE-BRING-UP note inline. If the 00x block actually lives in the .3x range on the router, .3 is wrong and both [connect].address and [connect].ssh need correcting before step 4.”)

Re-exports§

pub use crate::release_manifest::COSIGN_OIDC_ISSUER;

Structs§

RenderInput
Inputs needed to render mirror.yml for a single machine.

Enums§

CanonicalHome
Where the twin-drift guard should look for the canonical mirror.yml.

Constants§

COSIGN_SHA256_AMD64
sha256 of the cosign-linux-amd64 binary at COSIGN_VERSION, from cosign’s own published cosign_checksums.txt for that release. Cloud-init verifies the downloaded binary against this before chmod+exec. Pinning the verifier itself closes the bootstrap-trust gap: TLS to github.com proves origin, sha256 proves bytes, then cosign proves the yubaba tarball.
COSIGN_SHA256_ARM64
sha256 of the cosign-linux-arm64 binary at COSIGN_VERSION.
COSIGN_VERSION
Pinned cosign release used by the cloud-init verify-blob block. Bump in lockstep with COSIGN_SHA256_AMD64 / COSIGN_SHA256_ARM64 when upgrading the verifier — supply-chain hygiene (W203 §1.4, R330-F22). v3.x on purpose: cosign 3’s sigstore-bundle format (--bundle) is current best practice, replacing the old separate .sig/.cert file pair everywhere in this pipeline (install.sh, release_manifest.rs, yubaba_fetch.rs).
DEFAULT_TEMPLATE
Built-in fallback template, used when .yah/infra/cloud-init/mirror.yml is absent. Keeps the binary self-contained for tests + new workspaces.
DEFAULT_YUBABA_CHANNEL
Phase 1 defaults for new provisioning keys.
PLACEHOLDER_CLOUDFLARED_TOKEN
PLACEHOLDER_PREAUTH_KEY
PLACEHOLDER_YUBABA_SHA256
PLACEHOLDER_YUBABA_URL
Placeholders used when the user requests a dry-run without supplying real substitutes. Makes the rendered YAML obviously non-shippable while still preserving the structure for review.

Functions§

compute_yubaba_sha256
Read a yah-yubaba release tar.gz from disk and return its lowercase hex sha256. Used both as the cloud-init verification digest and for asserting that a local copy matches an expected --yubaba-sha256 value. The path should point to the release archive (matching what cloud-init downloads), not a bare binary.
load_template
Load the cloud-init template for a workspace.
locate_canonical_home
Walk start’s ancestors for the monorepo root that owns the canonical cloud-init template. See CanonicalHome for why .yah/infra/ is the marker rather than .yah/.
render
Substitute {{KEY}} placeholders. Fails loudly if any unsubstituted placeholder remains — better than silently shipping a broken cloud-init.