cloud/provision.rs
1//! Provisioning orchestrator: config → cloud-init render → MachineProvider call.
2//!
3//! Decoupled from the concrete provider so the CLI passes a `&dyn MachineProvider`
4//! (Hetzner in production, in-memory fakes in tests).
5//!
6//! `execute` here is called only from the operator-invoked `yah cloud machine
7//! provision` CLI (`app/yah/cli/src/cloud.rs`) — nothing in the yubaba server
8//! process calls it. `yubaba::headroom` (R737-T4) watches for when a fresh
9//! call here is warranted (N+1/N+2 warm-spare deficit) and surfaces that as a
10//! `GET /raft/status` field plus a log line; it does not call this module,
11//! deliberately — see its own doc for why.
12//!
13//! @yah:ticket(R737-T4, "N+1 headroom accounting + a background headroom-restore provisioning job that never sits on the failover path")
14//! @yah:status(review)
15//! @yah:at(2026-08-16T04:37:37Z)
16//! @yah:assignee(agent:bundle-anthropic-miravel)
17//! @yah:phase(P3)
18//! @yah:parent(R737)
19//! @yah:next("Tier: Cleric — an invariant plus an enqueue path; the hard constraint (never inline) is already the status quo to preserve.")
20//! @yah:next("VERIFIED ABSENT 2026-08-09: there is no N+1/N+2 warm-spare notion to fail over into. cloud::provision provisions on demand and nothing tracks headroom.")
21//! @yah:next("When headroom drops below N+1, ENQUEUE a background provisioning job. N+2 where a whole region can vanish.")
22//! @yah:depends_on(R737-F3)
23//! @yah:handoff("New oss/yubaba/crates/yubaba/src/headroom.rs: N+1 (N+2 across >=2 declared regions) warm-spare accounting, W253 §8. required_spare(members) derives the target from distinct MemberInfo.region values (2 once the fleet spans >=2 regions, else 1). is_spare(node,...) is the stricter predicate W253 asks for -- confirmed live (F2's TransitionTracker) + published capacity + ZERO currently-owned LIVE tenants (a node quietly carrying one small tenant is not 'held in reserve', even with room left -- that's node_headroom's job, not this). evaluate() counts spare nodes against the requirement -> HeadroomReport{spare, required}. 10 unit tests covering region-count thresholds, each spare-disqualifying condition individually (unconfirmed, unmeasured, carrying a live tenant), and the expired-lease-frees-the-node case that mirrors R737-F1's own node_load argument.")
24//! @yah:handoff("THE DESIGN CALL: does NOT call cloud::provision::execute, and does not build an in-process job queue with no consumer. Nothing in the yubaba server process has the MachineConfig/provider credentials that call needs -- those live with the operator-invoked `yah cloud machine provision` CLI (app/yah/cli/src/cloud.rs), confirmed by grep: provision::execute has exactly one call site and it is not in yubaba. Giving the control-plane daemon direct cloud-provisioning authority is a materially bigger architectural change than this ticket's own 'Cleric: an invariant plus an enqueue path' framing asks for. 'Enqueue' is scoped to: a leader-only background loop (paced at 10x election timeout -- a trend, not an emergency; scheduler.rs already owns the emergency path) evaluates the invariant, caches the verdict on ServerState.headroom (Mutex<Option<HeadroomReport>>), surfaces it as a new `headroom` section on GET /raft/status, and logs at warn on every tick a deficit persists. Wired in main.rs alongside scheduler/lease_renewal. Added a one-line pointer in provision.rs's own module doc so a future reader lands on the connection.")
25//! @yah:handoff("DISCOVERED WORK, and the reason this took far longer than the ticket's own Cleric-tier estimate: cluster_epoch_drift caught 2 UNDETECTED problems from this session's own earlier F2 and F3 work, not just from T4. raft_status (fn raft_*, a cluster_protocol surface input) had already gained F2's lease_liveness section and now T4's headroom section -- I had never run the epoch-drift gate after F2 or F3 landed. F3's three new store.rs read accessors (tenants/tenant_placement/node_admits) moved state_epoch's surface the same way the tenant_fencing_token()/tenant_ownership() precedents did (2026-08-09/10 entries). Verdict on both, matching the R734-F5 precedent exactly (same function, same additive-JSON-on-a-GET-surface shape): NOT BREAKING. Re-recorded via `cargo run -p xtask -- cluster-epochs --write` (cluster_protocol stays 5, state_epoch stays 4) and added a full surface_rerecords entry to cluster-epochs.json dated 2026-08-15 covering all three tickets' contribution to the drift in one entry.")
26//! @yah:handoff("Full verify, camp was under heavy concurrent-build load (10+ simultaneous cargo invocations across the camp against the same root workspace, confirmed via ps -- not a hang, rustc was genuinely CPU-bound, just deeply queued; several checks took 15-20 minutes to clear the lock): cargo check -p yubaba clean; cargo check -p yubaba --bin yubaba clean; cargo check -p yah-cloud clean (provision.rs doc-only edit); cargo test -p yubaba --lib = 489 passed / 0 failed (465 at F2's start -> 479 after F3+lease_renewal -> 489 after T4's 10 tests); cargo clippy -p yubaba --lib --bin yubaba: zero findings on headroom.rs or the main.rs wiring; cargo test -p xtask --test cluster_epoch_drift = 8 passed / 0 failed after re-recording (was 7/1, both axes red).")
27//! @yah:verify("cargo check -p yubaba (clean)")
28//! @yah:verify("cargo check -p yubaba --bin yubaba (clean)")
29//! @yah:verify("cargo check -p yah-cloud (clean)")
30//! @yah:verify("cargo test -p yubaba --lib = 489 passed, 0 failed, 0 ignored")
31//! @yah:verify("cargo clippy -p yubaba --lib --bin yubaba: no findings on headroom.rs, scheduler.rs, lease_renewal.rs, store.rs's new accessors, or main.rs")
32//! @yah:verify("cargo test -p xtask --test cluster_epoch_drift = 8 passed, 0 failed (cluster-epochs.json re-recorded, both axes NOT BREAKING, history entry added 2026-08-15)")
33//! @yah:gotcha("is_spare's metric is deliberately coarse: a node counts as spare only when it holds ZERO tenants, not merely 'has room for one more'. A cluster running near capacity with every node carrying at least one tenant reports 0 spare even if collectively there is plenty of headroom to absorb a single failure via partial rebalancing -- true bin-packing-aware headroom (can the union of remaining headroom absorb node X's specific tenant set) is more accurate but was out of this Cleric ticket's scope; flag this if a future ticket finds the invariant too conservative in practice.")
34//! @yah:gotcha("Nothing consumes the `headroom` GET /raft/status field yet -- no automation calls `yah cloud machine provision` off a deficit signal. That consumer (an operator runbook, or a `yah cloud auto-provision` watch command polling /raft/status) is explicitly out of this ticket's scope and unfiled; whoever wants headroom restoration to actually be automatic rather than merely visible should pick it up.")
35
36use crate::cloud_init::{self, RenderInput};
37use crate::config::MachineConfig;
38use crate::provider::{Location, MachineProvider, ProjectId, ServerId, ServerSpec};
39use anyhow::{Context, Result};
40use std::path::Path;
41
42/// A rendered provision payload, ready to send to a `MachineProvider`.
43#[derive(Debug)]
44pub struct ProvisionRequest {
45 pub machine_name: String,
46 pub server_type: String,
47 pub location: Location,
48 pub user_data: String,
49 /// Provider-side SSH-key IDs to authorize for `root` at create time.
50 /// Carried from `MachineConfig.ssh_keys`; empty defaults to no
51 /// per-key auth (Hetzner emails a random root password we discard).
52 pub ssh_keys: Vec<u64>,
53}
54
55/// Build a provision request: load the cloud-init template for the workspace and
56/// substitute per-machine values. The yubaba binary is fetched on the machine
57/// at first boot from `yubaba_url` and verified against `yubaba_sha256`
58/// (R040-F11) — base64-embedding it would blow past Hetzner's 32 KiB cap.
59///
60/// `headscale_preauth_key` decides mesh membership (R330-F28). `Some` ⟺ this
61/// machine is JOINING an existing mesh: the rendered cloud-init emits the
62/// tailscaled install + `tailscale up --auth-key=<key>` join block. `None` ⟺
63/// STANDALONE / coordinator-to-be — no mesh exists yet, so no join block is
64/// emitted; the node comes up as bare yubaba and becomes the coordinator later
65/// via `yah mesh bootstrap`. Membership is gated purely on this key's presence,
66/// independent of `machine.hosts_operator_bridge`.
67///
68/// `mesh_url` is the stable Headscale coordinator URL (R040-F18). When present
69/// (only meaningful alongside a preauth key), the rendered cloud-init passes
70/// `--login-server <url>` to `tailscale up` so the machine joins the camp's
71/// Headscale instead of Tailscale SaaS. When `None`, a joining machine uses the
72/// default Tailscale SaaS coordinator.
73///
74/// `yubaba_channel` selects the release channel (`"stable"` or `"beta"`);
75/// use [`cloud_init::DEFAULT_YUBABA_CHANNEL`] for Phase 1. containerd is
76/// installed unpinned (R330-T9 — an exact apt pin matched no Debian repo).
77pub fn build_request(
78 workspace_root: &Path,
79 machine: &MachineConfig,
80 yubaba_url: String,
81 yubaba_sha256: String,
82 yubaba_channel: String,
83 headscale_preauth_key: Option<String>,
84 mesh_url: Option<String>,
85 cloudflared_token: Option<String>,
86 yubaba_cosign_identity_regexp: Option<String>,
87) -> Result<ProvisionRequest> {
88 let template = cloud_init::load_template(workspace_root)?;
89 let input = RenderInput {
90 machine,
91 yubaba_url,
92 yubaba_sha256,
93 yubaba_channel,
94 headscale_preauth_key,
95 mesh_url,
96 cloudflared_token,
97 yubaba_cosign_identity_regexp,
98 };
99 let user_data = cloud_init::render(&template, &input)?;
100 machine.validate()?;
101 let location = Location::try_from(machine.location())
102 .with_context(|| format!("machine '{}' has unknown location", machine.name))?;
103 Ok(ProvisionRequest {
104 machine_name: machine.name.clone(),
105 server_type: machine.server_type().to_string(),
106 location,
107 user_data,
108 ssh_keys: machine.ssh_keys.clone(),
109 })
110}
111
112/// Execute a built request against a provider. Returns the new server ID on success.
113///
114/// Hostkey-fingerprint write-back lands with A8 (yah-yubaba `/identity` endpoint
115/// and `MachineConfig::save` are both already in place; the missing piece is the
116/// yubaba binary itself).
117pub async fn execute(
118 provider: &dyn MachineProvider,
119 project: &ProjectId,
120 req: &ProvisionRequest,
121) -> Result<ServerId> {
122 let spec = ServerSpec {
123 name: req.machine_name.clone(),
124 server_type: req.server_type.clone(),
125 image: "debian-12".into(),
126 location: req.location.clone(),
127 ssh_keys: req.ssh_keys.clone(),
128 };
129 provider.create_server(project, &spec, &req.user_data).await
130}
131
132#[cfg(test)]
133mod tests {
134 use super::*;
135 use crate::cloud_init;
136 use crate::config::{BucketSpec, MachineConfig};
137
138 fn sample_machine() -> MachineConfig {
139 MachineConfig {
140 name: "noisetable-pdx-1".into(),
141 provider: "hetzner".into(),
142 location: Some("pdx".into()),
143 server_type: Some("cpx22".into()),
144 hosts_mirrors: vec!["noisetable".into(), "yah".into()],
145 mesh_tags: vec!["tag:region-pdx".into(), "tag:tier-t2".into()],
146 region: None,
147 zone: None,
148 arch: None,
149 bucket: Some(BucketSpec {
150 name: "noisetable-assets-pdx-1".into(),
151 public_read: false,
152 }),
153 vendor: None,
154 nickname: None,
155 legacy_hostkey_fingerprint: None,
156 registration: Default::default(),
157 ssh_keys: vec![],
158 cloudflared: None,
159 hosts_operator_bridge: false,
160 connect: None,
161 allocatable: None,
162 taints: vec![],
163 sovereign_group: None,
164 sovereign_role: None,
165 }
166 }
167
168 fn build_req_defaults(
169 dir: &std::path::Path,
170 machine: &MachineConfig,
171 extra_url: Option<String>,
172 extra_cf: Option<String>,
173 ) -> crate::provision::ProvisionRequest {
174 build_request(
175 dir,
176 machine,
177 "https://example.com/yah-yubaba".into(),
178 "deadbeef".into(),
179 cloud_init::DEFAULT_YUBABA_CHANNEL.into(),
180 Some("KEY".into()),
181 extra_url,
182 extra_cf,
183 None,
184 )
185 .unwrap()
186 }
187
188 #[test]
189 fn build_request_renders_user_data_and_picks_location() {
190 // A preauth key present → join block emitted, so the key ("KEY") and tags appear.
191 let machine = sample_machine();
192 let dir = tempfile::tempdir().unwrap();
193 let req = build_req_defaults(dir.path(), &machine, None, None);
194 assert_eq!(req.machine_name, "noisetable-pdx-1");
195 assert_eq!(req.location, Location::Pdx);
196 assert!(req.user_data.contains("https://example.com/yah-yubaba"));
197 assert!(req.user_data.contains("deadbeef"));
198 assert!(req.user_data.contains("KEY"));
199 assert!(req.user_data.contains("tag:region-pdx,tag:tier-t2"));
200 }
201
202 #[test]
203 fn build_request_with_mesh_url_adds_login_server() {
204 let machine = sample_machine();
205 let dir = tempfile::tempdir().unwrap();
206 let req = build_req_defaults(
207 dir.path(),
208 &machine,
209 Some("https://mesh.example.com".into()),
210 None,
211 );
212 assert!(req
213 .user_data
214 .contains("--login-server https://mesh.example.com"));
215 }
216
217 #[test]
218 fn build_request_standalone_omits_join_block() {
219 // R330-F28: a standalone / coordinator-to-be node carries no preauth
220 // key (and no mesh_url). build_request must NOT emit the tailscale-up
221 // join block — the node comes up as bare yubaba.
222 let machine = sample_machine();
223 let dir = tempfile::tempdir().unwrap();
224 let req = build_request(
225 dir.path(),
226 &machine,
227 "https://example.com/yah-yubaba".into(),
228 "deadbeef".into(),
229 cloud_init::DEFAULT_YUBABA_CHANNEL.into(),
230 None, // standalone: no preauth
231 None, // standalone: no mesh_url
232 None,
233 None,
234 )
235 .unwrap();
236 assert!(
237 !req.user_data.contains("tailscale up --auth-key"),
238 "standalone node must not emit the tailscale-up join block"
239 );
240 // Prose in the template header mentions --login-server; the real arg
241 // form (`--login-server https://`) must be absent.
242 assert!(!req.user_data.contains("--login-server https://"));
243 }
244
245 #[test]
246 fn build_request_rejects_unknown_location() {
247 let mut machine = sample_machine();
248 machine.location = Some("moon".into());
249 let dir = tempfile::tempdir().unwrap();
250 let err = build_request(
251 dir.path(),
252 &machine,
253 "x".into(),
254 "y".into(),
255 "stable".into(),
256 Some("z".into()),
257 None,
258 None,
259 None,
260 )
261 .unwrap_err()
262 .to_string();
263 assert!(err.contains("unknown location"), "unexpected: {err}");
264 }
265
266 #[test]
267 fn build_request_threads_cosign_identity_into_render() {
268 // R330-F21: when an identity_regexp is passed, the rendered cloud-init
269 // emits the cosign verify-blob block. Without it (the existing
270 // build_req_defaults helper passes None) the block stays empty.
271 let machine = sample_machine();
272 let dir = tempfile::tempdir().unwrap();
273 let req = build_request(
274 dir.path(),
275 &machine,
276 "https://cdn.yah.dev/yubaba/0.9.0/x86_64-unknown-linux-musl/yah-yubaba-x86_64-unknown-linux-musl.tar.gz".into(),
277 "deadbeef".into(),
278 cloud_init::DEFAULT_YUBABA_CHANNEL.into(),
279 None,
280 None,
281 None,
282 Some(r"^https://github\.com/yah-ai/yah/".into()),
283 )
284 .unwrap();
285 assert!(
286 req.user_data
287 .contains("cosign verify-blob --certificate-identity-regexp"),
288 "verify-blob runcmd missing once identity_regexp is threaded"
289 );
290 assert!(
291 req.user_data.contains(r"^https://github\.com/yah-ai/yah/"),
292 "identity_regexp value missing from rendered output"
293 );
294 }
295}