Skip to main content

cloud/reconciler/
local_process.rs

1//! `[providers.compute] kind = "local-process"` — run a service's own compute
2//! component as a **kamaji-supervised host process** at the dev tier.
3//!
4//! ## Why this exists
5//!
6//! A component's `kind` says what the service *is*; the mirror's provider slot
7//! says how *this tier* runs it. That split already exists for storage —
8//! `mesofact-static` runs off `local-static` (files on disk) at dev and
9//! `miniflare-container` (containers) at pond, one component kind, two
10//! runtimes — and this is the same split for compute. A `kind = "container"`
11//! component now runs natively at dev and in docker at pond without the
12//! service declaring itself twice.
13//!
14//! The alternative — a `local-process` *component kind* — was rejected: it
15//! would make the artifact shape a property of the service rather than of the
16//! tier, which is exactly the fork that W265 exists to prevent.
17//!
18//! ## Why not just run the container at dev
19//!
20//! Because a container is a different filesystem, and at the dev tier that is
21//! a lie you have to maintain. yah-cloud-admin is the worked example: it reads
22//! `.yah/infra/machines/*.toml`, so its dev container has to bind-mount
23//! `.yah/infra` read-only into `/workspace/.yah/infra` and override
24//! `YAH_CLOUD_ADMIN_WORKSPACE_ROOT` to find it (see
25//! `crates/yah/cloud-admin/workload.toml`). Get the mount wrong and the
26//! service comes up *running, healthy, and describing a fleet of zero
27//! machines*. A host process inherits the operator's actual workspace and the
28//! whole class of failure disappears — along with a docker build per edit.
29//!
30//! ## Config
31//!
32//! The mirror opts in:
33//!
34//! ```toml
35//! # .yah/services/<svc>/mirrors/dev.toml
36//! shape = "local"
37//!
38//! [providers.compute]
39//! kind = "local-process"
40//! ```
41//!
42//! and the component's `workload.toml` describes the process:
43//!
44//! ```toml
45//! [process]
46//! # Cargo package to build before spawning. Omit for a prebuilt binary.
47//! cargo_package = "yah-cloud-admin"
48//! # Binary to exec, relative to the workspace root. Defaults to
49//! # target/<profile>/<cargo_package>.
50//! bin = "target/debug/yah-cloud-admin"
51//! # Extra argv after the binary.
52//! args = []
53//! # Port the process listens on. Required — it is the readiness signal.
54//! port = 4325
55//!
56//! [process.env]
57//! YAH_CLOUD_ADMIN_ADDR = "127.0.0.1:4325"
58//! ```
59//!
60//! @yah:relay(R715, "local-process compute provider: a dev tier that runs the binary, not a container")
61//! @yah:at(2026-08-03T22:21:13Z)
62//! @yah:status(open)
63//! @yah:assignee(agent:bundle-anthropic-ashguard)
64//! @yah:next("T1 (landed): LocalProcessReconciler + Provider::LocalProcess + native_support dedup + both dispatchers + the yah-cloud-admin dev/pond split. See the T1 handoff.")
65//! @yah:next("Follow-on: mesofact-dev's local-static arm and this reconciler now differ only in argv construction. W265 already plans to retire local-static when local-s3-fs lands - that is the moment to consider collapsing them.")
66//! @yah:next("Follow-on: the desktop mirror_run_up has no adopt-probe for local-process the way it does for mesofact-dev. Correct today because the reconciler reaps its predecessor, but an adopt path would let a re-click return the running URL without a rebuild.")
67//! @yah:gotcha("Naming rule this relay encodes, from the operator: if it runs in a container it is pond. Dev is the tier that runs against the operator's real filesystem. No service is required to have all three tiers - a service whose lowest tier is pond is fine.")
68//! @yah:gotcha("The runtime is a property of the MIRROR, not the component. A kind=container component runs natively at dev and in docker at pond, selected by the mirror's compute slot - the same split mesofact-static already had via providers.static (local-static at dev, miniflare-container at pond). Do NOT add a local-process COMPONENT kind: that makes artifact shape a property of the service and reintroduces the per-tier fork W265 exists to prevent.")
69//! @arch:see(.yah/docs/working/W265-service-capabilities-and-drivers.md)
70//!
71//! @yah:ticket(R715-T1, "LocalProcessReconciler + Provider::LocalProcess, wired into both dispatchers; yah-cloud-admin dev/pond split")
72//! @yah:status(review)
73//! @yah:at(2026-08-05T04:54:16Z)
74//! @yah:assignee(agent:bundle-anthropic-ashguard)
75//! @yah:phase(P1)
76//! @yah:parent(R715)
77//! @yah:next("R714-B1 (stop button no-ops on container mirrors) and R714-B2 (tier chip prints 'axum') were found during this work and filed separately - they predate it and are not regressions from it.")
78//! @yah:verify("cargo test --manifest-path oss/yubaba/Cargo.toml -p yah-cloud --lib - 688 passed, 0 failed")
79//! @yah:verify("cargo test -p yah --lib cloud:: - 108 passed")
80//! @yah:verify("cargo test -p xtask --test schema_drift --test workload_envelope - 4 passed")
81//! @yah:verify("MANUAL (done): yah cloud mirror up yah-cloud-admin --env dev binds 127.0.0.1:4325 with no container, and the process logs workspace=/Users/leif/ss/yah - the real checkout, not /workspace")
82//! @yah:verify("MANUAL (done): running mirror up twice in a row replaces the child (pid changes) instead of leaving a dead second instance behind an Address-already-in-use")
83//! @yah:gotcha("NativeRuntime's workload table is IN-MEMORY, so a second `mirror up` from a fresh process cannot see the first one's child. Without the owner.json sidecar + reap the new child dies on Address-already-in-use, wait_for_port sees the OLD listener answering, and the reconcile reports success while the just-edited binary is not what is running. This was observed live before the fix, not theorised. Do not remove the sidecar.")
84//! @yah:gotcha("yah-cloud-admin's pond tier moved to host port 4326 so both tiers can run at once. The container still listens on 4325 internally.")
85//! @yah:next("RESOLVED (R715-T1, 2026-08-04): the cloud.toml blocker comment was corrected by the peer whose uncommitted edits blocked it - lines 34-60 now carry the 2026-08-03 status (published image DONE, workload declaration DONE, cheers key STILL BLOCKED on both an issuer and a 0.8.21 yubaba roll). Nothing left to do there.")
86//! @yah:handoff("Re-verified independently on 2026-08-04 (session:5fe4274e, rune) rather than trusting the prior run's claims. All three automated suites green; counts are HIGHER than the ones recorded above because peers added tests since: yah-cloud lib 698 passed / 0 failed (was 688), yah lib cloud:: 111 passed (was 108), xtask schema_drift + workload_envelope 4 passed (unchanged).")
87//! @yah:verify("RE-VERIFIED 2026-08-04, sidecar survives across sessions: the owner.json left by the 2026-08-03 run ({pid:54869,port:4325}) still matched the live listener a day later - lsof -t on 4325 returned exactly 54869. That is the cross-process ownership link working in the wild, not in a test.")
88//! @yah:verify("RE-VERIFIED 2026-08-04, reap-and-replace: `cargo run -p yah -- cloud mirror up yah-cloud-admin --env dev` reaped pid 54869 (kill -0 confirms gone), spawned 71455, rewrote owner.json to {pid:71455,port:4325}, and 127.0.0.1:4325/ answers HTTP 200. No dead second instance, no Address-already-in-use.")
89//! @yah:verify("RE-VERIFIED 2026-08-04, the failure mode this tier exists to prevent: child cwd is /Users/leif/ss/yah (lsof -d cwd), `docker ps` shows NO cloud-admin container, and /partial/machines renders 9 machine rows against 9 files in .yah/infra/machines/ - us-west-001/002/003 + mac-builder among them. Not a fleet of zero.")
90//! @yah:verify("RE-VERIFIED 2026-08-04, both dispatcher arms compile: `cargo build -p yah` (app/yah/cli/src/cloud.rs:4705) and `cargo check -p desktop --lib` (app/yah/desktop/src/mirror_run.rs:624) are both clean.")
91//! @yah:gotcha("LEFT RUNNING: the re-verification above spawned yah-cloud-admin pid 71455 on 127.0.0.1:4325 and did not stop it - that matches the state found at session start (a dev-tier process was already up). Kill it with `kill $(lsof -nP -iTCP:4325 -sTCP:LISTEN -t)`, or just re-run `mirror up`, which reaps it.")
92//! @yah:gotcha("TRANSIENT, NOT THIS TICKET: `cargo run -p yah` failed once with E0599 mid-verification and compiled clean on immediate retry with no edit from this session - a peer's in-flight change on the shared tree. Same shape as the note in crates/yah/camp-service/tests/e2e.rs:35. If you see it, retry before investigating.")
93
94use std::collections::BTreeMap;
95use std::net::{IpAddr, Ipv4Addr, SocketAddr};
96use std::path::PathBuf;
97use std::sync::Arc;
98use std::time::Duration;
99
100use anyhow::{bail, Context, Result};
101use async_trait::async_trait;
102use kamaji::native::NativeRuntime;
103use kamaji::{Kamaji, MeshAssignment, MeshIdent};
104use serde::Deserialize;
105use tokio::sync::oneshot;
106use tracing::{info, warn};
107use workload_spec::EnvVar;
108
109use super::native_support::{
110    capture_paths, native_spec, sanitize_ident, spawn_native_log_supervisor,
111};
112use super::{into_running, wait_for_port, LogBuffer, ReconcileCtx, Reconciler, RunningWorkload};
113use crate::{MirrorProviderSlot, MirrorShape, Provider};
114
115/// The slot role a natively-run compute component occupies on its mirror.
116const SLOT: &str = "compute";
117
118/// How long to wait for the process to bind its declared port before calling
119/// the reconcile failed. Generous because the first `cargo build` of a cold
120/// target dir is included in the caller's patience, not in this window — the
121/// build finishes before the spawn.
122const READY_TIMEOUT: Duration = Duration::from_secs(20);
123
124/// True when `mirror` binds its compute slot to `local-process`. This is the
125/// dispatch predicate — callers use it to choose this reconciler over the
126/// container one for the same component kind.
127pub fn slot_declared(mirror: &crate::MirrorConfig) -> bool {
128    matches!(
129        mirror.providers.get(SLOT),
130        Some(MirrorProviderSlot::Inline {
131            kind: Provider::LocalProcess,
132            ..
133        })
134    )
135}
136
137/// Reconciler for components bound to a `local-process` compute slot.
138#[derive(Debug, Default)]
139pub struct LocalProcessReconciler;
140
141impl LocalProcessReconciler {
142    pub fn new() -> Self {
143        Self::default()
144    }
145}
146
147/// On-disk `workload.toml` shape — only the `[process]` section is read here.
148/// Other sections (`[build]`, `[run]`) belong to the container reconciler and
149/// are ignored, so one component file can describe both tiers.
150#[derive(Debug, Default, Deserialize)]
151struct ProcessComponent {
152    #[serde(default)]
153    process: Option<ProcessSpec>,
154}
155
156#[derive(Debug, Deserialize)]
157struct ProcessSpec {
158    /// Cargo package to `cargo build -p` before spawning. `None` → the binary
159    /// is expected to already exist.
160    #[serde(default)]
161    cargo_package: Option<String>,
162    /// Binary path relative to the workspace root (absolute taken as-is).
163    /// `None` → `target/<profile>/<cargo_package>`.
164    #[serde(default)]
165    bin: Option<String>,
166    /// Extra argv appended after the binary.
167    #[serde(default)]
168    args: Vec<String>,
169    /// Port the process listens on. Required: it is how readiness is decided,
170    /// and a component with no port has nothing for the Run tab to open.
171    port: u16,
172    /// Environment for the child.
173    #[serde(default)]
174    env: BTreeMap<String, String>,
175    /// Cargo profile used for both the build and the default binary path.
176    #[serde(default = "default_profile")]
177    profile: String,
178}
179
180fn default_profile() -> String {
181    "debug".to_string()
182}
183
184#[async_trait]
185impl Reconciler for LocalProcessReconciler {
186    fn kind(&self) -> &'static str {
187        "local-process"
188    }
189
190    async fn up(&self, ctx: ReconcileCtx<'_>) -> Result<RunningWorkload> {
191        ctx.materialize().await?;
192
193        // A host process is the operator's own machine by definition. Refusing
194        // non-local shapes here keeps a `local-process` slot from silently
195        // meaning "run it on my laptop" in a mirror that describes a fleet.
196        if !matches!(ctx.mirror.shape, MirrorShape::Local) {
197            bail!(
198                "component {}: `local-process` is a dev-tier compute slot — mirror shape is \
199                 {:?}, not `local`. Deploy the cloud tier via `yah cloud workload deploy`.",
200                ctx.component.id,
201                ctx.mirror.shape,
202            );
203        }
204
205        let spec = load_process_spec(&ctx)?;
206
207        // Build first, if asked. Doing this before the port probe means a
208        // compile error surfaces as a compile error, rather than as a bind
209        // timeout on a binary that was never rebuilt.
210        if let Some(pkg) = &spec.cargo_package {
211            cargo_build(ctx.workspace_root, pkg, &spec.profile).await?;
212        }
213
214        let bin = resolve_binary(&spec, ctx.workspace_root)?;
215
216        // Kamaji's native backend captures stdio under this dir; scope it per
217        // workspace so concurrent camps don't collide on the same ident.
218        let state_dir = ctx.workspace_root.join(".yah/jit/native");
219        let ident_str = sanitize_ident(&format!(
220            "local-process-{}-{}-{}",
221            ctx.service.name, ctx.env, ctx.component.id
222        ));
223        let ident = MeshIdent(ident_str.clone());
224
225        // Replace any predecessor before spawning. `ContainerReconciler` has
226        // the same semantics ("run clears any prior container of the same name
227        // first"), and here it is load-bearing rather than tidy: NativeRuntime
228        // keeps its workload table **in memory**, so a second `mirror up` from
229        // a fresh process knows nothing about the first. Without this the new
230        // child dies on `Address already in use`, `wait_for_port` sees the
231        // *old* listener still answering, and the reconcile reports success
232        // while the binary the operator just edited is not the one running.
233        // That is the R602-B4 foreign-listener failure with a friendlier
234        // disguise, and on the dev tier it is the single most costly thing
235        // this reconciler could get wrong.
236        let owner_path = state_dir.join(&ident_str).join("owner.json");
237        reap_predecessor(&owner_path).await;
238
239        // Anything still holding the port now is not ours. Refuse rather than
240        // adopt it: a `dev_url` that answers with someone else's service is
241        // worse than a failed bring-up.
242        let addr = SocketAddr::new(IpAddr::V4(Ipv4Addr::LOCALHOST), spec.port);
243        if tokio::net::TcpStream::connect(addr).await.is_ok() {
244            bail!(
245                "component {}: {addr} is already held by a process this reconciler did not \
246                 start. Stop it, or give [process] a different `port`.",
247                ctx.component.id,
248            );
249        }
250
251        let mut argv = vec![bin.display().to_string()];
252        argv.extend(spec.args.iter().cloned());
253
254        let env: Vec<EnvVar> = spec
255            .env
256            .iter()
257            .map(|(name, value)| EnvVar {
258                name: name.clone(),
259                value: workload_spec::EnvValue::Literal {
260                    value: value.clone(),
261                },
262            })
263            .collect();
264
265        let workload = native_spec(&ident_str, argv, env);
266        let runtime = Arc::new(NativeRuntime::new(&state_dir));
267        let mesh = MeshAssignment::inlined(Ipv4Addr::LOCALHOST);
268
269        info!(
270            binary = %bin.display(),
271            port = spec.port,
272            ident = %ident_str,
273            "spawning local-process component (kamaji native backend)",
274        );
275
276        let deployed = runtime
277            .deploy_workload(&workload, &mesh)
278            .await
279            .with_context(|| {
280                format!(
281                    "deploying component {} via kamaji native backend (binary {})",
282                    ctx.component.id,
283                    bin.display(),
284                )
285            })?;
286
287        let (stdout_path, stderr_path) = capture_paths(&state_dir, &ident_str);
288
289        // Record ownership so the *next* reconcile — in a different process,
290        // with a different in-memory NativeRuntime — can find and reap this
291        // child instead of colliding with it.
292        write_owner(&owner_path, deployed.task_pid as i32, spec.port).await;
293
294        // Wait for the bind. A dead process must not hand the Run tab a URL
295        // that will never answer.
296        if !wait_for_port(addr, READY_TIMEOUT).await {
297            warn!(addr = %addr, "local-process did not bind within timeout; tearing down");
298            let tail = read_capture_tail(&stdout_path, &stderr_path).await;
299            runtime.teardown_workload(&ident).await.ok();
300            bail!(
301                "component {} did not bind {addr} within {READY_TIMEOUT:?}{tail}",
302                ctx.component.id,
303            );
304        }
305
306        let dev_url = format!("http://{addr}");
307        info!(dev_url = %dev_url, pid = deployed.task_pid, "local-process ready");
308
309        let log_buf = LogBuffer::new();
310        let (shutdown_tx, shutdown_rx) = oneshot::channel::<()>();
311        let supervisor = spawn_native_log_supervisor(
312            runtime,
313            ident,
314            log_buf.clone(),
315            stdout_path,
316            stderr_path,
317            shutdown_rx,
318        );
319
320        Ok(into_running(
321            "local-process",
322            SLOT,
323            Some(dev_url),
324            None,
325            Some(log_buf),
326            shutdown_tx,
327            supervisor,
328        ))
329    }
330}
331
332/// Read `<workload_dir>/workload.toml` and pull out its `[process]` section.
333fn load_process_spec(ctx: &ReconcileCtx<'_>) -> Result<ProcessSpec> {
334    let path = ctx.workload_dir().join("workload.toml");
335    let src =
336        std::fs::read_to_string(&path).with_context(|| format!("reading {}", path.display()))?;
337    let parsed: ProcessComponent =
338        toml::from_str(&src).with_context(|| format!("parsing {}", path.display()))?;
339    parsed.process.with_context(|| {
340        format!(
341            "component {} is bound to a `local-process` compute slot but {} declares no \
342             [process] section — add one with at least `port`",
343            ctx.component.id,
344            path.display(),
345        )
346    })
347}
348
349/// Resolve the binary to exec: explicit `bin`, else the cargo convention.
350fn resolve_binary(spec: &ProcessSpec, workspace_root: &std::path::Path) -> Result<PathBuf> {
351    let rel = match (&spec.bin, &spec.cargo_package) {
352        (Some(bin), _) => PathBuf::from(bin),
353        (None, Some(pkg)) => PathBuf::from(format!("target/{}/{pkg}", spec.profile)),
354        (None, None) => bail!(
355            "[process] declares neither `bin` nor `cargo_package` — one is needed to know \
356             what to run"
357        ),
358    };
359    let path = if rel.is_absolute() {
360        rel
361    } else {
362        workspace_root.join(rel)
363    };
364    if !path.exists() {
365        bail!(
366            "[process] binary {} does not exist — set `cargo_package` to have it built, or \
367             point `bin` at an existing file",
368            path.display(),
369        );
370    }
371    Ok(path)
372}
373
374/// `cargo build -p <pkg>` in the workspace root, surfacing compiler output on
375/// failure. Inherits stdout/stderr is *not* an option here (the desktop has no
376/// console), so output is captured and the tail is folded into the error.
377async fn cargo_build(workspace_root: &std::path::Path, pkg: &str, profile: &str) -> Result<()> {
378    let mut cmd = tokio::process::Command::new("cargo");
379    cmd.arg("build").arg("-p").arg(pkg);
380    if profile == "release" {
381        cmd.arg("--release");
382    } else if profile != "debug" {
383        cmd.arg("--profile").arg(profile);
384    }
385    cmd.current_dir(workspace_root);
386
387    info!(
388        package = pkg,
389        profile, "cargo build for local-process component"
390    );
391    let out = cmd
392        .output()
393        .await
394        .with_context(|| format!("spawning `cargo build -p {pkg}`"))?;
395    if !out.status.success() {
396        let stderr = String::from_utf8_lossy(&out.stderr);
397        let tail: Vec<&str> = stderr.lines().rev().take(30).collect();
398        let tail: Vec<&str> = tail.into_iter().rev().collect();
399        bail!("cargo build -p {pkg} failed:\n{}", tail.join("\n"));
400    }
401    Ok(())
402}
403
404/// What a previous `up()` recorded about the child it left running.
405///
406/// This exists because [`NativeRuntime`]'s workload table is in-memory: a
407/// `yah cloud mirror up` from the shell and a ▶ click in the desktop are
408/// different processes, and neither can see the other's children. The sidecar
409/// is the only thing that makes "replace the predecessor" possible across them.
410#[derive(Debug, serde::Serialize, Deserialize)]
411struct OwnerRecord {
412    pid: i32,
413    port: u16,
414}
415
416/// Record the child we just spawned. Best-effort: failing to write the sidecar
417/// must not fail an otherwise-successful bring-up — the cost is a manual kill
418/// on the next re-run, not a broken mirror.
419async fn write_owner(path: &std::path::Path, pid: i32, port: u16) {
420    if let Some(dir) = path.parent() {
421        if tokio::fs::create_dir_all(dir).await.is_err() {
422            return;
423        }
424    }
425    if let Ok(json) = serde_json::to_vec(&OwnerRecord { pid, port }) {
426        if let Err(e) = tokio::fs::write(path, json).await {
427            warn!(path = %path.display(), error = %e, "could not record local-process owner");
428        }
429    }
430}
431
432/// Stop the child a previous `up()` left behind, if it is still alive.
433///
434/// SIGTERM, a short grace, then SIGKILL — the same ladder kamaji's own teardown
435/// uses. The sidecar is removed either way, including when the pid is already
436/// gone, so a stale record can't linger and make the next run think it has a
437/// predecessor to wait on.
438async fn reap_predecessor(owner_path: &std::path::Path) {
439    let Ok(bytes) = tokio::fs::read(owner_path).await else {
440        return;
441    };
442    let _ = tokio::fs::remove_file(owner_path).await;
443    let Ok(owner) = serde_json::from_slice::<OwnerRecord>(&bytes) else {
444        return;
445    };
446    if !pid_alive(owner.pid) {
447        return;
448    }
449
450    info!(
451        pid = owner.pid,
452        port = owner.port,
453        "replacing predecessor local-process"
454    );
455    signal_pid(owner.pid, libc::SIGTERM);
456    for _ in 0..50 {
457        if !pid_alive(owner.pid) {
458            return;
459        }
460        tokio::time::sleep(Duration::from_millis(100)).await;
461    }
462    warn!(
463        pid = owner.pid,
464        "predecessor ignored SIGTERM; sending SIGKILL"
465    );
466    signal_pid(owner.pid, libc::SIGKILL);
467    // Give the kernel a moment to release the port before we probe it.
468    tokio::time::sleep(Duration::from_millis(200)).await;
469}
470
471/// `kill(pid, 0)` — true when a process with this pid exists and we may signal
472/// it. Cannot prove the pid is still *our* child (pids are reused), which is
473/// why the sidecar is written next to this workload's capture dir and removed
474/// on every read: the window where a recycled pid could be signalled is one
475/// reconcile wide, and the alternative (never reaping) is a guaranteed
476/// collision rather than a theoretical one.
477fn pid_alive(pid: i32) -> bool {
478    // SAFETY: `kill` with signal 0 performs error checking only and never
479    // delivers a signal; any pid value is a defined input.
480    unsafe { libc::kill(pid, 0) == 0 }
481}
482
483fn signal_pid(pid: i32, sig: i32) {
484    // SAFETY: same contract as above — `kill` is defined for any pid/signal
485    // pair and reports failure through its return value, which we ignore
486    // because a vanished process is the outcome we wanted anyway.
487    unsafe {
488        libc::kill(pid, sig);
489    }
490}
491
492/// Last few lines of the capture files, formatted for an error message. A bind
493/// timeout with no output is nearly impossible to act on; the process almost
494/// always said why on the way down.
495async fn read_capture_tail(stdout_path: &std::path::Path, stderr_path: &std::path::Path) -> String {
496    let mut lines: Vec<String> = Vec::new();
497    for path in [stderr_path, stdout_path] {
498        if let Ok(s) = tokio::fs::read_to_string(path).await {
499            lines.extend(s.lines().rev().take(10).map(str::to_string));
500        }
501    }
502    if lines.is_empty() {
503        return String::new();
504    }
505    lines.reverse();
506    format!("\nlast output:\n{}", lines.join("\n"))
507}
508
509#[cfg(test)]
510mod tests {
511    use super::*;
512
513    fn spec(bin: Option<&str>, pkg: Option<&str>) -> ProcessSpec {
514        ProcessSpec {
515            cargo_package: pkg.map(str::to_string),
516            bin: bin.map(str::to_string),
517            args: vec![],
518            port: 4325,
519            env: BTreeMap::new(),
520            profile: "debug".to_string(),
521        }
522    }
523
524    #[test]
525    fn parses_a_process_section_alongside_container_sections() {
526        // One workload.toml describes both tiers; each reconciler reads its own
527        // section and ignores the other's.
528        let src = r#"
529schema_version = 1
530kind = "container"
531
532[build]
533image = "yah-local/x:dev"
534
535[run]
536port = 4325
537
538[process]
539cargo_package = "yah-cloud-admin"
540port = 4325
541args = ["--verbose"]
542
543[process.env]
544YAH_CLOUD_ADMIN_ADDR = "127.0.0.1:4325"
545"#;
546        let c: ProcessComponent = toml::from_str(src).unwrap();
547        let p = c.process.unwrap();
548        assert_eq!(p.cargo_package.as_deref(), Some("yah-cloud-admin"));
549        assert_eq!(p.port, 4325);
550        assert_eq!(p.args, vec!["--verbose".to_string()]);
551        assert_eq!(p.profile, "debug");
552        assert_eq!(
553            p.env.get("YAH_CLOUD_ADMIN_ADDR").map(String::as_str),
554            Some("127.0.0.1:4325")
555        );
556    }
557
558    #[test]
559    fn a_workload_without_a_process_section_parses_as_none() {
560        let c: ProcessComponent = toml::from_str("[run]\nport = 1\n").unwrap();
561        assert!(c.process.is_none());
562    }
563
564    #[test]
565    fn binary_defaults_to_the_cargo_convention() {
566        let tmp = tempfile::tempdir().unwrap();
567        std::fs::create_dir_all(tmp.path().join("target/debug")).unwrap();
568        std::fs::write(tmp.path().join("target/debug/yah-cloud-admin"), b"").unwrap();
569        let got = resolve_binary(&spec(None, Some("yah-cloud-admin")), tmp.path()).unwrap();
570        assert_eq!(got, tmp.path().join("target/debug/yah-cloud-admin"));
571    }
572
573    #[test]
574    fn explicit_bin_wins_over_the_cargo_convention() {
575        let tmp = tempfile::tempdir().unwrap();
576        std::fs::create_dir_all(tmp.path().join("bin")).unwrap();
577        std::fs::write(tmp.path().join("bin/custom"), b"").unwrap();
578        let got = resolve_binary(&spec(Some("bin/custom"), Some("pkg")), tmp.path()).unwrap();
579        assert_eq!(got, tmp.path().join("bin/custom"));
580    }
581
582    /// Spawning a path that isn't there fails with an exec error several
583    /// layers down; naming the missing file here is the actionable version.
584    #[test]
585    fn a_missing_binary_is_named_before_we_try_to_exec_it() {
586        let tmp = tempfile::tempdir().unwrap();
587        let err = resolve_binary(&spec(None, Some("nope")), tmp.path()).unwrap_err();
588        assert!(err.to_string().contains("does not exist"), "{err}");
589    }
590
591    #[test]
592    fn neither_bin_nor_package_is_an_error() {
593        let tmp = tempfile::tempdir().unwrap();
594        let err = resolve_binary(&spec(None, None), tmp.path()).unwrap_err();
595        assert!(err.to_string().contains("neither"), "{err}");
596    }
597
598    /// The bug this whole path exists to prevent: a second `up()` in a fresh
599    /// process must stop the first one's child. NativeRuntime's table is
600    /// in-memory, so the sidecar is the only link between the two runs.
601    ///
602    /// The assertion is on the child's exit status, not on [`pid_alive`],
603    /// because this test *is* the child's parent: a signalled child it has not
604    /// waited on stays a zombie, and `kill(pid, 0)` succeeds against a zombie.
605    /// In production the spawning process is either gone (CLI — init reaps) or
606    /// still holding kamaji's supervisor task (desktop — that reaps), so
607    /// neither leaves one behind.
608    #[tokio::test]
609    async fn reap_stops_a_live_predecessor_and_clears_the_record() {
610        let tmp = tempfile::tempdir().unwrap();
611        let owner_path = tmp.path().join("owner.json");
612
613        let mut child = tokio::process::Command::new("sleep")
614            .arg("120")
615            .spawn()
616            .unwrap();
617        let pid = child.id().unwrap() as i32;
618        assert!(pid_alive(pid));
619
620        write_owner(&owner_path, pid, 4325).await;
621        reap_predecessor(&owner_path).await;
622
623        let status = tokio::time::timeout(Duration::from_secs(5), child.wait())
624            .await
625            .expect("a signalled `sleep 120` must have exited well inside 5s")
626            .unwrap();
627        assert!(!status.success(), "predecessor should have been signalled");
628        assert!(!owner_path.exists(), "sidecar should be cleared");
629    }
630
631    /// A record left behind by a crash names a pid that is gone. Reaping it
632    /// must be a no-op that still clears the file — a stale record that
633    /// survives would make every later run wait on a corpse.
634    #[tokio::test]
635    async fn reap_clears_a_stale_record_without_signalling_anything() {
636        let tmp = tempfile::tempdir().unwrap();
637        let owner_path = tmp.path().join("owner.json");
638
639        // Exited process: spawn and wait, so the pid is definitely dead.
640        let mut child = tokio::process::Command::new("true").spawn().unwrap();
641        let pid = child.id().unwrap() as i32;
642        child.wait().await.unwrap();
643
644        write_owner(&owner_path, pid, 4325).await;
645        reap_predecessor(&owner_path).await;
646        assert!(!owner_path.exists());
647    }
648
649    #[tokio::test]
650    async fn reap_with_no_record_is_a_no_op() {
651        let tmp = tempfile::tempdir().unwrap();
652        reap_predecessor(&tmp.path().join("owner.json")).await;
653    }
654
655    #[test]
656    fn slot_declared_only_matches_the_local_process_compute_slot() {
657        let mut m = crate::MirrorConfig {
658            schema_version: 1,
659            shape: MirrorShape::Local,
660            ingress: Default::default(),
661            providers: Default::default(),
662            drivers: Default::default(),
663            asset_aliases: Default::default(),
664        };
665        assert!(!slot_declared(&m));
666        m.providers.insert(
667            SLOT.to_string(),
668            MirrorProviderSlot::Inline {
669                kind: Provider::LocalContainer,
670                fields: Default::default(),
671            },
672        );
673        assert!(!slot_declared(&m), "a container slot is not a process slot");
674        m.providers.insert(
675            SLOT.to_string(),
676            MirrorProviderSlot::Inline {
677                kind: Provider::LocalProcess,
678                fields: Default::default(),
679            },
680        );
681        assert!(slot_declared(&m));
682    }
683}