Skip to main content

MachineConfig

Struct MachineConfig 

Source
pub struct MachineConfig {
Show 23 fields pub name: String, pub provider: String, pub vendor: Option<String>, pub nickname: Option<String>, pub location: Option<String>, pub server_type: Option<String>, pub hosts_mirrors: Vec<String>, pub mesh_tags: Vec<String>, pub region: Option<String>, pub zone: Option<String>, pub arch: Option<String>, pub bucket: Option<BucketSpec>, pub legacy_hostkey_fingerprint: Option<String>, pub ssh_keys: Vec<u64>, pub cloudflared: Option<String>, pub ingress_floating_ip: Option<String>, pub hosts_operator_bridge: bool, pub connect: Option<ConnectSpec>, pub allocatable: Option<NodeAllocatable>, pub taints: Vec<String>, pub sovereign_group: Option<String>, pub sovereign_role: Option<SovereignRole>, pub registration: MachineRegistration,
}
Expand description

Per-machine TOML from .yah/infra/machines/<name>.toml.

Two halves, split by provenance (R707-T1): everything here is declaration — operator intent under review and blame — except [registration], which carries what the fleet observed. See MachineRegistration for why the boundary is drawn there and what depends on it.

@yah:ticket(R860-T5, “Model per-node native-exec capability as an admission axis (W338 §Placement consequences 3 / R858-T4 gap)”) @yah:status(review) @yah:phase(P1) @yah:at(2026-09-05T18:29:19Z) @yah:assignee(agent:bundle-anthropic-ashguard) @yah:parent(R860) @yah:next(“Cheapest defensible shape: express it on MachineConfig, which already has the two vocabularies — mesh_tags: Vec<String> (config.rs:246, superset match, already carries arch:/os:/tag:build-worker) and taints: Vec<String> (config.rs:337). A native-exec mesh tag required by any group member whose kind is native is a one-line admission axis in admission_spec(). Whichever is chosen, it must be declared in .yah/infra/machines/*.toml for the nodes that actually run kamaji with –native-exec-dir, and check_inert_taints (config.rs:703) lints unread taint keys dead — so a taint nobody reads will be flagged.”) @yah:verify(“cargo test -p cloud –lib config”) @arch:see(.yah/docs/working/W338-workload-dependencies-and-appliance-composition.md) @yah:depends_on(R860-T4) @yah:gotcha(“Verified 2026-09-04: native-exec capability is modelled NOWHERE in placement — rg \"native\" oss/yubaba/crates/cloud/src/config.rs returns zero hits, and the raft state machine models no member attributes, labels or taints at all (rg \"taint|capabilit|labels|mesh_tag\" over raft/{mod,store,network}.rs yields one unrelated comment at raft/store.rs:591). Native-exec is a node-local kamaji startup decision today: --native-exec-dir (oss/kamaji/crates/kamaji-bin/src/main.rs:152-156, :51-55) plus the native-exec cargo feature (kamaji-bin/src/server.rs:329-330). A node without it refuses the deploy at dispatch time and nothing upstream can see that in advance — which is exactly the deploy-time surprise W338 wants turned into a placement precondition.”) @yah:handoff(“NATIVE-EXEC IS NOW A PLACEMENT PRECONDITION, NOT A DISPATCH-TIME SURPRISE. New pub const NATIVE_EXEC_MESH_TAG: &str = \"cap:native-exec\" in oss/yubaba/crates/cloud/src/config.rs (declared just above node_selector_mesh_tags), and one axis in admission_spec() immediately after the R860-T4 group loop: if ANY member of placement_group(ws, declared) returns true from WorkloadSpec::wants_native_exec(), the tag is appended to the derived RequiredSpec.mesh_tags (deduped). No new field on RequiredSpec, no signature change anywhere, no wire or serde change — the mesh_tags axis is already an AND-ed superset check against machine.mesh_tags in matches and is already rendered by describe, so a refusal now reads required.mesh_tags=[...,cap:native-exec].”) @yah:handoff(“ITEM 1 — HOW A NATIVE WORKLOAD IS DETECTED, settled by opening the type rather than guessing. There is no kind on WorkloadSpec: on the wire a native workload is still Workload::Container(WorkloadSpec), and the ONLY difference is the annotation yah.exec = native, read through WorkloadSpec::wants_native_exec() (oss/yah-base/crates/workload-spec/src/lib.rs:2939; consts NATIVE_EXEC_ANNOTATION / NATIVE_EXEC_VALUE at :3402/:3407). That accessor is what the admission axis calls — matching kamaji, whose deploy_container checks the same marker first and routes to deploy_native_exec (oss/kamaji/crates/kamaji-bin/src/server.rs). The yah.exec key is a substrate selector with a second value, microvm (wants_microvm, same key, R605-F8), so per-node microVM capability is the obvious sibling axis and is NOT modelled here — see next-steps.”) @yah:handoff(“ITEM 2 — DECLARATIONS LANDED ON TWO NODES, FROM READINGS RECORDED IN-REPO, NOT INFERRED. cap:native-exec added to mesh_tags in .yah/infra/machines/us-west-001.toml and .yah/infra/machines/us-west-003.toml, each with a comment naming its evidence and its re-check condition. us-west-001: the R858 gotcha in its own header records a ps reading taken on the box 2026-09-05 — pid 515908 is /usr/local/bin/kamaji --native-exec-dir /var/lib/yah/kamaji/native, supervising headscale as a native child. us-west-003: its header’s ‘THE DEPLOYED KAMAJI PREDATES THE microVM BACKEND’ note quotes the box’s actual ExecStart, read over ssh 2026-09-01, carrying --native-exec-dir /var/lib/yah/kamaji/native (corroborated by .yah/docs/architecture/A043-yah-on-machine-daemons.md’s @yah:verify for the same probe). Both comments say plainly that the capability lives in the systemd unit’s ExecStart, not in the TOML, so it must be re-checked after any roll.”) @yah:handoff(“ITEM 2, THE NEGATIVES — TWO NODES ARE KNOWN NOT TO HAVE IT AND WERE DELIBERATELY LEFT UNSET. us-south-001: kamaji refused headscale there 2026-09-03 with ‘native backend not configured — start kamaji with –native-exec-dir’ (the R858 chain, quoted in .yah/infra/machines/us-west-001.toml and W267). I did NOT edit us-south-001.toml — it was already dirty in the working tree at the anchor SHA and @Ashguard:eclipse is live on R858, so I left it alone rather than race it; the mechanism fails closed there, which is the correct state. us-west-015 (the sole darwin builder): W254-darwin-build-nodes.md’s own next-step records that its kamaji is built/started --docker only. I added a comment to us-west-015.toml explaining that the tag is deliberately absent, that this is the node where the axis changes an error message (a darwin build row is native by construction, so it is now refused at ELECTION naming cap:native-exec instead of reaching the box and being refused by kamaji), and the exact enable sequence: rebuild with --features native-exec, restart with --native-exec-dir <dir>, THEN add the tag. us-west-002/011/013/014 are unestablished from the repo and left unset. THE OPERATOR-FACING ANSWER: the file is .yah/infra/machines/<node>.toml and the key is mesh_tags; add the literal string cap:native-exec to that array, and only after the roll.”) @yah:handoff(“DECISIONS THE BRIEF LEFT OPEN, all recorded in doc comments at the site. (1) MESH TAG, NOT TAINT — as recommended, and the doc says why in the terms the brief asked for: mesh tags are positive capability with superset matching (‘this node CAN’), which is the claim being made; a taint is repulsion and would have to be inverted to no-native-exec on every node LACKING the backend (declaration burden on the majority, and silently wrong for a node nobody has edited) AND taught to taint_effect, or check_inert_taints would correctly lint the key dead. (2) THE cap: NAMESPACE IS NEW. Live prefixes are tag: (operator-assigned role), arch:/os: (silicon and userland facts, emitted as requirements by qed::platform::build_worker_mesh_tags), and tier: which R763 RETIRED for architecture and reserved for the environment axis — so reusing any of them would have stated the wrong kind of fact. A capability the daemon was configured with is none of those. Nothing validates tag prefixes (only check_retired_arch_tags looks at one), so this costs no wiring. (3) COMPUTED OVER THE GROUP, not the requirer — that is literally W338’s sentence (‘supply = self specs must be placeable where their requirer lands’), and the second test proves it: an ordinary container requirer with a local edge to a native provider is pulled onto a capable node. (4) FAILS CLOSED, accepted deliberately: an undeclared node is simply not a candidate, so an undeclared fleet reports ‘no node admits’ at election rather than dispatching to a node that refuses. Nothing in .yah/infra/workloads/ is native-marked today (only yah-cloud-admin.toml exists there), so the only live consumer is the qed darwin build row, where failing closed is strictly the better error.”) @yah:handoff(“BLAST RADIUS, MEASURED. admission_spec is private and its callers are unchanged: admit_workload / admit_workload_candidates / admit_workload_in_group (config.rs), reached from app/yah/cli/src/cloud.rs (deploy, rolling, topology analyzer), app/yah/cli/src/yubaba_client.rs elect_node, and cloud/src/migrate.rs. The headscale appliance path inside yubaba (headscale_appliance.rs) does NOT go through admission — it is node-internal — so nothing eclipse holds on R858 is touched by this. Files edited, in full: oss/yubaba/crates/cloud/src/config.rs; .yah/infra/machines/{us-west-001,us-west-003,us-west-015}.toml. Nothing in oss/yubaba/crates/yubaba/ was opened, and oss/kamaji/crates/kamaji-bin/src/server.rs was READ ONLY (to confirm the marker check), per @Ashguard:hydra’s contention triage.”) @yah:handoff(“ONE SCOPE ADDITION, stated loudly rather than slipped in: MachineConfig::mesh_tags (config.rs:256) had NO doc comment at all — the operator-facing declaration key for four tag namespaces was undocumented. I gave it one enumerating tag: / arch:+os: / the new cap: / retired tier:, and noting that nothing validates the prefix (which is why the two lints exist). CONSEQUENCE TO KNOW: that field’s doc is the source of the mesh_tags description in the GENERATED .yah/schema/machine.toml.schema.json, so it is schema-drift-affecting — see the gotcha.”) @yah:handoff(“Tree anchor at handoff: 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2 — the shared tree as I found it and left it. Quote this SHA rather than ‘HEAD’ in any revert/restore instruction; to undo a hunk, read it with git show 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2:<path> and put it back with Edit, never git checkout/restore (they restore whole files and would delete peers’ uncommitted work).”) @yah:handoff(“Tree anchor at handoff: 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2 — the shared tree as I left it. Diff against it (git diff 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2..HEAD) to see what landed under you, and quote this SHA rather than ‘HEAD’ in any revert/restore instruction.”) @yah:next(“MICROVM IS THE IDENTICAL UNMODELLED GAP, one line away. yah.exec is a substrate selector with a second value: WorkloadSpec::wants_microvm() (workload-spec/src/lib.rs, R605-F8), and kamaji constructs MicroVmRuntime only when started with --microvm-dir — A043’s probe records that us-west-003’s deployed kamaji has --native-exec-dir but NOT --microvm-dir, so a microvm-marked deploy is refused there by exactly the same dispatch-time surprise this ticket removed for native. The shape is cap:microvm alongside NATIVE_EXEC_MESH_TAG in the same if in admission_spec. Not done here because no node in the fleet can host one yet (R605-F14 must land a guest kernel + rootfs first), so declaring the tag anywhere today would be the wrong fact.”) @yah:next(“us-south-001 needs cap:native-exec DECIDED, not defaulted, and it is the R858 node. It is the one machine the repo positively records as LACKING the backend (kamaji refused headscale there 2026-09-03), so leaving the tag off is correct TODAY — but if R858’s fix is ‘give us-south-001 a native-capable kamaji’ rather than ‘stop moving headscale’, then the roll and the tag must land together, in that order. I left .yah/infra/machines/us-south-001.toml untouched because it was already dirty at anchor 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2 and @Ashguard:eclipse is live on R858.”) @yah:next(“R860-T6 (supply = \"self\" provisioning) inherits this for free — admission_spec already requires the capability of the whole group, so a self-provisioned native member cannot be elected onto a node that cannot run it. What T6 must still not do is re-elect per member: reuse the node URL elect_node returned for the requirer, per R860-T4’s handoff.”) @yah:verify(“BASELINE RECORDED BEFORE EDITING, at tree anchor 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2: cargo test -p yah-cloud --lib from oss/yubaba = 1090 passed / 0 failed / 4 ignored, exit 0 — exactly the count the brief predicted. AFTER: 1093 passed / 0 failed / 4 ignored, exit 0 (+3, exactly the three tests added). cargo check -p yah-cloud --all-targets exit 0 and cargo check -p yubaba --all-targets exit 0 (yubaba consumes cloud, so it is where any signature change would surface — there is none). Every exit code echoed explicitly via an EXIT=$? / ${PIPESTATUS[0]} marker and read back, never inferred from an empty grep. The four yah-cloud warnings are all pre-existing and in other files (object-store r2.rs, reconciler/mesofact_static.rs unused imports, app_manifest.rs, reconciler/mod.rs non_snake_case); config.rs contributes none.”) @yah:verify(“NEW TESTS (config.rs mod tests, R860-T5 section at the end, after the R860-T4 block). (1) a_node_without_the_native_exec_capability_cannot_host_a_native_workload — a yah.exec = native spec is refused by a bare node with an error naming cap:native-exec, and admitted by a node declaring it, with both nodes in the same fleet so the choice is provably the tag. (2) a_local_edge_to_a_native_provider_makes_the_requirer_need_the_capability — an ordinary container requirer (asserted !wants_native_exec()) with a local edge to a native provider lands on the capable node, while the SAME spec without the edge still lands on the plain one, so the constraint provably comes from the group. (3) a_group_with_no_native_member_does_not_require_the_capability — the regression guard: the axis is absent from admission_spec’s mesh_tags and a group with a local edge between two ordinary specs still admits on a node declaring nothing. Helper native_spec() asserts the marker reads back through wants_native_exec() before the test uses it, so a typo cannot make the test pass vacuously.”) @yah:verify(“Machine-config lints were considered and are unaffected by construction: check_inert_taints reads taints (I touched none), and check_retired_arch_tags flags only the tier: prefix. cap: is a new namespace and nothing validates prefixes, so no lint fires and no lint needs teaching.”) @yah:gotcha(“SCHEMA DRIFT IS EXPECTED FROM THIS TICKET AND WAS ALREADY RED BEFORE IT. .yah/schema/machine.toml.schema.json is generated from cloud::config by cargo run -p xtask -- emit-schemas, and MachineConfig’s DOC COMMENT is what the generator emits as its description — which means (a) my new mesh_tags doc changes it, and (b) so does this very handoff, because R860-T5’s @yah: annotation block lives inside MachineConfig’s doc at config.rs:201. That file was ALSO already dirty in the working tree at anchor 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2, before I touched anything — scripts/check-schema-drift.sh compares the regenerated tree against git, so it is red for any uncommitted schema edit regardless of author. Regenerate with cargo run -p xtask -- emit-schemas (or scripts/check-schema-drift.sh --update) when the root target dir is not contended; the pre-commit hook no longer does it (disabled 2026-08-15, see CLAUDE.md).”) @yah:verify(“SCHEMA REGENERATED IN THIS SESSION, so the drift gate is not left for the next reader: cargo run --quiet -p xtask -- emit-schemas exit 0, run from the repo root after the handoff was written (so it captures the annotation text too). Two files moved. .yah/schema/machine.toml.schema.json: MachineConfig’s description grows by this ticket’s annotation block, plus a genuinely new mesh_tags.description from the doc comment I added. .yah/schema/workload.toml.schema.json: +104 lines that are NOT mine — the Locality / Requirement / Supply / WorkloadSpec.requires types R860-T1 landed had never been emitted, so the sibling ticket’s schema drift was still outstanding and my regen swept it in. Derived artifacts are not ownable (shared-tree doctrine), so this is deliberate rather than accidental; @Ashguard, whoever picks up R860-T1’s review should know the schema now describes requires.”) @yah:verify(“FINAL RE-RUN AFTER THE HANDOFF ANNOTATION WAS WRITTEN INTO config.rs (the board write edits MachineConfig’s doc block, so the file changed under the earlier green): cargo test -p yah-cloud --lib = 1093 passed / 0 failed / 4 ignored, exit 0. Unchanged. Note for anyone reading the camp build rail’s skew warnings on this session: the one SUSPECT RESULT it emitted names oss/yubaba/crates/cloud/src/config.rs as modified mid-run, and that modification was MY OWN board_handoff annotation write, not a peer — the two authoritative runs (full lib test, and both cargo checks) each came back Input closure unchanged across the whole run: no skew.”) @yah:verify(“All builds were run with CARGO_TARGET_DIR=/tmp/r860t5-target rather than the shared oss/yubaba/target, following R860-T4’s recorded gotcha — a peer (session:83093d9d) held the shared target lock for the entire session (20+ minutes of cargo check -p yubaba --lib). Costs one cold dep build, then every subsequent run is seconds. Worth reaching for immediately when the queue message says you are behind someone.”) @yah:handoff(“LEADER RE-VERIFIED (session:69b18855, independent of the courier’s self-report). cargo test -p yah-cloud --lib from oss/yubaba: 1093 passed / 0 failed / 4 ignored, exit 0, against the 1090/0/4 baseline this relay’s own T4 established — +3 = exactly its new tests. Axis confirmed by content: NATIVE_EXEC_MESH_TAG = \"cap:native-exec\" at config.rs:2304, appended to the derived RequiredSpec.mesh_tags at :2112-2114 when any placement_group member returns true from WorkloadSpec::wants_native_exec(). No new RequiredSpec field, no signature change, no wire change — it rides the existing AND-ed superset check, so a refusal now reads required.mesh_tags=[...,cap:native-exec].”) @yah:handoff(“Tree anchor at handoff: 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2 — the shared tree as I left it. Diff against it (git diff 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2..HEAD) to see what landed under you, and quote this SHA rather than ‘HEAD’ in any revert/restore instruction.”) @yah:verify(“MACHINE DECLARATIONS AUDITED FOR PROVENANCE, because a wrong capability declaration is worse than an absent one. Both are traceable to measurements ALREADY RECORDED IN-REPO, not inferred: us-west-001 from the ps reading at us-west-001.toml:21 (pid 517125, ppid 515908 = /usr/local/bin/kamaji --native-exec-dir /var/lib/yah/kamaji/native, cgroup 0::/yubaba.slice/kamaji.service/native, 2026-09-05); us-west-003 from the actual ExecStart read over ssh 2026-09-01 at us-west-003.toml:141. us-west-015 was deliberately left WITHOUT the tag and carries enable instructions at :207-217 — unknown fails closed, which is the correct direction. No node was guessed at and nothing was probed live.”) @yah:handoff(“THIS TICKET MODELS THE EXACT DRIFT THAT CAUSED THE 25-HOUR MESH OUTAGE, which is worth stating because it turns an abstract W338 bullet into a measured one. us-west-001.toml:8 records the root-cause chain: on 2026-09-03T06:03:03Z leadership moved to us-south-001, which tried to deploy headscale and kamaji refused — \"workload requests native host execution (yah.exec=native) but no native backend is available (native backend not configured — start kamaji with –native-exec-dir)\" — then the systemd fallback failed too, both at WARN, and the mesh had no coordination server for 25 hours. us-west-001.toml:10 names it explicitly as \"a silent per-node capability drift that placement does not model\". After this ticket, placement models it: a group needing native exec can no longer be admitted onto a node that has not declared cap:native-exec.”) @yah:handoff(“Tree anchor at handoff: 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2 — the shared tree as I left it. Diff against it (git diff 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2..HEAD) to see what landed under you, and quote this SHA rather than ‘HEAD’ in any revert/restore instruction.”) @yah:verify(“RE-VERIFIED AT HEAD 00ee20d1 (session:aa5e882d, 2026-09-05). NATIVE_EXEC_MESH_TAG present in oss/yubaba/crates/cloud/src/config.rs (declared above node_selector_mesh_tags, appended to the derived RequiredSpec.mesh_tags when any placement_group member wants native exec). cargo test --manifest-path oss/yubaba/Cargo.toml -p yah-cloud --lib = 1141 passed / 0 failed / 4 ignored, exit 0.”) @yah:cleanup(“cap:microvm remains the identical unmodelled axis, one line from done in the same if in admission_spec. Deliberately NOT taken: no node in the fleet can host a microvm until R605-F14 lands a guest kernel + rootfs, so declaring the tag today would assert a false fact. Do it when R605-F14 lands, not before.”)

Fields§

§name: String§provider: String§vendor: Option<String>

Who the hardware actually comes from ("ovh", "vultr", "on-prem").

Deliberately not provider, which selects the auto-provision driver: a box we rented by hand and brought up over SSH is provider = "static" for its whole life, and writing the vendor there instead would flip it driver-backed and make validate demand location + server_type it has no answer for. The two axes genuinely differ — vendor is who bills you, provider is who yah can call an API against.

Worth recording because vendor-scoped policy is invisible in every other field and decides real work: outbound port 25, rDNS/PTR control, IP reputation, egress billing. It survived only in TOML prose until now, which made it ungreppable at exactly the moment you need it.

§nickname: Option<String>

Human label for the box ("gamer", "the GEEKOM"). Free-form and never matched on — name stays the identity everywhere. This is only so operators and agents can say which box they mean out loud.

§location: Option<String>

Provider DC code (e.g. Hetzner "hil"). Provisioning-only: required iff the provider has an auto-provision driver (provider_has_machine_driver); a BYO static node we brought up over SSH has no such code. Optional at load time so static machine.tomls omit it; MachineConfig::validate enforces presence at the right moment for driver-backed providers.

§server_type: Option<String>

Provider SKU/size (e.g. Hetzner "ccx13"). Provisioning-only, same optionality contract as location.

§hosts_mirrors: Vec<String>

Deprecated (R330-F16). A machine should describe itself (region, zone, provider, mesh_tags); which mirrors run on it is derived by the reconciler from each mirror’s required placement spec, not declared here. Now optional + omitted-when-empty so new machine.tomls leave it out. The legacy resolve_mirror_machine topology fallback still reads it until yubaba’s reverse-index supersedes the topology.toml path; once that lands, this field and its readers are removed wholesale.

§mesh_tags: Vec<String>

Positive placement facts about this node, matched as a superset: a workload is admitted only where every tag it requires is present, so adding a tag can only ever make a machine match more, never fewer.

Four namespaces are live, and they are not interchangeable:

  • tag:<role> — a role the operator assigns (tag:build-worker, tag:qed, tag:cloud-runner, tag:mac-builder);
  • arch:<x86|arm> / os:<linux|darwin> — facts about the silicon and userland, emitted as requirements by [qed::platform::build_worker_mesh_tags];
  • cap:<capability> — something the node’s daemons were configured to be able to do. Today just NATIVE_EXEC_MESH_TAG (R860-T5);
  • tier: is retired for architecture (R763) and reserved for the environment axis — crate::validate::check_retired_arch_tags flags a machine still carrying tier:<arch>.

Nothing validates the prefix, which is why the lint above exists: a tag nobody requires is silently inert, and a stale one silently stops matching and reports “no node” rather than “wrong tag”.

§region: Option<String>

Canonical geo region label (latency axis), e.g. "us-west". F16’s three topology axes are orthogonal: region = geo (latency), zone = failure domain within a region (HA), provider = network/cost. region is distinct from location (the provider’s DC code, e.g. Hetzner "hil"): location is provider-scoped, region is our provider-neutral label. Optional for backward-compat; a machine without it never satisfies a required.regions constraint.

§zone: Option<String>

Failure-domain label within a region (HA axis), e.g. "hil". For single-DC Hetzner this typically mirrors location. F16 placement matches required.zones against this. Optional for backward-compat.

§arch: Option<String>

Declared CPU architecture ("x86_64" / "aarch64"). A machine has exactly one — it’s a first-class property of the box, not a reach detail and not a mesh tag. Drives the yubaba release triple. Optional only because there’s no provider API to probe it (static nodes declare it; a driver-backed provider may leave it unset until known).

§bucket: Option<BucketSpec>§legacy_hostkey_fingerprint: Option<String>

Legacy location, superseded by [registration].hostkey_fingerprint (R707-T1). Still deserialized so machine TOMLs written before the split keep parsing; never read directly — go through MachineConfig::hostkey_fingerprint, which prefers the registration block. MachineConfig::normalize folds this into registration, and MachineConfig::save normalizes before writing, so a load→save cycle migrates the file rather than dropping the value.

§ssh_keys: Vec<u64>

Provider-side SSH-key IDs (Hetzner: from GET /v1/ssh_keys) authorized for root at create time. Defaults to empty for backwards-compat with existing machine declarations; an empty list yields a Hetzner-emailed random root password (which the driver currently discards). Populate this when you want pre-mesh SSH access for bootstrap deploys or recovery.

§cloudflared: Option<String>

Cloudflare Tunnel ID this machine joins (e.g. abc123.cfargotunnel.com). None → no tunnel (mesh-only node, no public ingress). When set, yah cloud machine provision reads cloudflare-tunnel-token from the keys vault and injects the cloudflared install block into cloud-init so the new machine connects to CF edge on first boot.

§ingress_floating_ip: Option<String>

Provider-issued floating/reserved IP that follows public-ingress ownership onto this box — R859-F2 (W267 §Tier 1).

The value is the provider’s own identifier, opaque here and interpreted only by the matching adapter: a Hetzner numeric floating-IP id as a string, an OVH Additional-IP address ("51.81.85.200"), a Vultr reserved-IP UUID. Same “the adapter is the boundary” convention crate::envoy::floating_ip::FloatingIpAssignInput::ip_id documents.

§Why it lives on the machine

crate::envoy::floating_ip shipped the floating_ip.* verbs and three provider adapters with no config anywhere saying which floating IP is “the” ingress IP — the gap R594-F5 recorded and deliberately left. This is that field, and it sits beside cloudflared on purpose: that is already the per-node “how the world reaches this box” handle, and a floating IP is the sovereign-tier answer to the same question. [[ingress]]’s tunnel_id is the service side of ingress identity — which cohort a given service fronts through — and a floating IP is not per-service: one IP moves between boxes, so it cannot be partitioned by slot or hostname.

§Absent means “no floating-IP path”, never an error

Most machines have none, and that is the normal case: mesh-only nodes, boxes behind a Cloudflare tunnel, and every provider without a floating-IP adapter. The effector skips such a machine cleanly rather than refusing — see plan_ingress_owner_effect.

§The cohort has to agree

Every machine that can hold the same ingress IP must declare the same id: the IP is one resource that moves, so two ids inside one sovereign_group means an ownership flip silently reassigns a different IP than the one currently serving traffic. yah cloud validate refuses that (crate::validate::check_ingress_floating_ip) rather than leaving it to be discovered during a failover.

§hosts_operator_bridge: bool

When true, this machine hosts operator-bridge workloads (Tailscale operator access to mesh-internal services). yah cloud machine provision will install tailscaled and run tailscale up during cloud-init via the {{OPERATOR_BRIDGE_BLOCK}} placeholder. Defaults to false for backward-compat with existing machine declarations.

§connect: Option<ConnectSpec>

BYO static-node reach descriptor. Static nodes have no provider API to probe, so how the camp reaches them (SSH user@host + the yubaba URL, which is loopback until the WireGuard mesh lands) is declared here. None for driver-backed providers (Hetzner/Vultr), whose address is resolved from the provider API / mesh at provision time.

§allocatable: Option<NodeAllocatable>

Static node capacity (R572-F3). Declares the node’s total hardware budget; F5’s scheduler subtracts committed workload requests from this to check whether a new workload fits. Absent means unconstrained.

§taints: Vec<String>

Placement taint keys (R572-F3). There is no toleration — a no-<archetype> taint is an absolute block, not a preference (W305/R742-T4; the pre-2026-08-11 “repel-unless-tolerate” wording here described an unless that was never built).

A key in this list influences placement in exactly one of two ways, and taint_effect is the authority on which:

  • repulsion — "no-server" / "no-appliance" / "no-job" reject workloads of that LifecycleArchetype outright;
  • affinity — a key in AFFINITY_TAINT_KEYS (today just "public-ip") that a workload names in yah.placement.requires-taint, which then requires this node.

Anything else is inert: it parses, it round-trips, and no scheduler decision can ever read it. yah cloud validate rejects such keys (validate::check_inert_taints) rather than letting them sit looking load-bearing — which is how no-voter spent months asserting a falsehood on three nodes. Facts about a node that are not placement inputs belong in mesh_tags or a comment.

§sovereign_group: Option<String>

Which consensus group this node belongs to — W305/R742-F1. None means standalone: in no group at all, which is us-west-002 and us-west-015.

Membership is not by itself quorum eligibility; that is sovereign_role, added by R605-F12 because us-west-003 is in prod’s blast radius and must never vote in it.

Not a placement input. It is deliberately absent from RequiredSpec::matches, and adding it there would be a category error: a sovereign group is a blast radius, not a filter. Nothing about “which quorum does this box vote in” should decide where a workload runs — that is what made the fleet express three unrelated properties through one taint list and get all three wrong (W305).

What it is for is refusal. judge_join answers “may this node join that node’s cluster”, and the answer is no unless both declare the same group. Before this field the only guard was a comment in three machine TOMLs saying “never run a raft join against this box from a shell pointed at prod” — habit, with no mechanism behind it, which is the same class of guard W257 §8 admitted to.

§Why sovereign_group and not raft_group

Raft is today’s mechanism (operator, 2026-08-10). A field named for the mechanism goes stale the day the mechanism is swapped, and every consumer that reads it inherits the lie. sovereign names what the group has — its own authority, its own upgrade cadence, its own destruction — which stays true under any consensus protocol.

Note the word already appears in this tree as prose (W267’s title, the IngressProvider::Passway doc comment’s “sovereign edge”). That is an adjective meaning “self-hosted, not SaaS”; this is the first time it carries structure.

§sovereign_role: Option<SovereignRole>

Whether this node may hold a seat in its group’s quorum — R605-F12. Meaningless without sovereign_group: a standalone box has no quorum to be eligible for.

None is “not written”, not a third role. Read it through sovereign_membership, which resolves the absence to SovereignRole::Voter — what declaring a group has always meant, so the six nodes stamped before this field keep their seats without an edit. The distinction is kept only so crate::validate::check_unroled_sovereign_members can tell an operator who chose voter from one who never considered the question; no join decision reads the Option directly.

§Why this is not a taint

It was, once: no-voter sat in taints on three nodes for months, read by nothing, and R742-T4 removed it because the taint list is a placement vocabulary and this is not a placement input (see taint_effect). Nor is it a second group label. It is a modifier on the membership this node already declares, which is why it lives beside the group and is judged with it in one predicate, workload_spec::sovereign::join_permitted.

§registration: MachineRegistration

[registration] — the observed half (R707-T1). Empty until the box has been attached / mesh-joined. See MachineRegistration.

Implementations§

Source§

impl MachineConfig

Source

pub fn sovereign_membership(&self) -> Membership<'_>

This node’s declared place in a sovereign group, as the shared join rule wants it — R605-F12.

The one place sovereign_role’s None is resolved. Absence means SovereignRole::Voter, which is what declaring a group meant before the role existed; resolving it here rather than at each call site is what keeps the camp-side and node-side gates from disagreeing about a node that never wrote the field.

Source

pub fn location(&self) -> &str

Provider DC code, or "" when omitted (static nodes). Most readers want a &str; the driver-backed provision/status paths still go through validate which guarantees presence for those.

Source

pub fn server_type(&self) -> &str

Provider SKU, or "" when omitted (static nodes).

Source

pub fn validate(&self) -> Result<()>

Enforce the provisioning-only-field contract: a machine whose provider has an auto-provision driver MUST declare location + server_type (the driver can’t create a server without them). Static nodes may omit both. Call this before any provision/diff that assumes a driver.

Source

pub fn inert_taints(&self) -> Vec<&str>

Declared taints that no placement decision can read (W305/R742-T4).

Deliberately not folded into validate: that guard runs on the provision/diff hot path and answers a different question (can the driver create this server). An inert taint is a lint — it never breaks an operation in flight, it just means the file is asserting something the scheduler will not honour. yah cloud validate is where the operator asks for that judgement; see crate::validate::check_inert_taints.

Source

pub fn hostkey_fingerprint(&self) -> Option<&str>

Yubaba’s TOFU’d hostkey fingerprint, from [registration] and falling back to the pre-R707-T1 top-level field. The only read path — a caller that reaches for legacy_hostkey_fingerprint directly sees None on every migrated machine.

Source

pub fn set_hostkey_fingerprint(&mut self, fingerprint: Option<String>)

Record (or clear) the observed hostkey fingerprint. Writes [registration] and drops any pre-R707-T1 top-level value, so the two locations can never disagree after a writeback.

Source

pub fn mesh_ipv4(&self) -> Option<&str>

Mesh (tailnet) IPv4 for this node, or None pre-mesh.

Prefers [registration].mesh_ipv4; falls back to the host of a legacy [connect].yubaba URL when that host is in the 100.64.0.0/10 CGNAT range the mesh uses. A loopback placeholder (http://127.0.0.1:7443, meaning “pre-mesh, reachable only through an SSH tunnel”) is not a mesh address and yields None.

Source

pub fn yubaba_url(&self) -> Option<String>

Base URL for this node’s yubaba, or None when no reach resolves.

Thin wrapper over reach for the many call sites that only branch on presence. Prefer reach anywhere the operator sees the outcome — a None here throws away a refusal that names exactly which address is missing.

Source

pub fn reach(&self) -> Result<String, String>

The one address automation dials for this node — mesh-only.

Err is a named refusal, not an absence: a node with no mesh address is unresolvable to every automated path, and R605-T10’s whole complaint is that this used to surface as a connect timeout against an address the caller has no route to.

Resolution order:

  1. A declared [connect].yubaba on a private host (10/8, 172.16/12, 192.168/16) is not dialed — see below.
  2. Any other declared [connect].yubaba wins verbatim. That includes the pre-mesh loopback placeholder (http://127.0.0.1:7443, “I have no mesh address; reach me through the SSH tunnel to ssh”), which is a genuine declaration and stays honoured.
  3. Otherwise [registration].mesh_ipv4 composed with [connect].yubaba_port.

Why a LAN literal loses (R605-T10, operator 2026-08-19). The LAN address is an emergency break-glass route, never an official one, and automation must ALWAYS assume the caller is not on that LAN — this camp sits on 192.168.22.0/22 with no route to the fleet’s 192.168.10.0/24 at all. Writing one into the field every resolver dials does not sit beside the mesh route, it overrides it: R707-T6 made a declared literal beat mesh_ipv4 outright, so us-west-011 (mesh-joined, healthy) was elected for every aarch64 build and then dialed at an address that answers only from inside bldg-2506.

What R707-T6 wanted is preserved elsewhere. Its forcing case was identity, not reach: the dev raft group advertises LAN addrs (192.168.10.11:7443, verified live off /raft/status 2026-08-27), and rollout::yubaba::membership_to_nodes has to map those back to declared machines. That match now runs against lan_endpoint, which is composed from the break-glass [connect].address metadata and is never dialed — so the two concerns the old precedence rule fused are split, and the literal can stop squatting a dialed field.

The LAN address itself STAYS in the machine TOML. It is useful metadata and the manual ssh path is entitled to it; it is only disconnected from every automated process.

Source

pub fn lan_endpoint(&self) -> Option<String>

The LAN host:port this node’s yubaba answers on, composed from the break-glass [connect].address metadata plus the declared port.

Identity only — never dial this. It exists so a raft membership entry that names a node by its LAN address can be mapped back to the declared machine (rollout::yubaba::membership_to_nodes) without that address having to live in a field a resolver reads. None when the machine is unprovisioned.

Source

pub fn normalize(&mut self)

Fold the pre-R707-T1 top-level hostkey_fingerprint into [registration], and lift a mesh IP out of a legacy [connect].yubaba URL. Idempotent; a machine already on the split shape is untouched.

save calls this, so writing a machine TOML migrates it rather than round-tripping the old shape back out.

Source

pub fn save(&self, cloud_dir: &Path) -> Result<()>

Persist to <cloud_dir>/machines/<name>.toml, creating the dir if needed.

⚠ Serializes the struct, so operator comments in the target file are lost. Pre-existing behaviour, not introduced here, but it is why registration writeback (yah cloud machine attach) goes through crate::state::MachineState and the comment-preserving path in the CLI rather than calling this on a hand-authored inventory file.

Trait Implementations§

Source§

impl Clone for MachineConfig

Source§

fn clone(&self) -> MachineConfig

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Debug for MachineConfig

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more
Source§

impl<'de> Deserialize<'de> for MachineConfig

Source§

fn deserialize<__D>(__deserializer: __D) -> Result<Self, __D::Error>
where __D: Deserializer<'de>,

Deserialize this value from the given Serde deserializer. Read more
Source§

impl Serialize for MachineConfig

Source§

fn serialize<__S>(&self, __serializer: __S) -> Result<__S::Ok, __S::Error>
where __S: Serializer,

Serialize this value into the given Serde serializer. Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Allocation for T
where T: RefUnwindSafe + Send + Sync,

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> DeserializeOwned for T
where T: for<'de> Deserialize<'de>,

Source§

impl<T> Downcast for T
where T: Any,

Source§

fn into_any(self: Box<T>) -> Box<dyn Any>

Convert Box<dyn Trait> (where Trait: Downcast) to Box<dyn Any>. Box<dyn Any> can then be further downcast into Box<ConcreteType> where ConcreteType implements Trait.
Source§

fn into_any_rc(self: Rc<T>) -> Rc<dyn Any>

Convert Rc<Trait> (where Trait: Downcast) to Rc<Any>. Rc<Any> can then be further downcast into Rc<ConcreteType> where ConcreteType implements Trait.
Source§

fn as_any(&self) -> &(dyn Any + 'static)

Convert &Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot generate &Any’s vtable from &Trait’s.
Source§

fn as_any_mut(&mut self) -> &mut (dyn Any + 'static)

Convert &mut Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot generate &mut Any’s vtable from &mut Trait’s.
Source§

impl<T> Downcast for T
where T: Any,

Source§

fn into_any(self: Box<T>) -> Box<dyn Any>

Converts Box<dyn Trait> (where Trait: Downcast) to Box<dyn Any>, which can then be downcast into Box<dyn ConcreteType> where ConcreteType implements Trait.
Source§

fn into_any_rc(self: Rc<T>) -> Rc<dyn Any>

Converts Rc<Trait> (where Trait: Downcast) to Rc<Any>, which can then be further downcast into Rc<ConcreteType> where ConcreteType implements Trait.
Source§

fn as_any(&self) -> &(dyn Any + 'static)

Converts &Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot generate &Any’s vtable from &Trait’s.
Source§

fn as_any_mut(&mut self) -> &mut (dyn Any + 'static)

Converts &mut Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot generate &mut Any’s vtable from &mut Trait’s.
Source§

impl<T> DowncastSend for T
where T: Any + Send,

Source§

fn into_any_send(self: Box<T>) -> Box<dyn Any + Send>

Converts Box<Trait> (where Trait: DowncastSend) to Box<dyn Any + Send>, which can then be downcast into Box<ConcreteType> where ConcreteType implements Trait.
Source§

impl<T> DowncastSync for T
where T: Any + Send + Sync,

Source§

fn into_any_arc(self: Arc<T>) -> Arc<dyn Any + Sync + Send> ⓘ

Convert Arc<Trait> (where Trait: Downcast) to Arc<Any>. Arc<Any> can then be further downcast into Arc<ConcreteType> where ConcreteType implements Trait.
Source§

impl<T> DowncastSync for T
where T: Any + Send + Sync,

Source§

fn into_any_sync(self: Box<T>) -> Box<dyn Any + Sync + Send>

Converts Box<Trait> (where Trait: DowncastSync) to Box<dyn Any + Send + Sync>, which can then be downcast into Box<ConcreteType> where ConcreteType implements Trait.
Source§

fn into_any_arc(self: Arc<T>) -> Arc<dyn Any + Sync + Send> ⓘ

Converts Arc<Trait> (where Trait: DowncastSync) to Arc<Any>, which can then be downcast into Arc<ConcreteType> where ConcreteType implements Trait.
Source§

impl<T> ErasedDestructor for T
where T: 'static,

Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> FromRef<T> for T
where T: Clone,

Source§

fn from_ref(input: &T) -> T

Converts to this type from a reference to the input type.
Source§

impl<T> FromRef<T> for T
where T: Clone,

Source§

fn from_ref(input: &T) -> T

Converts to this type from a reference to the input type.
Source§

impl<T> Fruit for T
where T: Send + Downcast,

Source§

impl<A, B, T> HttpServerConnExec<A, B> for T
where B: Body,

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self> ⓘ

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self> ⓘ

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> IntoEither for T

Source§

fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ

Converts self into a Left variant of Either<Self, Self> if into_left is true. Converts self into a Right variant of Either<Self, Self> otherwise. Read more
Source§

fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
where F: FnOnce(&Self) -> bool,

Converts self into a Left variant of Either<Self, Self> if into_left(&self) returns true. Converts self into a Right variant of Either<Self, Self> otherwise. Read more
Source§

impl<T> Pointable for T

Source§

const ALIGN: usize

The alignment of pointer.
Source§

type Init = T

The type for initializers.
Source§

unsafe fn init(init: <T as Pointable>::Init) -> usize

Initializes a with the given initializer. Read more
Source§

unsafe fn deref<'a>(ptr: usize) -> &'a T

Dereferences the given pointer. Read more
Source§

unsafe fn deref_mut<'a>(ptr: usize) -> &'a mut T

Mutably dereferences the given pointer. Read more
Source§

unsafe fn drop(ptr: usize)

Drops the object pointed to by the given pointer. Read more
Source§

impl<T> PolicyExt for T
where T: ?Sized,

Source§

fn and<P, B, E>(self, other: P) -> And<T, P>
where T: Sized + Policy<B, E>, P: Policy<B, E>,

Create a new Policy that returns Action::Follow only if self and other return Action::Follow. Read more
Source§

fn or<P, B, E>(self, other: P) -> Or<T, P>
where T: Sized + Policy<B, E>, P: Policy<B, E>,

Create a new Policy that returns Action::Follow if either self or other returns Action::Follow. Read more
Source§

impl<T> Read<Exclusive, BecauseExclusive> for T
where T: ?Sized,

Source§

impl<T> Same for T

Source§

type Output = T

Should always be Self
Source§

impl<T> Serialize for T
where T: Serialize + ?Sized,

Source§

fn erased_serialize(&self, serializer: &mut dyn Serializer) -> Result<(), Error>

Source§

fn do_erased_serialize( &self, serializer: &mut dyn Serializer, ) -> Result<(), ErrorImpl>

Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, !>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<V, T> VZip<V> for T
where V: MultiLane<T>,

Source§

fn vzip(self) -> V

Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self> ⓘ
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self> ⓘ

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more