pub struct MachineConfig {Show 23 fields
pub name: String,
pub provider: String,
pub vendor: Option<String>,
pub nickname: Option<String>,
pub location: Option<String>,
pub server_type: Option<String>,
pub hosts_mirrors: Vec<String>,
pub mesh_tags: Vec<String>,
pub region: Option<String>,
pub zone: Option<String>,
pub arch: Option<String>,
pub bucket: Option<BucketSpec>,
pub legacy_hostkey_fingerprint: Option<String>,
pub ssh_keys: Vec<u64>,
pub cloudflared: Option<String>,
pub ingress_floating_ip: Option<String>,
pub hosts_operator_bridge: bool,
pub connect: Option<ConnectSpec>,
pub allocatable: Option<NodeAllocatable>,
pub taints: Vec<String>,
pub sovereign_group: Option<String>,
pub sovereign_role: Option<SovereignRole>,
pub registration: MachineRegistration,
}Expand description
Per-machine TOML from .yah/infra/machines/<name>.toml.
Two halves, split by provenance (R707-T1): everything here is declaration
— operator intent under review and blame — except [registration], which
carries what the fleet observed. See MachineRegistration for why the
boundary is drawn there and what depends on it.
@yah:ticket(R860-T5, “Model per-node native-exec capability as an admission axis (W338 §Placement consequences 3 / R858-T4 gap)”)
@yah:status(review)
@yah:phase(P1)
@yah:at(2026-09-05T18:29:19Z)
@yah:assignee(agent:bundle-anthropic-ashguard)
@yah:parent(R860)
@yah:next(“Cheapest defensible shape: express it on MachineConfig, which already has the two vocabularies — mesh_tags: Vec<String> (config.rs:246, superset match, already carries arch:/os:/tag:build-worker) and taints: Vec<String> (config.rs:337). A native-exec mesh tag required by any group member whose kind is native is a one-line admission axis in admission_spec(). Whichever is chosen, it must be declared in .yah/infra/machines/*.toml for the nodes that actually run kamaji with –native-exec-dir, and check_inert_taints (config.rs:703) lints unread taint keys dead — so a taint nobody reads will be flagged.”)
@yah:verify(“cargo test -p cloud –lib config”)
@arch:see(.yah/docs/working/W338-workload-dependencies-and-appliance-composition.md)
@yah:depends_on(R860-T4)
@yah:gotcha(“Verified 2026-09-04: native-exec capability is modelled NOWHERE in placement — rg \"native\" oss/yubaba/crates/cloud/src/config.rs returns zero hits, and the raft state machine models no member attributes, labels or taints at all (rg \"taint|capabilit|labels|mesh_tag\" over raft/{mod,store,network}.rs yields one unrelated comment at raft/store.rs:591). Native-exec is a node-local kamaji startup decision today: --native-exec-dir (oss/kamaji/crates/kamaji-bin/src/main.rs:152-156, :51-55) plus the native-exec cargo feature (kamaji-bin/src/server.rs:329-330). A node without it refuses the deploy at dispatch time and nothing upstream can see that in advance — which is exactly the deploy-time surprise W338 wants turned into a placement precondition.”)
@yah:handoff(“NATIVE-EXEC IS NOW A PLACEMENT PRECONDITION, NOT A DISPATCH-TIME SURPRISE. New pub const NATIVE_EXEC_MESH_TAG: &str = \"cap:native-exec\" in oss/yubaba/crates/cloud/src/config.rs (declared just above node_selector_mesh_tags), and one axis in admission_spec() immediately after the R860-T4 group loop: if ANY member of placement_group(ws, declared) returns true from WorkloadSpec::wants_native_exec(), the tag is appended to the derived RequiredSpec.mesh_tags (deduped). No new field on RequiredSpec, no signature change anywhere, no wire or serde change — the mesh_tags axis is already an AND-ed superset check against machine.mesh_tags in matches and is already rendered by describe, so a refusal now reads required.mesh_tags=[...,cap:native-exec].”)
@yah:handoff(“ITEM 1 — HOW A NATIVE WORKLOAD IS DETECTED, settled by opening the type rather than guessing. There is no kind on WorkloadSpec: on the wire a native workload is still Workload::Container(WorkloadSpec), and the ONLY difference is the annotation yah.exec = native, read through WorkloadSpec::wants_native_exec() (oss/yah-base/crates/workload-spec/src/lib.rs:2939; consts NATIVE_EXEC_ANNOTATION / NATIVE_EXEC_VALUE at :3402/:3407). That accessor is what the admission axis calls — matching kamaji, whose deploy_container checks the same marker first and routes to deploy_native_exec (oss/kamaji/crates/kamaji-bin/src/server.rs). The yah.exec key is a substrate selector with a second value, microvm (wants_microvm, same key, R605-F8), so per-node microVM capability is the obvious sibling axis and is NOT modelled here — see next-steps.”)
@yah:handoff(“ITEM 2 — DECLARATIONS LANDED ON TWO NODES, FROM READINGS RECORDED IN-REPO, NOT INFERRED. cap:native-exec added to mesh_tags in .yah/infra/machines/us-west-001.toml and .yah/infra/machines/us-west-003.toml, each with a comment naming its evidence and its re-check condition. us-west-001: the R858 gotcha in its own header records a ps reading taken on the box 2026-09-05 — pid 515908 is /usr/local/bin/kamaji --native-exec-dir /var/lib/yah/kamaji/native, supervising headscale as a native child. us-west-003: its header’s ‘THE DEPLOYED KAMAJI PREDATES THE microVM BACKEND’ note quotes the box’s actual ExecStart, read over ssh 2026-09-01, carrying --native-exec-dir /var/lib/yah/kamaji/native (corroborated by .yah/docs/architecture/A043-yah-on-machine-daemons.md’s @yah:verify for the same probe). Both comments say plainly that the capability lives in the systemd unit’s ExecStart, not in the TOML, so it must be re-checked after any roll.”)
@yah:handoff(“ITEM 2, THE NEGATIVES — TWO NODES ARE KNOWN NOT TO HAVE IT AND WERE DELIBERATELY LEFT UNSET. us-south-001: kamaji refused headscale there 2026-09-03 with ‘native backend not configured — start kamaji with –native-exec-dir’ (the R858 chain, quoted in .yah/infra/machines/us-west-001.toml and W267). I did NOT edit us-south-001.toml — it was already dirty in the working tree at the anchor SHA and @Ashguard:eclipse is live on R858, so I left it alone rather than race it; the mechanism fails closed there, which is the correct state. us-west-015 (the sole darwin builder): W254-darwin-build-nodes.md’s own next-step records that its kamaji is built/started --docker only. I added a comment to us-west-015.toml explaining that the tag is deliberately absent, that this is the node where the axis changes an error message (a darwin build row is native by construction, so it is now refused at ELECTION naming cap:native-exec instead of reaching the box and being refused by kamaji), and the exact enable sequence: rebuild with --features native-exec, restart with --native-exec-dir <dir>, THEN add the tag. us-west-002/011/013/014 are unestablished from the repo and left unset. THE OPERATOR-FACING ANSWER: the file is .yah/infra/machines/<node>.toml and the key is mesh_tags; add the literal string cap:native-exec to that array, and only after the roll.”)
@yah:handoff(“DECISIONS THE BRIEF LEFT OPEN, all recorded in doc comments at the site. (1) MESH TAG, NOT TAINT — as recommended, and the doc says why in the terms the brief asked for: mesh tags are positive capability with superset matching (‘this node CAN’), which is the claim being made; a taint is repulsion and would have to be inverted to no-native-exec on every node LACKING the backend (declaration burden on the majority, and silently wrong for a node nobody has edited) AND taught to taint_effect, or check_inert_taints would correctly lint the key dead. (2) THE cap: NAMESPACE IS NEW. Live prefixes are tag: (operator-assigned role), arch:/os: (silicon and userland facts, emitted as requirements by qed::platform::build_worker_mesh_tags), and tier: which R763 RETIRED for architecture and reserved for the environment axis — so reusing any of them would have stated the wrong kind of fact. A capability the daemon was configured with is none of those. Nothing validates tag prefixes (only check_retired_arch_tags looks at one), so this costs no wiring. (3) COMPUTED OVER THE GROUP, not the requirer — that is literally W338’s sentence (‘supply = self specs must be placeable where their requirer lands’), and the second test proves it: an ordinary container requirer with a local edge to a native provider is pulled onto a capable node. (4) FAILS CLOSED, accepted deliberately: an undeclared node is simply not a candidate, so an undeclared fleet reports ‘no node admits’ at election rather than dispatching to a node that refuses. Nothing in .yah/infra/workloads/ is native-marked today (only yah-cloud-admin.toml exists there), so the only live consumer is the qed darwin build row, where failing closed is strictly the better error.”)
@yah:handoff(“BLAST RADIUS, MEASURED. admission_spec is private and its callers are unchanged: admit_workload / admit_workload_candidates / admit_workload_in_group (config.rs), reached from app/yah/cli/src/cloud.rs (deploy, rolling, topology analyzer), app/yah/cli/src/yubaba_client.rs elect_node, and cloud/src/migrate.rs. The headscale appliance path inside yubaba (headscale_appliance.rs) does NOT go through admission — it is node-internal — so nothing eclipse holds on R858 is touched by this. Files edited, in full: oss/yubaba/crates/cloud/src/config.rs; .yah/infra/machines/{us-west-001,us-west-003,us-west-015}.toml. Nothing in oss/yubaba/crates/yubaba/ was opened, and oss/kamaji/crates/kamaji-bin/src/server.rs was READ ONLY (to confirm the marker check), per @Ashguard:hydra’s contention triage.”)
@yah:handoff(“ONE SCOPE ADDITION, stated loudly rather than slipped in: MachineConfig::mesh_tags (config.rs:256) had NO doc comment at all — the operator-facing declaration key for four tag namespaces was undocumented. I gave it one enumerating tag: / arch:+os: / the new cap: / retired tier:, and noting that nothing validates the prefix (which is why the two lints exist). CONSEQUENCE TO KNOW: that field’s doc is the source of the mesh_tags description in the GENERATED .yah/schema/machine.toml.schema.json, so it is schema-drift-affecting — see the gotcha.”)
@yah:handoff(“Tree anchor at handoff: 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2 — the shared tree as I found it and left it. Quote this SHA rather than ‘HEAD’ in any revert/restore instruction; to undo a hunk, read it with git show 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2:<path> and put it back with Edit, never git checkout/restore (they restore whole files and would delete peers’ uncommitted work).”)
@yah:handoff(“Tree anchor at handoff: 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2 — the shared tree as I left it. Diff against it (git diff 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2..HEAD) to see what landed under you, and quote this SHA rather than ‘HEAD’ in any revert/restore instruction.”)
@yah:next(“MICROVM IS THE IDENTICAL UNMODELLED GAP, one line away. yah.exec is a substrate selector with a second value: WorkloadSpec::wants_microvm() (workload-spec/src/lib.rs, R605-F8), and kamaji constructs MicroVmRuntime only when started with --microvm-dir — A043’s probe records that us-west-003’s deployed kamaji has --native-exec-dir but NOT --microvm-dir, so a microvm-marked deploy is refused there by exactly the same dispatch-time surprise this ticket removed for native. The shape is cap:microvm alongside NATIVE_EXEC_MESH_TAG in the same if in admission_spec. Not done here because no node in the fleet can host one yet (R605-F14 must land a guest kernel + rootfs first), so declaring the tag anywhere today would be the wrong fact.”)
@yah:next(“us-south-001 needs cap:native-exec DECIDED, not defaulted, and it is the R858 node. It is the one machine the repo positively records as LACKING the backend (kamaji refused headscale there 2026-09-03), so leaving the tag off is correct TODAY — but if R858’s fix is ‘give us-south-001 a native-capable kamaji’ rather than ‘stop moving headscale’, then the roll and the tag must land together, in that order. I left .yah/infra/machines/us-south-001.toml untouched because it was already dirty at anchor 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2 and @Ashguard:eclipse is live on R858.”)
@yah:next(“R860-T6 (supply = \"self\" provisioning) inherits this for free — admission_spec already requires the capability of the whole group, so a self-provisioned native member cannot be elected onto a node that cannot run it. What T6 must still not do is re-elect per member: reuse the node URL elect_node returned for the requirer, per R860-T4’s handoff.”)
@yah:verify(“BASELINE RECORDED BEFORE EDITING, at tree anchor 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2: cargo test -p yah-cloud --lib from oss/yubaba = 1090 passed / 0 failed / 4 ignored, exit 0 — exactly the count the brief predicted. AFTER: 1093 passed / 0 failed / 4 ignored, exit 0 (+3, exactly the three tests added). cargo check -p yah-cloud --all-targets exit 0 and cargo check -p yubaba --all-targets exit 0 (yubaba consumes cloud, so it is where any signature change would surface — there is none). Every exit code echoed explicitly via an EXIT=$? / ${PIPESTATUS[0]} marker and read back, never inferred from an empty grep. The four yah-cloud warnings are all pre-existing and in other files (object-store r2.rs, reconciler/mesofact_static.rs unused imports, app_manifest.rs, reconciler/mod.rs non_snake_case); config.rs contributes none.”)
@yah:verify(“NEW TESTS (config.rs mod tests, R860-T5 section at the end, after the R860-T4 block). (1) a_node_without_the_native_exec_capability_cannot_host_a_native_workload — a yah.exec = native spec is refused by a bare node with an error naming cap:native-exec, and admitted by a node declaring it, with both nodes in the same fleet so the choice is provably the tag. (2) a_local_edge_to_a_native_provider_makes_the_requirer_need_the_capability — an ordinary container requirer (asserted !wants_native_exec()) with a local edge to a native provider lands on the capable node, while the SAME spec without the edge still lands on the plain one, so the constraint provably comes from the group. (3) a_group_with_no_native_member_does_not_require_the_capability — the regression guard: the axis is absent from admission_spec’s mesh_tags and a group with a local edge between two ordinary specs still admits on a node declaring nothing. Helper native_spec() asserts the marker reads back through wants_native_exec() before the test uses it, so a typo cannot make the test pass vacuously.”)
@yah:verify(“Machine-config lints were considered and are unaffected by construction: check_inert_taints reads taints (I touched none), and check_retired_arch_tags flags only the tier: prefix. cap: is a new namespace and nothing validates prefixes, so no lint fires and no lint needs teaching.”)
@yah:gotcha(“SCHEMA DRIFT IS EXPECTED FROM THIS TICKET AND WAS ALREADY RED BEFORE IT. .yah/schema/machine.toml.schema.json is generated from cloud::config by cargo run -p xtask -- emit-schemas, and MachineConfig’s DOC COMMENT is what the generator emits as its description — which means (a) my new mesh_tags doc changes it, and (b) so does this very handoff, because R860-T5’s @yah: annotation block lives inside MachineConfig’s doc at config.rs:201. That file was ALSO already dirty in the working tree at anchor 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2, before I touched anything — scripts/check-schema-drift.sh compares the regenerated tree against git, so it is red for any uncommitted schema edit regardless of author. Regenerate with cargo run -p xtask -- emit-schemas (or scripts/check-schema-drift.sh --update) when the root target dir is not contended; the pre-commit hook no longer does it (disabled 2026-08-15, see CLAUDE.md).”)
@yah:verify(“SCHEMA REGENERATED IN THIS SESSION, so the drift gate is not left for the next reader: cargo run --quiet -p xtask -- emit-schemas exit 0, run from the repo root after the handoff was written (so it captures the annotation text too). Two files moved. .yah/schema/machine.toml.schema.json: MachineConfig’s description grows by this ticket’s annotation block, plus a genuinely new mesh_tags.description from the doc comment I added. .yah/schema/workload.toml.schema.json: +104 lines that are NOT mine — the Locality / Requirement / Supply / WorkloadSpec.requires types R860-T1 landed had never been emitted, so the sibling ticket’s schema drift was still outstanding and my regen swept it in. Derived artifacts are not ownable (shared-tree doctrine), so this is deliberate rather than accidental; @Ashguard, whoever picks up R860-T1’s review should know the schema now describes requires.”)
@yah:verify(“FINAL RE-RUN AFTER THE HANDOFF ANNOTATION WAS WRITTEN INTO config.rs (the board write edits MachineConfig’s doc block, so the file changed under the earlier green): cargo test -p yah-cloud --lib = 1093 passed / 0 failed / 4 ignored, exit 0. Unchanged. Note for anyone reading the camp build rail’s skew warnings on this session: the one SUSPECT RESULT it emitted names oss/yubaba/crates/cloud/src/config.rs as modified mid-run, and that modification was MY OWN board_handoff annotation write, not a peer — the two authoritative runs (full lib test, and both cargo checks) each came back Input closure unchanged across the whole run: no skew.”)
@yah:verify(“All builds were run with CARGO_TARGET_DIR=/tmp/r860t5-target rather than the shared oss/yubaba/target, following R860-T4’s recorded gotcha — a peer (session:83093d9d) held the shared target lock for the entire session (20+ minutes of cargo check -p yubaba --lib). Costs one cold dep build, then every subsequent run is seconds. Worth reaching for immediately when the queue message says you are behind someone.”)
@yah:handoff(“LEADER RE-VERIFIED (session:69b18855, independent of the courier’s self-report). cargo test -p yah-cloud --lib from oss/yubaba: 1093 passed / 0 failed / 4 ignored, exit 0, against the 1090/0/4 baseline this relay’s own T4 established — +3 = exactly its new tests. Axis confirmed by content: NATIVE_EXEC_MESH_TAG = \"cap:native-exec\" at config.rs:2304, appended to the derived RequiredSpec.mesh_tags at :2112-2114 when any placement_group member returns true from WorkloadSpec::wants_native_exec(). No new RequiredSpec field, no signature change, no wire change — it rides the existing AND-ed superset check, so a refusal now reads required.mesh_tags=[...,cap:native-exec].”)
@yah:handoff(“Tree anchor at handoff: 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2 — the shared tree as I left it. Diff against it (git diff 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2..HEAD) to see what landed under you, and quote this SHA rather than ‘HEAD’ in any revert/restore instruction.”)
@yah:verify(“MACHINE DECLARATIONS AUDITED FOR PROVENANCE, because a wrong capability declaration is worse than an absent one. Both are traceable to measurements ALREADY RECORDED IN-REPO, not inferred: us-west-001 from the ps reading at us-west-001.toml:21 (pid 517125, ppid 515908 = /usr/local/bin/kamaji --native-exec-dir /var/lib/yah/kamaji/native, cgroup 0::/yubaba.slice/kamaji.service/native, 2026-09-05); us-west-003 from the actual ExecStart read over ssh 2026-09-01 at us-west-003.toml:141. us-west-015 was deliberately left WITHOUT the tag and carries enable instructions at :207-217 — unknown fails closed, which is the correct direction. No node was guessed at and nothing was probed live.”)
@yah:handoff(“THIS TICKET MODELS THE EXACT DRIFT THAT CAUSED THE 25-HOUR MESH OUTAGE, which is worth stating because it turns an abstract W338 bullet into a measured one. us-west-001.toml:8 records the root-cause chain: on 2026-09-03T06:03:03Z leadership moved to us-south-001, which tried to deploy headscale and kamaji refused — \"workload requests native host execution (yah.exec=native) but no native backend is available (native backend not configured — start kamaji with –native-exec-dir)\" — then the systemd fallback failed too, both at WARN, and the mesh had no coordination server for 25 hours. us-west-001.toml:10 names it explicitly as \"a silent per-node capability drift that placement does not model\". After this ticket, placement models it: a group needing native exec can no longer be admitted onto a node that has not declared cap:native-exec.”)
@yah:handoff(“Tree anchor at handoff: 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2 — the shared tree as I left it. Diff against it (git diff 0a85122cdb33dbf97ebc04b84e07d9cfc049c0b2..HEAD) to see what landed under you, and quote this SHA rather than ‘HEAD’ in any revert/restore instruction.”)
@yah:verify(“RE-VERIFIED AT HEAD 00ee20d1 (session:aa5e882d, 2026-09-05). NATIVE_EXEC_MESH_TAG present in oss/yubaba/crates/cloud/src/config.rs (declared above node_selector_mesh_tags, appended to the derived RequiredSpec.mesh_tags when any placement_group member wants native exec). cargo test --manifest-path oss/yubaba/Cargo.toml -p yah-cloud --lib = 1141 passed / 0 failed / 4 ignored, exit 0.”)
@yah:cleanup(“cap:microvm remains the identical unmodelled axis, one line from done in the same if in admission_spec. Deliberately NOT taken: no node in the fleet can host a microvm until R605-F14 lands a guest kernel + rootfs, so declaring the tag today would assert a false fact. Do it when R605-F14 lands, not before.”)
Fields§
§name: String§provider: String§vendor: Option<String>Who the hardware actually comes from ("ovh", "vultr", "on-prem").
Deliberately not provider, which selects the
auto-provision driver: a box we rented by hand and brought up over SSH
is provider = "static" for its whole life, and writing the vendor
there instead would flip it driver-backed and make
validate demand location + server_type it has no
answer for. The two axes genuinely differ — vendor is who bills you,
provider is who yah can call an API against.
Worth recording because vendor-scoped policy is invisible in every other field and decides real work: outbound port 25, rDNS/PTR control, IP reputation, egress billing. It survived only in TOML prose until now, which made it ungreppable at exactly the moment you need it.
nickname: Option<String>Human label for the box ("gamer", "the GEEKOM"). Free-form and never
matched on — name stays the identity everywhere. This is
only so operators and agents can say which box they mean out loud.
location: Option<String>Provider DC code (e.g. Hetzner "hil"). Provisioning-only: required
iff the provider has an auto-provision driver (provider_has_machine_driver);
a BYO static node we brought up over SSH has no such code. Optional at
load time so static machine.tomls omit it; MachineConfig::validate
enforces presence at the right moment for driver-backed providers.
server_type: Option<String>Provider SKU/size (e.g. Hetzner "ccx13"). Provisioning-only, same
optionality contract as location.
hosts_mirrors: Vec<String>Deprecated (R330-F16). A machine should describe itself (region,
zone, provider, mesh_tags); which mirrors run on it is derived by the
reconciler from each mirror’s required placement spec, not declared
here. Now optional + omitted-when-empty so new machine.tomls leave it
out. The legacy resolve_mirror_machine topology fallback still reads
it until yubaba’s reverse-index supersedes the topology.toml path; once
that lands, this field and its readers are removed wholesale.
Positive placement facts about this node, matched as a superset: a workload is admitted only where every tag it requires is present, so adding a tag can only ever make a machine match more, never fewer.
Four namespaces are live, and they are not interchangeable:
tag:<role>— a role the operator assigns (tag:build-worker,tag:qed,tag:cloud-runner,tag:mac-builder);arch:<x86|arm>/os:<linux|darwin>— facts about the silicon and userland, emitted as requirements by [qed::platform::build_worker_mesh_tags];cap:<capability>— something the node’s daemons were configured to be able to do. Today justNATIVE_EXEC_MESH_TAG(R860-T5);tier:is retired for architecture (R763) and reserved for the environment axis —crate::validate::check_retired_arch_tagsflags a machine still carryingtier:<arch>.
Nothing validates the prefix, which is why the lint above exists: a tag nobody requires is silently inert, and a stale one silently stops matching and reports “no node” rather than “wrong tag”.
region: Option<String>Canonical geo region label (latency axis), e.g. "us-west". F16’s three
topology axes are orthogonal: region = geo (latency), zone = failure
domain within a region (HA), provider = network/cost. region is
distinct from location (the provider’s DC code, e.g. Hetzner "hil"):
location is provider-scoped, region is our provider-neutral label.
Optional for backward-compat; a machine without it never satisfies a
required.regions constraint.
zone: Option<String>Failure-domain label within a region (HA axis), e.g. "hil". For
single-DC Hetzner this typically mirrors location. F16 placement
matches required.zones against this. Optional for backward-compat.
arch: Option<String>Declared CPU architecture ("x86_64" / "aarch64"). A machine has
exactly one — it’s a first-class property of the box, not a reach
detail and not a mesh tag. Drives the yubaba release triple. Optional
only because there’s no provider API to probe it (static nodes declare
it; a driver-backed provider may leave it unset until known).
bucket: Option<BucketSpec>§legacy_hostkey_fingerprint: Option<String>Legacy location, superseded by [registration].hostkey_fingerprint
(R707-T1). Still deserialized so machine TOMLs written before the split
keep parsing; never read directly — go through
MachineConfig::hostkey_fingerprint, which prefers the registration
block. MachineConfig::normalize folds this into registration, and
MachineConfig::save normalizes before writing, so a load→save cycle
migrates the file rather than dropping the value.
ssh_keys: Vec<u64>Provider-side SSH-key IDs (Hetzner: from GET /v1/ssh_keys)
authorized for root at create time. Defaults to empty for
backwards-compat with existing machine declarations; an empty
list yields a Hetzner-emailed random root password (which the
driver currently discards). Populate this when you want pre-mesh
SSH access for bootstrap deploys or recovery.
cloudflared: Option<String>Cloudflare Tunnel ID this machine joins (e.g. abc123.cfargotunnel.com).
None → no tunnel (mesh-only node, no public ingress).
When set, yah cloud machine provision reads cloudflare-tunnel-token
from the keys vault and injects the cloudflared install block into
cloud-init so the new machine connects to CF edge on first boot.
ingress_floating_ip: Option<String>Provider-issued floating/reserved IP that follows public-ingress ownership onto this box — R859-F2 (W267 §Tier 1).
The value is the provider’s own identifier, opaque here and interpreted
only by the matching adapter: a Hetzner numeric floating-IP id as a
string, an OVH Additional-IP address ("51.81.85.200"), a Vultr
reserved-IP UUID. Same “the adapter is the boundary” convention
crate::envoy::floating_ip::FloatingIpAssignInput::ip_id documents.
§Why it lives on the machine
crate::envoy::floating_ip shipped the floating_ip.* verbs and
three provider adapters with no config anywhere saying which floating
IP is “the” ingress IP — the gap R594-F5 recorded and deliberately left.
This is that field, and it sits beside cloudflared
on purpose: that is already the per-node “how the world reaches this
box” handle, and a floating IP is the sovereign-tier answer to the same
question. [[ingress]]’s
tunnel_id is the service
side of ingress identity — which cohort a given service fronts through —
and a floating IP is not per-service: one IP moves between boxes, so it
cannot be partitioned by slot or hostname.
§Absent means “no floating-IP path”, never an error
Most machines have none, and that is the normal case: mesh-only nodes,
boxes behind a Cloudflare tunnel, and every provider without a
floating-IP adapter. The effector skips such a machine cleanly rather
than refusing — see
plan_ingress_owner_effect.
§The cohort has to agree
Every machine that can hold the same ingress IP must declare the same
id: the IP is one resource that moves, so two ids inside one
sovereign_group means an ownership flip
silently reassigns a different IP than the one currently serving
traffic. yah cloud validate refuses that
(crate::validate::check_ingress_floating_ip) rather than leaving it
to be discovered during a failover.
hosts_operator_bridge: boolWhen true, this machine hosts operator-bridge workloads (Tailscale
operator access to mesh-internal services). yah cloud machine provision
will install tailscaled and run tailscale up during cloud-init via the
{{OPERATOR_BRIDGE_BLOCK}} placeholder. Defaults to false for
backward-compat with existing machine declarations.
connect: Option<ConnectSpec>BYO static-node reach descriptor. Static nodes have no provider API to
probe, so how the camp reaches them (SSH user@host + the yubaba URL,
which is loopback until the WireGuard mesh lands) is declared here.
None for driver-backed providers (Hetzner/Vultr), whose address is
resolved from the provider API / mesh at provision time.
allocatable: Option<NodeAllocatable>Static node capacity (R572-F3). Declares the node’s total hardware budget; F5’s scheduler subtracts committed workload requests from this to check whether a new workload fits. Absent means unconstrained.
taints: Vec<String>Placement taint keys (R572-F3). There is no toleration — a
no-<archetype> taint is an absolute block, not a preference
(W305/R742-T4; the pre-2026-08-11 “repel-unless-tolerate” wording here
described an unless that was never built).
A key in this list influences placement in exactly one of two ways, and
taint_effect is the authority on which:
- repulsion —
"no-server"/"no-appliance"/"no-job"reject workloads of thatLifecycleArchetypeoutright; - affinity — a key in
AFFINITY_TAINT_KEYS(today just"public-ip") that a workload names inyah.placement.requires-taint, which then requires this node.
Anything else is inert: it parses, it round-trips, and no scheduler
decision can ever read it. yah cloud validate rejects such keys
(validate::check_inert_taints) rather than letting them sit looking
load-bearing — which is how no-voter spent months asserting a
falsehood on three nodes. Facts about a node that are not placement
inputs belong in mesh_tags or a comment.
sovereign_group: Option<String>Which consensus group this node belongs to — W305/R742-F1. None means
standalone: in no group at all, which is us-west-002 and us-west-015.
Membership is not by itself quorum eligibility; that is
sovereign_role, added by R605-F12 because
us-west-003 is in prod’s blast radius and must never vote in it.
Not a placement input. It is deliberately absent from
RequiredSpec::matches, and adding it there would be a category
error: a sovereign group is a blast radius, not a filter. Nothing
about “which quorum does this box vote in” should decide where a
workload runs — that is what made the fleet express three unrelated
properties through one taint list and get all three wrong (W305).
What it is for is refusal. judge_join answers “may this node join
that node’s cluster”, and the answer is no unless both declare the same
group. Before this field the only guard was a comment in three machine
TOMLs saying “never run a raft join against this box from a shell
pointed at prod” — habit, with no mechanism behind it, which is the
same class of guard W257 §8 admitted to.
§Why sovereign_group and not raft_group
Raft is today’s mechanism (operator, 2026-08-10). A field named for the
mechanism goes stale the day the mechanism is swapped, and every
consumer that reads it inherits the lie. sovereign names what the
group has — its own authority, its own upgrade cadence, its own
destruction — which stays true under any consensus protocol.
Note the word already appears in this tree as prose (W267’s title, the
IngressProvider::Passway doc comment’s “sovereign edge”). That is an
adjective meaning “self-hosted, not SaaS”; this is the first time it
carries structure.
sovereign_role: Option<SovereignRole>Whether this node may hold a seat in its group’s quorum — R605-F12.
Meaningless without sovereign_group: a
standalone box has no quorum to be eligible for.
None is “not written”, not a third role. Read it through
sovereign_membership, which resolves the
absence to SovereignRole::Voter — what declaring a group has always
meant, so the six nodes stamped before this field keep their seats
without an edit. The distinction is kept only so
crate::validate::check_unroled_sovereign_members can tell an
operator who chose voter from one who never considered the question;
no join decision reads the Option directly.
§Why this is not a taint
It was, once: no-voter sat in taints on three nodes
for months, read by nothing, and R742-T4 removed it because the taint
list is a placement vocabulary and this is not a placement input (see
taint_effect). Nor is it a second group label. It is a modifier on
the membership this node already declares, which is why it lives beside
the group and is judged with it in one predicate,
workload_spec::sovereign::join_permitted.
registration: MachineRegistration[registration] — the observed half (R707-T1). Empty until the box has
been attached / mesh-joined. See MachineRegistration.
Implementations§
Source§impl MachineConfig
impl MachineConfig
Sourcepub fn sovereign_membership(&self) -> Membership<'_>
pub fn sovereign_membership(&self) -> Membership<'_>
This node’s declared place in a sovereign group, as the shared join rule wants it — R605-F12.
The one place sovereign_role’s None is resolved. Absence means
SovereignRole::Voter, which is what declaring a group meant before
the role existed; resolving it here rather than at each call site is what
keeps the camp-side and node-side gates from disagreeing about a node
that never wrote the field.
Sourcepub fn location(&self) -> &str
pub fn location(&self) -> &str
Provider DC code, or "" when omitted (static nodes). Most readers want
a &str; the driver-backed provision/status paths still go through
validate which guarantees presence for those.
Sourcepub fn server_type(&self) -> &str
pub fn server_type(&self) -> &str
Provider SKU, or "" when omitted (static nodes).
Sourcepub fn validate(&self) -> Result<()>
pub fn validate(&self) -> Result<()>
Enforce the provisioning-only-field contract: a machine whose provider
has an auto-provision driver MUST declare location + server_type
(the driver can’t create a server without them). Static nodes may omit
both. Call this before any provision/diff that assumes a driver.
Sourcepub fn inert_taints(&self) -> Vec<&str>
pub fn inert_taints(&self) -> Vec<&str>
Declared taints that no placement decision can read (W305/R742-T4).
Deliberately not folded into validate: that
guard runs on the provision/diff hot path and answers a different
question (can the driver create this server). An inert taint is a lint
— it never breaks an operation in flight, it just means the file is
asserting something the scheduler will not honour. yah cloud validate
is where the operator asks for that judgement; see
crate::validate::check_inert_taints.
Sourcepub fn hostkey_fingerprint(&self) -> Option<&str>
pub fn hostkey_fingerprint(&self) -> Option<&str>
Yubaba’s TOFU’d hostkey fingerprint, from [registration] and falling
back to the pre-R707-T1 top-level field. The only read path — a
caller that reaches for legacy_hostkey_fingerprint directly sees
None on every migrated machine.
Sourcepub fn set_hostkey_fingerprint(&mut self, fingerprint: Option<String>)
pub fn set_hostkey_fingerprint(&mut self, fingerprint: Option<String>)
Record (or clear) the observed hostkey fingerprint. Writes
[registration] and drops any pre-R707-T1 top-level value, so the two
locations can never disagree after a writeback.
Sourcepub fn mesh_ipv4(&self) -> Option<&str>
pub fn mesh_ipv4(&self) -> Option<&str>
Mesh (tailnet) IPv4 for this node, or None pre-mesh.
Prefers [registration].mesh_ipv4; falls back to the host of a legacy
[connect].yubaba URL when that host is in the 100.64.0.0/10 CGNAT
range the mesh uses. A loopback placeholder (http://127.0.0.1:7443,
meaning “pre-mesh, reachable only through an SSH tunnel”) is not a
mesh address and yields None.
Sourcepub fn yubaba_url(&self) -> Option<String>
pub fn yubaba_url(&self) -> Option<String>
Base URL for this node’s yubaba, or None when no reach resolves.
Thin wrapper over reach for the many call sites that
only branch on presence. Prefer reach anywhere the operator sees the
outcome — a None here throws away a refusal that names exactly which
address is missing.
Sourcepub fn reach(&self) -> Result<String, String>
pub fn reach(&self) -> Result<String, String>
The one address automation dials for this node — mesh-only.
Err is a named refusal, not an absence: a node with no mesh address
is unresolvable to every automated path, and R605-T10’s whole complaint
is that this used to surface as a connect timeout against an address the
caller has no route to.
Resolution order:
- A declared
[connect].yubabaon a private host (10/8, 172.16/12, 192.168/16) is not dialed — see below. - Any other declared
[connect].yubabawins verbatim. That includes the pre-mesh loopback placeholder (http://127.0.0.1:7443, “I have no mesh address; reach me through the SSH tunnel tossh”), which is a genuine declaration and stays honoured. - Otherwise
[registration].mesh_ipv4composed with[connect].yubaba_port.
Why a LAN literal loses (R605-T10, operator 2026-08-19). The LAN
address is an emergency break-glass route, never an official one, and
automation must ALWAYS assume the caller is not on that LAN — this camp
sits on 192.168.22.0/22 with no route to the fleet’s 192.168.10.0/24 at
all. Writing one into the field every resolver dials does not sit beside
the mesh route, it overrides it: R707-T6 made a declared literal beat
mesh_ipv4 outright, so us-west-011 (mesh-joined, healthy) was elected
for every aarch64 build and then dialed at an address that answers only
from inside bldg-2506.
What R707-T6 wanted is preserved elsewhere. Its forcing case was
identity, not reach: the dev raft group advertises LAN addrs
(192.168.10.11:7443, verified live off /raft/status 2026-08-27), and
rollout::yubaba::membership_to_nodes has to map those back to declared
machines. That match now runs against lan_endpoint,
which is composed from the break-glass [connect].address metadata and
is never dialed — so the two concerns the old precedence rule fused are
split, and the literal can stop squatting a dialed field.
The LAN address itself STAYS in the machine TOML. It is useful metadata
and the manual ssh path is entitled to it; it is only disconnected
from every automated process.
Sourcepub fn lan_endpoint(&self) -> Option<String>
pub fn lan_endpoint(&self) -> Option<String>
The LAN host:port this node’s yubaba answers on, composed from the
break-glass [connect].address metadata plus the declared port.
Identity only — never dial this. It exists so a raft membership
entry that names a node by its LAN address can be mapped back to the
declared machine (rollout::yubaba::membership_to_nodes) without that
address having to live in a field a resolver reads. None when the
machine is unprovisioned.
Sourcepub fn normalize(&mut self)
pub fn normalize(&mut self)
Fold the pre-R707-T1 top-level hostkey_fingerprint into
[registration], and lift a mesh IP out of a legacy [connect].yubaba
URL. Idempotent; a machine already on the split shape is untouched.
save calls this, so writing a machine TOML migrates it
rather than round-tripping the old shape back out.
Sourcepub fn save(&self, cloud_dir: &Path) -> Result<()>
pub fn save(&self, cloud_dir: &Path) -> Result<()>
Persist to <cloud_dir>/machines/<name>.toml, creating the dir if needed.
⚠ Serializes the struct, so operator comments in the target file are
lost. Pre-existing behaviour, not introduced here, but it is why
registration writeback (yah cloud machine attach) goes through
crate::state::MachineState and the comment-preserving path in the
CLI rather than calling this on a hand-authored inventory file.
Trait Implementations§
Source§impl Clone for MachineConfig
impl Clone for MachineConfig
Source§fn clone(&self) -> MachineConfig
fn clone(&self) -> MachineConfig
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read moreSource§impl Debug for MachineConfig
impl Debug for MachineConfig
Source§impl<'de> Deserialize<'de> for MachineConfig
impl<'de> Deserialize<'de> for MachineConfig
Source§fn deserialize<__D>(__deserializer: __D) -> Result<Self, __D::Error>where
__D: Deserializer<'de>,
fn deserialize<__D>(__deserializer: __D) -> Result<Self, __D::Error>where
__D: Deserializer<'de>,
Auto Trait Implementations§
impl Freeze for MachineConfig
impl RefUnwindSafe for MachineConfig
impl Send for MachineConfig
impl Sync for MachineConfig
impl Unpin for MachineConfig
impl UnsafeUnpin for MachineConfig
impl UnwindSafe for MachineConfig
Blanket Implementations§
impl<T> Allocation for T
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> DeserializeOwned for Twhere
T: for<'de> Deserialize<'de>,
Source§impl<T> Downcast for Twhere
T: Any,
impl<T> Downcast for Twhere
T: Any,
Source§fn into_any(self: Box<T>) -> Box<dyn Any>
fn into_any(self: Box<T>) -> Box<dyn Any>
Box<dyn Trait> (where Trait: Downcast) to Box<dyn Any>. Box<dyn Any> can
then be further downcast into Box<ConcreteType> where ConcreteType implements Trait.Source§fn into_any_rc(self: Rc<T>) -> Rc<dyn Any>
fn into_any_rc(self: Rc<T>) -> Rc<dyn Any>
Rc<Trait> (where Trait: Downcast) to Rc<Any>. Rc<Any> can then be
further downcast into Rc<ConcreteType> where ConcreteType implements Trait.Source§fn as_any(&self) -> &(dyn Any + 'static)
fn as_any(&self) -> &(dyn Any + 'static)
&Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot
generate &Any’s vtable from &Trait’s.Source§fn as_any_mut(&mut self) -> &mut (dyn Any + 'static)
fn as_any_mut(&mut self) -> &mut (dyn Any + 'static)
&mut Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot
generate &mut Any’s vtable from &mut Trait’s.Source§impl<T> Downcast for Twhere
T: Any,
impl<T> Downcast for Twhere
T: Any,
Source§fn into_any(self: Box<T>) -> Box<dyn Any>
fn into_any(self: Box<T>) -> Box<dyn Any>
Box<dyn Trait> (where Trait: Downcast) to Box<dyn Any>, which can then be
downcast into Box<dyn ConcreteType> where ConcreteType implements Trait.Source§fn into_any_rc(self: Rc<T>) -> Rc<dyn Any>
fn into_any_rc(self: Rc<T>) -> Rc<dyn Any>
Rc<Trait> (where Trait: Downcast) to Rc<Any>, which can then be further
downcast into Rc<ConcreteType> where ConcreteType implements Trait.Source§fn as_any(&self) -> &(dyn Any + 'static)
fn as_any(&self) -> &(dyn Any + 'static)
&Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot
generate &Any’s vtable from &Trait’s.Source§fn as_any_mut(&mut self) -> &mut (dyn Any + 'static)
fn as_any_mut(&mut self) -> &mut (dyn Any + 'static)
&mut Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot
generate &mut Any’s vtable from &mut Trait’s.Source§impl<T> DowncastSend for T
impl<T> DowncastSend for T
Source§impl<T> DowncastSync for T
impl<T> DowncastSync for T
Source§impl<T> DowncastSync for T
impl<T> DowncastSync for T
impl<T> ErasedDestructor for Twhere
T: 'static,
impl<T> Fruit for T
impl<A, B, T> HttpServerConnExec<A, B> for Twhere
B: Body,
Source§impl<T> Instrument for T
impl<T> Instrument for T
Source§fn instrument(self, span: Span) -> Instrumented<Self> ⓘ
fn instrument(self, span: Span) -> Instrumented<Self> ⓘ
Source§fn in_current_span(self) -> Instrumented<Self> ⓘ
fn in_current_span(self) -> Instrumented<Self> ⓘ
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more