pub struct RunConfig {Show 22 fields
pub mecha_version: String,
pub provider: String,
pub model: String,
pub workspace: PathBuf,
pub system_prompt: Option<String>,
pub tools: Vec<String>,
pub tools_hash: Option<String>,
pub effort: Option<Effort>,
pub temperature: Option<f64>,
pub seed: Option<u64>,
pub thinking: bool,
pub cache_prompt: bool,
pub max_tokens: u32,
pub max_turns: u32,
pub max_output_tokens: Option<u64>,
pub max_cost_usd: Option<f64>,
pub compact_at_tokens: Option<u64>,
pub compact_keep_recent: usize,
pub permission_mode: PermissionMode,
pub trifecta: TrifectaPolicy,
pub sandbox: String,
pub sandbox_network: bool,
}Expand description
What a run was configured with, recorded so it can be replayed.
The rule behind the field list: anything that shapes the request or constrains the run is a confound if it is not recorded. That is not theoretical here — compaction on versus off measured 1/5 against 5/5 on the same task, so a replay that did not know whether compaction was enabled would compare two incomparable runs and report a model regression.
The system prompt is stored in full rather than hashed. A hash tells you only that something differed; the text lets a replay rebuild the request. It is no more sensitive than the transcript sitting beside it.
The sampler is recorded only as far as it is pinned: temperature and
seed hold what this process sent, and None means the server chose.
Replay against an unpinned run has to be pass@k-shaped rather than
exact-match-shaped; against a pinned, seeded run driven sequentially it can
expect to match. (Not greedy — temperature 0.0 walks qwen3.6 into verbatim
repetition loops. And only sequentially: llama-server’s continuous batching
makes concurrent requests perturb each other’s numerics, seed or no seed.)
Fields§
§mecha_version: StringWhich harness produced this. The axis every replay diff is measured on.
provider: String§model: String§workspace: PathBuf§system_prompt: Option<String>The resolved text, not the path it may have come from.
tools: Vec<String>Tool names in registry order — which is the order they are sent, and the front of the cached prefix. A tool added, removed or renamed between recording and replay changes what the model could have done.
tools_hash: Option<String>The surface those names actually described, by hash.
Names were never enough, and the comment above says why without seeing it. Add, remove and rename are the three that almost never happen; re-describe happens constantly — 49 commits touched tool definitions in three weeks of this store — and a list of names cannot see it. Tools render before the system prompt, so a replay was rebuilding the second half of the prefix byte-exactly and the first half from whatever the registry says today. Measured consequence: 12 of 13 counterfactual probes inconclusive, deterministically, median divergence one tool call in.
The specs themselves live in crate::surface::SurfaceStore — 69 KB
against a 25 KB average session is why this is a citation and not the
text, where system_prompt above is the text.
None is a recording from before this existed, and must never read as
a match — crate::surface::Fidelity is the three-state answer, and
its Unknown arm is the one every session on disk today lands in.
Scope: registry().specs(), unfiltered — not necessarily what this
turn’s request actually sent. The wire request goes through
registry.specs_for(cx.phase), which also applies Phase::Plan’s
read-only filter and a loaded skill’s tools: narrowing (matching this
struct’s own tools field, so this is not a new gap, only a named
one). A run under Plan, or one that had a narrowing skill loaded,
sent fewer specs than this hash covers — and since the surface can
narrow mid-run, no single hash can describe every turn’s request
exactly. Fidelity::Matches here means “the full registry is
unchanged since this was recorded”, which is what makes a replay
worth attempting; it is not a claim that the request bytes were
identical.
effort: Option<Effort>§temperature: Option<f64>The temperature and seed actually sent, when the provider config pins them. Unset means the server chose, and the run is not repeatable.
seed: Option<u64>§thinking: bool§cache_prompt: boolNo effect on semantics; large effect on the token counts a replay diffs.
max_tokens: u32§max_turns: u32§max_output_tokens: Option<u64>§max_cost_usd: Option<f64>§compact_at_tokens: Option<u64>§compact_keep_recent: usize§permission_mode: PermissionModeA denied call redirects the whole trajectory, so replaying a read-only
session under --yes compares nothing.
trifecta: TrifectaPolicy§sandbox: Stringnone | bwrap | docker | landlock. Load-bearing beyond the
obvious: shell declares narrower capabilities when confined, and
the interlock believes them, so the same prompt can be refused in one
and allowed in the other. (landlock never narrows external_send —
see the sandbox module — so it patterns with none for the interlock
while still confining files.)
sandbox_network: bool