pub struct RunConfig {Show 21 fields
pub mecha_version: String,
pub provider: String,
pub model: String,
pub workspace: PathBuf,
pub system_prompt: Option<String>,
pub tools: Vec<String>,
pub effort: Option<Effort>,
pub temperature: Option<f64>,
pub seed: Option<u64>,
pub thinking: bool,
pub cache_prompt: bool,
pub max_tokens: u32,
pub max_turns: u32,
pub max_output_tokens: Option<u64>,
pub max_cost_usd: Option<f64>,
pub compact_at_tokens: Option<u64>,
pub compact_keep_recent: usize,
pub permission_mode: PermissionMode,
pub trifecta: TrifectaPolicy,
pub sandbox: String,
pub sandbox_network: bool,
}Expand description
What a run was configured with, recorded so it can be replayed.
The rule behind the field list: anything that shapes the request or constrains the run is a confound if it is not recorded. That is not theoretical here — compaction on versus off measured 1/5 against 5/5 on the same task, so a replay that did not know whether compaction was enabled would compare two incomparable runs and report a model regression.
The system prompt is stored in full rather than hashed. A hash tells you only that something differed; the text lets a replay rebuild the request. It is no more sensitive than the transcript sitting beside it.
The sampler is recorded only as far as it is pinned: temperature and
seed hold what this process sent, and None means the server chose.
Replay against an unpinned run has to be pass@k-shaped rather than
exact-match-shaped; against a pinned, seeded run driven sequentially it can
expect to match. (Not greedy — temperature 0.0 walks qwen3.6 into verbatim
repetition loops. And only sequentially: llama-server’s continuous batching
makes concurrent requests perturb each other’s numerics, seed or no seed.)
Fields§
§mecha_version: StringWhich harness produced this. The axis every replay diff is measured on.
provider: String§model: String§workspace: PathBuf§system_prompt: Option<String>The resolved text, not the path it may have come from.
tools: Vec<String>Tool names in registry order — which is the order they are sent, and the front of the cached prefix. A tool added, removed or renamed between recording and replay changes what the model could have done.
effort: Option<Effort>§temperature: Option<f64>The temperature and seed actually sent, when the provider config pins them. Unset means the server chose, and the run is not repeatable.
seed: Option<u64>§thinking: bool§cache_prompt: boolNo effect on semantics; large effect on the token counts a replay diffs.
max_tokens: u32§max_turns: u32§max_output_tokens: Option<u64>§max_cost_usd: Option<f64>§compact_at_tokens: Option<u64>§compact_keep_recent: usize§permission_mode: PermissionModeA denied call redirects the whole trajectory, so replaying a read-only
session under --yes compares nothing.
trifecta: TrifectaPolicy§sandbox: Stringnone | bwrap | docker | landlock. Load-bearing beyond the
obvious: shell declares narrower capabilities when confined, and
the interlock believes them, so the same prompt can be refused in one
and allowed in the other. (landlock never narrows external_send —
see the sandbox module — so it patterns with none for the interlock
while still confining files.)
sandbox_network: bool