pub struct EvalRun {
pub run_id: String,
pub passed: u64,
pub failed: u64,
pub summary: Map<String, Value>,
pub recorded_ms: i64,
pub spend: Option<RunSpend>,
}Expand description
One recorded execution of an evalset.
Fields§
§run_id: StringThe eval- run id the cases were journaled under.
passed: u64§failed: u64§summary: Map<String, Value>The whole summary object, so host-defined fields (category_accuracy,
…) are reachable without this module having to know them.
recorded_ms: i64When the summary grain was written.
spend: Option<RunSpend>What the run spent, when the runtime journaled it: the run_outcome
Observation areev run writes for the same run id. Read only for a
cost key the summary does not carry, so an eval run that IS an
areev run run quotes one spend to both the Verify gate and the
run_outcome analyzer.
Implementations§
Source§impl EvalRun
impl EvalRun
Sourcepub fn field(&self, name: &str) -> Option<f64>
pub fn field(&self, name: &str) -> Option<f64>
A numeric field of the summary. passed/failed are promoted to typed
fields but stay readable here too, so a metric string may name any of
them uniformly.
Sourcepub fn effects(&self) -> Option<u64>
pub fn effects(&self) -> Option<u64>
Tool calls the cases made. Summary key effects only — the runtime’s
Observation counts supersteps, not effects, so there is no fallback.