Skip to main content

Module eval

Module eval 

Source
Expand description

Evalset runs — the one reader for the evalset:<hash> mg:eval_run summaries that areev eval run journals into agent:harness.

Two edges consume these, and they deliberately share this code:

  • the gating edge (areev loop apply --gating-run <id>), which names one run and reads its numbers rather than trusting the command line, and
  • the outcome edge (measure_metric’s evalset: kind), which re-reads the newest run at each checkpoint.

Sharing matters because the alternative is two parsers of the same JSON drifting apart, so a rule could be admitted on one reading of an evalset and judged on another.

Structs§

EvalRun
One recorded execution of an evalset.
RunSpend
The spent figures of a terminal run_outcome Observation (spent_input_tokens … spent_wall_ms), integers as the runtime wrote them.

Constants§

COST_FIELDS
The cost fields run_value promotes, beside the quality ones. tokens is input + output; usd is usd_micros / 1e6; cost_per_pass is usd / passed, undefined when nothing passed.
EVAL_RUN_RELATION
The relation an eval-run summary is recorded under.
HARNESS_NS
The namespace those summaries live in.

Functions§

eval_run_by_id
One recorded run by id, over all of history.
eval_runs
Every recorded run of evalset_hash, oldest first.
newest_eval_run
The newest recorded run of evalset_hash at or after since_ms.
parse_evalset_metric
Parse an evalset:<hash>:<field> metric string into its parts.
run_value
The value a metric field takes on one run. failed/passed/total are promoted so a metric can be written against any evalset without the host having to add fields; error_rate is derived (undefined, not zero, when no case ran); the cost fields (COST_FIELDS) read the integer cost keys fail-closed, with cost_per_pass undefined — not zero, not a division — when nothing passed; anything else is read from the summary the host did write. The proposal-time baseline and every later measurement go through this one reader, so a rule cannot be admitted on one reading of a run and judged on another.