Skip to main content

Module eval

Module eval 

Source
Expand description

Evalset runs — the one reader for the evalset:<hash> mg:eval_run summaries that areev eval run journals into agent:harness.

Two edges consume these, and they deliberately share this code:

  • the gating edge (areev loop apply --gating-run <id>), which names one run and reads its numbers rather than trusting the command line, and
  • the outcome edge (measure_metric’s evalset: kind), which re-reads the newest run at each checkpoint.

Sharing matters because the alternative is two parsers of the same JSON drifting apart, so a rule could be admitted on one reading of an evalset and judged on another.

Structs§

EvalRun
One recorded execution of an evalset.

Constants§

EVAL_RUN_RELATION
The relation an eval-run summary is recorded under.
HARNESS_NS
The namespace those summaries live in.

Functions§

eval_run_by_id
One recorded run by id, over all of history.
eval_runs
Every recorded run of evalset_hash, oldest first.
newest_eval_run
The newest recorded run of evalset_hash at or after since_ms.
newest_eval_run_before
The newest recorded run of evalset_hash strictly before before_ms — the state of the world when something was about to be applied.
parse_evalset_metric
Parse an evalset:<hash>:<field> metric string into its parts.
run_value
The value a metric field takes on one run. failed/passed/total are promoted so a metric can be written against any evalset without the host having to add fields; error_rate is derived (undefined, not zero, when no case ran); anything else is read from the summary the host did write. The proposal-time baseline and every later measurement go through this one reader, so a rule cannot be admitted on one reading of a run and judged on another.