Expand description
Evalset runs — the one reader for the evalset:<hash> mg:eval_run
summaries that areev eval run journals into agent:harness.
Two edges consume these, and they deliberately share this code:
- the gating edge (
areev loop apply --gating-run <id>), which names one run and reads its numbers rather than trusting the command line, and - the outcome edge (
measure_metric’sevalset:kind), which re-reads the newest run at each checkpoint.
Sharing matters because the alternative is two parsers of the same JSON drifting apart, so a rule could be admitted on one reading of an evalset and judged on another.
Structs§
- EvalRun
- One recorded execution of an evalset.
Constants§
- EVAL_
RUN_ RELATION - The relation an eval-run summary is recorded under.
- HARNESS_
NS - The namespace those summaries live in.
Functions§
- eval_
run_ by_ id - One recorded run by id, over all of history.
- eval_
runs - Every recorded run of
evalset_hash, oldest first. - newest_
eval_ run - The newest recorded run of
evalset_hashat or aftersince_ms. - newest_
eval_ run_ before - The newest recorded run of
evalset_hashstrictly beforebefore_ms— the state of the world when something was about to be applied. - parse_
evalset_ metric - Parse an
evalset:<hash>:<field>metric string into its parts. - run_
value - The value a metric field takes on one run.
failed/passed/totalare promoted so a metric can be written against any evalset without the host having to add fields;error_rateis derived (undefined, not zero, when no case ran); anything else is read from the summary the host did write. The proposal-time baseline and every later measurement go through this one reader, so a rule cannot be admitted on one reading of a run and judged on another.