Expand description
Comparing a run against a previous one — and refusing to, when the two runs stopped being the same experiment.
Two things look like “the evaluation got worse” and must not be treated the
same. A deterministic regression is an item that used to satisfy its
assertions and no longer does: something the agent does changed, it is
reproducible, and it can be a merge blocker. A judge drift is the same
behaviour graded differently: a signal about the measurement, and gating a
merge on it is how a team learns to ignore its own evaluation. ChangeKind
keeps them in separate variants.
The third thing is an item that changed because the code under test generates
part of it, which silently unpairs the comparison. Exclusion is then as loud
as the score, there is no headline to read past
ComparisonPolicy::max_excluded_share, and the three kinds of item change
have three different names. NoiseFloor makes a difference news only when
it is bigger than the one the unchanged system produces.
The reasoning, and what each of the three names means, is in
docs/evaluation.md.
Structs§
- Change
- One item’s change.
- Comparison
- What changed between two runs of the same suite.
- Comparison
Policy - Everything
compareneeds beyond the two reports. - Drift
Tolerance - How much movement is noise rather than news.
- Excluded
Item - One item that appeared in both runs and could not be paired.
- Headline
Figures - The figures of a comparison that stood.
- Withheld
Headline - Why a comparison produced no figure.
Enums§
- Change
Kind - The kinds of change a comparison reports.
- Exclusion
Reason - Why an item could not be paired between two runs.
- Headline
- The figures of a comparison, or the reason there are none.
- Noise
Verdict - How a movement compares with what the unchanged system produced against itself.
Functions§
- compare
- Compares a current run against a baseline.