Expand description
A control run: the same corpus, twice, against the same code.
§Why a harness needs one
Every other number in this crate is a difference between two runs, and a difference is only news if it is bigger than the difference the harness produces when nothing has changed at all. Agents are sampled from a distribution; judges are too. Without a control, “the pass rate fell from 1.00 to 0.75” and “the pass rate wobbles by a quarter between any two runs” are indistinguishable, and a team that cannot tell them apart eventually learns to ignore its own evaluation.
So the affordance is first class. Runner::run_control
runs a suite twice against the same harness and hands back a ControlRun;
ControlRun::noise_floor turns the two reports into a NoiseFloor,
which is the largest movement the unchanged system produced against
itself. A Comparison built with that floor
then labels every change it reports as
WithinNoise or
ExceedsNoise, so a reader
never has to guess.
§What a control run cannot check for you
That both passes really saw the same code and the same corpus. Running them
back to back in one call is the strongest guarantee a library can offer;
deploying a change between the two passes would produce a “noise floor” that
is a measurement of the change, and nothing here can detect that. The
NoiseFloor::suite name and the fingerprints on both reports are what a
reviewer checks.
§A floor is a ceiling on credulity, not a licence
A large noise floor is itself the finding. A corpus whose control run moves by half is not a corpus that tolerates movement of half; it is a corpus whose items are too flaky to measure anything, and the honest next step is more samples per item rather than a wider tolerance.
Structs§
- Control
Run - Two runs of one suite against the same code.
- Item
Noise - How much one item moved between the two passes of a control run.
- Noise
Floor - How much the unchanged system moved when it was measured twice.