Skip to main content

Module config

Module config 

Source
Expand description

How a run is configured (spec §27.6).

The whole point of this module is one distinction the specification insists on, and it is worth stating before any type appears:

  • A sample is one execution of the agent under test. More samples measure how much the model varies. Ten samples of the same item are ten chances for understanding to read something different, and the spread between them is the number a release gate cares about.
  • A vote is one judge opinion about one sample. More votes measure how much the judge varies. Three votes on one sample tell you nothing about the agent; they tell you whether the judge would have said the same thing twice.

Averaging the two together produces a number that moves when either the model or the judge wobbles and cannot say which — which is exactly the failure §26.3 warns about. So they are two settings, they are counted separately, and they are reported separately.

The file format is the one printed in the specification:

[execution]
samples_per_item = 10

[judging]
votes_per_sample = 3

Structs§

EvalConfig
Everything one evaluation run needs to know.
ExecutionConfig
How the agent under test is exercised.
JudgingConfig
How the judge is polled.
SelectionConfig
Which items of a suite a run selects.

Enums§

ConfigError
Why a configuration could not be loaded or run.