Expand description
Running a corpus against a real orchestrator (spec §27.6).
§What the runner does, and what it refuses to do
For every selected item it asks the harness for a fresh world, builds the
turn the item describes, runs it through a real Orchestrator, reads back
what happened from the stores, and checks the item’s deterministic
expectations. Then it does it again, samples_per_item times.
It never turns a failing item into an error. A scenario that passes seven
times out of ten is a measurement — the most valuable one in the whole
harness, because it is the one a single run cannot see. So the sample
records its failures, the item records its variance, and the run continues.
The only thing that stops a run early is an explicit
stop_after_failures.
§The seam
EvalHarness is the one thing an application implements. It owns the
domain: it knows how to turn the item’s seeded JSON state into a trip or a
traveler, which providers to configure, and how much autonomy to grant.
The runner owns everything that must not vary between applications — the
order, the sampling, the assertions and the report.
The sample index is handed to EvalHarness::prepare on purpose: a
harness that wants to exercise model variance without a network can script a
different answer per sample, and a harness talking to a real endpoint can
simply ignore it.
Structs§
- Prepared
Run - One item’s world, freshly built for one sample.
- Runner
- Runs a corpus.
- Sample
Index - Which execution of an item this is, zero-based.
Enums§
- Harness
Error - Why a sample could not be set up or started.
Traits§
- Eval
Harness - Builds one world per sample.