Expand description
What a run produced: per item, per suite, and in a form a build can gate on (spec §26.3, §27.6).
§One number would be a lie
§26.3 is blunt about it: “a single ‘agent accuracy’ percentage hides the most important distinctions”. A corpus in which the assistant sent a rebooking it should not have, but phrased three replies beautifully, can average to a very healthy figure. So the report keeps the dashboard’s categories apart — side-effect integrity, operational claim integrity, semantic interpretation, clarification, abandonment, provider failure and the judge’s user-experience scores — and refuses to combine them into one.
§Three numbers, and each one sees what the others cannot
The pass rate is about effects: what the turn did and did not do. It is
the number a release blocks on, and on its own it flatters the design,
because every reading the runtime refused to carry out reads as a clean
turn. ItemReport::refused_proposals is the correction: the model
proposed an act and nothing was journaled, which is a fact about the model’s
reading rather than about the effect.
Neither of them can see a turn whose answer never became a plan at all.
ItemReport::samples_with_discards is that one: the runtime read what the
model produced and threw it away whole — an invented citation, a question
quoting nothing, a document of the wrong shape — and asked again. It is
usually invisible, since the repair round recovers and the effects come out
right, and when it is not invisible the turn simply has no effects, which
looks exactly like a turn that correctly had nothing to do.
§Samples and votes stay apart too
ItemReport::deterministic_pass_rate counts samples, never votes. A
judge score never enters it. CriterionSummary carries the judge’s
numbers with their vote spread, so a reader can see whether a low score is
the model’s fault or the judge’s disagreement with itself.
Structs§
- Behaviour
Count - One distinct behaviour and how often it happened.
- Category
Stats - Samples counted for one category.
- Criterion
Summary - The judge’s numbers for one criterion of one item.
- Eval
Report - One whole run.
- Gate
Outcome - The gate’s verdict.
- Gate
Thresholds - What a continuous integration gate refuses to merge.
- Gate
Violation - One unmet threshold.
- Item
Report - Everything one item produced.
- Reliability
- The dashboard of §26.3, with nothing averaged across its rows.
- Sample
Report - One execution of one item.
- Variance
- How much the agent under test varied across the samples of one item.
Enums§
- Reliability
Category - The reliability categories of the specification’s dashboard (§26.3).
- Report
Error - Why a report could not be rendered or read back.