Skip to main content

Module report

Module report 

Source
Expand description

What a run produced: per item, per suite, and in a form a build can gate on (spec §26.3, §27.6).

§One number would be a lie

§26.3 is blunt about it: “a single ‘agent accuracy’ percentage hides the most important distinctions”. A corpus in which the assistant sent a rebooking it should not have, but phrased three replies beautifully, can average to a very healthy figure. So the report keeps the dashboard’s categories apart — side-effect integrity, operational claim integrity, semantic interpretation, clarification, abandonment, provider failure and the judge’s user-experience scores — and refuses to combine them into one.

§Three numbers, and each one sees what the others cannot

The pass rate is about effects: what the turn did and did not do. It is the number a release blocks on, and on its own it flatters the design, because every reading the runtime refused to carry out reads as a clean turn. ItemReport::refused_proposals is the correction: the model proposed an act and nothing was journaled, which is a fact about the model’s reading rather than about the effect.

Neither of them can see a turn whose answer never became a plan at all. ItemReport::samples_with_discards is that one: the runtime read what the model produced and threw it away whole — an invented citation, a question quoting nothing, a document of the wrong shape — and asked again. It is usually invisible, since the repair round recovers and the effects come out right, and when it is not invisible the turn simply has no effects, which looks exactly like a turn that correctly had nothing to do.

§Samples and votes stay apart too

ItemReport::deterministic_pass_rate counts samples, never votes. A judge score never enters it. CriterionSummary carries the judge’s numbers with their vote spread, so a reader can see whether a low score is the model’s fault or the judge’s disagreement with itself.

Structs§

BehaviourCount
One distinct behaviour and how often it happened.
CategoryStats
Samples counted for one category.
CriterionSummary
The judge’s numbers for one criterion of one item.
EvalReport
One whole run.
GateOutcome
The gate’s verdict.
GateThresholds
What a continuous integration gate refuses to merge.
GateViolation
One unmet threshold.
ItemReport
Everything one item produced.
Reliability
The dashboard of §26.3, with nothing averaged across its rows.
SampleReport
One execution of one item.
Variance
How much the agent under test varied across the samples of one item.

Enums§

ReliabilityCategory
The reliability categories of the specification’s dashboard (§26.3).
ReportError
Why a report could not be rendered or read back.