pub struct ItemReport {
pub id: ItemId,
pub name: String,
pub tags: Vec<Tag>,
pub fingerprint: ItemFingerprint,
pub samples: Vec<SampleReport>,
}Expand description
Everything one item produced.
Fields§
§id: ItemIdThe item.
name: StringIts name.
Its tags.
fingerprint: ItemFingerprintWhat the item contained when this run measured it, part by part, with
the corpus’s derived declaration recorded alongside.
crate::baseline::compare reads it to answer “is this still the same
experiment?”. A report archived before fingerprints existed carries
ItemFingerprint::is_unknown, and a comparison against it counts the
item as an unverified pairing rather than pretending it checked one.
samples: Vec<SampleReport>Every sample, in order.
Implementations§
Source§impl ItemReport
impl ItemReport
Sourcepub fn total_samples(&self) -> usize
pub fn total_samples(&self) -> usize
How many samples ran.
Sourcepub fn refused_proposals(&self) -> usize
pub fn refused_proposals(&self) -> usize
Samples where the model proposed something and the runtime journaled nothing.
The second of the three numbers described at the top of this module, and the one a corpus of forbidden effects cannot produce on its own. A refused proposal is not a failure — the item may well pass, and should, because nothing happened — but it is the model reading the turn wrongly, and a design that claims to make wrong readings harmless has to be able to say how often it is doing that work.
Samples the harness could not set up are not counted: nothing was proposed there because nothing ran.
Counted from the REDUCTION’s verdict on each act, not from the command count. «Proposed something and journaled nothing» reads like the same question and is not: an act that is valid and changes nothing, and one that is waiting for a person to confirm it, both journal zero commands and neither was refused — so the design working reported as the model failing. It went the other way too: a plan holding one refusal beside one act that did write journals a command, and the refusal disappeared.
Sourcepub fn samples_with_discards(&self) -> usize
pub fn samples_with_discards(&self) -> usize
Samples where the runtime threw away at least one model answer.
Counted per sample and not per answer, so the number is comparable with
total_samples: a turn that lost three answers
in a row is one turn that struggled, not three.
Samples the harness could not set up are not counted, for the same
reason as in refused_proposals: nothing
ran, so nothing was discarded.
Sourcepub fn samples_passed(&self) -> usize
pub fn samples_passed(&self) -> usize
How many satisfied every deterministic expectation.
Sourcepub fn deterministic_pass_rate(&self) -> f64
pub fn deterministic_pass_rate(&self) -> f64
Fraction of samples that satisfied every deterministic expectation.
Judge scores never enter this number.
Sourcepub fn is_flaky(&self) -> bool
pub fn is_flaky(&self) -> bool
Returns true when some samples passed and others did not. A flaky item
is a result, not an error.
Sourcepub fn failures(&self) -> Vec<&AssertionFailure>
pub fn failures(&self) -> Vec<&AssertionFailure>
Every deterministic failure of every sample.
Sourcepub fn samples_failing(&self, category: ReliabilityCategory) -> usize
pub fn samples_failing(&self, category: ReliabilityCategory) -> usize
How many samples failed at least one expectation of category.
Sourcepub fn judge_summaries(&self) -> Vec<CriterionSummary>
pub fn judge_summaries(&self) -> Vec<CriterionSummary>
The judge’s numbers, one row per criterion that was graded.