pub struct SampleReport {Show 14 fields
pub sample: u32,
pub failures: Vec<AssertionFailure>,
pub harness_error: Option<String>,
pub signature: String,
pub judge: Vec<CriterionOutcome>,
pub provider_failures: usize,
pub cards_created: usize,
pub acts_proposed: usize,
pub acts_refused: usize,
pub commands_journaled: usize,
pub abandoned: bool,
pub discarded_answers: Vec<String>,
pub answer: String,
pub tasks: TaskScores,
}Expand description
One execution of one item.
Fields§
§sample: u321-based position in the item’s samples.
failures: Vec<AssertionFailure>Deterministic expectations that did not hold.
harness_error: Option<String>The harness could not even set the sample up. Distinct from a failing assertion: nothing was measured.
signature: StringCanonical rendering of what this run did, for counting distinct behaviours across samples.
judge: Vec<CriterionOutcome>Judge opinions about this sample, when the item asked for any.
provider_failures: usizeHow many provider attempts failed or fell back.
cards_created: usizeHow many cards the turn created.
acts_proposed: usizeHow many acts the message was understood to ask for.
The pair below is the whole point of recording it: a turn where the model proposed something and nothing was journaled is a turn the runtime REFUSED, and that is a fact about the model’s reading rather than about the effect. A report that shows only the effect says an assistant is perfect exactly where its reading is worst — every refusal reads as a clean turn — which flatters a design whose whole claim is that it makes bad readings harmless.
acts_refused: usizeHow many of them the reduction refused outright.
Only rejected: an act the domain turned down. Not the one that
changes nothing, not the one awaiting a confirmation, not the one a
later act superseded — see
refused_proposals.
commands_journaled: usizeHow many commands the turn journaled.
Not comparable one-to-one with Self::acts_proposed: one act can
produce several commands and a command type is not an operation name. It
is here to be read against zero — nothing journaled after something was
proposed — and not as a ratio of the two.
abandoned: boolThe turn produced no usable answer at all.
discarded_answers: Vec<String>The stable code of every answer the runtime threw away whole this turn, in order and with repeats.
The third number of a reliability report, and the one neither of the other two can reach. The pass rate is about effects and the refused proposals are about a plan the runtime declined to carry out; this is about a plan that never became one, because the runtime read the model’s answer and refused it whole — a citation the user’s message does not contain, a question that does not quote what it answers, a document that does not match its schema.
Codes rather than the full records, because a report is read in
aggregate and the runtime’s own wording names positions inside one
turn’s plan. The full records, reasons included, are on
Observation::discarded_answers.
answer: StringWhat the turn actually said, when the run was configured to keep it.
Empty unless
ExecutionConfig::record_answers
is on — see there for why that is the default. It is here for a person
curating a corpus, never for an assertion: nothing in this crate reads
it, and an expectation that did would be measuring prose.
tasks: TaskScoresHow each understanding task did, where the item says what it expects of it.