pub struct EvalItem {
pub id: ItemId,
pub name: String,
pub description: Option<String>,
pub tags: Vec<Tag>,
pub setup: Setup,
pub before: Vec<TurnSpec>,
pub turn: TurnSpec,
pub expect: Expectations,
pub judge: Vec<JudgeCriterion>,
pub provenance: Provenance,
}Expand description
One named scenario.
Fields§
§id: ItemIdStable identifier, unique within a suite. Reports and baselines join on it, so renaming one is renaming a measurement.
name: StringOne sentence a human reads in a report.
description: Option<String>Longer prose, when the scenario needs it.
Labels a run selects on.
setup: SetupThe state the world is in before the turn.
before: Vec<TurnSpec>Turns the person takes first, in the same conversation. Only the last turn is observed; what the conversation built is read back after it.
turn: TurnSpecThe turn the person takes.
expect: ExpectationsWhat must be true afterwards, checked without a model.
judge: Vec<JudgeCriterion>Which linguistic qualities a judge grades. Empty means no judge runs, which is the right answer for most items.
provenance: ProvenanceWhere each part of this item came from: authored, derived or recorded.
Declared once here, or once for a whole directory in a
SuiteManifest. It changes nothing about how the item runs and
everything about how crate::baseline::compare reads a difference — an
intended projection change, a broken pairing, or a corpus defect.
Implementations§
Source§impl EvalItem
impl EvalItem
Sourcepub fn validate(&self) -> Result<(), CorpusError>
pub fn validate(&self) -> Result<(), CorpusError>
Sourcepub fn carries(&self, part: ItemPart) -> bool
pub fn carries(&self, part: ItemPart) -> bool
Returns true when the item actually has something in part.
A turn always has something in it — TurnSpec::validate refuses one
that does not — so ItemPart::Turn is always carried.
Sourcepub fn fingerprint(&self) -> ItemFingerprint
pub fn fingerprint(&self) -> ItemFingerprint
Digests the item part by part, recording where the corpus said each part came from.
A run stores this on ItemReport so a later
comparison can tell whether the two runs measured the same thing. The
digest is over a canonical rendering with object keys sorted, so an item
whose seeded state was written out with its fields in a different order
still fingerprints the same.
use turnframe_eval::corpus::{EvalItem, ItemPart, PartProvenance};
let item: EvalItem = toml::from_str(
r#"
id = "a"
name = "A scenario"
[provenance]
setup = "derived"
[turn]
text = "hello"
[[setup.cases]]
workflow = "trip"
case_id = "trip-1"
label = "Trip 1"
state = { status = "draft" }
"#,
)?;
item.validate()?;
let fingerprint = item.fingerprint();
assert_eq!(fingerprint.provenance_of(ItemPart::Setup), PartProvenance::Derived);
assert_eq!(fingerprint.provenance_of(ItemPart::Turn), PartProvenance::Authored);