Expand description
The corpus: named scenarios loaded from files (spec §27.6).
An item is one scenario — a starting state, the turn a person takes, and what must be true afterwards. A suite is a set of items with tags, so a run can say “only the trip write paths” without a second list of file names.
The loader refuses what it does not understand. Every structure carries
deny_unknown_fields, and EvalItem::validate rejects combinations that
parse but cannot mean anything. A loader that skipped a key it did not
recognise would quietly turn a typo into a weaker test, and a weaker test
into a green build.
An item also says where each of its parts came from — authored, derived
or recorded (PartProvenance) — declared per item or per directory
(SuiteManifest). EvalItem::fingerprint records what was declared, and
crate::baseline::compare uses it to tell an intended change from a broken
pairing from a corpus defect. What each name means, and why no tooling may
ever regenerate a recorded part, is in
docs/evaluation.md.
§An item file
id = "trip.set_name"
name = "Setting the subject commits exactly one command"
tags = ["trip", "write"]
[[setup.cases]]
workflow = "trip"
case_id = "trip-1"
label = "Trip 1"
revision = 3
state = { status = "draft" }
[turn]
text = "Set the name to Lisbon"
[expect]
commands = ["trip.set_name"]
events = ["trip.name_set"]
blocks = ["receipt"]
[[expect.acts]]
kind = "apply_operation"
operation = "trip.set_name"
[expect.forbid]
commands = ["trip.rebook"]
judge = ["language_quality"]Structs§
- ActExpectation
- One expected normalized act.
- Card
Reply Spec - A click, named by what it answers rather than by an identifier.
- Case
Count Expectation - How many cases of a workflow must exist at the end.
- Case
Seed - One case seeded before the turn.
- Eval
Item - One named scenario.
- Expectations
- What must be true after the turn, checked without a model (spec §27.6).
- External
Spec - A change to a record made outside the conversation, as its own system makes it: an airline re-quoting a fare. The command is the domain’s own, as the workflow serializes it, and never passes through understanding.
- Forbidden
Effects - Effects an item asserts must not happen.
- Interaction
Status Expectation - The statuses the cards of a case must have, in creation order (spec §15.4).
- Item
Fingerprint - What an item contained when a run measured it.
- ItemId
- Stable identifier of one corpus item.
- Origin
Spec - A server-issued origin token (spec §12.4).
- Part
Digest - One part’s digest, and where the corpus said that part came from.
- Prior
Exchange - An exchange that already happened in this conversation.
- Provenance
- Where each part of an item came from.
- Revision
Expectation - The revision a case must end the turn at.
- Seeded
Record - Something the world holds that is not a case.
- Setup
- The world before the turn.
- State
Expectation - What a case’s state must say after the turn.
- Suite
- A set of items with a name.
- Suite
Manifest - The optional
suite.toml(orsuite.json) beside a directory of items. - Tag
- A label a run can select on.
- Target
Expectation - How one act’s target must resolve (spec §12.2).
- Turn
Spec - The turn the person takes (spec §9: text and a card answer may coexist).
- Workflow
State Expectation - A value some case of a workflow must hold at the end.
Enums§
- ActKind
- The act variants an item can expect, named exactly as they appear in JSON.
- Block
Kind - The kinds of response block, as an item file names them (spec §18.1).
- Corpus
Error - Why a corpus could not be loaded.
- Item
Part - A part of an item, named exactly as the item file names it.
- Outcome
Expectation - Whether the turn is expected to complete.
- Part
Provenance - Where a part of an item came from.
- Resolution
Kind - The resolutions of
TargetResolution, as an item file names them.