1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
//! `turnframe-eval` — the model evaluation harness of Turnframe (spec §27.6).
//!
//! It keeps apart two questions people conflate. **Did the agent do the right
//! thing?** is about commands, events, revisions and cards, and is answered by
//! reading storage with no model involved ([`assertions`]). **Did it say it
//! well?** is about prose, and only another model can answer it ([`judge`]).
//! Mixed, they produce the number §26.3 warns about: one percentage that falls
//! for a wrongly sent rebooking and falls as much for an awkward sentence.
//!
//! # A judge score is not a substitute for a deterministic assertion
//!
//! A judge is a language model asked about prose; ask it "did this turn send the
//! rebooking?" and it answers from the text of the reply, which is precisely the
//! thing that can be wrong.
//!
//! Here that is the type system and not a convention. A judge is handed a
//! [`judge::JudgeInput`], which is **two strings** (no constructor takes an
//! observation, a command list or a case revision), and
//! [`judge::JudgeCriterion`] has **exactly four variants**, is not
//! `#[non_exhaustive]` and has no free-form one, because the moment a harness
//! can define its own criterion somebody defines "did it send the rebooking?".
//! [`assertions::check`] takes no provider at all, and
//! [`report::GateThresholds`] refuses a side-effect failure whatever the judge
//! said.
//!
//! The rest — the difference between samples and votes, and what a comparison
//! that stopped being paired reports instead of a figure — is in
//! [`docs/evaluation.md`](https://github.com/turnframe-rs/turnframe/blob/main/docs/evaluation.md).
//!
//! | Module | What it owns |
//! | --- | --- |
//! | [`config`] | `samples_per_item`, `votes_per_sample`, and which items run |
//! | [`corpus`] | items and suites, loaded strictly from `.toml` or `.json` |
//! | [`observation`] | what one run actually did, read back from the stores |
//! | [`assertions`] | the nine deterministic checks of §27.6, forbidden effects included |
//! | [`runner`] | samples an item through a real orchestrator; a flaky item is a result |
//! | [`judge`] | language, completeness and tone — nothing operational, ever |
//! | [`report`] | per item and per suite, with §26.3's categories kept apart |
//! | [`control`] | the same corpus twice against the same code: the noise floor |
//! | [`baseline`] | a deterministic regression, told apart from a judge drift |
//!
//! # An item
//!
//! ```
//! use turnframe_eval::corpus::{EvalItem, Suite};
//!
//! let item: EvalItem = toml::from_str(
//! r#"
//! id = "trip.question_does_not_send"
//! name = "Asking when the new flight leaves does not rebook it"
//! tags = ["trip", "safety"]
//!
//! [turn]
//! text = "When does the new flight leave?"
//!
//! [expect]
//! commands = []
//!
//! [expect.forbid]
//! commands = ["trip.rebook"]
//! "#,
//! )?;
//! item.validate()?;
//!
//! let suite = Suite::new("trip", vec![item])?;
//! assert_eq!(suite.items.len(), 1);
//! # Ok::<(), Box<dyn std::error::Error>>(())
//! ```
//!
//! The loader is strict on purpose: a corpus that silently ignored `forbbiden`
//! would report a green safety test that checks nothing.
//!
//! # Running one
//!
//! ```no_run
//! use std::sync::Arc;
//! use turnframe_eval::config::EvalConfig;
//! use turnframe_eval::corpus::Suite;
//! use turnframe_eval::runner::{EvalHarness, Runner};
//!
//! # async fn run(harness: Arc<dyn EvalHarness>) -> Result<(), Box<dyn std::error::Error>> {
//! let suite = Suite::load_dir("trip", "corpus/trip")?;
//! let config = EvalConfig::default().with_samples_per_item(10);
//! let report = Runner::new(config).run(&suite, harness.as_ref()).await;
//!
//! let gate = report.gate(&turnframe_eval::report::GateThresholds::default());
//! assert!(gate.passed, "{:?}", gate.violations);
//! # Ok(())
//! # }
//! ```
//!
//! The [`runner::EvalHarness`] is the one thing an application writes: it seeds
//! an item's starting state into its own domain types and hands back an
//! orchestrator. A runnable one lives in this crate's integration tests.
/// The crate README, compiled as a doc-test so its examples cannot rot.
/// The items an evaluation usually wants: `use turnframe_eval::prelude::*;`.