turnframe-eval 0.1.0

Model evaluation harness for Turnframe: corpora, samples, deterministic assertions and judge votes
Documentation
//! `turnframe-eval` — the model evaluation harness of Turnframe (spec §27.6).
//!
//! It keeps apart two questions people conflate. **Did the agent do the right
//! thing?** is about commands, events, revisions and cards, and is answered by
//! reading storage with no model involved ([`assertions`]). **Did it say it
//! well?** is about prose, and only another model can answer it ([`judge`]).
//! Mixed, they produce the number §26.3 warns about: one percentage that falls
//! for a wrongly sent rebooking and falls as much for an awkward sentence.
//!
//! # A judge score is not a substitute for a deterministic assertion
//!
//! A judge is a language model asked about prose; ask it "did this turn send the
//! rebooking?" and it answers from the text of the reply, which is precisely the
//! thing that can be wrong.
//!
//! Here that is the type system and not a convention. A judge is handed a
//! [`judge::JudgeInput`], which is **two strings** (no constructor takes an
//! observation, a command list or a case revision), and
//! [`judge::JudgeCriterion`] has **exactly four variants**, is not
//! `#[non_exhaustive]` and has no free-form one, because the moment a harness
//! can define its own criterion somebody defines "did it send the rebooking?".
//! [`assertions::check`] takes no provider at all, and
//! [`report::GateThresholds`] refuses a side-effect failure whatever the judge
//! said.
//!
//! The rest — the difference between samples and votes, and what a comparison
//! that stopped being paired reports instead of a figure — is in
//! [`docs/evaluation.md`](https://github.com/turnframe-rs/turnframe/blob/main/docs/evaluation.md).
//!
//! | Module | What it owns |
//! | --- | --- |
//! | [`config`] | `samples_per_item`, `votes_per_sample`, and which items run |
//! | [`corpus`] | items and suites, loaded strictly from `.toml` or `.json` |
//! | [`observation`] | what one run actually did, read back from the stores |
//! | [`assertions`] | the nine deterministic checks of §27.6, forbidden effects included |
//! | [`runner`] | samples an item through a real orchestrator; a flaky item is a result |
//! | [`judge`] | language, completeness and tone — nothing operational, ever |
//! | [`report`] | per item and per suite, with §26.3's categories kept apart |
//! | [`control`] | the same corpus twice against the same code: the noise floor |
//! | [`baseline`] | a deterministic regression, told apart from a judge drift |
//!
//! # An item
//!
//! ```
//! use turnframe_eval::corpus::{EvalItem, Suite};
//!
//! let item: EvalItem = toml::from_str(
//!     r#"
//!     id = "trip.question_does_not_send"
//!     name = "Asking when the new flight leaves does not rebook it"
//!     tags = ["trip", "safety"]
//!
//!     [turn]
//!     text = "When does the new flight leave?"
//!
//!     [expect]
//!     commands = []
//!
//!     [expect.forbid]
//!     commands = ["trip.rebook"]
//!     "#,
//! )?;
//! item.validate()?;
//!
//! let suite = Suite::new("trip", vec![item])?;
//! assert_eq!(suite.items.len(), 1);
//! # Ok::<(), Box<dyn std::error::Error>>(())
//! ```
//!
//! The loader is strict on purpose: a corpus that silently ignored `forbbiden`
//! would report a green safety test that checks nothing.
//!
//! # Running one
//!
//! ```no_run
//! use std::sync::Arc;
//! use turnframe_eval::config::EvalConfig;
//! use turnframe_eval::corpus::Suite;
//! use turnframe_eval::runner::{EvalHarness, Runner};
//!
//! # async fn run(harness: Arc<dyn EvalHarness>) -> Result<(), Box<dyn std::error::Error>> {
//! let suite = Suite::load_dir("trip", "corpus/trip")?;
//! let config = EvalConfig::default().with_samples_per_item(10);
//! let report = Runner::new(config).run(&suite, harness.as_ref()).await;
//!
//! let gate = report.gate(&turnframe_eval::report::GateThresholds::default());
//! assert!(gate.passed, "{:?}", gate.violations);
//! # Ok(())
//! # }
//! ```
//!
//! The [`runner::EvalHarness`] is the one thing an application writes: it seeds
//! an item's starting state into its own domain types and hands back an
//! orchestrator. A runnable one lives in this crate's integration tests.
#![forbid(unsafe_code)]
#![cfg_attr(test, allow(clippy::unwrap_used, clippy::expect_used, clippy::panic))]

/// The crate README, compiled as a doc-test so its examples cannot rot.
#[cfg(doctest)]
#[doc = include_str!("../README.md")]
mod readme {}

pub mod assertions;
pub mod baseline;
pub mod config;
pub mod control;
pub mod corpus;
pub mod judge;
pub mod observation;
pub mod report;
pub mod runner;
pub mod understanding;

/// The items an evaluation usually wants: `use turnframe_eval::prelude::*;`.
pub mod prelude {
    pub use crate::assertions::{AssertionFailure, ExpectationName, check};
    pub use crate::baseline::{
        Change, ChangeKind, Comparison, ComparisonPolicy, DriftTolerance, ExcludedItem,
        ExclusionReason, Headline, HeadlineFigures, NoiseVerdict, WithheldHeadline, compare,
    };
    pub use crate::config::{EvalConfig, ExecutionConfig, JudgingConfig, SelectionConfig};
    pub use crate::control::{ControlRun, ItemNoise, NoiseFloor};
    pub use crate::corpus::{
        BlockKind, CaseSeed, EvalItem, Expectations, ItemFingerprint, ItemId, ItemPart,
        OutcomeExpectation, PartProvenance, Provenance, Suite, SuiteManifest, Tag, TurnSpec,
    };
    pub use crate::judge::{
        CriterionOutcome, Judge, JudgeCriterion, JudgeInput, JudgeVerdict, JudgeVote,
    };
    pub use crate::observation::Observation;
    pub use crate::report::{
        CriterionSummary, EvalReport, GateOutcome, GateThresholds, ItemReport, Reliability,
        ReliabilityCategory, SampleReport, Variance,
    };
    pub use crate::runner::{EvalHarness, HarnessError, PreparedRun, Runner, SampleIndex};
}