Expand description
eval-core — pytest/jest, but for LLM agents: a batteries-included agent testing framework
where a test case is a prompt and the assertions are built-in checks on what the agent DID —
which tools it called, with which parameters, and what it finally said or computed — scored over
the universal RunArtifacts.
Prompts in, assertions on behavior out; bring your own harness. For the common case a host
implements ONE method (Agent::run), authors expect::Expectation predicates (in RON or
inline), and calls run_suite — no World, no Setup, no Scorer impl. It is
game-agnostic, so it doubles as a generic result/metric data model plus a self-contained HTML
comparison report (a small “Weights & Biases for evals”).
§Quickstart
Implement Agent::run over your harness (run one prompt, return what the agent did via the
with_* builders + ToolCall::new), author Expectation cases, and call run_suite:
use eval_core::{run_suite, Agent, EvalCase, EvalError, Expectation, RunArtifacts, ToolCall};
use serde_json::json;
// A toy agent (no real LLM): for an "add" prompt it emits a calculator tool call and ends with
// the sum; for anything else it just greets, making no tool call.
struct MyAgent;
impl Agent for MyAgent {
fn run(&self, instruction: &str) -> Result<RunArtifacts, EvalError> {
if instruction.contains("add") {
Ok(RunArtifacts::new()
.with_tool_calls(vec![ToolCall::new(
"calculator",
json!({ "op": "add", "a": 2, "b": 2 }),
)])
.with_final_text("The answer is 4."))
} else {
Ok(RunArtifacts::new().with_final_text("Hello!"))
}
}
}
let cases: Vec<EvalCase<(), Expectation>> = vec![
EvalCase {
name: "adds-two-numbers".to_owned(),
instruction: "please add 2 and 2".to_owned(),
setup: (), // no `setup` on the easy path — it is `()`
expect: vec![
Expectation::CalledToolWith {
tool: "calculator".to_owned(),
args: json!({ "op": "add" }),
},
Expectation::FinalNumberEquals { value: 4.0, tolerance: 0.0 },
],
},
EvalCase {
name: "no-tools-for-chitchat".to_owned(),
instruction: "hello there".to_owned(),
setup: (),
expect: vec![Expectation::NoToolCalls],
},
];
let report = run_suite(&MyAgent, &cases);
assert_eq!(report.total(), 2);
assert_eq!(report.passed(), 2); // both cases pass
// `println!("{report}")` prints the human-readable summary table.In practice cases are usually authored as RON and loaded with load_cases:
(
name: "adds-two-numbers",
instruction: "what is 2 + 2?",
expect: [
CalledToolWith(tool: "calculator", args: { "op": "add" }),
FinalNumberEquals(value: 4.0),
],
)eval-core also ships a ready-to-run baseline() suite (arithmetic / language / tool-use, 18
cases) you can hand straight to run_suite, and baseline_files to dump it as a template.
See examples/calculator.rs for a complete, dependency-free agent-framework example, and the
crate README.md for the full assertion catalog.
§Isolation guarantee
This crate depends only on small, well-scoped third-party crates (serde, serde_json, ron,
regex, thiserror, anyhow, tracing, chrono (local timestamps on auto-persisted runs),
include_dir to embed the shipped baseline suite, and ureq (a small blocking HTTP client with
rustls TLS, used to upload finished runs to the EvalForge dashboard)). It has ZERO dependency on any
host engine/game crate, so it can be lifted into a standalone public repository unchanged. The
dependency arrow points one way: a host harness depends on eval-core, never the reverse.
§Modules
report— the result/metric data model:report::RunRecord,report::EvalReport,report::CaseOutcome, with a readableDisplaysummary and the aggregate statistics (accuracy, latency percentiles, token totals).report_html— the self-contained HTML report generator (report_html::generate_report): loads persistedreport::RunRecords from a directory and writes a single offlinereport.html.persist— automatic run persistence (persist::save_and_report/persist::save_record): write a run as a JSONreport::RunRecordand regeneratereport.html. Driven automatically when aRunMetacarries apersist::Persisttarget (seeRunMeta::persist_to).upload— automatic upload of a finished run to the EvalForge API (evalforge.ai), configured at runtime with a project id + API key viaRunMeta::upload_to/RunMeta::upload_from_env(envEVALFORGE_API_KEY); reuses the samereport::RunRecordas the request body.case— the generic, RON-authored case containerEvalCase+ the fail-loudload_casesloader (andparse_cases_from_strfor one-or-many cases per file), both generic over the host’sSetup/Expecttypes.baseline— a shipped, ready-to-run baseline capability suite (baseline()/baseline_files): basic arithmetic / language / tool-use checks, embedded into the crate, that a user runs against their agent in one call or copies as a template.harness— theHarnesstrait (the thing being benchmarked), the easy-pathAgenttrait,RunArtifacts(what one run produced, minus scoring), and the structuredharness::ToolCall.expect— the built-in assertion libraryexpect::Expectation(tool use / text / math / health checks overRunArtifacts), serde/RON-authored.scorer— theScorertrait (score one expectation against the post-run world + artifacts) and the batteries-includedBuiltinScorer.runner— the generic engine:run_eval/run_eval_with_metatie aHarness+ aScorerover a shared world;run_suite/run_suite_with_metaare the easy path (Agent+BuiltinScorer). Both run every case, time each, isolate panics, and assemble anreport::EvalReport.error— the publicEvalErrorsurfaced byload_casesandAgent::run.
§Advanced — the full path (custom world)
When scoring needs post-run WORLD state, implement Harness over your agent + world, implement
Scorer over the same world, and call run_eval. See examples/minimal.rs.
Re-exports§
pub use baseline::baseline;pub use baseline::baseline_files;pub use case::EvalCase;pub use case::load_cases;pub use case::parse_cases_from_str;pub use error::EvalError;pub use expect::Expectation;pub use harness::Agent;pub use harness::Harness;pub use harness::RunArtifacts;pub use harness::ToolCall;pub use persist::Persist;pub use persist::build_record;pub use persist::save_record;pub use persist::write_record_and_report;pub use runner::AgentHarness;pub use runner::RunMeta;pub use runner::run_eval;pub use runner::run_eval_with_meta;pub use runner::run_suite;pub use runner::run_suite_with_meta;pub use scorer::BuiltinScorer;pub use scorer::Scorer;pub use upload::Upload;pub use upload::UploadResponse;pub use upload::upload_record;
Modules§
- baseline
- A shipped, ready-to-run baseline capability suite — basic checks any agent can be measured against in one call.
- case
- The generic, RON-authored eval case schema + loader.
- error
- The crate’s public error type,
EvalError. - expect
- The built-in assertion library:
Expectation, a serde/RON-authored predicate over a run’s universalRunArtifacts. - harness
- The thing being benchmarked: the host’s agent harness, behind the
Harnesstrait, plusRunArtifacts— everything a single run produced EXCEPT scoring — and the structuredToolCallthe artifacts carry. - persist
- Automatic persistence of an eval run: write it as a JSON
RunRecordinto a results directory and (re)generate the self-containedreport.htmlover every run saved there. - report
- The eval report types: one
CaseOutcomeper case and the aggregateEvalReportwith a readableDisplaysummary table. - report_
html - Self-contained HTML comparison report over the persisted eval runs.
- runner
- The generic benchmark runner: ties a
Harness+Scorerover a sharedWorld, runs everyEvalCase, times each, isolates panics, and assembles anEvalReport. - scorer
- Scoring one expectation against a run’s result, behind the
Scorertrait — plus the batteries-includedBuiltinScorerthat needs NO host scoring code at all. - upload
- Automatic upload of a finished eval run to the EvalForge API (evalforge.ai) so results show up in the online dashboard with no manual export/import.