vtcode-eval
Agent evaluation framework for VT Code — defines eval tasks, runs them through an agent executor, grades outcomes, and reports pass@k / pass^k metrics split by capability and regression categories.
The crate is deliberately small and I/O-free at its core: run_suite orchestrates the loop of tasks × attempts,
computes per-task metrics, and assembles a report. Everything that touches the filesystem, config, or the concrete agent
runner is pushed behind the EvalExecutor trait, so the harness is fully unit-testable with an in-memory fake executor.
Layout
| Module | Responsibility |
|---|---|
task |
Data model: EvalTask, EvalCategory, RunOutcome, EvalRunResult |
suite |
EvalSuite — a named set of tasks with an attempts count |
metric |
EvalMetric and combinatorial pass@k, independent pass^k, and aggregate_metrics |
executor |
EvalExecutor trait + bounded run_suite_with_options orchestration |
environment |
EnvironmentProbe checks: CommandProbe, FileExistsProbe, GitCleanProbe |
report |
EvalReport / SuiteReport / TaskReport + to_markdown renderer |
trace_analyzer |
Privacy-preserving JSONL summaries for DeepSeek and VT Code harness traces |
Concepts
EvalTask— a prompt plusverify_commandsand an optionaltimeout_secs.categoryisCapabilityorRegression.RunOutcome—Pass,Fail, orErrorfor a single task attempt.EvalMetric— combinatorialpass_at_k, independent reliabilitypass_power_k((passed / attempts)^k), andpass_all_k, averaged per task. It also carries the selectedkand rawpassed_runs/total_runs.EvalExecutor— the trait boundary. Implementors own "run this task" semantics (drive the agent, apply environment probes, grade the result).run_suiteschedules attempts with a default concurrency of two;run_suite_with_optionsmakes the bound and metrickexplicit.
HarnessTraceSummary provides aggregate-only trace facts: turns, bounded tool/error counts, latency, output byte
totals, repetition, and token/cache usage. Raw prompts, arguments, file contents, and output text are not retained. Use
analyze_jsonl_file for buffered file analysis or analyze_jsonl_reader when the caller already owns a stream.
Usage
use ;
// Implement EvalExecutor to drive your agent + grade outcomes, then:
let report = run_suite.await?;
println!;
Trace analysis can be kept out of the agent hot path and run against a persisted JSONL session:
use analyze_jsonl_file;
let summary = analyze_jsonl_file?;
println!;
The analyzer recognizes both DeepSeek-style records and serialized vtcode-exec-events::ThreadEvent shapes.
Thread-level aggregate usage is used only when per-turn usage is absent, preventing double counting.
Notes
run_suiteperforms no file I/O or trust checks; the caller owns configuration, while the scheduler rejectsattempts == 0and invalid metrickvalues.- Reports sum known per-attempt costs and count unknown-pricing attempts separately; unknown cost is never silently treated as a known zero.
SuiteReportcarries cost-efficiency aggregates —cost_per_solve,mean_cost_per_attempt,mean_tokens_per_attempt,mean_turns_per_attempt.cost_per_solveisNonewhen any attempt is unpriced or nothing passed; the means cover priced/token/tracked attempts only. The markdown renderer prints them as anEfficiency: cost/solve … · per attempt · tokens · turnsline.- Environment verification (
EnvironmentProbe) is a separate concern from outcome grading — executor implementations decide whether and how to apply probes before returning aRunOutcome.