vtcode-eval 0.172.1

Agent evaluation framework for VT Code: pass@k / pass^k metrics, capability and regression evals, and environment-based outcome verification.
Documentation

vtcode-eval

Agent evaluation framework for VT Code — defines eval tasks, runs them through an agent executor, grades outcomes, and reports pass@k / pass^k metrics split by capability and regression categories.

The crate is deliberately small and I/O-free at its core: run_suite orchestrates the loop of tasks × attempts, computes per-task metrics, and assembles a report. Everything that touches the filesystem, config, or the concrete agent runner is pushed behind the EvalExecutor trait, so the harness is fully unit-testable with an in-memory fake executor.

Layout

Module Responsibility
task Data model: EvalTask, EvalCategory, RunOutcome, EvalRunResult
suite EvalSuite — a named set of tasks with an attempts count
metric EvalMetric and combinatorial pass@k, independent pass^k, and aggregate_metrics
executor EvalExecutor trait + bounded run_suite_with_options orchestration
environment EnvironmentProbe checks: CommandProbe, FileExistsProbe, GitCleanProbe
report EvalReport / SuiteReport / TaskReport + to_markdown renderer
trace_analyzer Privacy-preserving JSONL summaries for DeepSeek and VT Code harness traces

Concepts

  • EvalTask — a prompt plus verify_commands and an optional timeout_secs. category is Capability or Regression.
  • RunOutcome — Pass, Fail, or Error for a single task attempt.
  • EvalMetric — combinatorial pass_at_k, independent reliability pass_power_k ((passed / attempts)^k), and pass_all_k, averaged per task. It also carries the selected k and raw passed_runs / total_runs.
  • EvalExecutor — the trait boundary. Implementors own "run this task" semantics (drive the agent, apply environment probes, grade the result). run_suite schedules attempts with a default concurrency of two; run_suite_with_options makes the bound and metric k explicit.

HarnessTraceSummary provides aggregate-only trace facts: turns, bounded tool/error counts, latency, output byte totals, repetition, and token/cache usage. Raw prompts, arguments, file contents, and output text are not retained. Use analyze_jsonl_file for buffered file analysis or analyze_jsonl_reader when the caller already owns a stream.

Usage

use vtcode_eval::{EvalExecutor, run_suite, EvalSuite};

// Implement EvalExecutor to drive your agent + grade outcomes, then:
let report = run_suite(&my_executor, &suite).await?;
println!("{}", report.to_markdown());

Trace analysis can be kept out of the agent hot path and run against a persisted JSONL session:

use vtcode_eval::analyze_jsonl_file;

let summary = analyze_jsonl_file("session.jsonl")?;
println!("{} tool calls, {} input tokens", summary.tool_calls, summary.token_usage.input_tokens);

The analyzer recognizes both DeepSeek-style records and serialized vtcode-exec-events::ThreadEvent shapes. Thread-level aggregate usage is used only when per-turn usage is absent, preventing double counting.

Notes

  • run_suite performs no file I/O or trust checks; the caller owns configuration, while the scheduler rejects attempts == 0 and invalid metric k values.
  • Reports sum known per-attempt costs and count unknown-pricing attempts separately; unknown cost is never silently treated as a known zero.
  • SuiteReport carries cost-efficiency aggregates — cost_per_solve, mean_cost_per_attempt, mean_tokens_per_attempt, mean_turns_per_attempt. cost_per_solve is None when any attempt is unpriced or nothing passed; the means cover priced/token/tracked attempts only. The markdown renderer prints them as an Efficiency: cost/solve … · per attempt · tokens · turns line.
  • Environment verification (EnvironmentProbe) is a separate concern from outcome grading — executor implementations decide whether and how to apply probes before returning a RunOutcome.