Skip to main content

Module benchagent

Module benchagent 

Source
Expand description

Agent-run recorder and baseline harness (SCC-002, docs/TEST_PLAN.md §9).

Runs each ground-truth task through an external agent command and records outcome metrics. When the agent emits a JSON event stream (codex exec --json emits JSONL: {type:item.completed, item:{type:command_execution| mcp_tool_call|...}}), the recorder additionally extracts tool-level exploration metrics: files opened, search/read tool calls, wrong-first locations opened before the first ground-truth file, and the wall time until the first ground-truth file is touched. When no JSON events are present the recorder falls back to the portable layer: wall time, exit status, output size, and per-task pass/fail localization.

Structs§

AgentBenchSummary
AgentGateResult
A-vs-E agent-behavior gate result: does the atlas variant (E) reduce exploration vs the baseline (A)?
AgentTaskResult
VariantTask
One ground-truth task definition used by the variant runners (run_variant_tasks). The files are the localization ground truth (the benchagent protocol compares tool events against them); plan_keys are the first-plan correctness keys (files + symbols).

Enums§

SurfaceAblation
Harness-level surface ranking modes for the PPR ablation matrix. Every mode runs the PRODUCTION surface pipeline (render_ablation_surface → build_surface_staged) with exactly one stage toggled — never a harness reimplementation of ranking. lexical switches every stage off (lexical-only); global-ppr disables task PPR; task-ppr is the full task pipeline; ppr-mmr/ppr-quotas/ppr-optimizer disable the MMR/quotas/optimizer tail stages respectively. The ablation rows are the production rows with one stage removed. See benchmarks/external/README.md.

Functions§

evaluate_gate
Evaluate the gate clauses from two summaries (A = baseline, E = atlas variant). The gate FAILS when the atlas variant does not reduce exploration: requires E.search_tool_calls < A.search_tool_calls AND E.files_opened <= A.files_opened + 1 AND E.first_correct_ms <= A.first_correct_ms (means).
metrics_from_jsonl
Score an in-process JSONL tool stream the same way [run_task] scores an agent. Event index is used as first_correct_ms so deterministic pack-consumers have a stable ordinal (not wall-clock noise).
paid_opt_in_allowed
Paid-model safety interlock: paid benchmark entrypoints refuse execution by default and spend no model quota unless the operator opted in with SCC_ALLOW_PAID_BENCHMARKS=1. The pure predicate is unit-testable without touching the process environment.
parse_surface_paths
Parse repo-relative file paths from a rendered surface — the PRODUCTION format: groups headed by an uppercased component line, a blank line, then the group’s path line (<path> or <path> [<sub>]), then blank-line-separated <kind> <name> entry blocks. Order preserved, deduped. This is the harness’s file-selection oracle for the scc-full structural section: selection comes from the rendered surface (task PPR + lexical), NEVER from ground-truth task.files.
print_agent_gate
print_agent_summary
render_ablation_surface
Render one ablation-mode task surface for goal under budget_tokens (the chars/4 estimate) through the PRODUCTION surface pipeline (build_surface_staged) with the mode’s stage toggle — the ablation rows are the production rows with one stage removed, never a harness reimplementation of ranking. A one-line mode label prefixes the production render so artifacts stay identifiable; the label’s token cost is subtracted from the budget (equal-token discipline: the final artifact never exceeds the cap).
require_paid_opt_in
Read the process environment for the paid-benchmark opt-in.
run_agent_benchmark
scc bench agent --cmd "<command>" — the command receives the task goal via the SCC_GOAL env var and the repo path as its working directory (like claude -p "$SCC_GOAL" or codex exec -- "$SCC_GOAL").
run_agent_gate
Run the agent-behavior release gate: score the baseline (A) command and the atlas variant (E) command over the same corpus with the same harness, then evaluate the exploration-reduction clauses (see evaluate_gate). min_files applies to BOTH runs.
run_variant_benchmark
Wave-15 variant benchmark: same protocol as run_agent_benchmark over the full benchmarks/tasks.json corpus, recording the variant name in the summary. cmd is a shell command; the variant context artifact is expected to be generated inside it (see benchmarks/run_agent_bench.sh).
run_variant_tasks
The shared variant runner: for each task, copy the fixture repo, index it, ask cmd_for for the per-task shell command (this is where a variant generates its context artifact into the freshly indexed repo and returns the agent command that consumes it), then run the benchagent protocol. Aggregates the same metrics as bench agent plus first-plan accuracy and graph-query counts, and records variant.
structural_source_for_goal
Build the structural-source section for a goal with NO ground truth: file selection comes from the task-personalized surface (task PPR + lexical), and per-file structural units are rendered while the token estimate fits budget_tokens. When the running binary supports it, the production scc context structural --task "<goal>" --budget N CLI performs the same selection and render in one shot. task.files NEVER enters context construction — ground truth is scoring-only.