Expand description
Agent-run recorder and baseline harness (SCC-002, docs/TEST_PLAN.md §9).
Runs each ground-truth task through an external agent command and records
outcome metrics. When the agent emits a JSON event stream (codex exec --json emits JSONL: {type:item.completed, item:{type:command_execution| mcp_tool_call|...}}), the recorder additionally extracts tool-level
exploration metrics: files opened, search/read tool calls, wrong-first
locations opened before the first ground-truth file, and the wall time
until the first ground-truth file is touched. When no JSON events are
present the recorder falls back to the portable layer: wall time, exit
status, output size, and per-task pass/fail localization.
Structs§
- Agent
Bench Summary - Agent
Gate Result - A-vs-E agent-behavior gate result: does the atlas variant (E) reduce exploration vs the baseline (A)?
- Agent
Task Result - Variant
Task - One ground-truth task definition used by the variant runners
(
run_variant_tasks). Thefilesare the localization ground truth (the benchagent protocol compares tool events against them);plan_keysare the first-plan correctness keys (files + symbols).
Enums§
- Surface
Ablation - Harness-level surface ranking modes for the PPR ablation matrix. Every
mode runs the PRODUCTION surface pipeline (
render_ablation_surface→build_surface_staged) with exactly one stage toggled — never a harness reimplementation of ranking.lexicalswitches every stage off (lexical-only);global-pprdisables task PPR;task-ppris the full task pipeline;ppr-mmr/ppr-quotas/ppr-optimizerdisable the MMR/quotas/optimizer tail stages respectively. The ablation rows are the production rows with one stage removed. See benchmarks/external/README.md.
Functions§
- evaluate_
gate - Evaluate the gate clauses from two summaries (A = baseline, E = atlas variant). The gate FAILS when the atlas variant does not reduce exploration: requires E.search_tool_calls < A.search_tool_calls AND E.files_opened <= A.files_opened + 1 AND E.first_correct_ms <= A.first_correct_ms (means).
- metrics_
from_ jsonl - Score an in-process JSONL tool stream the same way [
run_task] scores an agent. Event index is used asfirst_correct_msso deterministic pack-consumers have a stable ordinal (not wall-clock noise). - paid_
opt_ in_ allowed - Paid-model safety interlock: paid benchmark entrypoints refuse execution
by default and spend no model quota unless the operator opted in with
SCC_ALLOW_PAID_BENCHMARKS=1. The pure predicate is unit-testable without touching the process environment. - parse_
surface_ paths - Parse repo-relative file paths from a rendered surface — the
PRODUCTION format: groups headed by an uppercased component line, a
blank line, then the group’s path line (
<path>or<path> [<sub>]), then blank-line-separated<kind> <name>entry blocks. Order preserved, deduped. This is the harness’s file-selection oracle for the scc-full structural section: selection comes from the rendered surface (task PPR + lexical), NEVER from ground-truthtask.files. - print_
agent_ gate - print_
agent_ summary - render_
ablation_ surface - Render one ablation-mode task surface for
goalunderbudget_tokens(the chars/4 estimate) through the PRODUCTION surface pipeline (build_surface_staged) with the mode’s stage toggle — the ablation rows are the production rows with one stage removed, never a harness reimplementation of ranking. A one-line mode label prefixes the production render so artifacts stay identifiable; the label’s token cost is subtracted from the budget (equal-token discipline: the final artifact never exceeds the cap). - require_
paid_ opt_ in - Read the process environment for the paid-benchmark opt-in.
- run_
agent_ benchmark scc bench agent --cmd "<command>"— the command receives the task goal via theSCC_GOALenv var and the repo path as its working directory (likeclaude -p "$SCC_GOAL"orcodex exec -- "$SCC_GOAL").- run_
agent_ gate - Run the agent-behavior release gate: score the baseline (A) command and
the atlas variant (E) command over the same corpus with the same
harness, then evaluate the exploration-reduction clauses (see
evaluate_gate).min_filesapplies to BOTH runs. - run_
variant_ benchmark - Wave-15 variant benchmark: same protocol as
run_agent_benchmarkover the fullbenchmarks/tasks.jsoncorpus, recording the variant name in the summary.cmdis a shell command; the variant context artifact is expected to be generated inside it (see benchmarks/run_agent_bench.sh). - run_
variant_ tasks - The shared variant runner: for each task, copy the fixture repo, index
it, ask
cmd_forfor the per-task shell command (this is where a variant generates its context artifact into the freshly indexed repo and returns the agent command that consumes it), then run the benchagent protocol. Aggregates the same metrics asbench agentplus first-plan accuracy and graph-query counts, and recordsvariant. - structural_
source_ for_ goal - Build the structural-source section for a goal with NO ground truth:
file selection comes from the task-personalized surface (task PPR +
lexical), and per-file structural units are rendered while the token
estimate fits
budget_tokens. When the running binary supports it, the productionscc context structural --task "<goal>" --budget NCLI performs the same selection and render in one shot.task.filesNEVER enters context construction — ground truth is scoring-only.