Expand description
Atlas recall benchmark v2 (Wave 8 §57): structured recall of independently documented ground truth against the startup System Atlas, with precision and token-density metrics and per-gap diagnosis.
Ground truth is organized into seven layers (the v2 ontology):
architecture (components/subsystems the agent must know at startup),
entrypoints (invokable surfaces), behavior (flows/lifecycles),
state_authority (state owners), contracts (HTTP/CLI/API contracts),
landmarks (symbols one zoom level deeper — informational), and
tests (informational). The quality gate is the equal-weighted mean of
the FIVE startup-required layers only (architecture, entrypoints,
behavior, state_authority, contracts) — landmarks/tests are excluded,
which is the anti-bloat guarantee: dumping implementation symbols into
the atlas can no longer inflate the score.
Scoring is STRUCTURED, not text-substring: each layer is matched against
the machine model (scc_context::atlas::build_atlas), not the rendered
text. Item/haystack normalization applies the documented aliases
(:: -> ., fn X -> X, ./p -> p) so e.g. Controller::run
matches a flow step rendered as Controller.run.
v3 metrics:
precision(startup_required_precision): per startup-required layer, |atlas entries in that layer that match a ground-truth item| / |atlas entries in that layer| — a too-much-architecture detector: an atlas bloated with facts the ground truth never describes scores low.RepoRecall.precisionis the equal-weighted mean over the five startup-required layers.F2per layer: (5 * P * R) / (4 * P + R) from that layer’s precision P and recall R (zero when P + R == 0); the gate still uses recall.density(architecture_density): matched startup-required items per 1000 atlas tokens;atlas_tokensper repo is reported too.
The v1/v2 --holdout protocol is now labelled development vs
validation (the “holdout” corpus has been inspected and tuned
against — calling it blind would be dishonest; the on-disk dirs
benchmarks/holdout stay as they are).
--blind scores the NEW frozen corpus (benchmarks/blind-test +
benchmarks/blind-test-ground-truth), never used by tuning, and prints
ONLY aggregates (overall, per-section means, the validation-vs-blind
generalization gap, precision, density) — no per-repo rows, no missed
keys, no filenames. blind-test failures are never shown to tuning agents.
--diagnose classifies every missed item by WHERE it disappeared
(PARSER/EXTRACTOR/WRITER/RESOLUTION/COMPILER/PROJECTION/ALIAS) via a
deterministic store->flows->components->text ladder, and prints a
per-kind histogram plus per-repo gap lines (the regeneration source for
benchmarks/results/ground-truth-gaps.md).
When benchmarks/corpus/ is absent (or empty), the harness falls back to
the golden fixtures/: ground truth is synthesized from
benchmarks/tasks.json, fixture copies are indexed in a temp dir (the
golden fixtures are never written into), and the same recall pipeline runs.
Modules§
- sha256
- Minimal pure-Rust SHA-256 (FIPS 180-4) for the blind-test manifest hash. Deterministic, dependency-free (scc-cli has no crypto dep), panic-free. Public only so the roundtrip unit test can exercise it directly.
Structs§
- Atlas
Recall Report - Blind
Comparison - Blind-test protocol comparison: validation-vs-blind generalization.
- Blind
Lock Entry - One pinned blind-test clone in
benchmarks/blind-lock.json: the upstream URL plus the exact commit the on-disk clone must sit at. - Blind
Manifest - Blind-test manifest (Wave 11 — GENERALIZATION II): a sha256 fingerprint
of the frozen blind set — the ground-truth answer keys
(
benchmarks/blind-test-ground-truth/**), the clone list (the committedbenchmarks/blind-test/README.mdmanifest plus the on-disk repo dirs — the git-ls-files equivalent for the gitignored clones), and the commit pins frombenchmarks/blind-lock.json(alock <name> <sha>line per repo, so the digest covers the pinned commits). Written into the blind results header;--blindverifies the hash matches the previous run before scoring and errors on mismatch, so a changed blind set can never silently re-score different keys. - Compare
Report - Wave-11 gate report over two saved holdout result files (JSON
HoldoutComparisons): the GE gate (--gate-ge MIN, default 0.0 — fails whenGE <= MIN; semantic waves must generalize) and the per-section validation regression guard (--guard-section-delta MAX, default 0.05 — fails when ANY startup-required section regresses by more than MAX between the two runs, in development or validation). - GapFinding
- Ground
Truth Doc - Ground-truth sections parsed from
benchmarks/ground-truth/<name>.md(one- <key string>bullet per item). The v2 ontology; legacy section names (components/flows/ownership) are accepted and normalized. - Holdout
Comparison - Dev-vs-holdout comparison for
scc bench atlas --holdout. - Repo
Recall
Enums§
- GapKind
- Gap-kind classification for a missed ground-truth item (
--diagnose): where the fact disappeared between source and the rendered atlas. - Holdout
Verdict - Overfit verdict over the dev-vs-holdout overall gap.
Constants§
- ATLAS_
GATE - Quality gate: overall mean recall must be >= this floor (Wave 8 §57). The floor is over the five startup-required layers ONLY.
- DEFAULT_
SECTION_ GUARD - Default per-section regression guard (
--guard-section-delta): any startup-required section dropping by more than this between two compared runs fails the Wave-11 guard. - HOLDOUT_
TOLERANCE - Holdout verdict tolerance: the validation corpus may lag the development corpus by up to this much (overall recall, absolute) before the run is called OVERFIT. The band absorbs corpus-difficulty, LOC-mix, and ground-truth-strictness differences; a lag beyond it means the development-tuned rules do not generalize to unseen repos.
- STARTUP_
SECTIONS - The five startup-required sections the regression guard watches.
Functions§
- blind_
manifest - Deterministic manifest text + sha256 over the blind set under
root: every file inbenchmarks/blind-test-ground-truth/**(path + content hash), the clone list (the committedbenchmarks/blind-test/README.mdcontent hash + the sorted top-level repo dir names — the git-ls-files equivalent for the gitignored clones), and alock <name> <sha>line per repo from the committedbenchmarks/blind-lock.json— the digest covers the pinned commits. Missing ground-truth dir is an error (the protocol requires it); a missing README is tolerated (the clone list then reduces to the repo dirs); a missing lock is an error (the commit pins are a protocol artifact). - compare_
runs - Compare two saved holdout result files (JSON
HoldoutComparisons, e.g.scc bench atlas --holdout --jsonoutput) and apply the Wave-11 gates.oldis the earlier run (the pre-wave baseline),newthe current one; deltas are new - old. - generalization_
efficiency - Generalization efficiency (GE): how much of the development improvement between two runs transferred to validation.
- holdout_
verdict - Verdict over the dev-vs-holdout overall gap (
dev,holdoutare the equal-weighted mean recalls of the five startup-required layers). - load_
holdout_ result - Load a saved holdout JSON result file into a
HoldoutComparison. - locate_
fixtures_ dir - Locate the fixtures directory: walk up from cwd; fall back to the workspace-relative path (dev tooling).
- parse_
ground_ truth - Parse a ground-truth markdown doc into per-section key strings.
- print_
blind_ report - Print the blind protocol: aggregates ONLY (no per-repo rows, no missed keys, no filenames) plus the validation-vs-blind generalization gap.
- print_
compare_ report - Print the Wave-11 compare report (deltas, GE, per-section guard).
- print_
holdout_ report - Print the development and validation reports side by side plus the gap summary.
- print_
report - run_
atlas_ bench - Top-level entry for
scc bench atlas: locate the workspace, resolve the corpus/ground-truth directories (or the fixtures fallback), and run. - run_
atlas_ blind - Run the blind protocol: verify the frozen blind clones are at their
pinned commits (
benchmarks/blind-lock.json), score the validation corpus (benchmarks/holdout) and the blind-test corpus (benchmarks/blind-test) with the same recall pipeline, keep ONLY aggregates (per-repo rows, missed keys, and filenames are stripped — blind-test failures are never shown to tuning agents), compute the validation-vs-blind generalization gap, writebenchmarks/results/blind-v1.txt, and return the comparison. - run_
atlas_ holdout - Run the holdout protocol: score the development corpus and the
validation corpus with the same recall pipeline, compute per-layer gaps,
write
benchmarks/results/holdout-v3.txt, and return the comparison. - run_
atlas_ recall - Run the recall benchmark over
repo_names(sorted for deterministic output). Repos whose corpus dir is missing, whose ground-truth doc is missing, or whose index/atlas fails are recorded withskipped_reason— this function never panics on missing dirs. - score_
repo - Index one repo in place and score its ground truth against the atlas.