Skip to main content

Module benchatlas

Module benchatlas 

Source
Expand description

Atlas recall benchmark v2 (Wave 8 §57): structured recall of independently documented ground truth against the startup System Atlas, with precision and token-density metrics and per-gap diagnosis.

Ground truth is organized into seven layers (the v2 ontology): architecture (components/subsystems the agent must know at startup), entrypoints (invokable surfaces), behavior (flows/lifecycles), state_authority (state owners), contracts (HTTP/CLI/API contracts), landmarks (symbols one zoom level deeper — informational), and tests (informational). The quality gate is the equal-weighted mean of the FIVE startup-required layers only (architecture, entrypoints, behavior, state_authority, contracts) — landmarks/tests are excluded, which is the anti-bloat guarantee: dumping implementation symbols into the atlas can no longer inflate the score.

Scoring is STRUCTURED, not text-substring: each layer is matched against the machine model (scc_context::atlas::build_atlas), not the rendered text. Item/haystack normalization applies the documented aliases (:: -> ., fn X -> X, ./p -> p) so e.g. Controller::run matches a flow step rendered as Controller.run.

v3 metrics:

  • precision (startup_required_precision): per startup-required layer, |atlas entries in that layer that match a ground-truth item| / |atlas entries in that layer| — a too-much-architecture detector: an atlas bloated with facts the ground truth never describes scores low. RepoRecall.precision is the equal-weighted mean over the five startup-required layers.
  • F2 per layer: (5 * P * R) / (4 * P + R) from that layer’s precision P and recall R (zero when P + R == 0); the gate still uses recall.
  • density (architecture_density): matched startup-required items per 1000 atlas tokens; atlas_tokens per repo is reported too.

The v1/v2 --holdout protocol is now labelled development vs validation (the “holdout” corpus has been inspected and tuned against — calling it blind would be dishonest; the on-disk dirs benchmarks/holdout stay as they are).

--blind scores the NEW frozen corpus (benchmarks/blind-test + benchmarks/blind-test-ground-truth), never used by tuning, and prints ONLY aggregates (overall, per-section means, the validation-vs-blind generalization gap, precision, density) — no per-repo rows, no missed keys, no filenames. blind-test failures are never shown to tuning agents.

--diagnose classifies every missed item by WHERE it disappeared (PARSER/EXTRACTOR/WRITER/RESOLUTION/COMPILER/PROJECTION/ALIAS) via a deterministic store->flows->components->text ladder, and prints a per-kind histogram plus per-repo gap lines (the regeneration source for benchmarks/results/ground-truth-gaps.md).

When benchmarks/corpus/ is absent (or empty), the harness falls back to the golden fixtures/: ground truth is synthesized from benchmarks/tasks.json, fixture copies are indexed in a temp dir (the golden fixtures are never written into), and the same recall pipeline runs.

Modules§

sha256
Minimal pure-Rust SHA-256 (FIPS 180-4) for the blind-test manifest hash. Deterministic, dependency-free (scc-cli has no crypto dep), panic-free. Public only so the roundtrip unit test can exercise it directly.

Structs§

AtlasRecallReport
BlindComparison
Blind-test protocol comparison: validation-vs-blind generalization.
BlindLockEntry
One pinned blind-test clone in benchmarks/blind-lock.json: the upstream URL plus the exact commit the on-disk clone must sit at.
BlindManifest
Blind-test manifest (Wave 11 — GENERALIZATION II): a sha256 fingerprint of the frozen blind set — the ground-truth answer keys (benchmarks/blind-test-ground-truth/**), the clone list (the committed benchmarks/blind-test/README.md manifest plus the on-disk repo dirs — the git-ls-files equivalent for the gitignored clones), and the commit pins from benchmarks/blind-lock.json (a lock <name> <sha> line per repo, so the digest covers the pinned commits). Written into the blind results header; --blind verifies the hash matches the previous run before scoring and errors on mismatch, so a changed blind set can never silently re-score different keys.
CompareReport
Wave-11 gate report over two saved holdout result files (JSON HoldoutComparisons): the GE gate (--gate-ge MIN, default 0.0 — fails when GE <= MIN; semantic waves must generalize) and the per-section validation regression guard (--guard-section-delta MAX, default 0.05 — fails when ANY startup-required section regresses by more than MAX between the two runs, in development or validation).
GapFinding
GroundTruthDoc
Ground-truth sections parsed from benchmarks/ground-truth/<name>.md (one - <key string> bullet per item). The v2 ontology; legacy section names (components/flows/ownership) are accepted and normalized.
HoldoutComparison
Dev-vs-holdout comparison for scc bench atlas --holdout.
RepoRecall

Enums§

GapKind
Gap-kind classification for a missed ground-truth item (--diagnose): where the fact disappeared between source and the rendered atlas.
HoldoutVerdict
Overfit verdict over the dev-vs-holdout overall gap.

Constants§

ATLAS_GATE
Quality gate: overall mean recall must be >= this floor (Wave 8 §57). The floor is over the five startup-required layers ONLY.
DEFAULT_SECTION_GUARD
Default per-section regression guard (--guard-section-delta): any startup-required section dropping by more than this between two compared runs fails the Wave-11 guard.
HOLDOUT_TOLERANCE
Holdout verdict tolerance: the validation corpus may lag the development corpus by up to this much (overall recall, absolute) before the run is called OVERFIT. The band absorbs corpus-difficulty, LOC-mix, and ground-truth-strictness differences; a lag beyond it means the development-tuned rules do not generalize to unseen repos.
STARTUP_SECTIONS
The five startup-required sections the regression guard watches.

Functions§

blind_manifest
Deterministic manifest text + sha256 over the blind set under root: every file in benchmarks/blind-test-ground-truth/** (path + content hash), the clone list (the committed benchmarks/blind-test/README.md content hash + the sorted top-level repo dir names — the git-ls-files equivalent for the gitignored clones), and a lock <name> <sha> line per repo from the committed benchmarks/blind-lock.json — the digest covers the pinned commits. Missing ground-truth dir is an error (the protocol requires it); a missing README is tolerated (the clone list then reduces to the repo dirs); a missing lock is an error (the commit pins are a protocol artifact).
compare_runs
Compare two saved holdout result files (JSON HoldoutComparisons, e.g. scc bench atlas --holdout --json output) and apply the Wave-11 gates. old is the earlier run (the pre-wave baseline), new the current one; deltas are new - old.
generalization_efficiency
Generalization efficiency (GE): how much of the development improvement between two runs transferred to validation.
holdout_verdict
Verdict over the dev-vs-holdout overall gap (dev, holdout are the equal-weighted mean recalls of the five startup-required layers).
load_holdout_result
Load a saved holdout JSON result file into a HoldoutComparison.
locate_fixtures_dir
Locate the fixtures directory: walk up from cwd; fall back to the workspace-relative path (dev tooling).
parse_ground_truth
Parse a ground-truth markdown doc into per-section key strings.
print_blind_report
Print the blind protocol: aggregates ONLY (no per-repo rows, no missed keys, no filenames) plus the validation-vs-blind generalization gap.
print_compare_report
Print the Wave-11 compare report (deltas, GE, per-section guard).
print_holdout_report
Print the development and validation reports side by side plus the gap summary.
print_report
run_atlas_bench
Top-level entry for scc bench atlas: locate the workspace, resolve the corpus/ground-truth directories (or the fixtures fallback), and run.
run_atlas_blind
Run the blind protocol: verify the frozen blind clones are at their pinned commits (benchmarks/blind-lock.json), score the validation corpus (benchmarks/holdout) and the blind-test corpus (benchmarks/blind-test) with the same recall pipeline, keep ONLY aggregates (per-repo rows, missed keys, and filenames are stripped — blind-test failures are never shown to tuning agents), compute the validation-vs-blind generalization gap, write benchmarks/results/blind-v1.txt, and return the comparison.
run_atlas_holdout
Run the holdout protocol: score the development corpus and the validation corpus with the same recall pipeline, compute per-layer gaps, write benchmarks/results/holdout-v3.txt, and return the comparison.
run_atlas_recall
Run the recall benchmark over repo_names (sorted for deterministic output). Repos whose corpus dir is missing, whose ground-truth doc is missing, or whose index/atlas fails are recorded with skipped_reason — this function never panics on missing dirs.
score_repo
Index one repo in place and score its ground truth against the atlas.