Expand description
Context benchmark (docs/TEST_PLAN.md §84–87, M0/M4): score task packs
against the ground-truth corpus in benchmarks/tasks.json.
Metrics:
- recall: ground-truth entities surfaced in the pack / total ground truth
- precision: ground-truth entities / pack entities of ground-truth kinds
- localization: ground-truth files mentioned in the pack content
- hallucination: nonexistent entities that must NOT surface
- budget: pack under the configured token budget
Structs§
Functions§
- locate_
fixtures_ dir - Locate the fixtures directory: walk up from cwd; fall back to the build-time manifest path (dev tooling).
- print_
summary - run_
context_ benchmark - Run the full context benchmark against
benchmarks/tasks.json. - score_
task_ public