Expand description
Structured evaluator with scoring dimensions and hard thresholds.
Following the long-running harness pattern from Anthropic’s research: “When using a judge, we cannot ask abstractly whether this is ‘good.’ We need to break ‘good’ into multiple checkable dimensions.”
The evaluator scores a sprint across multiple dimensions, each with a hard threshold. If any dimension falls below its threshold, the sprint fails and the generator must revise based on concrete feedback. This prevents the common failure mode where an agent sees a button render and declares the feature complete.
Key principles:
- Evaluate outcomes, not claims (check actual test results, not agent assertions)
- Every dimension has a hard threshold (below = sprint fails)
- Feedback is specific and actionable (not “looks good”)
Structs§
- Dimension
Score - Score for a single dimension in an evaluation result.
- Evaluation
Result - The complete evaluation result for a sprint.
- Evaluation
Rubric - The rubric that defines how a sprint is evaluated.
- Scoring
Dimension - A scoring dimension with a hard threshold.
Functions§
- default_
code_ rubric - Default rubric for code tasks.
- default_
code_ rubric_ with_ hypothesis_ revision - Default code rubric plus the hypothesis-revision dimension.
- evaluation_
to_ markdown - Render an EvaluationResult as a markdown report suitable for harness artifacts.
- hypothesis_
revision_ dimension - Scoring dimension for the hypothesis loop (additive;
default_code_rubricis unchanged). - score_
hypothesis_ revision - Score hypothesis-revision behavior from harness-observed counts.