Skip to main content

Module evaluator

Module evaluator 

Source
Expand description

Structured evaluator with scoring dimensions and hard thresholds.

Following the long-running harness pattern from Anthropic’s research: “When using a judge, we cannot ask abstractly whether this is ‘good.’ We need to break ‘good’ into multiple checkable dimensions.”

The evaluator scores a sprint across multiple dimensions, each with a hard threshold. If any dimension falls below its threshold, the sprint fails and the generator must revise based on concrete feedback. This prevents the common failure mode where an agent sees a button render and declares the feature complete.

Key principles:

  • Evaluate outcomes, not claims (check actual test results, not agent assertions)
  • Every dimension has a hard threshold (below = sprint fails)
  • Feedback is specific and actionable (not “looks good”)

Structs§

DimensionScore
Score for a single dimension in an evaluation result.
EvaluationResult
The complete evaluation result for a sprint.
EvaluationRubric
The rubric that defines how a sprint is evaluated.
ScoringDimension
A scoring dimension with a hard threshold.

Functions§

default_code_rubric
Default rubric for code tasks.
default_code_rubric_with_hypothesis_revision
Default code rubric plus the hypothesis-revision dimension.
evaluation_to_markdown
Render an EvaluationResult as a markdown report suitable for harness artifacts.
hypothesis_revision_dimension
Scoring dimension for the hypothesis loop (additive; default_code_rubric is unchanged).
score_hypothesis_revision
Score hypothesis-revision behavior from harness-observed counts.