Available on crate feature
evaluation only.Expand description
Evaluation framework: judge runs over production AI traffic, failure replay with repair hints, and the auto-improve loop.
Modules§
- error
- Typed error boundary for the
systemprompt-evaluationcrate. - extension
- Extension registration — wires the evaluation schemas (runs, cases, results, pairs, judge calls, rubrics) and their reconcile migrations into the extension framework.
- models
- Data model for evaluation runs, cases, results, rubrics, and sampling.
- repository
- Repositories over the
eval_*tables plus the sampling reader over theai_requeststrace owned bysystemprompt-ai. - services
- Evaluation services: sampling, judging, replay, and the auto-improve loop.
Structs§
- Auto
Improve Loop - Sample → judge → repair-hint → replay → re-judge, one pass.
- Canonical
Message - Canonical
Prompt - Provider-neutral reconstruction of an AI request.
- Dimension
Score - Eval
Case - Eval
Case Repository - Eval
Judge Call Repository - Eval
Repositories - Eval
Result - Eval
Result Repository - Eval
Rubric Repository - EvalRun
- Eval
RunRepository - Evaluation
Extension - Evaluation
Service - Judge
Service - Judge
Verdict - Structured output the judge model is constrained to produce.
- Loop
Limits - Loop
Report - NewCase
Params - NewResult
Params - NewRun
Params - Replay
Service - Rubric
- Rubric
Dimension - RunRequest
- Sample
Filter - Sampled
Request - A completed production request hydrated with everything the judge needs.
- Sampler
Service - Sampling
Repository