What DISCOVER optimizes for (docs/loop-reflection.md §5.1). Host config
like everything else here: it changes the scoring rule the proposer is
given, never the gates — every draft still has to survive GROUND, VERIFY,
the confidence floor and a human review with a BECAUSE.
The review-queue objective: “nothing to report” is a zero-penalty
answer and a wrong finding costs twice a right one. Right for a queue
a person triages — it keeps the queue clean at the price of drafts
the model was not sure enough about.
The learner objective: the agent has to improve from THIS pass, so
abstaining in the face of a recurring failure, repeated rejections or
a person’s instruction is penalized like a wrong lesson. Measured
need: under the review-queue rule a cheap model authored a lesson on
fewer than half of its passes over evidence that plainly held one.