pub const REFUSED: &str = "this action was not permitted";Expand description
What a model may be told about a refusal.
§A denial reason is an oracle
Every message in PolicyError is written for an operator reading a
journal, and each one is precise on purpose: which principal, which sink,
what sensitivity, which ceiling. That precision is exactly what makes it
unsafe to hand back to a model.
An agent loop that feeds the refusal into its next prompt turns the policy
into a queryable service. Injected content steering the agent can probe it:
vary the request, watch which variants come back refused, and read the
boundary off the answers. EgressCeiling is the sharpest case — it reports
the sensitivity of the data and the sink’s ceiling, so a few probes
classify data the run was never allowed to reveal, without any of it ever
crossing the boundary.
So the split is deliberate: the journal keeps everything, the model is told one uniform sentence. An auditor needs to know why; the thing that might be attacking the policy must not learn anything it can differentiate.
This does not remove the denied/allowed bit itself. Nothing can, short of
fabricating success. What bounds that channel is
Budget::max_denials, which counts the
refusals a model is shown — these — as well as the engine denials it is not:
a run that keeps being refused is probing, and it is stopped.