pub struct RunOutcome {Show 20 fields
pub text: String,
pub stop_reason: StopReason,
pub usage: Usage,
pub turns: u32,
pub refusal: Option<Refusal>,
pub exhausted: bool,
pub tool_calls: Vec<ToolCallTrace>,
pub malformed_tool_args: u32,
pub blocked_sends: u32,
pub taint: Taint,
pub homeostat: Option<Homeostat>,
pub stop_cause: StopCause,
pub cost_usd: Option<f64>,
pub ended_on_failed_call: bool,
pub compactions: u32,
pub context_overflows: u32,
pub boredom_notices: u32,
pub step_escalations_attempted: u32,
pub step_escalations_revised: u32,
pub usage_complete: bool,
}Fields§
§text: StringText of the final assistant turn.
stop_reason: StopReason§usage: Usage§turns: u32§refusal: Option<Refusal>§exhausted: boolTrue when the loop stopped because it hit max_turns, not because the
model was finished. The answer is probably incomplete.
tool_calls: Vec<ToolCallTrace>Every tool call attempted, in order.
malformed_tool_args: u32Calls whose arguments did not parse as JSON.
blocked_sends: u32Outbound calls refused because the trifecta was armed.
taint: TaintTaint state when the run ended.
homeostat: Option<Homeostat>Conditions the run happened under, when the caller asked for them.
stop_cause: StopCause§cost_usd: Option<f64>Cost of this run, when the provider has prices configured.
ended_on_failed_call: boolThe model said it was finished, and the last thing it did was fail.
The silent-failure shape: an agent that stops on its own after a failed call may have understood the failure and said so, or may be reporting success over it. Measured elsewhere, 75.8% of self-assessing AppWorld runs are false successes and no LLM-judge configuration exceeds AUROC 0.65 at catching one — while this signal is free, deterministic, and visible nowhere in the answer text.
Deliberately an observation rather than a verdict, which is why it is
named for what it saw. “Read this file” answered with “that file does
not exist” is a correct run that ends on a failed call, so this is not
an error condition; it is a flag a case or a human can gate on, and a
false positive costs one read. Only Completed runs can set it: a run
the harness cut short already says so through stop_cause and
exhausted.
The last call only. One failure among successes is ordinary recovery — what this names is a run whose final act failed and which then declared itself done.
compactions: u32How many times the transcript was summarised to keep it sendable.
Reported because compaction is lossy: an answer produced after four compactions is a different claim about the harness than the same answer produced without any, and only one of them tests that summaries carry the task forward.
context_overflows: u32How many times a prompt was refused as too large.
Named for the observation, not the response. An earlier spelling
counted recoveries, which left the one overflow that is never recovered
— the forced final-answer turn, whose failure is swallowed so the run
can still return its text — recorded as Some(0): sensor present, saw
nothing. The question this field exists to answer is whether the
threshold failed, and whether the harness got out of it afterwards is a
separate fact.
Distinct from compactions, and the distinction is the whole reason
this exists. compactions counts summaries, so an overflow the
recovery answered with eviction and thinning alone — which is the
common shape, because those cost no request — incremented nothing and
was invisible in every store. The harness caught a 400, rebuilt the
transcript and retried, and no counter anywhere said so.
What it measures is the reactive threshold failing: compact_at is
checked between turns against the previous prompt’s size, so a turn’s
parallel tool results can take the next request over the window from
under the threshold. Every recovery is one instance of that, and the
count is the baseline any change claiming to predict the overflow has
to be measured against.
A retry that overflows again propagates and ends the run, so it leaves no outcome to be recorded on — the count on a row that exists is always of overflows the run survived.
boredom_notices: u32Times this run was told an approach had stopped teaching it anything
(docs/GOAL-SYSTEM-DESIGN.md §9.1).
Here so the mechanism is falsifiable. Every threshold in boredom.rs
is a number chosen from argument rather than from measurement, and a
detector nobody can count fires either constantly or never with no way
to tell which — the silent failure this project keeps naming. The
notice is in the transcript verbatim, so this could in principle be
recovered by matching prose; that is what is_context_overflow has to
do because no backend gives it a code, and it is not something to
choose when the count is right here.
step_escalations_attempted: u32How many step-escalation candidates (docs/GOAL-SYSTEM-DESIGN.md
§5.5) actually spent a quarantined call this run — todo may flag
more, but MAX_STEP_ESCALATIONS_PER_RUN and stopping_now both
silently drop candidates without spending anything, so this counts
what happened, not what was offered.
The feature’s own off-by-default posture is explicitly pending a
measurement the pre-filter’s thresholds have never had (span ≥3× the
mean, floor of 6 calls — argued, not measured). Without a counter
recorded per run, that measurement can never be taken from the store:
mecha sessions health cannot say whether the mechanism ever fired,
and candidate.rs’s gate has no metric to move. boredom_notices
just above is the same argument already accepted for a sibling
mechanism.
step_escalations_revised: u32Of those, how many came back revise_plan — the run-level shape of
the same “not every fired check was right” question the appraiser’s
sign/agency split asks elsewhere.
usage_complete: boolFalse when usage is a lower bound rather than a measurement.
A run cancelled mid-stream keeps the input tokens, which arrive in the first frame, but not the output tokens of the cut turn, which arrive in a frame that never comes. Reporting the shortfall as zero would be a quiet lie in the same field a budget reads; saying the number is partial costs one bool.
Trait Implementations§
Source§impl Clone for RunOutcome
impl Clone for RunOutcome
Source§fn clone(&self) -> RunOutcome
fn clone(&self) -> RunOutcome
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read more