Skip to main content

RunOutcome

Struct RunOutcome 

Source
pub struct RunOutcome {
Show 20 fields pub text: String, pub stop_reason: StopReason, pub usage: Usage, pub turns: u32, pub refusal: Option<Refusal>, pub exhausted: bool, pub tool_calls: Vec<ToolCallTrace>, pub malformed_tool_args: u32, pub blocked_sends: u32, pub taint: Taint, pub homeostat: Option<Homeostat>, pub stop_cause: StopCause, pub cost_usd: Option<f64>, pub ended_on_failed_call: bool, pub compactions: u32, pub context_overflows: u32, pub boredom_notices: u32, pub step_escalations_attempted: u32, pub step_escalations_revised: u32, pub usage_complete: bool,
}

Fields§

§text: String

Text of the final assistant turn.

§stop_reason: StopReason§usage: Usage§turns: u32§refusal: Option<Refusal>§exhausted: bool

True when the loop stopped because it hit max_turns, not because the model was finished. The answer is probably incomplete.

§tool_calls: Vec<ToolCallTrace>

Every tool call attempted, in order.

§malformed_tool_args: u32

Calls whose arguments did not parse as JSON.

§blocked_sends: u32

Outbound calls refused because the trifecta was armed.

§taint: Taint

Taint state when the run ended.

§homeostat: Option<Homeostat>

Conditions the run happened under, when the caller asked for them.

§stop_cause: StopCause§cost_usd: Option<f64>

Cost of this run, when the provider has prices configured.

§ended_on_failed_call: bool

The model said it was finished, and the last thing it did was fail.

The silent-failure shape: an agent that stops on its own after a failed call may have understood the failure and said so, or may be reporting success over it. Measured elsewhere, 75.8% of self-assessing AppWorld runs are false successes and no LLM-judge configuration exceeds AUROC 0.65 at catching one — while this signal is free, deterministic, and visible nowhere in the answer text.

Deliberately an observation rather than a verdict, which is why it is named for what it saw. “Read this file” answered with “that file does not exist” is a correct run that ends on a failed call, so this is not an error condition; it is a flag a case or a human can gate on, and a false positive costs one read. Only Completed runs can set it: a run the harness cut short already says so through stop_cause and exhausted.

The last call only. One failure among successes is ordinary recovery — what this names is a run whose final act failed and which then declared itself done.

§compactions: u32

How many times the transcript was summarised to keep it sendable.

Reported because compaction is lossy: an answer produced after four compactions is a different claim about the harness than the same answer produced without any, and only one of them tests that summaries carry the task forward.

§context_overflows: u32

How many times a prompt was refused as too large.

Named for the observation, not the response. An earlier spelling counted recoveries, which left the one overflow that is never recovered — the forced final-answer turn, whose failure is swallowed so the run can still return its text — recorded as Some(0): sensor present, saw nothing. The question this field exists to answer is whether the threshold failed, and whether the harness got out of it afterwards is a separate fact.

Distinct from compactions, and the distinction is the whole reason this exists. compactions counts summaries, so an overflow the recovery answered with eviction and thinning alone — which is the common shape, because those cost no request — incremented nothing and was invisible in every store. The harness caught a 400, rebuilt the transcript and retried, and no counter anywhere said so.

What it measures is the reactive threshold failing: compact_at is checked between turns against the previous prompt’s size, so a turn’s parallel tool results can take the next request over the window from under the threshold. Every recovery is one instance of that, and the count is the baseline any change claiming to predict the overflow has to be measured against.

A retry that overflows again propagates and ends the run, so it leaves no outcome to be recorded on — the count on a row that exists is always of overflows the run survived.

§boredom_notices: u32

Times this run was told an approach had stopped teaching it anything (docs/GOAL-SYSTEM-DESIGN.md §9.1).

Here so the mechanism is falsifiable. Every threshold in boredom.rs is a number chosen from argument rather than from measurement, and a detector nobody can count fires either constantly or never with no way to tell which — the silent failure this project keeps naming. The notice is in the transcript verbatim, so this could in principle be recovered by matching prose; that is what is_context_overflow has to do because no backend gives it a code, and it is not something to choose when the count is right here.

§step_escalations_attempted: u32

How many step-escalation candidates (docs/GOAL-SYSTEM-DESIGN.md §5.5) actually spent a quarantined call this run — todo may flag more, but MAX_STEP_ESCALATIONS_PER_RUN and stopping_now both silently drop candidates without spending anything, so this counts what happened, not what was offered.

The feature’s own off-by-default posture is explicitly pending a measurement the pre-filter’s thresholds have never had (span ≥3× the mean, floor of 6 calls — argued, not measured). Without a counter recorded per run, that measurement can never be taken from the store: mecha sessions health cannot say whether the mechanism ever fired, and candidate.rs’s gate has no metric to move. boredom_notices just above is the same argument already accepted for a sibling mechanism.

§step_escalations_revised: u32

Of those, how many came back revise_plan — the run-level shape of the same “not every fired check was right” question the appraiser’s sign/agency split asks elsewhere.

§usage_complete: bool

False when usage is a lower bound rather than a measurement.

A run cancelled mid-stream keeps the input tokens, which arrive in the first frame, but not the output tokens of the cut turn, which arrive in a frame that never comes. Reporting the shortfall as zero would be a quiet lie in the same field a budget reads; saying the number is partial costs one bool.

Trait Implementations§

Source§

impl Clone for RunOutcome

Source§

fn clone(&self) -> RunOutcome

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Debug for RunOutcome

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more
Source§

impl From<&RunOutcome> for RunStats

Source§

fn from(o: &RunOutcome) -> Self

Converts to this type from the input type.

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self>

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self>

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> PolicyExt for T
where T: ?Sized,

Source§

fn and<P, B, E>(self, other: P) -> And<T, P>
where T: Sized + Policy<B, E>, P: Policy<B, E>,

Create a new Policy that returns Action::Follow only if self and other return Action::Follow. Read more
Source§

fn or<P, B, E>(self, other: P) -> Or<T, P>
where T: Sized + Policy<B, E>, P: Policy<B, E>,

Create a new Policy that returns Action::Follow if either self or other returns Action::Follow. Read more
Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, <T as TryFrom<U>>::Error>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self>
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self>

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more