Skip to main content

Expectation

Enum Expectation 

Source
pub enum Expectation {
    CalledTool {
        tool: String,
    },
    DidNotCallTool {
        tool: String,
    },
    CalledToolWith {
        tool: String,
        args: Value,
    },
    ToolCallCount {
        tool: Option<String>,
        min: Option<usize>,
        max: Option<usize>,
    },
    CalledToolsInOrder {
        tools: Vec<String>,
    },
    NoToolCalls,
    FinalTextContains {
        text: String,
        case_insensitive: bool,
    },
    FinalTextEquals {
        text: String,
    },
    FinalTextMatches {
        regex: String,
    },
    FinalNumberEquals {
        value: f64,
        tolerance: f64,
    },
    NoError,
}
Expand description

A single built-in assertion over a run’s RunArtifacts.

A case PASSES iff the run did not error AND every Expectation holds. Evaluate one against a run via Expectation::evaluate (or let BuiltinScorer + run_suite do it). Each variant produces a clear, report-ready label via Expectation::label.

Variants§

§

CalledTool

The agent called tool at least once (any args).

Fields

§tool: String

The tool/function name that must appear among the calls.

§

DidNotCallTool

The agent did NOT call tool at all.

Fields

§tool: String

The tool/function name that must be absent from the calls.

§

CalledToolWith

The agent called tool at least once with args that SUPERSET the given args (subset match — see the module docs). The canonical “called X with these parameters” assertion.

Fields

§tool: String

The tool/function name that must have been called.

§args: Value

The expected argument subset (objects recurse; other values match exactly).

§

ToolCallCount

The number of tool calls is within [min, max] (each optional). When tool is Some, only calls to that tool are counted; when None, ALL calls are counted.

Fields

§tool: Option<String>

Restrict the count to this tool; None counts every call.

§min: Option<usize>

Inclusive lower bound on the count (optional).

§max: Option<usize>

Inclusive upper bound on the count (optional).

§

CalledToolsInOrder

The named tools appear as a SUBSEQUENCE of the call order (in order, but not necessarily contiguous — other calls may be interleaved). Empty tools trivially holds.

Fields

§tools: Vec<String>

The tools that must appear in this relative order among the calls.

§

NoToolCalls

The agent made NO tool calls at all (a pure-reasoning / refusal assertion).

§

FinalTextContains

final_text contains text (optionally case-insensitively). Fails when there is no final text.

Fields

§text: String

The substring that must be present in the final reply.

§case_insensitive: bool

Compare case-insensitively when true (default false, an exact-case substring match).

§

FinalTextEquals

final_text equals text exactly (after trimming surrounding whitespace on both sides). Fails when there is no final text.

Fields

§text: String

The exact (trimmed) final reply expected.

§

FinalTextMatches

final_text matches the regex (anywhere, via regex::Regex::is_match). A malformed regex is a hard EvalError::Regex from Expectation::evaluate (NOT a silent failure). Fails when there is no final text.

Fields

§regex: String

The regular expression to match against the final reply.

§

FinalNumberEquals

The LAST number in final_text equals value within tolerance (see the module docs for the extraction rule). Fails when there is no final text or it contains no number.

Fields

§value: f64

The expected numeric answer.

§tolerance: f64

Allowed absolute difference (default 0.0 = exact match).

§

NoError

The run reported no error (RunArtifacts::error is None). The runner already fails a case on any run error, so this is mostly for an explicit, labeled “the run was clean” predicate.

Implementations§

Source§

impl Expectation

Source

pub fn evaluate( &self, artifacts: &RunArtifacts, ) -> Result<(String, bool), EvalError>

Evaluate this expectation against a run’s artifacts, returning (label, passed).

The label identifies WHICH predicate it is (for the report’s per-predicate diagnostics), matching the labeling style of a hand-rolled scorer. The ONLY fallible case is FinalTextMatches with a malformed regex, which returns EvalError::Regex; every other variant is infallible.

Source

pub fn label(&self) -> String

A short, human-readable label identifying this expectation, for the report’s per-predicate lines.

Trait Implementations§

Source§

impl Clone for Expectation

Source§

fn clone(&self) -> Expectation

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Debug for Expectation

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more
Source§

impl<'de> Deserialize<'de> for Expectation

Source§

fn deserialize<__D>(__deserializer: __D) -> Result<Self, __D::Error>
where __D: Deserializer<'de>,

Deserialize this value from the given Serde deserializer. Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> DeserializeOwned for T
where T: for<'de> Deserialize<'de>,

Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self>

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self>

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = Infallible

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, <T as TryFrom<U>>::Error>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self>
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self>

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more