pub struct Observation {Show 17 fields
pub turn_id: TurnId,
pub error_code: Option<String>,
pub acts: Vec<ObservedAct>,
pub target_resolutions: Vec<ObservedResolution>,
pub understanding: Option<Understanding>,
pub commands: Vec<String>,
pub events: Vec<String>,
pub events_truncated: bool,
pub revisions: Vec<(CaseKey, u64)>,
pub states: Vec<(CaseKey, ObservedState)>,
pub interactions: Vec<ObservedInteraction>,
pub blocks: Vec<BlockKind>,
pub phase: Option<TurnPhase>,
pub provider_failures: usize,
pub discarded_answers: Vec<DiscardedAnswer>,
pub cards_created: usize,
pub answer: String,
}Expand description
Everything one execution of an item left behind.
Fields§
§turn_id: TurnIdThe turn that was run.
error_code: Option<String>The orchestrator’s error code, when the turn failed outright.
acts: Vec<ObservedAct>The acts the message was understood to ask for, in order. Empty for a turn with no text, whose acts come from the card it answers.
target_resolutions: Vec<ObservedResolution>How each of those acts’ targets resolved.
understanding: Option<Understanding>The whole understanding of the message, for scoring each task that made it.
commands: Vec<String>Command types journaled this turn, in admission order.
events: Vec<String>Event types committed this turn, in append order.
events_truncated: boolWhether events was cut short by a configured bound.
It is false for every observation collected with
collect, which pages the journal to the end. It can
only be true when a run set
ExecutionConfig::max_observed_events,
and when it is, every assertion that reads the event list fails: an
expectation checked against half a ledger is not a measurement, and a
forbidden event hiding in the half nobody read would otherwise report a
green safety row.
revisions: Vec<(CaseKey, u64)>The revision each seeded case ended the turn at.
states: Vec<(CaseKey, ObservedState)>What each seeded case held before the turn and after it.
Both sides, because the two questions a corpus asks about state are different and one of them has no answer without the before: «this field ends at Lisbon» is about the after, and «this field did not move» is about the pair.
interactions: Vec<ObservedInteraction>Cards known for the seeded cases and cards this turn created.
blocks: Vec<BlockKind>Response block kinds, in order.
phase: Option<TurnPhase>Phase the turn finished in.
provider_failures: usizeHow many provider attempts failed or fell back.
discarded_answers: Vec<DiscardedAnswer>Every answer a model produced this turn that the runtime threw away whole, in the order it discarded them.
§Why a measurement wants this
It is the only channel for a turn that did nothing and had nothing to
show for it. Every other field here reports an effect, and the failure
this answers — the assistant closing a turn without proposing anything
or asking anything — leaves no effect by definition, so an empty
commands and an empty events look exactly like a turn that correctly
had nothing to do. This list separates the two: a turn whose plan was
refused for citing words the user never wrote is a defect, and a turn
that quietly agreed there was nothing to do is not.
Deliberately not part of signature — see there.
cards_created: usizeHow many cards this turn created.
answer: StringThe model-authored text of the turn, which is all a judge ever sees.
Implementations§
Source§impl Observation
impl Observation
Sourcepub async fn with_conversation_cases(
self,
stores: &Stores,
workflows: &WorkflowRegistry,
account: &AccountId,
earlier: &[TurnId],
) -> Self
pub async fn with_conversation_cases( self, stores: &Stores, workflows: &WorkflowRegistry, account: &AccountId, earlier: &[TurnId], ) -> Self
Adds the cases every turn of the conversation wrote, read back now, beside the seeded ones. A case the conversation created was seeded as nothing.
Sourcepub async fn collect(
stores: &Stores,
workflows: &WorkflowRegistry,
account: &AccountId,
turn_id: TurnId,
cases: &[CaseSeed],
outcome: Result<&AssistantTurn, String>,
) -> Self
pub async fn collect( stores: &Stores, workflows: &WorkflowRegistry, account: &AccountId, turn_id: TurnId, cases: &[CaseSeed], outcome: Result<&AssistantTurn, String>, ) -> Self
Reads back everything one turn left in the stores, paging the event journal to the end.
outcome is the turn as the orchestrator returned it, or the stable
error code of the failure. A failed turn is still observed: what it
journaled and committed before it failed is exactly what a safety
assertion is about.
The ledger is read through the journal’s sequence cursor, one page after another, until the journal reports it has caught up. A case with nine thousand prior events is therefore observed exactly like a case with nine, which is the property a seeded corpus depends on: an item whose fixture carries a long history must still see the events its own turn committed, and must still see a forbidden one.
Sourcepub async fn collect_bounded(
stores: &Stores,
workflows: &WorkflowRegistry,
account: &AccountId,
turn_id: TurnId,
cases: &[CaseSeed],
outcome: Result<&AssistantTurn, String>,
max_events: Option<usize>,
) -> Self
pub async fn collect_bounded( stores: &Stores, workflows: &WorkflowRegistry, account: &AccountId, turn_id: TurnId, cases: &[CaseSeed], outcome: Result<&AssistantTurn, String>, max_events: Option<usize>, ) -> Self
collect, with an optional cap on how many of the
turn’s events are recorded.
max_events is None for the complete ledger, which is what every
reproducible run wants. A Some(limit) is a deliberate ceiling for a
corpus run against a real endpoint where one item could commit an
unbounded number of events, and it is never silent: reaching it sets
events_truncated, and
check turns that into a failure of every
expectation that reads the event list. A bound that made an assertion
pass would be worse than no assertion at all.
Sourcepub fn revision_of(&self, case: &CaseKey) -> Option<u64>
pub fn revision_of(&self, case: &CaseKey) -> Option<u64>
The revision a case ended the turn at, when it was observed.
Sourcepub fn interactions_of(&self, case: &CaseKey) -> Vec<&ObservedInteraction>
pub fn interactions_of(&self, case: &CaseKey) -> Vec<&ObservedInteraction>
The cards of one case, oldest first.
Sourcepub fn signature(&self) -> String
pub fn signature(&self) -> String
A canonical rendering of the deterministic facts of this run.
Two samples with the same signature behaved the same way; two with different signatures did not. It deliberately excludes the model’s wording, block identifiers and timestamps, because a run that phrases the same receipt differently is not a different behaviour — and counting it as one would make every item look flaky.
Sourcepub fn discard_codes(&self) -> Vec<&str>
pub fn discard_codes(&self) -> Vec<&str>
The stable codes of the answers this turn lost, in order and with repeats.
The reason strings are for a person reading one turn; these are what a run groups by.
Trait Implementations§
Source§impl Clone for Observation
impl Clone for Observation
Source§impl Debug for Observation
impl Debug for Observation
impl Eq for Observation
Source§impl PartialEq for Observation
impl PartialEq for Observation
impl StructuralPartialEq for Observation
Auto Trait Implementations§
impl Freeze for Observation
impl RefUnwindSafe for Observation
impl Send for Observation
impl Sync for Observation
impl Unpin for Observation
impl UnsafeUnpin for Observation
impl UnwindSafe for Observation
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
Source§impl<Q, K> Equivalent<K> for Q
impl<Q, K> Equivalent<K> for Q
Source§fn equivalent(&self, key: &K) -> bool
fn equivalent(&self, key: &K) -> bool
key and return true if they are equal.