#[non_exhaustive]pub struct ExecutionConfig {
pub samples_per_item: u32,
pub stop_after_failures: Option<u32>,
pub sample_concurrency: u32,
pub max_observed_events: Option<usize>,
pub record_answers: bool,
}Expand description
How the agent under test is exercised.
#[non_exhaustive]: this 0.1 grows a field whenever a run needs to say
something new, and each one would be source-breaking for anybody constructing
this with a struct literal — #[serde(default)] keeps a FILE readable and
does nothing for Rust. Build it from Default and the with_* methods,
which is what the next added field will not break.
Fields (Non-exhaustive)§
This struct is marked as non-exhaustive
Struct { .. } syntax; cannot be matched against without a wildcard ..; and struct update syntax will not work.samples_per_item: u32How many times each item runs end to end.
This is the only knob that measures model variance. Raising it makes
a flaky item visible; raising JudgingConfig::votes_per_sample never
will.
stop_after_failures: Option<u32>Stop the whole run after this many samples have failed their
deterministic assertions. None runs the corpus to the end, which is
what a nightly job wants; a pull-request gate may prefer to stop early.
sample_concurrency: u32How many samples may be in flight at once.
One — the default — runs the corpus strictly one sample at a time, which is what a reproducible in-memory run wants: the harness sees the samples in index order, and a scripted provider keyed on call order sees exactly the sequence it was written for.
Raising it is for a corpus against a real endpoint, where a hundred samples in series is hours of waiting on a network. What it changes is the execution order: samples start and finish interleaved, so a harness that shares anything between them — a counter, a queue of scripted answers, a rate limit — will see a different order every run. The report does not move: samples are still reported in index order, and the same set of results comes back whatever this is set to. Zero is read as one.
max_observed_events: Option<usize>How many of a turn’s events one observation may record before it stops and says so.
None — the default — reads the ledger to the end, paging the journal
by its sequence cursor. A Some(limit) is a deliberate ceiling for a
run where one item could commit an unbounded number of events, and it is
never silent: an observation that hit it is marked truncated, and every
assertion that reads the event list then fails loudly instead of passing
over the half of the ledger nobody read.
record_answers: boolKeep the turn’s own words on each sample of the report.
Off by default, and the default is the careful one: the reply is the only model-authored prose a run produces, it is the one field that can carry whatever a person typed, and a report is a file that gets attached to things. A measurement does not need it — every assertion reads storage, and the judge is handed the text directly whether this is on or off.
Turn it on to CURATE. Writing the expectations of an item means deciding what the right reply would have been, and that cannot be done from a signature: two runs whose effects are identical can differ entirely in whether the assistant asked the question the turn needed. Reading them is how a corpus of scenes carried over from another engine gets its assertions.
Trait Implementations§
Source§impl Clone for ExecutionConfig
impl Clone for ExecutionConfig
Source§impl Debug for ExecutionConfig
impl Debug for ExecutionConfig
Source§impl Default for ExecutionConfig
impl Default for ExecutionConfig
Source§impl<'de> Deserialize<'de> for ExecutionConfigwhere
ExecutionConfig: Default,
impl<'de> Deserialize<'de> for ExecutionConfigwhere
ExecutionConfig: Default,
Source§fn deserialize<__D>(__deserializer: __D) -> Result<Self, __D::Error>where
__D: Deserializer<'de>,
fn deserialize<__D>(__deserializer: __D) -> Result<Self, __D::Error>where
__D: Deserializer<'de>,
impl Eq for ExecutionConfig
Source§impl PartialEq for ExecutionConfig
impl PartialEq for ExecutionConfig
Source§impl Serialize for ExecutionConfig
impl Serialize for ExecutionConfig
impl StructuralPartialEq for ExecutionConfig
Auto Trait Implementations§
impl Freeze for ExecutionConfig
impl RefUnwindSafe for ExecutionConfig
impl Send for ExecutionConfig
impl Sync for ExecutionConfig
impl Unpin for ExecutionConfig
impl UnsafeUnpin for ExecutionConfig
impl UnwindSafe for ExecutionConfig
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> DeserializeOwned for Twhere
T: for<'de> Deserialize<'de>,
Source§impl<Q, K> Equivalent<K> for Q
impl<Q, K> Equivalent<K> for Q
Source§fn equivalent(&self, key: &K) -> bool
fn equivalent(&self, key: &K) -> bool
key and return true if they are equal.