pub struct ToolFeedback {
pub tool_success_rates: HashMap<String, f64>,
pub tool_dispatch_counts: HashMap<String, (u64, u64)>,
}Expand description
Historical tool success rates computed from the trajectory store.
Pass to rank_with_feedback() to bias scoring based on past outcomes.
Fields§
§tool_success_rates: HashMap<String, f64>Per-tool success rate (0.0–1.0). Tools not in the map are assumed 0.5 (no data).
tool_dispatch_counts: HashMap<String, (u64, u64)>Per-tool (succeeded, dispatched) counts backing each rate, when the
feedback came from ToolFeedback::dispatched_from_trajectories.
A rate is not self-describing: 1.0 from one observation and 1.0 from
five hundred are the same number and very different evidence. Consumers
that price risk on these rates need the sample size to say how much to
trust them, so it travels alongside. Empty for feedback built by
ToolFeedback::from_trajectories, which does not track it.
Implementations§
Source§impl ToolFeedback
impl ToolFeedback
Sourcepub fn dispatched_from_trajectories(trajectories: &[Trajectory]) -> ToolFeedback
pub fn dispatched_from_trajectories(trajectories: &[Trajectory]) -> ToolFeedback
Compute dispatch-conditional tool feedback — P(tool succeeds | the action was actually dispatched).
This differs from ToolFeedback::from_trajectories in what lands in
the denominator, and the difference is not cosmetic. That constructor
counts every event carrying a tool name, including action_rejected
(the validator or policy refused the action) and action_skipped (an
upstream failure meant it never ran). Neither is evidence about the
tool: nothing was dispatched, so the tool had no opportunity to fail.
For scoring that conflation is defensible — a tool whose actions keep
getting rejected genuinely is a worse bet, and that is the signal
rank_with_feedback wants. For Monte Carlo rollout it is a
double-count: car_verify::simulate_monte_carlo already models
rejection structurally, deriving it from the dependency cascade, so a
rate that has rejections baked in penalizes the tool twice and reports
a plan as more fragile than the evidence supports.
Rates are Laplace-smoothed — (succeeded + 1) / (dispatched + 2),
the posterior mean under a uniform prior. One success becomes 0.67
rather than a categorical 1.0, while 100/100 becomes 0.99; evidence
still dominates once there is any. This avoids the alternative of a
hard minimum-sample cliff, where a tool’s rate would lurch from the
0.5 default to 1.0 on crossing some arbitrary count. Raw counts are
preserved in Self::tool_dispatch_counts for callers that want to
smooth differently or report the underlying evidence.
Tools that were never dispatched in the window are absent from the map entirely rather than present at 0.5 — “no evidence” and “evidence of a coin flip” are different claims, and the consumer applies its own default.
Sourcepub fn from_trajectories(trajectories: &[Trajectory]) -> ToolFeedback
pub fn from_trajectories(trajectories: &[Trajectory]) -> ToolFeedback
Compute tool feedback from a trajectory store.
Sourcepub fn rate(&self, tool: &str) -> f64
pub fn rate(&self, tool: &str) -> f64
Get the success rate for a tool, defaulting to 0.5 (no data).
Sourcepub fn proposal_tool_confidence(&self, proposal: &ActionProposal) -> f64
pub fn proposal_tool_confidence(&self, proposal: &ActionProposal) -> f64
Average success rate across all tools in a proposal.
Trait Implementations§
Source§impl Clone for ToolFeedback
impl Clone for ToolFeedback
Source§fn clone(&self) -> ToolFeedback
fn clone(&self) -> ToolFeedback
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read moreSource§impl Debug for ToolFeedback
impl Debug for ToolFeedback
Source§impl Default for ToolFeedback
impl Default for ToolFeedback
Source§fn default() -> ToolFeedback
fn default() -> ToolFeedback
Auto Trait Implementations§
impl Freeze for ToolFeedback
impl RefUnwindSafe for ToolFeedback
impl Send for ToolFeedback
impl Sync for ToolFeedback
impl Unpin for ToolFeedback
impl UnsafeUnpin for ToolFeedback
impl UnwindSafe for ToolFeedback
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> ErasedDestructor for Twhere
T: 'static,
Source§impl<T> Instrument for T
impl<T> Instrument for T
Source§fn instrument(self, span: Span) -> Instrumented<Self> ⓘ
fn instrument(self, span: Span) -> Instrumented<Self> ⓘ
Source§fn in_current_span(self) -> Instrumented<Self> ⓘ
fn in_current_span(self) -> Instrumented<Self> ⓘ
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more