pub struct Usage {Show 14 fields
pub prompt_tokens: usize,
pub completion_tokens: usize,
pub total_tokens: usize,
pub completion_tokens_details: Option<CompletionTokensDetails>,
pub prompt_per_second: Option<f64>,
pub predicted_per_second: Option<f64>,
pub prompt_eval_duration_ms: Option<f64>,
pub generation_duration_ms: Option<f64>,
pub time_to_first_token_ms: Option<f64>,
pub cached_tokens: Option<usize>,
pub acceptance_length: Option<f64>,
pub draft_tokens: Option<usize>,
pub accepted_draft_tokens: Option<usize>,
pub draft_accept_rate_per_position: Option<Vec<f64>>,
}Fields§
§prompt_tokens: usize§completion_tokens: usize§total_tokens: usize§completion_tokens_details: Option<CompletionTokensDetails>OpenAI’s nested completion breakdown. Absent unless the reasoning split actually ran, because a zero here is a claim about the model and not about this server.
prompt_per_second: Option<f64>Prefill throughput (prompt tokens / prefill seconds), when timed.
predicted_per_second: Option<f64>Decode throughput (completion tokens / decode seconds), when timed.
prompt_eval_duration_ms: Option<f64>Wall time spent processing the prompt, in milliseconds.
generation_duration_ms: Option<f64>Wall time spent in the decode loop, in milliseconds. Kept
separate from prompt_eval_duration_ms on purpose (see the
module docs).
time_to_first_token_ms: Option<f64>Time to first token: from the start of prefill to the moment the
first token was produced. None when no token was produced at
all (an immediate EOS), because a zero there would read as an
instantaneous response.
cached_tokens: Option<usize>Prompt tokens served from the KV prefix cache instead of being
recomputed. Some(0) means “the cache was consulted and missed”;
None means “no prefix cache is configured” – a distinction the
UI needs to decide whether to show the row at all.
acceptance_length: Option<f64>Completion tokens per verification step when speculative
decoding ran: the published acceptance length. None means
speculation did not run, which is not the same as an acceptance
length of 1.0 (speculation ran and never helped).
draft_tokens: Option<usize>Draft tokens the target actually evaluated. Positions after a rejection are not counted, so the ratio below tracks the drafter’s accuracy rather than its block size.
accepted_draft_tokens: Option<usize>Draft tokens accepted.
draft_accept_rate_per_position: Option<Vec<f64>>Accept rate at each position within the draft block, each conditional on that position having been reached.
Reported alongside the mean and not folded into it: a drafter that is right at position 0 and useless by position 7 has the same mean as one that is uniformly mediocre, and the two want opposite block sizes. Suffix decay is only visible per position.
Implementations§
Source§impl Usage
impl Usage
pub fn new(prompt_tokens: usize, completion_tokens: usize) -> Self
Sourcepub fn with_timings(self, prompt_secs: f64, predicted_secs: f64) -> Self
pub fn with_timings(self, prompt_secs: f64, predicted_secs: f64) -> Self
Records the two phase durations, in seconds, and the rates they imply. A zero-length phase leaves the rate unset rather than dividing by zero into infinity.
Sourcepub fn with_ttft(self, secs: f64) -> Self
pub fn with_ttft(self, secs: f64) -> Self
Time-to-first-token, in seconds, measured from the start of prefill. Ignored when no token was generated.
Sourcepub fn with_reasoning_tokens(self, reasoning: usize) -> Self
pub fn with_reasoning_tokens(self, reasoning: usize) -> Self
How many of the completion’s tokens were spent reasoning.
OpenAI’s own field, and the only part of a reasoning model’s
accounting that IS in their spec – reasoning_content is a
DeepSeek convention this server also speaks, but the token
count is standard, and it is how a caller prices or budgets a
thinking model.
None, not zero, when this build cannot know: a checkpoint
whose family emits no reasoning at all, and any path that did
not run the split. /v1/responses used to report a hardcoded
0 here, which reads as “this model did not think” rather than
“nobody counted” – the exact confusion this module’s header
rules out for timings.
Sourcepub fn with_cached_tokens(self, cached: usize) -> Self
pub fn with_cached_tokens(self, cached: usize) -> Self
Prompt tokens that came from the prefix cache. Call this only
when a prefix cache actually exists (see cached_tokens).
Sourcepub fn with_speculation(
self,
verification_steps: usize,
accepted: usize,
drafted: usize,
per_position: Vec<f64>,
) -> Self
pub fn with_speculation( self, verification_steps: usize, accepted: usize, drafted: usize, per_position: Vec<f64>, ) -> Self
Records what speculative decoding actually achieved for this request. Call this only when speculation ran: leaving the fields unset is how a non-speculative request says so, and a zero would read as “speculation ran and failed”.
accepted and drafted are token counts, per_position the
accept rate at each position inside the draft block. A zero
verification_steps leaves acceptance_length unset rather
than dividing by zero.