pub struct Usage {
pub input_tokens: u32,
pub output_tokens: u32,
pub cache_read_tokens: Option<u32>,
pub cache_creation_tokens: Option<u32>,
}Expand description
Token usage for one model response.
Convention, uniform across every provider mapping: input_tokens is the
full, cache-inclusive prompt size. It always contains cache_read_tokens
as a subset (verified live: OpenAI/OpenRouter/DeepSeek report it inside
prompt_tokens; the Anthropic mapping adds it back since that API reports
input net of cache). cache_creation_tokens is likewise inside input_tokens
for providers that report cache writes (Anthropic); the OpenAI-style
providers we use don’t report writes at all. So Usage::total
(input + output) is the context-window occupancy.
For cost, the cache buckets bill at different rates, so subtract them
from the input rate instead of charging the full rate twice:
cost = (input - cache_read - cache_creation)·p_in + cache_read·p_cache_read + cache_creation·p_cache_write + output·p_out.
Fields§
§input_tokens: u32Full prompt tokens, cache-inclusive (a superset of the two cache fields).
output_tokens: u32Generated output (completion) tokens.
cache_read_tokens: Option<u32>Subset of input_tokens served from the prompt cache (cheaper rate).
cache_creation_tokens: Option<u32>Subset of input_tokens written into the prompt cache (surcharge rate).
Implementations§
Trait Implementations§
Source§impl AddAssign for Usage
impl AddAssign for Usage
Source§fn add_assign(&mut self, other: Usage)
fn add_assign(&mut self, other: Usage)
+= operation. Read more