pub struct TokenUsage {
pub input: u64,
pub cache_write_5m: u64,
pub cache_write_1h: u64,
pub cache_write_unsplit: u64,
pub cache_read: u64,
pub output: u64,
}Expand description
Token counts, split the way the Anthropic and OpenAI usage objects split them.
Fields§
§input: u64§cache_write_5m: u64§cache_write_1h: u64§cache_write_unsplit: u64Cache creation a harness recorded without the TTL that decides its rate. No price table can cover it: charging either TTL’s rate would be a guess about a lifetime nobody wrote down. Only a harness that reports its own costs can put a number against these tokens.
cache_read: u64§output: u64Implementations§
Source§impl TokenUsage
impl TokenUsage
pub fn cache_write(&self) -> u64
Sourcepub fn total(&self) -> u64
pub fn total(&self) -> u64
Everything the model consumed or produced. This is the “TOKENS” column.
Sourcepub fn prompt(&self) -> u64
pub fn prompt(&self) -> u64
The input side: everything sent to the model as the prompt, fresh input plus cache writes plus cache reads, but not the output it produced.
Sourcepub fn cache_hit_rate(&self) -> Option<f64>
pub fn cache_hit_rate(&self) -> Option<f64>
The share of the prompt served from cache (the cheap reads), in 0..=1.
A high number means most of the re-sent conversation was billed at the
cache-read rate rather than full input; a low one on a long session is
money left on the table. None when there was no prompt to judge, or
the model does not cache at all.