Expand description
Provider-normalized usage accumulation and cache-aware session cost estimation. Provider-normalized usage accumulation and cache-aware session cost estimation.
Different providers report prompt_tokens with different cache semantics:
Anthropic and Minimax report prompt_tokens exclusive of cache-read and
cache-creation tokens, while every other supported provider (OpenAI, Gemini,
etc.) reports prompt_tokens as a total that already includes cached
tokens. This module normalizes per-turn provider usage into the canonical
harness Usage shape, where input_tokens always means the total prompt
tokens (uncached + cached + cache-creation), and provides a shared cost
estimator used by both the interactive and headless runloops.
Structs§
- Session
Budget - Durable per-session cost budget for long-running (full-auto) sessions.
- Session
Cost Accumulator - Accumulates independently priced turns without repricing earlier model routes. Once a turn cannot be priced, a complete session total remains unknown.
- Session
Cost Estimate - Cache-aware and conservative session cost estimates in USD.
Enums§
- Budget
Status - Outcome of classifying cumulative spend against a budget cap.
Constants§
- DEFAULT_
BUDGET_ WARNING_ RATIO - Default fraction of the budget at which the harness warns before hard
exhaustion. Mirrors
agent.harness.budget_warning_threshold’s default so aSessionBudgetbuilt without an explicit threshold behaves like the harness default.
Functions§
- estimate_
decisions_ cost - Base price for the bounded, standard-endpoint Decisions probe, not generation. OpenAI charges $0.10 per million input tokens and no output/cache surcharge. Regional and long-context premiums are outside this bounded probe route.
- estimate_
session_ costs - Resolve pricing for
provider/modeland estimate session costs from accumulated harness usage. ReturnsNonewhen the model cannot be resolved or pricing metadata is unavailable. - estimate_
session_ costs_ with_ pricing - Estimate session costs from an already-resolved
ModelPricing. - estimate_
tool_ definition_ tokens - Estimate the token overhead of sending
toolsin the request payload. - normalized_
turn_ usage - Build a per-turn harness
Usagesample from raw provider usage, applying the provider-specific normalization documented onprovider_reports_exclusive_inputsoinput_tokensalways represents the total prompt token count across every provider. - provider_
reports_ exclusive_ input - Returns true when
providerreportsprompt_tokensexclusive of cache-read and cache-creation tokens. - require_
budget_ pricing - Reject a priced session budget when its selected route cannot be priced. Call before any inference, including automatic compaction.