Skip to main content

Module usage_cost

Module usage_cost 

Source
Expand description

Provider-normalized usage accumulation and cache-aware session cost estimation. Provider-normalized usage accumulation and cache-aware session cost estimation.

Different providers report prompt_tokens with different cache semantics: Anthropic and Minimax report prompt_tokens exclusive of cache-read and cache-creation tokens, while every other supported provider (OpenAI, Gemini, etc.) reports prompt_tokens as a total that already includes cached tokens. This module normalizes per-turn provider usage into the canonical harness Usage shape, where input_tokens always means the total prompt tokens (uncached + cached + cache-creation), and provides a shared cost estimator used by both the interactive and headless runloops.

Structs§

SessionBudget
Durable per-session cost budget for long-running (full-auto) sessions.
SessionCostAccumulator
Accumulates independently priced turns without repricing earlier model routes. Once a turn cannot be priced, a complete session total remains unknown.
SessionCostEstimate
Cache-aware and conservative session cost estimates in USD.

Enums§

BudgetStatus
Outcome of classifying cumulative spend against a budget cap.

Constants§

DEFAULT_BUDGET_WARNING_RATIO
Default fraction of the budget at which the harness warns before hard exhaustion. Mirrors agent.harness.budget_warning_threshold’s default so a SessionBudget built without an explicit threshold behaves like the harness default.

Functions§

estimate_decisions_cost
Base price for the bounded, standard-endpoint Decisions probe, not generation. OpenAI charges $0.10 per million input tokens and no output/cache surcharge. Regional and long-context premiums are outside this bounded probe route.
estimate_session_costs
Resolve pricing for provider/model and estimate session costs from accumulated harness usage. Returns None when the model cannot be resolved or pricing metadata is unavailable.
estimate_session_costs_with_pricing
Estimate session costs from an already-resolved ModelPricing.
estimate_tool_definition_tokens
Estimate the token overhead of sending tools in the request payload.
normalized_turn_usage
Build a per-turn harness Usage sample from raw provider usage, applying the provider-specific normalization documented on provider_reports_exclusive_input so input_tokens always represents the total prompt token count across every provider.
provider_reports_exclusive_input
Returns true when provider reports prompt_tokens exclusive of cache-read and cache-creation tokens.
require_budget_pricing
Reject a priced session budget when its selected route cannot be priced. Call before any inference, including automatic compaction.