1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
//! **How full is this conversation's context?** (DESIGN §5.1 #35, §11's
//! settings rows.) The context-window percentage is shown per chat.
//!
//! Derived, never stored, and pure: a query over the worker's already-walked
//! [`StepBill`]s (§3.5, bl-9dd4) and the windows `models.yaml` declares. Yog
//! keeps no counter and no cache — asking twice re-derives, like every other
//! §5.1 fact.
//!
//! **Fullness is not spend.** [`crate::budgets`] sums every counter of every
//! attempt of every step of a whole descent — that is what exhausts
//! `max_total_tokens`, and it grows without bound while a context is compacted
//! back down. Fullness is one number off **one** step: the prompt the *latest*
//! step of the root agent actually sent. Two questions, two derivations, one
//! walk.
//!
//! **The latest step of the root, not of the descent.** A dispatched child runs
//! its own context in its own `steps/<root>-<child>/` tree; folding the descent
//! in would answer no question anyone asked. The conversation's context is the
//! conversation's.
//!
//! **Why the prompt is `max(input, cache_read + cache_write)`.** brazen's
//! canonical `Usage` is deliberately *unnormalized* about overlap, because its
//! providers disagree: Anthropic reports `input_tokens`,
//! `cache_read_input_tokens` and `cache_creation_input_tokens` as three
//! **disjoint** slices of one prompt, while OpenAI's `prompt_tokens` and
//! Google's `promptTokenCount` already **contain** the cached slice they report
//! beside it. So summing all three over-states OpenAI by the cached prefix, and
//! taking `input_tokens` alone under-states Anthropic by nearly the whole
//! prompt — brazen marks Anthropic prompts for caching unconditionally
//! (`protocol/anthropic/encode/cache.rs`: *"caching is brazen-owned POLICY with
//! zero canonical surface"*), so mid-conversation `input_tokens` is only the
//! uncached tail. The maximum of the two readings is exact wherever the cached
//! slice is contained (OpenAI, Google, Ollama, and any provider reporting no
//! cache counters at all — where it degrades to plain `input_tokens`) and a
//! **floor** where the slices are disjoint (Anthropic, short by that same
//! uncached tail). It never over-states, and yog already renders a figure it
//! cannot complete as a floor rather than as an answer (§3.5's
//! `+N tok unpriced`). One formula, no per-provider branch — normalizing that
//! overlap is brazen's to do, not yog's to guess at.
/// The egui line the settings rows paint.
use crate;
use BTreeMap;
/// One conversation's context as of its latest step (§5.1 #35).
/// The prompt one `Usage` reading describes — the module header's rule, at its
/// one home.
/// One conversation's fullness, or `None` when nothing honest can be said: the
/// root has taken no step, its latest step names no model, or that model
/// declares no window. **Render nothing, never an estimate** — the whole point
/// of the figure is that it is measured.