1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
//! The shared token estimator (SPEC.md C9). The repo has no tokenizer — only
//! provider-reported `Usage` (`provider.rs:153-164`) and the turn/output
//! counters `agent.rs` already tracks. Every other UX figure (the banner,
//! `/status`, `/tokens`, `show-reductions`, notices) derives from one
//! documented heuristic here, and is always printed with a `~` prefix so it
//! reads as an estimate, never a measurement. Acceptance criteria never
//! assert against these numbers directly — ACs assert exact byte counts
//! (`SessionInfo::full_bytes`/`view_bytes`, C9/D15) and only *derive* a
//! token/dollar figure from them for the test log.
//!
//! Provider-reported figures (`Usage.completion_tokens`, B7's optional
//! `cached_tokens`, `agent.total_output_tokens()`) are real counts and are
//! never routed through this module — they print un-tilded.
use ;
use crateToolSchema;
pub use ;
/// Deterministic token estimate: `ceil(utf8_bytes / 4)`. Documented
/// heuristic — all UX figures derived from it are printed with a `~`
/// prefix (see [`fmt_approx_tokens`]). Never used in acceptance criteria
/// (ACs assert exact byte counts).
/// PARITY-18 D2 — conservative safety margin folded into the context-guard
/// boundary only (never into the plain `~`-prefixed UX estimates
/// themselves, which stay the documented `ceil(bytes/4)` heuristic
/// unmodified). `ceil(utf8_bytes/4)` under-counts real tokenizer output on
/// CJK text, base64/binary-ish blobs, and dense code — informally by 25%+
/// against common tokenizers for those corpora, since multi-byte UTF-8
/// sequences and non-whitespace-delimited runs pack more real tokens per
/// byte than the heuristic assumes. 25% is a round number comfortably above
/// that observed skew: applying it can only make the guard MORE
/// conservative (refuse sooner), never let an over-context request through
/// that a real tokenizer would also have refused.
const GUARD_MARGIN_NUM: u64 = 5;
const GUARD_MARGIN_DEN: u64 = 4;
/// Apply the runtime's 5/4 context-guard safety margin to a raw token
/// estimate. Rounds up (`div_ceil`), never down —
/// the margin only ever pushes the boundary check to be more cautious.
/// PARITY-18 — headroom reserved for the model's own completion, folded
/// into the context-guard boundary alongside [`with_guard_margin`]. Neither
/// [`estimate_view_tokens`] nor [`estimate_request_tokens`] counts anything
/// for the reply the model is about to generate — this is a flat token
/// budget carved out of the model's context window for it, since the
/// completion shares the same window as the request on every provider this
/// crate targets.
///
/// PARITY-18 v3 NOTE-AND-DECIDE (NF6, owner-recorded, not fixed here): this
/// is a FLAT reserve — it does not read `Config::max_tokens`
/// (`crates/harness/src/config.rs`, user-settable via `--max-tokens`, wired
/// into the actual provider request at `provider.rs`'s
/// `ChatRequest::max_tokens`). A user who passes `--max-tokens` greater than
/// 16,384 can still pass this guard (`projected + 16_384 <= context_limit`)
/// and then draw a provider-side `input_tokens + max_tokens > context_window`
/// rejection the guard never anticipated — i.e. the guard's margin can be
/// smaller than what the user actually asked the provider to reserve for the
/// completion. Making the reserve `max(CONTEXT_RESPONSE_RESERVE_TOKENS,
/// config.max_tokens)` would close this, but `context_guard` doesn't
/// currently receive `Config` at all (only `messages`/`tools`/
/// `context_limit`) — threading it through is a small but real signature
/// change touching every call site (`resume_cmd`, `Agent::run_loop`, and
/// this pass's new `reduce_to_fit` `fits` closures) that's out of scope for
/// this pass's reducer/guard boundary fix. Recorded for the owner.
pub const CONTEXT_RESPONSE_RESERVE_TOKENS: u64 = 16_384;
/// PARITY-18 D1 — the full wire-request token estimate: every message in
/// `messages` (including the system prompt at index 0) plus the serialized
/// `tools` schema array, which is a real part of the provider request but — before
/// PARITY-18's re-fix — was never counted by the preflight guard at all.
/// A session whose messages alone fit comfortably could still carry a fat
/// builtin/MCP tool-schema array that blows the real wire request; this is
/// the fix.
/// PARITY-18 D1/D2/D4 — the single context-guard decision, shared by every
/// call site that must decide whether a request is safe to send: the CLI's
/// `resume_cmd` preflight check AND `Agent::run_loop`'s per-send check
/// (D4 — the guard is a session invariant, not a one-shot preflight, so
/// turn 2+ and `/expand all` are covered too). Because both call through
/// this one function, a request can never pass one gate and fail the
/// other — there is only one formula.
///
/// `fits` is true iff the [`with_guard_margin`]-adjusted
/// [`estimate_request_tokens`] estimate, plus the
/// [`CONTEXT_RESPONSE_RESERVE_TOKENS`] completion reserve, is still within
/// `context_limit` — i.e. gates on the reduce TARGET
/// (`context_limit - CONTEXT_RESPONSE_RESERVE_TOKENS`), not the raw limit,
/// closing the "blind band between reduce target and pass/fail boundary"
/// gap. Returns the margin-adjusted projected total either way so callers
/// can report it (dev/02) regardless of verdict.
/// Schema-token estimate for the B6 "tools" banner line: the estimate over
/// the serialized `full` [`ToolSchema`] list minus the estimate over the
/// serialized `advertised` list — i.e. the token cost of what's currently
/// deferred (hidden behind `tool_search`) rather than eagerly advertised.
/// Saturates to `0` rather than underflow if `advertised` somehow estimates
/// larger than `full` (e.g. formatting differences), since "negative
/// deferred tokens" has no meaning for the banner.
/// Render an estimated token count in the shared UX style: `~21,904 tok`
/// (tilde prefix + comma-grouped thousands, matching the stub-line comma
/// style in `reduce.rs`). Every figure that flows through
/// [`estimate_tokens`]/[`estimate_view_tokens`]/[`estimate_deferred_schema_tokens`]
/// should be rendered through this helper so the `~` discipline (D11) is
/// applied uniformly rather than ad hoc at each call site.