1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
//! What one provider call actually took: every attempt at it, and every move to
//! a different provider when one would not serve.
//!
//! A run bills one [`RunRecord::InferenceUsage`](super::RunRecord::InferenceUsage)
//! per call that *worked*, which is the right shape for an invoice and the wrong
//! shape for a post-mortem. A call that was refused three times and answered on
//! the fourth is journaled identically to one that was answered at once, and a
//! call that moved from one provider to another leaves nothing behind at all. The
//! records here are the missing half: one per trip to a provider, plus one per
//! failover, so "why did this turn take ninety seconds" has an answer that does
//! not depend on the daemon's log still being around.
//!
//! Small by default. Timing, classification, how the answer ended, and enough
//! identity ([`RequestDigest`]) to answer whether two attempts sent the same
//! thing: no response body, no error message, and no request body unless the
//! operator asked for one. A record that grew with the prompt would put a copy of the
//! whole window in the journal once per retry, which is why [`ModelInput`]
//! carries a body only where [`CaptureStatus::Retained`] says it does.
use BTreeMap;
use ;
/// Cheap identity for a request: enough to tell whether two attempts sent the
/// same thing, and nothing more.
///
/// Every field is something the assembly already computed, so producing this
/// costs no hashing of its own. `system_hash` is the digest the prefix-cache
/// decision is made from, which is the one number that moves when the system
/// blocks change; the counts and the sampling knobs cover the rest of what an
/// attempt could differ by.
///
/// The per-block digests are left out on purpose. They are a vector as long as
/// the stage has blocks, and this record is written once per attempt: the whole
/// point of a digest here is that it is a fixed hundred bytes whatever the run
/// is doing.
/// Whether an attempt's exact request is in the journal, and where it went if
/// not.
///
/// Every variant describes the state of the *record*, not the intent behind it,
/// because that is what a reader can act on: a body that is here can be read, a
/// body that was never taken cannot be recovered, and a body that was taken and
/// then removed is a different fact from one that never existed.
/// What one attempt sent, and what the request was assembled from.
///
/// Written per attempt whether or not capture is on, because everything here
/// except `request` is cheap and answers questions a digest cannot: which
/// sampling knobs were really in force after resolution, which tools the model
/// was offered, and which build of the assembly produced the shape.
///
/// `request` is the whole prompt. It holds whatever the run's context held -
/// file contents, command output, credentials a tool read - so it is written
/// only for a run whose operator asked for it, and there is no size cap on it:
/// every call re-sends the window, so a captured run's journal grows by roughly
/// the context size per attempt.
/// What the retry loop did after an attempt failed.
///
/// The distinction a reader needs is "was the same thing sent again, and why":
/// three records with the same [`RequestDigest`] are a provider that kept
/// refusing, while three with a different one are a request that kept changing
/// underneath the run. A move to a *different* provider is not here, because the
/// loop never makes one: see [`FailoverRecord`].
/// How one attempt ended.
/// One attempt at one provider call.
///
/// Written per trip to the provider, including the first, and including the ones
/// that failed. This is the record that was missing: a retried call and a
/// first-time success were indistinguishable in the journal, so the time a run
/// spent being refused was invisible and a failover looked like a run that had
/// simply always used the second provider.
/// One provider was unusable, so the next configured model is being tried.
///
/// Its own record rather than a field on [`AttemptRecord`], because the decision
/// is made somewhere else and later: the job reports its failure, the tick loop
/// collects it, and only then does the stage look at what else it was given. By
/// that point the attempt that failed has already been journaled, and an
/// append-only journal cannot go back and amend it.
///
/// Paired with the attempts it sits between, this is what separates "the same
/// provider refused four times" from "four providers each refused once": the
/// attempt records say what was tried, and these say when the target moved.