1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
//! Phase-12A read-cost sampling: the Hot-DAG terminalization oracle's
//! instrumentation.
//!
//! # Purpose
//!
//! Collect a bounded [`ReadCostSample`] per materialization so the 12A
//! oracle can answer the phase's one question: **for a reference DAG that
//! is legal and compact, when does its repeated read/materialization cost
//! exceed the storage savings it provides?** The first oracle constructs
//! controlled DAG families (depth 0–4, fanout 1/2/4, cold/warm/hot cache
//! states, ExactRef / BaseResidual / SequenceDict / mixed diamonds) and
//! measures random-read p50/p95/p99 against the samples — the
//! "depth != latency" distinction (a depth-4 chain whose dependencies are
//! hot in memory may be cheaper than a depth-1 representation requiring a
//! large cold fetch) is the hypothesis under test.
//!
//! # Model
//!
//! One sample per materialization, filled by the two halves of the
//! batched read (`Store::materialize_prepare` and
//! `Store::materialize_decode`), carried inside `PreparedRead` so the
//! two-phase FUSE read (guard-held prepare, guard-released decode)
//! completes the same sample. Fields:
//!
//! ```text
//! family the representation family of the batch's first
//! descriptor (the oracle's families are homogeneous,
//! so the first is representative)
//! reference_depth the descriptor's own contribution
//! (core::cost::reference_depth: 0 or 1)
//! max_path_depth the deepest reference-chain level actually walked
//! by collect_read_deps (the real chain length)
//! dag_nodes distinct DAG nodes traversed (top-level descriptors
//! + distinct nested descriptors resolved)
//! fanout max nested children of any single node in the walk
//! referenced_objects distinct object ids prefetched for the closure
//! bytes_fetched physical bytes of those objects
//! read_many_submissions backend fetch submissions (1 in the batched
//! path; more if decode-time fallbacks fire)
//! cache_hits/misses decoded-model cache events (snapshotted as a DELTA
//! of the store's atomic counters across the
//! materialization — exact for sequential reads,
//! approximate under concurrent decodes, which the
//! oracle's single-threaded probe never triggers)
//! decode_cpu_ns decode thread-CPU (exact for the inline single-
//! extent path the oracle reads; the requesting
//! thread's share for scoped/worker decodes)
//! io_wait_ns the prefetch's wall time (the read_many)
//! read_latency_ns total materialization wall (prepare + decode)
//! logical_bytes materialized output length
//! ```
//!
//! The ring is bounded (FIFO, [`READ_COST_RING`]) so the instrumentation
//! is memory-bounded regardless of read volume.
//!
//! # Boundary
//!
//! Samples are diagnostic and **never persisted, never an authority**:
//! deleting every sample changes nothing. The hotness tracker (below) is
//! the same — advisory read-frequency evidence that a FUTURE
//! terminalizer would consult, kept strictly in-memory until the oracle
//! proves depth/hotness predict latency.
//!
//! # Hotness tracker
//!
//! Exponentially decayed per-chunk-id read-frequency counters
//! (`h ← h·0.9 + 1` per touch), diagnostic and non-persistent. Touched
//! once per distinct referenced object per materialization, so a hot base
//! shared by many consumers accumulates evidence while one-off chunks
//! decay toward zero. The 12A terminalization design (if adopted) would
//! use this to distinguish "cache-first" (retain a hot shared base) from
//! "rewrite-first" (terminalize a costly DAG); the oracle only reports it.
//!
//! # Correctness invariants
//!
//! - A sample never affects read bytes, scheduling, or persisted state.
//! - The ring never grows beyond [`READ_COST_RING`] entries.
//! - Hotness values stay in `[0, ∞)`; decay keeps stale entries bounded in
//! count by the distinct-chunk set.
//!
//! # Concurrency
//!
//! The ring and the hotness map are each behind their own store-level
//! mutex (bounded, short critical sections — a push/read is a HashMap
//! op). The cache counters are per-store lock-free atomics (bumped by
//! `Store::decode_rans` from any thread, sampled as a delta by the
//! materialization — see the field note above). No thread-local state:
//! the sample travels inside `PreparedRead`, so worker-thread decodes
//! complete the same sample their requesting thread opened.
//!
//! # Resource bounds
//!
//! Ring: [`READ_COST_RING`] fixed-size samples. Hotness: one entry per
//! distinct chunk id touched (bounded by the store's object set; no
//! growth vector beyond distinct content).
//!
//! # Performance
//!
//! Per materialization: one ring push (mutex + Vec push, amortized), one
//! thread-CPU read in the decode half (the same `CLOCK_THREAD_CPUTIME`
//! the 11D worker clock uses), and the hotness touches (one HashMap
//! lookup per distinct dep — already O(deps) work the read must do
//! anyway). The oracle measured no material impact on the read rows it
//! reports (the samples are the measurement, not the perturbation).
//!
//! # Failure modes
//!
//! Infallible by construction. A poisoned ring/hotness mutex panics
//! (loudly, like every other store mutex); samples are advisory, so a
//! panic would be an instrumentation defect, not data corruption.
//!
//! # History / evidence
//!
//! Phase 12A oracle (`docs/performance/dag-read-cost.md`, sealed
//! `evidence/performance/dag-read-cost-probe-*/`): the first oracle
//! measured whether depth predicts read latency once family and cache
//! state are controlled, and decided terminalization's fate (adopted only
//! on a measured latency penalty; `depth > N => RAW` was never a
//! candidate — the brief's explicit rejection).
use HashMap;
use Instant;
use crateChunkId;
/// Bounded sample ring length (FIFO; the per-materialization sample
/// stream, capped so read volume cannot grow memory).
pub const READ_COST_RING: usize = 4096;
/// Hotness decay per touch. 0.9: a chunk read once per sweep stays warm
/// for ~20 sweeps; the tracker is clocked by READS (like the DSFB
/// observer is clocked by writes), not wall time.
pub const HOTNESS_DECAY: f64 = 0.9;
/// One materialization's read-cost sample (Phase-12A; fields documented
/// in the module doc).
/// The bounded sample ring (FIFO; [`READ_COST_RING`] entries).
/// Exponentially decayed per-chunk-id read-frequency evidence
/// (`h ← h·0.9 + 1` per touch). Diagnostic and non-persistent; a future
/// terminalizer's cache-first/rewrite-first input (module doc).
/// True thread-CPU time (`CLOCK_THREAD_CPUTIME`) on Linux; `None`
/// elsewhere (the same clock the 11D worker oracle uses; wall is the
/// documented fallback where the kernel lacks it).
pub
/// A start marker for thread-CPU measurement (the decode half's clock).
pub