1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
//! Storage transport abstraction (Phase 10F, ADR-0021).
//!
//! `Store` / transactions / epoch checkpoint sit above an `IoBackend`:
//!
//! ```text
//! Store / transactions / epoch checkpoint
//! │
//! ▼
//! IoBackend
//! / \
//! SyncIo UringIo
//! reference path performance path
//! ```
//!
//! - [`sync::SyncIo`] is the pre-10F synchronous engine, preserved
//! byte-for-byte as the crash-consistency oracle (the default).
//! - [`uring::UringIo`] implements the same record format and the exact
//! same durability ordering with the syscalls issued through an
//! io_uring ring.
//!
//! Every backend call completes its durability work before returning, so
//! the store's orchestration — and its crash-court injection points
//! between calls (`CrashPoint`) — is unchanged. The acceptance test for
//! `UringIo` is crash-court parity: at every injection point the store
//! directory must be byte-identical to the `SyncIo` state, and recovery
//! must produce the same admissible state.
//!
//! # PURPOSE
//!
//! Define the transport seam below the store and make the syscall
//! *issuing* strategy replaceable without touching format, durability
//! ordering, or recovery: `SyncIo` issues each syscall directly and
//! synchronously; `UringIo` issues the same logical operations through
//! one io_uring ring. The on-disk format is untouched by the choice — a
//! store is equally mountable with either backend.
//!
//! # BOUNDARY
//!
//! KNOWS: the store directory layout, segment file naming and open mode,
//! the record header size (payloads begin at `offset +
//! RECORD_HEADER_SIZE`), and the durability primitives the store needs.
//! NEVER KNOWS: record format semantics, transaction / epoch
//! orchestration, recovery, or any policy. This module is safe Rust
//! (`#![forbid(unsafe_code)]`); the crate's one `unsafe` surface is
//! [`crate::platform::io_uring`], with a ledger entry and a
//! walk-the-src enforcement test.
//!
//! # MODEL
//!
//! A backend is a byte-addressed transport: every operation addresses
//! `(segment_seq, offset, length)` where offsets and lengths are in
//! bytes within a segment file (or the superblock file, byte offsets).
//! Backends are `Send + Sync` handles; the store holds one `Arc<dyn
//! IoBackend>` for its lifetime. `open_segment_common` and
//! `find_clean_end_bytes` are backend-agnostic and shared so the
//! open-time torn-tail state machine cannot drift between engines.
//!
//! # PERSISTENT AUTHORITY
//!
//! Yes — this seam writes the persistent-data surface: segment bytes,
//! torn-tail truncation, superblock slots + `fsync`, GC unlinks. The
//! contract per call is identical across backends, and the acceptance
//! test for `UringIo` is canonically byte-identical store directories at
//! every crash injection point (inode wall-clock times canonicalized),
//! with recovery producing the same admissible state.
//!
//! # CORRECTNESS INVARIANTS
//!
//! - Every call completes its durability work before returning (the
//! ADR-0008 recovery contract holds for both backends by
//! construction).
//! - A fresh segment is the 4-byte `SEGMENT_MAGIC` made durable before
//! `open_segment` returns; an existing segment has its torn tail
//! truncated and made durable; a truncated magic (< 4 bytes) or a
//! wrong magic is `Malformed`, never silently accepted.
//! - `write_at` is pwrite semantics (page-cache accept); short writes
//! loop, 0-byte completions are errors.
//! - `read_many` returns results in request order (the i-th result
//! corresponds to the i-th request) — the parallel-decode consumer
//! relies on this.
//! - `delete_segment` is idempotent (missing file tolerated) and only
//! runs after the new root is durable (GC ordering lives above).
//!
//! # CONCURRENCY
//!
//! The seam itself adds no locks: `SyncIo` guards only its fd map
//! (10E/10E1 discipline — clone the `Arc`, never hold across an op);
//! `UringIo` guards only its ring. The write path is serialized above
//! (`commit_lock` + segment mutex); `read_many` is the read-path
//! parallelism unit (one submission for `UringIo`, sequential preads for
//! `SyncIo`).
//!
//! # DURABILITY
//!
//! Acknowledgment semantics are spelled out per method: `write_at` =
//! page-cache accept; `sync_segment_file` = full fsync (fresh-magic
//! durability); `fdatasync_segment` = record durability;
//! `sync_segments_dir` = new segment directory entry durable;
//! `write_superblock_slot` = page cache; `fsync_superblock` = commit
//! durable. The store composes these into checkpoints and barriers.
//!
//! # RESOURCE BOUNDS
//!
//! `read_payload` / `read_many` allocate `stored_len` bytes per request
//! (record `stored_len` is a `u32` field; the read path validates via
//! `Limits` above this seam). `uring_entries` bounds the submission
//! queue capacity of `UringIo` only.
//!
//! # PERFORMANCE
//!
//! The two implementations exist because the synchronous syscall-per-op
//! shape dominated the read path (a materialization fetches a model, an
//! encoded stream, a dictionary and B-tree nodes individually) and the
//! commit durability sequence. `read_many` batches those fetches into
//! one ring submission for `UringIo`. The sealed 10F court pair
//! (tmpfs-backed, `fuse-court-*-10f-sync/uring`) measured `UringIo`
//! trailing by 5–27% on writes and 7–12% on reads — the ~2.3 µs ring
//! submit/wait floor on sub-µs tmpfs I/O; the default stays `sync`
//! (the oracle) until real-device evidence flips it.
//!
//! # FAILURE MODES
//!
//! `StoreError::Io` for syscall failures; `StoreError::Limit` for
//! arithmetic overflow in payload offsets; `SegmentError::Malformed` for
//! a torn/wrong magic; `SegmentError::Overflow` in `find_clean_end_bytes`
//! for a record whose size overflows. A record that fails decode is a
//! clean-end boundary, never an error, at open time (torn tail).
//!
//! # HISTORY / EVIDENCE
//!
//! Phase 10F (v0.6.2, ADR-0021): the seam was introduced with `SyncIo`
//! as the preserved oracle and `UringIo` as the opt-in performance path;
//! crash and durability courts are parameterized over both backends
//! (`src/tests/io_backend_parity.rs`); the sealed pair is
//! `fuse-court-*-10f-sync/uring`. The `Arc<File>` fd-cache shape came
//! from Phase 10E1 (`fuse-court-*-10e1-before/after`). The write-path
//! hunt during 10F also found and fixed `apply_sorted_batch` walking the
//! whole tree per tiny batch (empty-batch short-circuit, ~50× win on
//! both backends).
use Path;
use Arc;
use crateSEGMENT_MAGIC;
use crateStoreError;
use crateSegmentError;
/// Which transport the store uses.
/// One payload read request for [`IoBackend::read_many`]: the record
/// payload at `(segment_seq, offset)` with `stored_len` bytes.
/// The storage transport. Every method has the same semantics in both
/// implementations; the difference is only *how* the syscalls are issued.
/// Build the backend for a store directory (mkfs / mount).
/// Backend-agnostic `open_segment`: create (magic, made durable) or
/// validate + truncate the torn tail. Both backends share this; the
/// primitives (`segment_len`, `write_at`, `truncate_segment`,
/// `sync_segment_file`) are backend-specific.
///
/// All offsets and lengths here are byte units within the segment file.
pub
/// Find the last clean record boundary in segment bytes (the offset at
/// which sequential record validation first fails, or EOF). The pure
/// version of `segment::find_clean_end`, shared by both backends.
pub