1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
//! The crate's ONLY file containing `unsafe`: one io_uring ring behind a
//! small, safe submit-and-collect API (Phase 10F, ADR-0021). This is the
//! narrowly isolated Linux ABI boundary the unsafe ledger designates
//! (`docs/security/unsafe-ledger.md`). The rest of the crate — including
//! the storage transport [`crate::store::io::uring::UringIo`] — calls the
//! safe [`Uring::submit_and_wait`]; no other module needs `unsafe`.
//!
//! # PURPOSE
//!
//! Own the crate's single unsafe primitive — the SQE push into the
//! kernel-shared submission ring — and present it as a safe, synchronous
//! "submit all, wait all, collect all" API. The completion queue needs no
//! `unsafe` here: the `io-uring` crate exposes it through a safe
//! `Iterator`.
//!
//! # BOUNDARY
//!
//! KNOWS: the io_uring ABI through the `io-uring` crate (0.7.14, already
//! a transitive dependency of `libublk` — no new dependency tree),
//! `user_data` completion tokens, and the file descriptors and buffers
//! callers hand in. NEVER KNOWS: the record format, segment layout, store
//! orchestration, or durability ordering — those live above
//! [`crate::store::io`]. Ledger policy additionally forbids this code
//! being reachable from persistent-data parsing (parsers are
//! `forbid(unsafe_code)` and cannot call platform code).
//!
//! # MODEL
//!
//! A batch in, exactly one completion per op out. Callers submit
//! `(token, Sqe)` pairs; the kernel completes each op independently and
//! MAY do so in any order, so completions are correlated by the caller's
//! token, never by position. The API does not return until every
//! completion is collected, so callers observe the same blocking
//! semantics as the synchronous syscalls (`pread`/`pwrite`/`fsync`) the
//! reference backend uses — which is what lets `UringIo` reproduce the
//! sync engine's crash-court injection-point sequence exactly.
//!
//! # PERSISTENT AUTHORITY
//!
//! None for the on-disk format: this is a transport, not a format (ADR
//! 0021: a store is equally mountable with either backend). Persistence
//! semantics are the CALLER's choice of ops — a WRITE completion means
//! the kernel accepted the bytes into its page cache (crash-unsafe
//! alone); an FSYNC completion means the data is durable
//! (power-loss-safe). This module never decides what to write or when to
//! fsync; the store's orchestration above does, and the crash courts
//! check that ordering.
//!
//! # CORRECTNESS INVARIANTS
//!
//! - Exactly one completion per submitted op, in arbitrary order.
//! - Every buffer/fd referenced by a submitted SQE remains valid until
//! the call returns (the buffer-lifetime contract on
//! [`Uring::submit_and_wait`]).
//! - Distinct `user_data` tokens within a batch; no aliasing between the
//! buffers of different ops in a batch.
//! - Submit and collect are atomic under the mutex, so one caller's
//! completions are never interleaved with another's.
//!
//! # CONCURRENCY
//!
//! One `Mutex<IoUring>`; concurrent callers serialize only on the short
//! submit/collect critical section. ONE ring is correct for the store:
//! the write path is serialized above (`commit_lock` + segment mutex),
//! and batches are the read-path parallelism unit. Deliberately out of
//! scope (ADR-0021): registered fixed buffers/files, `SQPOLL`, `IOPOLL`,
//! and async submission from multiple threads.
//!
//! # DURABILITY
//!
//! The ring itself provides none — it is a syscall transport. What
//! survives a crash is decided by which ops the store lets complete
//! before acknowledging, with exactly the sync backend's syscall
//! semantics (page-cache accept vs fdatasync vs fsync vs dir fsync). The
//! crash courts parameterize over both backends and require the same
//! recoverable state per injection point.
//!
//! # RESOURCE BOUNDS
//!
//! Ring depth: `entries` submission slots (u32; clamped to ≥ 8; the
//! kernel rounds up to a power of two). The `ops` slice may exceed the
//! ring depth; it is chunked internally, and buffers must stay valid
//! across all chunks (the whole call). Buffer sizes are chosen by the
//! caller — the store allocates `stored_len` bytes per read from
//! persistent descriptors, bounded by the decode limits the hostile-media
//! court exercises, never by this module.
//!
//! # PERFORMANCE
//!
//! The ring exists because a materialization's dependencies (models,
//! streams, dictionaries, tree nodes) and a commit's durability sequence
//! can be ONE submission instead of a syscall per op. Measured ring
//! economics (Phase 10F, tmpfs): ~2.3 µs per submit-and-wait cycle (the
//! kernel's submit+wait+wake floor) vs ~0.1 µs per `pread`, amortizing
//! to ~0.34 µs/op at a 32-op batch. The 10F-sync/10F-uring court pair
//! (evidence `fuse-court-*-10f-sync/uring`) shows uring trailing sync by
//! 5–27% (reads −7–12%) on tmpfs, so the default stays `sync` (the
//! oracle) until real-device (NVMe, queue-depth) evidence flips it.
//!
//! # FAILURE MODES
//!
//! Expected: `SubmissionQueue::push` failure (no room in the SQ) and
//! `io_uring_enter` failures surface as `io::Error` from
//! [`Uring::submit_and_wait`]; per-op failures arrive as a negative
//! `Completion::result` (`-errno`), which the caller maps to a typed
//! store error. MUST NEVER HAPPEN: returning without exactly one
//! completion per op — the caller would free (or reuse) a buffer the
//! kernel may still read or write: a use-after-free of kernel-shared
//! memory.
//!
//! # HISTORY / EVIDENCE
//!
//! Phase 10F, ADR-0021. The unsafe isolation is one of the repo's
//! custodial items: `SubmissionQueue::push` is the `io-uring` crate's
//! sole unsafe primitive (the CQ is consumed through its safe `Iterator`),
//! and the ledger pins it HERE with exact preconditions (buffer lifetime,
//! fd validity, alignment, no aliasing), enforced by
//! `unsafe_files_match_ledger` (`src/tests/unsafe_ledger.rs`) and the
//! crate-root `#![deny(unsafe_code)]`. Acceptance is crash-court parity:
//! `src/tests/io_backend_parity.rs` runs the crash matrix against BOTH
//! backends and asserts the store directories are canonically
//! byte-identical at every crash point (evidence
//! `fuse-court-*-10f-sync/uring`).
use Mutex;
use IoUring;
use Entry as Sqe;
/// One completion: the caller's token plus the operation's result — the
/// io_uring CQE contract. A batch in, exactly one completion per op out;
/// completions are correlated by `user_data` because the kernel may
/// complete ops in arbitrary order.
/// A mutex-guarded io_uring ring: the crate's single `unsafe`-owning
/// type (see the module docs for the ledger story).
///
/// ROLE: the one place SQEs are pushed into the kernel-shared submission
/// ring and completions are collected, behind a safe synchronous API.
///
/// INVARIANTS: submission and completion collection are atomic per call,
/// so a caller that submits a batch receives exactly that batch's
/// completions before the call returns — the ordering semantics the
/// storage transport needs; exactly one completion per submitted op;
/// every op's buffers/fd stay valid until the call returns (the safety
/// contract on [`Uring::submit_and_wait`]).
///
/// One ring is correct for the store: the write path is serialized
/// (`commit_lock` + segment mutex) and batches are the read-path
/// parallelism unit.