1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
//! Phase-12B durability generations: the group-commit coordinator.
//!
//! # Purpose
//!
//! Amortize concurrent `fsync` barriers without weakening the durability
//! contract. The 12B model (the brief):
//!
//! ```text
//! logical_seq monotonically identifies acknowledged mutation state
//! durable_seq highest logical sequence known to survive power loss
//!
//! fsync(required_seq = N) may return success iff durable_seq >= N
//! ```
//!
//! N concurrent fsyncs coalesce onto ONE physical barrier whose cut
//! covers every waiter registered before the physical work starts; each
//! waiter completes only after a cut that includes its writes (the
//! linearizability requirement the crash courts pin). A mutation
//! acknowledged AFTER the cut is chosen must NOT inherit that barrier.
//!
//! EntropyFS's acknowledged-mutation state has TWO monotonic coordinates:
//!
//! - `seq` — the epoch's mutation-log sequence (`Epoch::seq`; envelopes
//! `> root.log_seq` are replayed at recovery). Covers staged epoch ops.
//! - `gen` — the published root's `generation` (bumped by EVERY commit:
//! epoch checkpoints AND direct transaction commits). Covers direct
//! (non-epoch) writes, which never touch `seq`.
//!
//! A barrier makes durable everything ≤ its cut `(seq, gen)` componentwise;
//! the group's durable state is the pair of store atomics advanced to the
//! last completed cut.
//!
//! # Model: the coordinator
//!
//! ```text
//! DurabilityGroup
//! waiters: (required_seq, required_gen, joined_gen)
//! owner_cut: Option<(seq, gen)> the in-flight physical barrier
//! owner_error: Option<(gen, String)> a FAILED generation's error,
//! tagged with its generation so only
//! ITS covered waiters surface it
//! next_gen: u64 the generation counter
//! ```
//!
//! A caller registers with `joined_gen = next_gen` — the generation that
//! will cover it (the in-flight one if it registered before the takeover,
//! otherwise the next). The first waiter when idle becomes the OWNER:
//! it fixes the cut at the componentwise max of the CURRENT waiters'
//! requirements, increments `next_gen`, releases the group lock, and runs
//! the physical barrier (epoch checkpoint + commit-lock-held
//! fdatasync → dir sync → superblock write → superblock fsync). On
//! success the durable atomics advance to the CUT (conservative: never
//! beyond what the generation was required to cover, even if the
//! checkpoint flushed more); on failure the error is stored tagged with
//! the generation. Every waiter is woken and re-checks:
//!
//! ```text
//! durable covers required -> Ok
//! my generation failed -> Err(the generation's error)
//! otherwise -> park again (or become the next owner)
//! ```
//!
//! # Correctness invariants
//!
//! - **Linearizability (write→fsync):** a waiter returns Ok only after a
//! completed physical barrier whose cut covers its requirement — the
//! checkpoint consumed every envelope ≤ its `seq` into the root (which
//! recovery replays nothing beyond) and the superblock fsync'd a root
//! at generation ≥ its `gen`.
//! - **No barrier inheritance across cuts:** a mutation acknowledged
//! after the owner fixed the cut is not covered by `durable`'s advance
//! to that cut (its `seq`/`gen` exceed it), so its fsync stays pending.
//! - **Failed generations are surfaced, never silently retried forever:**
//! each waiter returns exactly one outcome; a failure is tagged with
//! the generation so late arrivals (joined the NEXT generation) retry
//! rather than inheriting the previous generation's error.
//! - **The physical barrier is unchanged:** same steps, same crash hooks,
//! same commit-lock hold — the group only decides WHO runs it and WHO
//! waits.
//!
//! # Concurrency
//!
//! One short mutex critical section at registration/takeover/wake; the
//! physical barrier runs outside the group lock (the owner holds the
//! COMMIT lock across the fsync window, exactly like the pre-12B
//! barrier). Waiters park on the condition variable; the durable atomics
//! are lock-free (Relaxed is sufficient — the store state they certify is
//! synchronized by the store's own locks, and the atomics are only
//! readiness markers).
//!
//! # Resource bounds
//!
//! One `(u64, u64, u64)` entry per concurrent fsync; the epoch bounds
//! outstanding requests. No allocations on the fast path.
//!
//! # Failure modes
//!
//! A poisoned group mutex panics (like every store mutex). A physical
//! barrier failure (I/O, ENOSPC) is stored and surfaced to the covered
//! waiters as `StoreError::Io` — durability is never silently claimed.
//!
//! # History / evidence
//!
//! Phase-11B/11C measured the fsync convoy (`commit_lock_wait` 34.7% of
//! 16-thread request time pre-11C; the barrier's own comment named the
//! group-durability future). The 12B oracle (`src/tests/fsync_group_probe.rs`,
//! sealed `evidence/performance/fsync-group-probe-*/`) sealed the
//! baseline: amplification 1.00 at every concurrency, fsync p99 45 µs →
//! 7.9 ms at 32 callers. This coordinator is the 12B-1 fix; the re-run
//! seals the amplification reduction (CHANGELOG v0.7.9).
use VecDeque;
/// One registered fsync waiter.
/// The group-commit coordinator state (behind the store's
/// `durability_group` mutex; the condition variable lives beside it in
/// the store).