Skip to main content

turso_backup/
stream.rs

1//! Tier 2 — WAL-frame streaming sink (R005-F2).
2//!
3//! Tee WAL frames from a live turso connection to an object store, anchored
4//! to a tier-1a base snapshot. Restore (R005-F3) replays the frames onto the
5//! snapshot. Near-zero RPO; engine-coupled but the public seam is small —
6//! see the spike findings in `.yah/docs/working/turso-s3-backup.md`.
7//!
8//! ## Object layout under `BackupTarget::prefix`
9//!
10//! ```text
11//! frames/{epoch:020}/{checkpoint_seq:010}/{first:020}-{last:020}  batch of consecutive WAL frames
12//! frames/{checkpoint_seq:010}/{first:020}-{last:020}              ditto, epoch 0 (unfenced)
13//! frames/{epoch:020}/{checkpoint_seq:010}/{frame_no:020}          one raw frame, pre-R761-F2 layout
14//! frames/{checkpoint_seq:010}/{frame_no:020}                      ditto, epoch 0 (pre-fencing)
15//! generations/gen-{unix_nanos:020}.manifest                one per `tail_frames` call that uploaded
16//! latest.stream-watermark    text sidecar: "<checkpoint_seq> <last_frame> <written_at_nanos> <epoch> <pointer_generation> <salt1> <salt2>"
17//! ```
18//!
19//! R858-B19: the trailing salt pair is the WAL generation the `last_frame`
20//! position belongs to. `checkpoint_seq` alone does not identify a generation —
21//! a writer-process restart recreates the WAL back at sequence 0 — so the salt
22//! is what `tail_frames` compares and what a generation manifest stamps. See
23//! [`WalGeneration`].
24//!
25//! A frame object holds `n` consecutive frames, each `24 + page_size` bytes,
26//! concatenated in ascending frame order — so a batch is exactly the bytes the
27//! old per-frame objects held, glued together, and offset `i * frame_size`
28//! within it is frame `first + i`.
29//!
30//! The generation manifest names the base snapshot key, the page size, the
31//! frame range covered, which batch objects cover it (since R761-F2), and
32//! (since R732-F2) the fencing epoch and owner label. Object keys are
33//! zero-padded so lexical order matches chronological order (same convention as
34//! tier 1a snapshots / tier 1b manifests; clock-skew-immune).
35//!
36//! ## Frame batching (R761-F2, W248/W313 §9)
37//!
38//! One object per WAL frame made per-write Class A ops scale with *frames*,
39//! and a frame is one page — at 4 KB pages, every 4 KB of changed data was a
40//! billed op, which is a pathologically small object. A tail call's frames now
41//! go up as one ranged object per drain batch (bounded by
42//! [`BackpressureConfig::spill_buffer_frames`], the same bound that already
43//! caps how much this sink holds in memory), so an uploading call costs
44//! `ceil(frames / spill_buffer_frames) + 2` PUTs instead of `frames + 2` —
45//! typically **3, flat**, whatever the write volume in the interval.
46//!
47//! Reading stays compatible in the direction that matters: a manifest without
48//! a `frame_batch` list is a pre-R761-F2 generation and its frames are fetched
49//! one object each, so backups written before this change (and chains that
50//! straddle it) still restore. The reverse is a loud refusal by construction —
51//! the manifest header moved to `v3`, so an older binary reading a batched
52//! generation says "unexpected manifest header" rather than mis-reading it.
53//!
54//! ## Fencing (R732-F2 / R732-T3, W245; R736-T2, W250)
55//!
56//! [`StreamConfig::epoch`] is a per-tenant fencing token minted by yubaba's
57//! raft state machine. It is checked against the sidecar before any frame is
58//! uploaded, and the sidecar advance is a compare-and-swap on the version that
59//! check read — so a stale owner is *rejected* (with
60//! [`StreamOutcome::Fenced`]) both when it arrives late and when it races. An
61//! epoch of `0` means unfenced, which preserves single-writer behaviour but is
62//! still fenced *by* a claimed sink.
63//!
64//! That epoch is a **local** raft counter — it fences ownership moves within
65//! one cell, but two cells are independent raft groups that share no epoch
66//! counter, so it is blind to a tenant moving to a *different* cell.
67//! [`StreamConfig::pointer_generation`] is the second, cross-cell fence: the
68//! writer's belief about the global tenant→cell pointer's generation
69//! (`yah_tenant_pointer::PointerRecord`). It is checked alongside `epoch` at
70//! the same two points — up front against the sidecar, and again on a lost
71//! watermark CAS — so a stale generation bounces exactly like a stale epoch.
72//! A node is the real owner of a tenant only when **both** fences pass.
73//!
74//! ## Seam isolation
75//!
76//! Every call into `turso_core::Connection`'s `feature = "conn_raw_api"`
77//! surface goes through the [`WalSeam`] trait. If a future turso release
78//! renames or reshapes those calls, the delta is a single impl block — not a
79//! sed across the whole sink. The spike measured one breaking rename in three
80//! months (`wal_auto_checkpoint_disable` → `wal_auto_actions_disable`); the
81//! trait makes that a one-file fix.
82//!
83//! @yah:relay(R005, "Tier 2 — WAL-frame streaming (deferred)")
84//! @yah:at(2026-05-26T22:28:31Z)
85//! @yah:status(open)
86//! @yah:phase(P3)
87//! @yah:parent(Q002)
88//! @arch:see(.yah/docs/working/turso-s3-backup.md)
89//!
90//! @arch:see(.yah/docs/working/turso-s3-backup.md)
91//!
92//! @yah:ticket(R005-F2, "Frame-streaming sink: base snapshot + incremental frames + generation tracking")
93//! @yah:assignee(agent:claude)
94//! @yah:at(2026-05-26T22:30:09Z)
95//! @yah:status(review)
96//! @yah:phase(P3)
97//! @yah:parent(R005)
98//! @arch:see(.yah/docs/working/turso-s3-backup.md)
99//! @yah:handoff("Implemented in src/stream.rs (~590 LOC). Public API: WalSeam trait (3 methods: wal_state, wal_get_frame, wal_auto_actions_disable) + real CoreWalSeam impl over turso_core::Connection + StreamConfig{base_snapshot_key, page_size} + tail_frames(seam, target, cfg) -> StreamOutcome (Empty | Streamed | Restarted) + GenerationManifest format/parse + Watermark sidecar.")
100//! @yah:handoff("Object layout under prefix: frames/{checkpoint_seq:010}/{frame_no:020} for raw frames (24-byte header + page), generations/gen-{nanos:020}.manifest for per-call manifests, latest.stream-watermark for the (checkpoint_seq, last_frame) sidecar. Same zero-padded-key convention as tier 1a snapshots and tier 1b manifests — clock-skew-immune.")
101//! @yah:handoff("Tier-2 invariants from the R005-T1 spike are baked in: (checkpoint_seq, frame_no) compound key (not raw frame_no), so a WAL restart -> Restarted outcome under a new seq; page_size is a required cfg parameter (no 4096 hardcode like sync_server.rs); WalSeam isolates all turso_core::Connection calls (one-file delta if upstream renames again); CoreWalSeam::open() takes WAL ownership via wal_auto_actions_disable() at construction.")
102//! @yah:handoff("Cargo.toml: added turso_core = '0.6.1' with features = ['conn_raw_api'] as sibling dep to turso='0.6.1'. The friendly turso wrapper does NOT re-export the raw WAL API; pinned in lockstep — if either bumps, bump both.")
103//! @yah:handoff("Verified: 6 new stream unit tests (empty-on-empty-wal, initial-tail-records-watermark, second-tail-no-new-frames-is-Empty, second-tail-uploads-only-new, wal-restart-emits-Restarted-under-new-seq, manifest-roundtrip+rejection) using a mockable WalSeam. Full crate: 18/18 tests green. cargo clippy --all-targets -- --deny=warnings clean.")
104//! @yah:handoff("What's NOT verified at F2 level: live-DB ping-pong of CoreWalSeam (seed rows -> wal_state -> wal_get_frame -> assert is_commit_frame on the last frame). F3's restore path will be the natural end-to-end exercise. Optional sanity test could be added under F2 if you'd rather catch a CoreWalSeam regression here vs in F3.")
105//! @yah:handoff("Generation manifest text format: 'TURSO-BACKUP STREAM v1' header, then `base_snapshot <key>`, `page_size <n>`, `checkpoint_seq <n>`, `first_frame <n>`, `last_frame <n>`. Dependency-free, same convention as dedup::Manifest. parse_generation_manifest fails loudly on any other shape.")
106//! @yah:next("User: review/approve F2. If you want a live-DB CoreWalSeam sanity test before signoff, say so and I'll add it under F2; otherwise F3 picks it up naturally.")
107//! @yah:next("On approval: archive F2, claim R005-F3 (Restore via frame replay onto snapshot). F3 fetches latest gen-manifest, downloads referenced base_snapshot + frames in (checkpoint_seq, frame_no) order, replays into a writable DB via wal_insert_begin/wal_insert_frame/wal_insert_end (the same WalSeam trait, extended with the insert side).")
108//! @yah:next("Optional independent of F3: file a tiny upstream PR to re-export conn_raw_api from the `turso` wrapper crate so the sibling turso_core dep collapses to one.")
109//!
110//! @yah:ticket(R005-F3, "Restore via frame replay onto snapshot + restart/crash-consistency handling")
111//! @yah:assignee(agent:claude)
112//! @yah:at(2026-05-26T22:30:10Z)
113//! @yah:status(review)
114//! @yah:phase(P3)
115//! @yah:parent(R005)
116//! @arch:see(.yah/docs/working/turso-s3-backup.md)
117//! @yah:depends_on(R005-F2)
118//! @yah:handoff("Built restore_latest_stream(&BackupTarget, dest_path) + the WalInsertSeam trait extension on the existing WalSeam pattern. Public surface: WalInsertSeam{wal_insert_begin, wal_insert_frame, wal_insert_end} + impl for CoreWalSeam (wraps turso_core::Connection::wal_insert_*); RestoreOutcome{base_snapshot_key, checkpoint_seq, generation_count, frames_replayed, last_frame}; restore_latest_stream(target, dest_path)->RestoreOutcome.")
119//! @yah:handoff("Flow: list+sort all generations/*.manifest keys lexicographically → parse each → validate_generation_chain checks all share one base, one page_size, one checkpoint_seq, frames start at 1 and are contiguous → download base_snapshot to dest_path → CoreWalSeam::open(dest) → wal_insert_begin → replay_frames_into walks (seq, frame_no) in order, downloads frames/{seq:010}/{frame_no:020}, asserts byte length = 24+page_size, calls wal_insert_frame → wal_insert_end(force_commit=false). Crash-consistency story = the engine's own truncate-to-last-commit-frame on insert_end(false).")
120//! @yah:handoff("Restart handling = REFUSE: a chain spanning two checkpoint_seqs means the source engine folded WAL into main between generations, so the post-restart frames don't replay onto our pre-restart base. validate_generation_chain bails with 'WAL restart between generations, restore needs a fresh tier-1a snapshot'. Same refusal for cross-base chains and frame gaps. This is the right semantics — heroic restart-spanning replay would silently corrupt.")
121//! @yah:handoff("Verified with 13 new stream tests on top of F2's 6: validate_chain accepts single/contiguous-multi, rejects empty/gap/non-one-start/restart/base-mismatch/page_size-mismatch (7 tests); replay_walks_manifests_in_frame_order with MockInsertSeam; restore_errors_when_no_generations; replay_rejects_wrong_size_frame; live_db_seed_snapshot_tail_restore_round_trips (real turso + CoreWalSeam: seed 50 → snapshot → checkpoint → seed 25 more → tail → restore → assert dest has 75 rows); manifest_with_uncommitted_tail_rolls_back_via_insert_end (live: stage a phantom non-commit frame in the sink, extend the manifest, assert restore drops it and dest=3 rows = 2 base + 1 committed, NOT 4).")
122//! @yah:handoff("cargo test -p turso-backup = 31/31 green (up from 18); cargo clippy --all-targets -- --deny=warnings = clean. One clippy lint fixed in the new test code: manual_is_multiple_of (Rust 1.95 lint). The live tests use turso::Builder for writes/reads + CoreWalSeam for tailing — exclusive WAL lock means we drop the high-level conn before opening the low-level seam (matched by the snapshot tests' pattern).")
123//! @yah:handoff("Two known design decisions worth flagging to the reviewer: (1) replay starts at frame 1 in the dest's fresh WAL but the source's frames could overlap content already in the base snapshot (since VACUUM INTO point-in-time includes WAL state). turso's wal_insert_frame compares-and-returns-OK on identical content, so redundant frames are no-ops — clean orchestration via checkpoint-then-snapshot-then-stream avoids them entirely (which is what the live test does). (2) on replay error we attempt a best-effort wal_insert_end(false) before bubbling the error up, so we don't leave the dest's WAL with an uncommitted suffix half-open.")
124//! @yah:next("User: review/approve F3. Once green, archive F3 and the F2/F3-dependent state of R005 collapses to F4 only (concurrent-writer-safe raw copy).")
125//! @yah:verify("cargo test -p turso-backup")
126//! @yah:verify("cargo clippy --all-targets -- --deny=warnings")
127//!
128//! @yah:ticket(R005-F4, "Concurrent-writer-safe raw copy: read-only main + WAL-frame replay (no TRUNCATE-checkpoint dependency), for backing up under a live writer")
129//! @yah:assignee(agent:claude)
130//! @yah:at(2026-05-27T03:05:00Z)
131//! @yah:status(review)
132//! @yah:phase(P3)
133//! @yah:parent(R005)
134//! @arch:see(.yah/docs/working/turso-s3-backup.md)
135//! @yah:gotcha("R004's dedup::raw_consistent_copy assumes no live writer: it folds WAL->main via PRAGMA wal_checkpoint(TRUNCATE), which a concurrent writer can make return 'busy'. The live-writer alternative copies the main file under a read txn and replays WAL frames itself (wal_get_frame seam) — tier-2 territory, gated on R005-T1's seam assessment. Filed as an R004-T4 followup.")
136//! @yah:handoff("Implemented raw_consistent_copy_live(db_path, page_size) + pub(crate) replay_wal_onto_main inner. Algorithm: open CoreWalSeam (auto-actions disabled on our conn) -> wal_state for (cp_seq, max_frame) -> std::fs::read main -> walk frames 1..=max_frame collecting (page_no, db_size, page_bytes), locate the last is_commit_frame -> grow image to fit the largest page slot in the commit prefix -> overwrite each page at (page_no-1)*page_size -> truncate to db_size*page_size. Frames past the last commit are dropped wholesale (crash-consistency, matches restore's wal_insert_end(false)).")
137//! @yah:next("User: review/approve F4. Once green, archive F4 and R005 collapses to closed (R005-F2/F3/F4 all in review).")
138//! @yah:verify("cargo test -p turso-backup: 38/38 green (up from 31). New: 5 unit tests on replay_wal_onto_main with MockWal (empty-WAL/single-commit/uncommitted-tail-dropped/multi-commit-prefix-grows-image/only-uncommitted-frames-returns-main) + 2 live tests on raw_consistent_copy_live (live_db_consistent_copy_without_truncate_round_trips: seed -> checkpoint_truncate -> seed more uncheckpointed -> copy-live -> reopen = 35 rows; live_db_consistent_copy_reflects_new_writes: repeatable copy after subsequent writes).")
139//! @yah:verify("cargo clippy --all-targets -- --deny=warnings: clean.")
140//!
141//! ## R574-F2 — explicit R2 backpressure
142//!
143//! `tail_frames`'s upload loop drains frames through
144//! [`crate::backpressure::put_with_backoff`] against a bounded spill buffer
145//! (see [`drain_frames_with_backpressure`]) instead of putting each frame
146//! directly and bubbling any error raw. Full design in the `backpressure`
147//! module doc; ticket tracked in the W248 relay, not in-source.
148//! @arch:see(.yah/docs/working/W248-wal-streamer-hardening.md)
149//!
150//! ## R574-T4 — explicit RPO knob
151//!
152//! Cadence was entirely implicit in the caller's `tail_frames` invocation
153//! interval (doc §10). `StreamConfig::rpo_target` states that interval as a
154//! number; every `tail_frames` call reports [`RpoStatus`] (age of the last
155//! durably-persisted watermark, and whether that age has drifted past the
156//! target) on its [`StreamOutcome`] so an orchestrator can alert without
157//! reimplementing the bookkeeping. This crate still does not schedule
158//! anything itself — cadence/retention policy stays caller-driven per the
159//! doc's "Not in scope" — `rpo_target` documents the contract the caller's
160//! own scheduler is expected to uphold, and `RpoStatus` is the receipt.
161//! @arch:see(.yah/docs/working/W248-wal-streamer-hardening.md)
162//!
163//! ## R761-T1 — a measured default tail cadence
164//!
165//! [`DEFAULT_TAIL_INTERVAL`] (60 s) and [`DEFAULT_RPO_TARGET`] (120 s) are
166//! the cadence this crate recommends, derived from
167//! `examples/tail_sweep_harness.rs`'s 2026-08-13 sweep rather than chosen:
168//! every uploading `tail_frames` call writes two fixed objects (generation
169//! manifest + watermark CAS) on top of the frames, so per-write cost was
170//! `frames/write + 2/writes_per_tail` — 4.03 PUTs/write at one write per
171//! tail, 2.05 at a hundred. The constants' docs carry the full table and
172//! the reasoning for landing at 60 s instead of chasing the last 3 %. Still
173//! no scheduler in this crate; these are numbers for the caller's.
174//!
175//! R761-F2 then removed the `frames/write` term those numbers were floored
176//! by (see *Frame batching* above), so the cost is now `3/writes_per_tail`
177//! for a call whose frames fit one batch — the cadence lever and the layout
178//! lever compose, and the second is worth more the longer the interval.
179//! @arch:see(.yah/docs/working/W248-wal-streamer-hardening.md)
180//!
181//! ## R574-F3 — one-puller-per-box fan-out + warm applier
182//!
183//! [`crate::puller::WalPuller`] is the read side's counterpart to
184//! `tail_frames`'s write side: it pulls each newly-uploaded frame from R2
185//! exactly once per box and fans it out in-process to every attached warm
186//! applier (a [`WalInsertSeam`] with its page cache trimmed hard via
187//! [`CoreWalSeam::trim_page_cache_kb`]), so R2 read ops are O(boxes), not
188//! O(replicas). See the `puller` module doc for the full design and its v1
189//! scope cut (attach is cold-start-only; no mid-stream backlog replay).
190//! @arch:see(.yah/docs/working/W248-wal-streamer-hardening.md)
191
192use anyhow::{Context, Result};
193use object_store::path::Path as ObjPath;
194use object_store::{ObjectStore, ObjectStoreExt, PutMode, PutOptions, UpdateVersion};
195use std::collections::{HashSet, VecDeque};
196use std::sync::Arc;
197use std::time::{Duration, SystemTime, UNIX_EPOCH};
198
199use crate::backpressure::{put_with_backoff, BackpressureConfig, BackpressurePolicy, BackpressureReport};
200use crate::snapshot::BackupTarget;
201
202/// WAL frame layout: 24-byte frame header (page_no big-endian at 0..4,
203/// db_size big-endian at 4..8, salts/checksums in the remainder) followed by
204/// `page_size` bytes of page data. Constant per the SQLite WAL format; the
205/// page size is read from the base snapshot's header (page 1, byte 16, `u16`
206/// with the value `1` meaning 65536). `sync_server.rs` hardcodes 4096 — we
207/// don't, see [`StreamConfig::page_size`].
208pub const WAL_FRAME_HEADER_SIZE: usize = 24;
209
210/// Position in the WAL: a `(checkpoint_seq, last_frame)` pair. `max_frame`
211/// resets to 0 every time the WAL header restarts (`WalAutoActions::Restart`
212/// fires), and `checkpoint_seq_no` increments alongside it — so the pair is
213/// the right primary key for sink objects, not raw frame_no.
214///
215/// **R858-B19 — `checkpoint_seq` alone does NOT identify a WAL generation, and
216/// this doc used to claim it did ("increments monotonically across restarts").
217/// That claim is false and it cost a silent wrong restore.** It holds only for
218/// an *in-process* restart, where the same WAL file is reused. When the last
219/// connection to a SQLite database closes, the engine checkpoints and DELETES
220/// the `-wal` file; the next writer creates a fresh WAL back at
221/// checkpoint-sequence `0` with a brand-new random salt. Across that fold
222/// `checkpoint_seq` goes `0 -> 0` over two completely unrelated WALs (measured
223/// 2026-09-06 against system sqlite3 3.51.0 by
224/// `examples/foreign_checkpoint_probe.rs`, probes B and F).
225///
226/// Generation identity therefore lives in [`WalGeneration`], which pairs the
227/// sequence with the WAL header's salt. Never compare two `Watermark`s'
228/// `checkpoint_seq` to decide "same WAL" — that is exactly the inference this
229/// ticket exists to delete. `last_frame` is only meaningful *relative to a
230/// known generation*.
231#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
232pub struct Watermark {
233    pub checkpoint_seq: u32,
234    pub last_frame: u64,
235}
236
237/// R858-B19 — the WAL header's `(salt1, salt2)`, the two 32-bit values SQLite
238/// re-rolls on every WAL reset. Stored big-endian at bytes 16..24 of the
239/// 32-byte WAL header, and copied verbatim into bytes 8..16 of **every frame
240/// header** written under that header — which is where this crate reads it
241/// from, since `turso_core::WalState` exposes only `checkpoint_seq_no` and
242/// `max_frame`.
243///
244/// Why the salt and not the sequence: the salt moves in *both* fold regimes,
245/// and the sequence moves in only one.
246///
247/// - WAL recreated by a writer restart: fresh randomness (measured
248///   `b83c03f5 -> 9ca89e39`, unrelated), while `checkpoint_seq` resets `0 -> 0`.
249/// - In-process autocheckpoint: `salt1` increments in lockstep with the
250///   sequence (measured `d492ea8a -> ... -> d492ea92` alongside seq `0 -> 8`).
251#[derive(Debug, Clone, Copy, PartialEq, Eq)]
252pub struct WalSalt {
253    pub salt1: u32,
254    pub salt2: u32,
255}
256
257impl WalSalt {
258    /// Read the salt out of a WAL **frame** header (bytes 8..16 of the 24-byte
259    /// header, big-endian). Panics on a short slice — every caller here sizes
260    /// its buffer at `WAL_FRAME_HEADER_SIZE + page_size`.
261    pub(crate) fn from_frame_header(frame: &[u8]) -> Self {
262        let be = |o: usize| u32::from_be_bytes(frame[o..o + 4].try_into().unwrap());
263        Self { salt1: be(8), salt2: be(12) }
264    }
265
266    /// R858-B18 — read the salt out of the **32-byte `-wal` file header**
267    /// (bytes 16..24, big-endian), the copy SQLite writes once per WAL
268    /// generation and duplicates into every frame header.
269    ///
270    /// The two constructors differ only in offset, and they live together
271    /// deliberately: one type, one place, so the two ways this crate can reach
272    /// the same value can never disagree. [`from_frame_header`](Self::from_frame_header)
273    /// is the seam-only path ([`read_wal_salt`] — works through a
274    /// [`CoreWalSeam::from_conn`] that has no path); this one is the
275    /// path-only path ([`SourceFingerprint`] — works with no engine open at
276    /// all, which is what makes it usable as an independent check *on* a copy
277    /// the engine took).
278    ///
279    /// Panics on a slice shorter than 24 bytes; [`WalFileHeader::parse`] is the
280    /// length-checked entry point every caller here actually uses.
281    pub(crate) fn from_wal_file_header(hdr: &[u8]) -> Self {
282        let be = |o: usize| u32::from_be_bytes(hdr[o..o + 4].try_into().unwrap());
283        Self { salt1: be(16), salt2: be(20) }
284    }
285}
286
287impl std::fmt::Display for WalSalt {
288    fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
289        write!(f, "{:08x}/{:08x}", self.salt1, self.salt2)
290    }
291}
292
293/// R858-B19 — the identity of one WAL generation: the checkpoint sequence
294/// **and** the header salt that actually distinguishes it. This is what
295/// [`tail_frames`] compares across calls, what the watermark sidecar persists,
296/// and what a generation manifest stamps.
297///
298/// `salt: None` means **unknown generation**, and unknown is a value here, not
299/// a missing one: it arises from a sidecar or manifest written before this
300/// field existed, or from a WAL with no frames to read a salt out of. An
301/// unknown generation is never provably equal to anything — see
302/// [`WalGeneration::is_provably_same_as`] — so it forces a restart on the write
303/// side and a refusal on the restore side rather than a guess.
304#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
305pub struct WalGeneration {
306    pub checkpoint_seq: u32,
307    pub salt: Option<WalSalt>,
308}
309
310impl WalGeneration {
311    /// True only when both sides carry a **known** salt and every component
312    /// agrees — i.e. only when these are *provably* the same WAL.
313    ///
314    /// The asymmetry is the whole point. "Not provably the same" is treated as
315    /// "different", which costs a redundant re-upload (or a loud restore
316    /// refusal) in the worst case. The opposite default — "no evidence of a
317    /// change, so assume it is the same WAL" — is what spliced two WAL
318    /// generations into one chain and restored a plausible wrong image.
319    pub fn is_provably_same_as(&self, other: &WalGeneration) -> bool {
320        match (self.salt, other.salt) {
321            (Some(a), Some(b)) => a == b && self.checkpoint_seq == other.checkpoint_seq,
322            _ => false,
323        }
324    }
325
326    /// Render for an error message: `seq 4 salt 1109ca5e/7acf42a3`, or
327    /// `seq 4 salt <unknown>` when the salt was never recorded.
328    pub(crate) fn describe(&self) -> String {
329        match self.salt {
330            Some(s) => format!("seq {} salt {s}", self.checkpoint_seq),
331            None => format!("seq {} salt <unknown>", self.checkpoint_seq),
332        }
333    }
334}
335
336/// R858-B18 — the 32-byte `-wal` file header, read straight off disk with no
337/// engine open. Ground truth about which WAL generation is on disk and how far
338/// it has been written, independent of anything turso caches.
339///
340/// Only the three fields that move are kept. The rest of the header (magic,
341/// format version, the two header checksums) is either constant for a given
342/// build or a function of these; a change to any of it that did *not* move one
343/// of these three would not be a change this crate can act on.
344#[derive(Debug, Clone, Copy, PartialEq, Eq)]
345pub struct WalFileHeader {
346    /// Page size the WAL's frames carry (bytes 8..12).
347    pub page_size: u32,
348    /// Checkpoint sequence (bytes 12..16) — the field `tail_frames` used to key
349    /// restart detection on before [`WalGeneration`] paired it with the salt.
350    pub checkpoint_seq: u32,
351    /// The generation's salt (bytes 16..24).
352    pub salt: WalSalt,
353}
354
355impl WalFileHeader {
356    /// SQLite's WAL header is exactly this many bytes, ahead of frame 1.
357    pub const SIZE: usize = 32;
358
359    /// Parse a WAL header out of the first [`Self::SIZE`] bytes of a `-wal`
360    /// file. `None` for anything shorter — a truncated or freshly-created WAL
361    /// has no generation to name yet, which is an honest unknown rather than an
362    /// error (see [`WalGeneration`]'s `salt: None`).
363    pub fn parse(bytes: &[u8]) -> Option<Self> {
364        if bytes.len() < Self::SIZE {
365            return None;
366        }
367        let be = |o: usize| u32::from_be_bytes(bytes[o..o + 4].try_into().unwrap());
368        Some(Self {
369            page_size: be(8),
370            checkpoint_seq: be(12),
371            salt: WalSalt::from_wal_file_header(bytes),
372        })
373    }
374
375    /// Read `{db_path}-wal`'s header. `Ok(None)` when the sidecar is absent
376    /// (WAL folded and deleted, or journal_mode != WAL) or too short to parse.
377    pub fn read(db_path: &str) -> Result<Option<Self>> {
378        match std::fs::File::open(format!("{db_path}-wal")) {
379            Ok(mut f) => {
380                let mut buf = [0u8; Self::SIZE];
381                let mut filled = 0usize;
382                loop {
383                    match std::io::Read::read(&mut f, &mut buf[filled..]) {
384                        Ok(0) => break,
385                        Ok(n) => filled += n,
386                        Err(e) if e.kind() == std::io::ErrorKind::Interrupted => continue,
387                        Err(e) => {
388                            return Err(e).with_context(|| format!("reading {db_path}-wal header"))
389                        }
390                    }
391                    if filled == Self::SIZE {
392                        break;
393                    }
394                }
395                Ok(Self::parse(&buf[..filled]))
396            }
397            Err(e) if e.kind() == std::io::ErrorKind::NotFound => Ok(None),
398            Err(e) => Err(e).with_context(|| format!("opening {db_path}-wal")),
399        }
400    }
401}
402
403/// R858-B18 — everything about a source database's **files** that must hold
404/// still for a copy of them to be a point-in-time image, sampled with no engine
405/// open and no lock taken.
406///
407/// This is the other half of reading a foreign-written database. A ReadOnly
408/// open ([`CoreWalSeam::open_reader`]) gets us in the door without stealing the
409/// whole-file lock, but it buys **no shared locking protocol** with upstream C
410/// SQLite: turso locks whole files with `fcntl`, C SQLite uses byte-range locks
411/// plus the `-shm` WAL index, and neither engine observes the other's. So a
412/// foreign checkpoint landing in the middle of our copy can fold WAL frames
413/// into the main file we have already half-read, and the result is a torn image
414/// that still passes `PRAGMA integrity_check`.
415///
416/// Rather than reimplement SQLite's reader protocol (registering a read-mark in
417/// the `-shm` WAL index — a research project and a permanent compatibility
418/// liability against an engine we do not control), [`raw_consistent_copy_live`]
419/// uses textbook **optimistic validation**: sample this before the copy, sample
420/// it again after, and accept the copy only if nothing moved. That converts
421/// "may silently read torn state" into "detects torn state and refuses", needs
422/// no cooperation from the foreign engine, and costs two stats and a 132-byte
423/// read per attempt.
424///
425/// ## Why these fields
426///
427/// - `wal` (salt + checkpoint_seq) moves on **every** WAL reset, in both fold
428///   regimes — fresh randomness on a writer restart, `salt1` incrementing on an
429///   in-process autocheckpoint (both measured; see [`WalSalt`]).
430/// - `main_len` moves when a checkpoint grows the main database.
431/// - `change_counter` (main header bytes 24..28) moves on every write to the
432///   main file — i.e. on every checkpoint — even one that leaves its length
433///   alone. It is meaningful here **only because the foreign writer is C
434///   SQLite**: turso does not maintain this field (it stays `1` in every
435///   journal mode, verified — see `snapshot.rs`'s two-gate rationale), which is
436///   exactly why the WAL salt carries the weight and this one is corroboration.
437/// - `wal_len` is recorded for the report but deliberately **not** part of the
438///   accept/reject test — see [`Self::stable_across`], which is the comparison
439///   to use. There is no `PartialEq` on this type on purpose: a bare `==` would
440///   silently include `wal_len` and refuse every copy taken while the
441///   application was merely writing.
442#[derive(Debug, Clone, Copy)]
443pub struct SourceFingerprint {
444    /// Length of the main database file.
445    pub main_len: u64,
446    /// The main header's change counter, or `None` when the file is too short
447    /// to carry a SQLite header at all (a database whose page 1 still lives
448    /// only in the WAL). Unknown-and-unknown compares equal, which is safe
449    /// because `main_len` participates in the same comparison.
450    pub change_counter: Option<u32>,
451    /// The `-wal` header, or `None` when there is no WAL sidecar.
452    pub wal: Option<WalFileHeader>,
453    /// Length of the `-wal` file (`0` when absent).
454    pub wal_len: u64,
455}
456
457impl SourceFingerprint {
458    /// Offset of the change counter in SQLite's 100-byte database header.
459    const CHANGE_COUNTER_OFFSET: usize = 24;
460
461    /// Sample the fingerprint of the database at `db_path`. Touches nothing:
462    /// two metadata calls plus a 28-byte and a 32-byte read.
463    pub fn read(db_path: &str) -> Result<Self> {
464        let mut hdr = [0u8; Self::CHANGE_COUNTER_OFFSET + 4];
465        let main_len = match std::fs::File::open(db_path) {
466            Ok(mut f) => {
467                let len = f
468                    .metadata()
469                    .with_context(|| format!("stat {db_path}"))?
470                    .len();
471                if len as usize >= hdr.len() {
472                    std::io::Read::read_exact(&mut f, &mut hdr)
473                        .with_context(|| format!("reading {db_path} header"))?;
474                }
475                len
476            }
477            Err(e) => return Err(e).with_context(|| format!("opening source db {db_path}")),
478        };
479        let change_counter = (main_len as usize >= hdr.len()).then(|| {
480            u32::from_be_bytes(hdr[Self::CHANGE_COUNTER_OFFSET..].try_into().unwrap())
481        });
482        Ok(Self {
483            main_len,
484            change_counter,
485            wal: WalFileHeader::read(db_path)?,
486            wal_len: std::fs::metadata(format!("{db_path}-wal"))
487                .map(|m| m.len())
488                .unwrap_or(0),
489        })
490    }
491
492    /// True when nothing that can **tear** a copy moved between `self` (sampled
493    /// before) and `after` (sampled after). This is the accept test in
494    /// [`raw_consistent_copy_live`], and it is narrower than field equality on
495    /// purpose.
496    ///
497    /// ## What can tear the copy, and what cannot
498    ///
499    /// The copy reads the main file, then replays WAL frames `1..=max_frame`
500    /// captured when the seam opened. Against that algorithm:
501    ///
502    /// - **A checkpoint tears it.** It rewrites pages of the main file *and*
503    ///   resets the WAL, so our already-read main bytes and our frame reads can
504    ///   straddle the fold — replaying pre-fold frames over post-fold pages
505    ///   rolls pages backwards. Caught: a checkpoint bumps `change_counter`
506    ///   and/or `main_len`, and a WAL restart re-rolls the salt and the sequence
507    ///   (measured in both regimes, `examples/foreign_checkpoint_probe.rs`
508    ///   probes A and E).
509    /// - **A plain append does NOT tear it.** SQLite only ever appends frames
510    ///   within a generation, and only a reset (which re-rolls `salt1`) lets it
511    ///   overwrite an existing frame. So frames `1..=max_frame` are immutable
512    ///   for as long as the salt holds, and a writer that commits during our
513    ///   copy just means our image is a slightly earlier point in time — which
514    ///   is what a point-in-time copy *is*.
515    ///
516    /// Which is why `wal_len` and the frame count are excluded. Including them
517    /// buys no additional safety and costs a refusal on every copy taken while
518    /// the application is writing at all — turning a working backup into one
519    /// that only succeeds against an idle database. R858-B18 measured that
520    /// difference rather than assuming it; see probe H.
521    pub fn stable_across(&self, after: &Self) -> bool {
522        self.main_len == after.main_len
523            && self.change_counter == after.change_counter
524            && self.wal == after.wal
525    }
526
527    /// One-line rendering for the refusal message, so a failure names what
528    /// actually moved instead of asserting "something did".
529    pub(crate) fn describe(&self) -> String {
530        let wal = match self.wal {
531            Some(h) => format!("seq {} salt {}", h.checkpoint_seq, h.salt),
532            None => "<no WAL>".to_string(),
533        };
534        format!(
535            "main {}B change_counter {} | wal {}B {wal}",
536            self.main_len,
537            self.change_counter.map_or("<unknown>".to_string(), |c| c.to_string()),
538            self.wal_len,
539        )
540    }
541}
542
543/// Subset of `turso_core::Connection`'s `feature = "conn_raw_api"` surface
544/// we use. Every WAL call into turso goes through this trait. A future
545/// signature shift becomes a one-impl delta.
546pub trait WalSeam {
547    /// Snapshot the WAL position (checkpoint seq + max_frame).
548    fn wal_state(&self) -> Result<Watermark>;
549
550    /// Fetch frame `frame_no` (1-based) into `buf` (must be
551    /// `WAL_FRAME_HEADER_SIZE + page_size` bytes). Returns the page number
552    /// the frame applies to and the post-frame DB size (in pages) — non-zero
553    /// `db_size` marks a commit frame.
554    fn wal_get_frame(&self, frame_no: u64, buf: &mut [u8]) -> Result<FrameInfo>;
555
556    /// Take WAL ownership: turn off both the auto-checkpoint and the
557    /// auto-WAL-restart so our watermark stays meaningful across calls. The
558    /// in-tree consumer (`cli/sync_server.rs`) makes the same move at startup.
559    fn wal_auto_actions_disable(&self);
560}
561
562/// Restore-side counterpart to [`WalSeam`]: the three `wal_insert_*` calls
563/// `cli/sync_server.rs` clients use to replay frames into a fresh DB. Same
564/// rationale — one impl block to update if upstream renames.
565///
566/// Ordering contract: `begin` → N × `insert_frame(monotonic frame_no)` → `end`.
567/// `end(false)` is the crash-safe default: the engine rolls back any suffix of
568/// frames written after the last commit frame (`db_size > 0`) in the session.
569pub trait WalInsertSeam {
570    /// Open a write transaction with auto-checkpoint/restart suppressed so our
571    /// monotonically-numbered inserts aren't reshuffled mid-session.
572    fn wal_insert_begin(&self) -> Result<()>;
573
574    /// Insert `frame` (a 24-byte header + `page_size` bytes of page data) at
575    /// position `frame_no` (1-based, must be exactly `prev + 1` — gaps error).
576    /// Identical content at an already-written position is a no-op (the engine
577    /// compares and returns OK).
578    fn wal_insert_frame(&self, frame_no: u64, frame: &[u8]) -> Result<()>;
579
580    /// Close the session. With `force_commit = false` (the restore default) the
581    /// engine drops any frames after the last commit frame — automatic
582    /// crash-consistency for a tail captured mid-transaction. `force_commit = true`
583    /// commits even an uncommitted suffix; not used by restore.
584    fn wal_insert_end(&self, force_commit: bool) -> Result<()>;
585}
586
587/// Frame-header projection (what we need for streaming + commit detection).
588/// Mirrors `turso_core::types::WalFrameInfo` without re-exporting the type.
589#[derive(Debug, Clone, Copy, PartialEq, Eq)]
590pub struct FrameInfo {
591    pub page_no: u32,
592    /// Number of pages in the DB after this frame's commit, or `0` if the
593    /// frame is mid-transaction.
594    pub db_size: u32,
595}
596
597impl FrameInfo {
598    pub fn is_commit_frame(&self) -> bool {
599        self.db_size > 0
600    }
601}
602
603/// Real WAL seam backed by a live `turso_core::Connection`.
604pub struct CoreWalSeam {
605    conn: Arc<turso_core::Connection>,
606}
607
608impl CoreWalSeam {
609    /// Open a **writable** connection at `path` and disable auto-checkpoint /
610    /// auto-restart so the caller owns WAL maintenance. Requires `turso_core`
611    /// with `features = ["conn_raw_api"]` (set in this crate's Cargo.toml).
612    ///
613    /// This takes turso's **whole-file exclusive `fcntl` lock**, so it is
614    /// mutually exclusive with any other process holding the database open —
615    /// including upstream C SQLite (measured: `examples/foreign_checkpoint_probe.rs`
616    /// probes D and G). That is correct for the restore/apply direction, which
617    /// owns the destination file outright, and wrong for reading a live source:
618    /// use [`Self::open_reader`] there.
619    pub fn open(path: &str) -> Result<Self> {
620        let io: Arc<dyn turso_core::IO> =
621            Arc::new(turso_core::PlatformIO::new().context("creating turso_core PlatformIO")?);
622        let db = turso_core::Database::open_file(io, path)
623            .with_context(|| format!("opening turso_core db {path}"))?;
624        let conn = db.connect().context("connecting to turso_core db")?;
625        // Take WAL ownership — same first move sync_server.rs makes.
626        conn.wal_auto_actions_disable();
627        Ok(Self { conn })
628    }
629
630    /// R858-B18 — open a **read-only** connection at `path`, taking **no**
631    /// whole-file lock. This is the constructor the backup direction wants: a
632    /// backup is a reader, and it must not lock out the application whose
633    /// database it is reading.
634    ///
635    /// `OpenFlags::ReadOnly` is what buys that, per handle and with no
636    /// process-wide effect. `turso_core-0.7.2/io/unix.rs:67-72` takes the
637    /// exclusive lock only when
638    /// `env::var(ENV_DISABLE_FILE_LOCK).is_err() && !flags.intersects(ReadOnly | NoLock)`;
639    /// `io_uring.rs:463` and `windows.rs:305` carry the identical condition, so
640    /// this is not a unix-only accident. The `LIMBO_DISABLE_FILE_LOCK=1` escape
641    /// hatch reaches the same no-lock state but does it for **every** open in
642    /// the process, including the writable ones — never use it here.
643    ///
644    /// ## Two things this does NOT give you
645    ///
646    /// 1. **No shared locking protocol.** Getting in without a lock is not
647    ///    coordination: a foreign checkpoint can still land mid-read. That is
648    ///    what [`SourceFingerprint`] validation is for, and why
649    ///    [`raw_consistent_copy_live`] pairs the two rather than shipping this
650    ///    flag alone — the flag alone converts "refuses to open" into "may
651    ///    return a torn image", which is strictly worse.
652    /// 2. **No escape from turso's process-global registry.**
653    ///    `Database::open_file_with_flags` consults `DATABASE_MANAGER`, keyed by
654    ///    file id, *before* it looks at the flags (`lib.rs:940`), and hands back
655    ///    an already-open `Database` with its own flags discarded. So in a
656    ///    process that already holds a writable handle on this exact file, this
657    ///    call returns that writable handle — same as `snapshot.rs`'s
658    ///    `upload_base_snapshot` doc records. Cross-process (the headscale case)
659    ///    is unaffected: the registry is per-process.
660    pub fn open_reader(path: &str) -> Result<Self> {
661        let io: Arc<dyn turso_core::IO> =
662            Arc::new(turso_core::PlatformIO::new().context("creating turso_core PlatformIO")?);
663        let db = turso_core::Database::open_file_with_flags(
664            io,
665            path,
666            turso_core::OpenFlags::ReadOnly,
667            turso_core::DatabaseOpts::new(),
668            None,
669        )
670        .with_context(|| format!("opening turso_core db {path} read-only"))?;
671        let conn = db.connect().context("connecting to turso_core db")?;
672        // Belt and braces: a read-only connection cannot checkpoint anyway, but
673        // the seam contract is that nothing we hold folds the WAL under us.
674        conn.wal_auto_actions_disable();
675        Ok(Self { conn })
676    }
677
678    /// Wrap an already-open connection. The caller is responsible for having
679    /// called `wal_auto_actions_disable()` on it.
680    pub fn from_conn(conn: Arc<turso_core::Connection>) -> Self {
681        Self { conn }
682    }
683}
684
685impl WalSeam for CoreWalSeam {
686    fn wal_state(&self) -> Result<Watermark> {
687        let s = self.conn.wal_state().context("turso_core wal_state")?;
688        Ok(Watermark {
689            checkpoint_seq: s.checkpoint_seq_no,
690            last_frame: s.max_frame,
691        })
692    }
693
694    fn wal_get_frame(&self, frame_no: u64, buf: &mut [u8]) -> Result<FrameInfo> {
695        let info = self
696            .conn
697            .wal_get_frame(frame_no, buf)
698            .with_context(|| format!("turso_core wal_get_frame({frame_no})"))?;
699        Ok(FrameInfo {
700            page_no: info.page_no,
701            db_size: info.db_size,
702        })
703    }
704
705    fn wal_auto_actions_disable(&self) {
706        self.conn.wal_auto_actions_disable();
707    }
708}
709
710impl WalInsertSeam for CoreWalSeam {
711    fn wal_insert_begin(&self) -> Result<()> {
712        self.conn
713            .wal_insert_begin()
714            .context("turso_core wal_insert_begin")
715    }
716
717    fn wal_insert_frame(&self, frame_no: u64, frame: &[u8]) -> Result<()> {
718        self.conn
719            .wal_insert_frame(frame_no, frame)
720            .with_context(|| format!("turso_core wal_insert_frame({frame_no})"))?;
721        Ok(())
722    }
723
724    fn wal_insert_end(&self, force_commit: bool) -> Result<()> {
725        self.conn
726            .wal_insert_end(force_commit)
727            .context("turso_core wal_insert_end")
728    }
729}
730
731impl CoreWalSeam {
732    /// Resize this connection's page cache toward `target_kb` kilobytes via
733    /// the standard `PRAGMA cache_size` surface (negative value = KB, per
734    /// SQLite's own convention — see `turso_core::translate::pragma`'s
735    /// `update_cache_size`) rather than reaching into `turso_core`'s
736    /// `Pager::change_page_cache_size` / `CacheResizeResult` directly: the
737    /// latter are technically reachable (`Connection::get_pager()` and
738    /// `Pager`/`Page`/`PageRef` are all `pub use`d at the crate root) but
739    /// `CacheResizeResult` itself is not re-exported, so calling it from
740    /// here would mean handling a value of an unnameable type. The pragma
741    /// path exercises the exact same resize logic through turso_core's own
742    /// public, documented SQL surface instead.
743    ///
744    /// R574-F3: a warm applier trims its cache hard immediately on attach —
745    /// it never serves reads, so cached pages are pure standing RSS cost
746    /// (see the R574-T1 measurement this sizes against).
747    pub(crate) fn trim_page_cache_kb(&self, target_kb: i64) -> Result<()> {
748        anyhow::ensure!(
749            target_kb > 0,
750            "trim_page_cache_kb: target_kb must be positive, got {target_kb}"
751        );
752        self.conn
753            .execute(format!("PRAGMA cache_size = -{target_kb}"))
754            .with_context(|| format!("PRAGMA cache_size = -{target_kb}"))
755    }
756}
757
758/// R761-T1: the tail cadence this crate recommends when a caller has no
759/// reason of its own to pick a different one — **60 s**, and the number is
760/// a measurement result, not a guess.
761///
762/// ## The measurement
763///
764/// `examples/tail_sweep_harness.rs`, run 2026-08-13 against a casual-app
765/// workload (single-row transactions into a table with a secondary index),
766/// sweeping `writes_per_tail` — how many application writes accumulate
767/// between two `tail_frames` calls:
768///
769/// | writes_per_tail | PUTs/write | frames/write |
770/// |----------------:|-----------:|-------------:|
771/// |               1 |       4.03 |         2.03 |
772/// |               5 |       2.43 |         2.03 |
773/// |              25 |       2.11 |         2.03 |
774/// |             100 |       2.05 |         2.03 |
775///
776/// Those four points are not four independent facts. Every `tail_frames`
777/// call that uploads anything writes exactly **two fixed objects** beyond
778/// the frames themselves — one generation manifest and one watermark CAS —
779/// so
780///
781/// ```text
782/// PUTs/write = frames/write + 2 / writes_per_tail
783/// ```
784///
785/// which reproduces all four measured rows to the last digit: `2.03 + 2/1`,
786/// `2.03 + 2/5`, `2.03 + 2/25`, `2.03 + 2/100`. frames/write is FLAT across
787/// the sweep — it is a property of the schema and transaction shape, and
788/// cadence cannot touch it. The entire lever this default pulls is the
789/// `2 / writes_per_tail` term.
790///
791/// ## Why 60 s and not longer
792///
793/// That term has sharply diminishing returns. Of the total reduction
794/// available (4.03 → 2.05), moving from `writes_per_tail` 1 → 5 captures
795/// 81 %, 1 → 25 captures 97 %, and everything from 25 → 100 is the last
796/// 3 %. So the target is the 25 region, not 100.
797///
798/// Converting that writes axis into a time interval needs a write RATE,
799/// and the honest one is the rate DURING an active session — W313 §3.1's
800/// "~10 writes/day" arrives clustered in bursts, not spread evenly, and an
801/// idle tenant costs nothing at any cadence (a tail call with no new frames
802/// uploads no objects at all, so a long interval only ever helps a tenant
803/// that is actively writing). Across the burst rates a casual app produces,
804/// ~0.1–1 write/s:
805///
806/// - 15 s (what a 30 s RPO bound derives today) spans 1.5–15 writes/tail
807///   → 3.36–2.16 PUTs/write. The slow-burst end is the worst case the
808///   ticket is named after, and it is nearly the full 4.03.
809/// - **60 s spans 6–60 writes/tail → 2.36–2.06 PUTs/write.**
810/// - 300 s spans 30–300 → 2.10–2.04: at most 0.26 PUTs/write better than
811///   60 s, bought with a 5× wider data-loss window on exactly the tenants
812///   that are actively writing. Not a trade worth making for 3 % of a
813///   bill.
814///
815/// 60 s is where the curve has flattened but the exposure window is still
816/// something an operator can say out loud.
817///
818/// ## What it did not fix, and what did
819///
820/// The 2.03 frames/write floor was untouchable from here — it was one object
821/// per WAL frame, and cadence cannot amortize a per-frame cost. **R761-F2
822/// removed it** by uploading a tail call's frames as one ranged object, so
823/// the curve above is now
824///
825/// ```text
826/// PUTs/write = 3 / writes_per_tail
827/// ```
828///
829/// (one batch object + manifest + watermark), i.e. 3.00 / 0.60 / 0.12 / 0.03
830/// at the same four points — a 26 % cut at `writes_per_tail = 1` and 98.5 %
831/// at 100. **That does not move this default**, and the reasoning above is
832/// why rather than an accident: the 60 s choice was made against the shape
833/// of the curve, and what batching changes is its scale. Over the same
834/// 0.1–1 write/s burst band, going 60 s → 300 s now buys 0.50 → 0.10
835/// PUTs/write at the slow end and 0.05 → 0.01 at the fast one — against a
836/// pre-batching worst case of 3.36, an absolute difference small enough
837/// that RPO exposure is the only term still worth optimizing here. The
838/// measured points above are kept as the pre-batching baseline the
839/// reduction is stated against.
840///
841/// Re-run the harness and revisit both numbers whenever schema shape, page
842/// size, or turso's WAL behaviour changes:
843/// `cargo run -p turso-backup --example tail_sweep_harness`.
844///
845/// This crate still schedules nothing itself (R574-T4's "document the
846/// contract" choice stands) — the constant is the number a caller's
847/// scheduler should adopt absent a reason not to, and
848/// [`DEFAULT_RPO_TARGET`] is the bound that goes with it.
849pub const DEFAULT_TAIL_INTERVAL: Duration = Duration::from_secs(60);
850
851/// R761-T1: the RPO bound that goes with [`DEFAULT_TAIL_INTERVAL`] — twice
852/// it, because tailing *at* the bound makes ordinary scheduling jitter read
853/// as a breach, and tailing at half of it means one missed tick still lands
854/// inside the promise. See [`DEFAULT_TAIL_INTERVAL`] for why the cadence is
855/// 60 s; this is that number expressed as the promise rather than the
856/// mechanism, and it is what belongs in [`StreamConfig::rpo_target`].
857pub const DEFAULT_RPO_TARGET: Duration = Duration::from_secs(120);
858
859/// Configuration for a streaming session.
860pub struct StreamConfig<'a> {
861    /// Object-store key of the base tier-1a snapshot the frames replay onto.
862    /// Recorded in every generation manifest; restore re-fetches it.
863    pub base_snapshot_key: &'a str,
864    /// Page size of the base snapshot — read it from the snapshot header
865    /// (offset 16, `u16` big-endian; the on-disk value `1` means 65 536).
866    /// Required because the seam does not return it and `sync_server.rs`'s
867    /// 4 KB hardcode is the wrong default to inherit.
868    pub page_size: usize,
869    /// R574-F2: bounded spill buffer + overflow policy + 429/503 backoff
870    /// for the R2 upload side of the drain loop. `Default` is
871    /// behavior-preserving (`BackpressurePolicy::Fail` — errors bubble
872    /// immediately, same as before this field existed).
873    pub backpressure: BackpressureConfig,
874    /// R574-T4: the stated RPO bound — the caller's own scheduler is
875    /// expected to invoke `tail_frames` often enough that the watermark
876    /// never goes stale past this. `None` (the default) means no target is
877    /// asserted; [`RpoStatus::breached`] is always `false` in that case,
878    /// but [`RpoStatus::watermark_age`] is still reported so a caller can
879    /// observe the real gap before picking a number.
880    ///
881    /// R761-T1: [`DEFAULT_RPO_TARGET`] (120 s, tailed at
882    /// [`DEFAULT_TAIL_INTERVAL`] = 60 s) is the number to put here absent a
883    /// reason to pick another — its doc carries the measured write-op table
884    /// that justifies it. `None` stays the default so that a caller who has
885    /// not thought about RPO never gets a breach flag it did not ask for.
886    pub rpo_target: Option<Duration>,
887    /// R732-F2 (W245): this writer's **fencing token** — the per-tenant epoch
888    /// handed out by yubaba's raft state machine
889    /// (`YubabaState::tenant_fencing_token`). It is stamped into every frame
890    /// key, every generation manifest, and the watermark sidecar, and it is
891    /// checked before a single frame is uploaded: a writer whose epoch is
892    /// *lower* than the one already recorded at the sink is a stale owner and
893    /// bounces with [`StreamOutcome::Fenced`] rather than interleaving its
894    /// frames with the real owner's.
895    ///
896    /// **`0` means unfenced**, and is the behaviour-preserving default for a
897    /// single-writer deployment that has no ownership authority to ask. Note
898    /// that unfenced is not exempt: a `0` writer is still fenced by any sink
899    /// already stamped with a real epoch, which is exactly what should happen
900    /// when a tenant has been claimed and a legacy streamer is still running.
901    pub epoch: u64,
902    /// R732-F2: opaque label for *who* holds `epoch` — a yubaba node id, a
903    /// hostname, whatever the caller finds useful. Recorded in the manifest
904    /// and never interpreted here. Purely diagnostic: when you are staring at
905    /// a fenced stream at 3am, "which owner wrote generation 7" is the first
906    /// question, and the epoch alone does not answer it.
907    pub owner: Option<&'a str>,
908    /// R736-T2 (W250): the **cross-cell** fence — this writer's belief about
909    /// the global tenant→cell pointer's generation
910    /// (`yah_tenant_pointer::PointerRecord::generation`). `epoch` alone
911    /// fences *within* one raft group; it says nothing when ownership moves
912    /// to a different cell, because the two cells run independent raft
913    /// groups that share no epoch counter. Checked alongside `epoch` before
914    /// any frame is uploaded: a writer whose generation is *lower* than the
915    /// one already recorded at the sink is a stale cell and bounces with
916    /// [`StreamOutcome::Fenced`], exactly like a stale epoch.
917    ///
918    /// **`0` means unfenced** — the same behaviour-preserving default as
919    /// `epoch`, for a single-cell deployment that has no global pointer to
920    /// ask. This mirrors `yah_tenant_pointer`'s `FIRST_GENERATION = 1`: `0`
921    /// is what an unfenced writer defaults to and must never be mistakable
922    /// for a live cross-cell owner. The caller (yubaba's control plane)
923    /// resolves the pointer and hands the generation in as a plain `u64` —
924    /// this crate does not link `yah_tenant_pointer` to read it itself.
925    pub pointer_generation: u64,
926}
927
928/// What a [`tail_frames`] call did.
929#[derive(Debug, Clone, PartialEq, Eq)]
930pub enum StreamOutcome {
931    /// No new frames since the last tail — sink is already current.
932    Empty {
933        watermark: Watermark,
934        /// R574-T4: staleness of the last durably-persisted watermark —
935        /// still meaningful here, since "nothing new to stream" and "the
936        /// caller's scheduler stopped invoking us" look identical from the
937        /// engine's side and only this field tells them apart.
938        rpo: RpoStatus,
939    },
940    /// Uploaded a contiguous range of frames and wrote a generation manifest.
941    Streamed {
942        generation_key: String,
943        first_frame: u64,
944        last_frame: u64,
945        checkpoint_seq: u32,
946        frame_count: u64,
947        /// R574-F2: backpressure activity during this call (policy,
948        /// high-water spill-buffer occupancy, shed count, throttle retries).
949        backpressure: BackpressureReport,
950        /// R574-T4: see [`StreamOutcome::Empty::rpo`].
951        rpo: RpoStatus,
952    },
953    /// The live WAL is not provably the one the sidecar watermark was taken
954    /// from, so this call re-uploaded frames `1..N` from the top rather than
955    /// resuming.
956    ///
957    /// R858-B19 widened this from "`checkpoint_seq` advanced" to "the
958    /// [`WalGeneration`] did not prove itself unchanged", which is why both
959    /// fields are now generations rather than bare sequence numbers: the
960    /// motivating case is a writer restart where the sequence reads `0 -> 0`
961    /// and only the salt moved. `previous_generation.salt == None` names the
962    /// third case — a sidecar written before the salt existed, restarted
963    /// because it cannot be checked, not because it was seen to change.
964    Restarted {
965        generation_key: String,
966        previous_generation: WalGeneration,
967        new_generation: WalGeneration,
968        first_frame: u64,
969        last_frame: u64,
970        frame_count: u64,
971        /// R574-F2: see [`StreamOutcome::Streamed::backpressure`].
972        backpressure: BackpressureReport,
973        /// R574-T4: see [`StreamOutcome::Empty::rpo`].
974        rpo: RpoStatus,
975    },
976    /// R574-F2: `BackpressurePolicy::Shed` dropped every buffered frame in
977    /// this call before any of them persisted (R2 was throttling harder
978    /// than the spill buffer + backoff could absorb). No manifest/watermark
979    /// was written — the next `tail_frames` call re-attempts the same
980    /// range from the unchanged prior watermark. Distinct from `Empty`,
981    /// which means the engine itself had nothing new.
982    Shed {
983        checkpoint_seq: u32,
984        first_frame: u64,
985        last_frame: u64,
986        backpressure: BackpressureReport,
987        /// R574-T4: see [`StreamOutcome::Empty::rpo`].
988        rpo: RpoStatus,
989    },
990    /// R732-F2 (W245) / R736-T2 (W250): **this writer is a stale owner and
991    /// wrote nothing.** Either the sink's watermark is stamped with an epoch
992    /// higher than [`StreamConfig::epoch`] (ownership moved within the cell),
993    /// or with a pointer generation higher than
994    /// [`StreamConfig::pointer_generation`] (ownership moved to a different
995    /// cell) — the two-level fence bounces on either. Detected before the
996    /// first frame upload, so a fenced call is a pure read — no frames, no
997    /// manifest, no watermark write.
998    ///
999    /// This is the outcome the whole fencing design exists to produce. Without
1000    /// it a partitioned old master and a freshly-promoted new master both
1001    /// stream into the same prefix and silently corrupt each other; with it
1002    /// the loser finds out on its very next tail and can stop.
1003    ///
1004    /// Deliberately carries no [`RpoStatus`]: a fenced writer's view of
1005    /// watermark staleness is not its stream's RPO any more, and reporting one
1006    /// here would page the wrong operator about the wrong node.
1007    Fenced {
1008        /// The epoch recorded at the sink — the real owner's token.
1009        current_epoch: u64,
1010        /// The (lower) epoch this writer tried to stream under.
1011        our_epoch: u64,
1012        /// R736-T2: the pointer generation recorded at the sink.
1013        current_pointer_generation: u64,
1014        /// R736-T2: the (lower) generation this writer tried to stream under.
1015        our_pointer_generation: u64,
1016    },
1017}
1018
1019impl StreamOutcome {
1020    /// This call's RPO snapshot, or `None` for [`StreamOutcome::Fenced`] —
1021    /// see that variant's doc for why it deliberately carries none.
1022    ///
1023    /// R782: the accessor a caller (`tenant-streamer`'s tail loop) uses to
1024    /// push `watermark_age` onward without re-deriving this match on every
1025    /// call site that needs it.
1026    pub fn rpo(&self) -> Option<&RpoStatus> {
1027        match self {
1028            StreamOutcome::Empty { rpo, .. }
1029            | StreamOutcome::Streamed { rpo, .. }
1030            | StreamOutcome::Restarted { rpo, .. }
1031            | StreamOutcome::Shed { rpo, .. } => Some(rpo),
1032            StreamOutcome::Fenced { .. } => None,
1033        }
1034    }
1035
1036    /// This call's backpressure activity, or `None` for the two outcomes that
1037    /// never reached the drain loop ([`StreamOutcome::Empty`] had nothing to
1038    /// send, [`StreamOutcome::Fenced`] was refused before the first frame).
1039    ///
1040    /// R760-B8: the accessor a *multi-tenant* caller needs. `Streamed` is not
1041    /// the same thing as "the sink took everything" — under
1042    /// [`BackpressurePolicy::Shed`] a call that persisted a partial prefix and
1043    /// dropped the rest reports `Streamed` with a nonzero
1044    /// [`BackpressureReport::frames_shed`], and a caller that only matches the
1045    /// variant cannot tell that apart from a clean tail. roadcase's shard
1046    /// flusher reads this on every arm so a struggling cell is nameable from
1047    /// its own metrics rather than by bisecting tenants.
1048    pub fn backpressure(&self) -> Option<&BackpressureReport> {
1049        match self {
1050            StreamOutcome::Streamed { backpressure, .. }
1051            | StreamOutcome::Restarted { backpressure, .. }
1052            | StreamOutcome::Shed { backpressure, .. } => Some(backpressure),
1053            StreamOutcome::Empty { .. } | StreamOutcome::Fenced { .. } => None,
1054        }
1055    }
1056}
1057
1058/// R574-T4: watermark-staleness snapshot for RPO drift alerting, computed
1059/// fresh on every `tail_frames` call against [`StreamConfig::rpo_target`].
1060#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
1061pub struct RpoStatus {
1062    /// The stated bound from `StreamConfig::rpo_target`, echoed back for
1063    /// convenience (so a caller reading only the outcome still knows what
1064    /// was being enforced).
1065    pub target: Option<Duration>,
1066    /// Elapsed time since the watermark last durably advanced, measured at
1067    /// the start of this call. `None` when the sink has never persisted a
1068    /// watermark — there is nothing yet to measure staleness against.
1069    pub watermark_age: Option<Duration>,
1070    /// `true` when both `target` and `watermark_age` are set and
1071    /// `watermark_age > target` — the RPO bound is breached and an
1072    /// orchestrator should alert. Always `false` when no target is
1073    /// configured.
1074    pub breached: bool,
1075}
1076
1077impl BackupTarget {
1078    pub(crate) fn watermark_key(&self) -> ObjPath {
1079        join_key(&self.prefix, "latest.stream-watermark")
1080    }
1081
1082    /// R732-F2: frames are namespaced by the writer's fencing epoch, so two
1083    /// owners at different epochs cannot land on the same key even if they
1084    /// somehow both get as far as uploading. The epoch check in
1085    /// [`tail_frames`] is the guard; this key shape is the backstop that makes
1086    /// a guard failure recoverable (both owners' frames survive and the
1087    /// manifests say who wrote what) instead of a silent overwrite.
1088    ///
1089    /// Epoch `0` keeps the original two-level layout. That is not cosmetic:
1090    /// backups written before fencing existed are still restorable because
1091    /// their manifests say `epoch 0` and land back on this branch.
1092    ///
1093    /// R761-F2: this is the *read-side legacy* key shape now — nothing writes
1094    /// one-object-per-frame any more. See [`Self::frame_batch_key`].
1095    pub(crate) fn frame_key(&self, epoch: u64, checkpoint_seq: u32, frame_no: u64) -> ObjPath {
1096        let suffix = if epoch == 0 {
1097            format!("frames/{checkpoint_seq:010}/{frame_no:020}")
1098        } else {
1099            format!("frames/{epoch:020}/{checkpoint_seq:010}/{frame_no:020}")
1100        };
1101        join_key(&self.prefix, &suffix)
1102    }
1103
1104    /// R761-F2: key of a batch object holding frames `first..=last`
1105    /// concatenated. Same epoch namespacing and zero-padding as
1106    /// [`Self::frame_key`], so lexical order still matches frame order — one
1107    /// stream's batches never overlap, because each is a slice of a single
1108    /// monotonic drain.
1109    ///
1110    /// The range is in the key rather than only in the manifest so the sink
1111    /// stays self-describing: a `ls` of the prefix tells an operator exactly
1112    /// which frames are present, which is what made the per-frame layout easy
1113    /// to reason about and is worth keeping.
1114    ///
1115    /// Distinguishable from a legacy per-frame key by construction (that one
1116    /// has no `-`), so a stream that straddles the layout change can hold both
1117    /// shapes under the same directory without collision.
1118    pub(crate) fn frame_batch_key(
1119        &self,
1120        epoch: u64,
1121        checkpoint_seq: u32,
1122        first: u64,
1123        last: u64,
1124    ) -> ObjPath {
1125        let suffix = if epoch == 0 {
1126            format!("frames/{checkpoint_seq:010}/{first:020}-{last:020}")
1127        } else {
1128            format!("frames/{epoch:020}/{checkpoint_seq:010}/{first:020}-{last:020}")
1129        };
1130        join_key(&self.prefix, &suffix)
1131    }
1132
1133    /// Every object holding generation `m`'s frames, as `(key, first, last)`
1134    /// in ascending frame order.
1135    ///
1136    /// The single place that knows how a generation's frames are laid out:
1137    /// R761-F2 batches when the manifest carries a `frame_batch` list, the
1138    /// pre-R761-F2 one-object-per-frame layout when it does not. Restore's
1139    /// replay and [`crate::puller::WalPuller`] both go through here, so the
1140    /// two can never drift apart about where a frame lives.
1141    pub(crate) fn frame_objects_of(
1142        &self,
1143        m: &OwnedGenerationManifest,
1144    ) -> Vec<(ObjPath, u64, u64)> {
1145        if m.frame_batches.is_empty() {
1146            (m.first_frame..=m.last_frame)
1147                .map(|n| (self.frame_key(m.epoch, m.checkpoint_seq, n), n, n))
1148                .collect()
1149        } else {
1150            m.frame_batches
1151                .iter()
1152                .map(|&(first, last)| {
1153                    (
1154                        self.frame_batch_key(m.epoch, m.checkpoint_seq, first, last),
1155                        first,
1156                        last,
1157                    )
1158                })
1159                .collect()
1160        }
1161    }
1162
1163    pub(crate) fn generation_key(&self, unix_nanos: u128) -> ObjPath {
1164        join_key(
1165            &self.prefix,
1166            &format!("generations/gen-{unix_nanos:020}.manifest"),
1167        )
1168    }
1169}
1170
1171/// R858-B19 — identify the WAL generation `seam` is currently reading, by
1172/// pulling the salt out of **frame 1's** header.
1173///
1174/// Frame 1 rather than the 32-byte WAL header because this goes through the
1175/// [`WalSeam`] trait, which is the crate's one boundary against `turso_core`:
1176/// `turso_core::WalState` reports only `checkpoint_seq_no` and `max_frame`, and
1177/// reading the `-wal` file directly would need a path that
1178/// [`CoreWalSeam::from_conn`] does not have. Every frame header carries a
1179/// verbatim copy of the WAL header's salt, so frame 1 answers the same question
1180/// through machinery every seam already implements — including the mocks, and
1181/// including roadcase's connection-backed seam.
1182///
1183/// `Ok(None)` means the WAL holds no frames, so there is no generation to name
1184/// yet. That is not a failure: it is the honest "unknown", and
1185/// [`WalGeneration::is_provably_same_as`] treats it as such.
1186///
1187/// Costs one page-sized read per [`tail_frames`] call — negligible next to the
1188/// frames that call is about to upload, and it is the only thing standing
1189/// between a WAL recreate and a silently spliced generation chain.
1190pub(crate) fn read_wal_salt<S: WalSeam>(
1191    seam: &S,
1192    page_size: usize,
1193    last_frame: u64,
1194) -> Result<Option<WalSalt>> {
1195    if last_frame == 0 {
1196        return Ok(None);
1197    }
1198    let mut buf = vec![0u8; WAL_FRAME_HEADER_SIZE + page_size];
1199    seam.wal_get_frame(1, &mut buf)
1200        .context("reading WAL frame 1 to identify the WAL generation (R858-B19)")?;
1201    Ok(Some(WalSalt::from_frame_header(&buf)))
1202}
1203
1204/// Tail new WAL frames from `seam` into `target`, anchored to a base
1205/// snapshot. Idempotent and resumable: on the second call only frames after
1206/// the recorded watermark are uploaded.
1207///
1208/// Ordering within a single call:
1209/// 1. Read [`Watermark`] from the seam (snapshot the current `(checkpoint_seq,
1210///    max_frame)`) and the current [`WalGeneration`] (that sequence plus the
1211///    WAL salt, via [`read_wal_salt`]).
1212/// 2. Read the prior watermark sidecar (if any).
1213/// 3. Unless the sidecar's generation is *provably* the same WAL as the live
1214///    one, treat this as a restart: upload frames `1..=max_frame`. R858-B19 —
1215///    the test is [`WalGeneration::is_provably_same_as`], not a `checkpoint_seq`
1216///    comparison, because a writer-process restart recreates the WAL at
1217///    sequence `0` and an unchanged sequence therefore proves nothing.
1218/// 4. Otherwise upload frames `prior.last_frame+1..=max_frame`.
1219/// 5. Advance the watermark sidecar with a compare-and-swap on the version
1220///    read in step 2, then write a generation manifest pointing at the base
1221///    snapshot + frame range + the batch objects covering it. Frames precede
1222///    both, so a manifest never references a missing frame; the manifest
1223///    follows the CAS, so a writer that loses the sidecar race publishes
1224///    nothing (R732-T3).
1225///
1226/// Returns [`StreamOutcome::Empty`] if there is nothing to do (max_frame
1227/// hasn't advanced and checkpoint_seq is unchanged), or
1228/// [`StreamOutcome::Fenced`] if this writer's [`StreamConfig::epoch`] has been
1229/// superseded — either observed up front in step 2 or discovered by losing the
1230/// CAS in step 5.
1231pub async fn tail_frames<S: WalSeam>(
1232    seam: &S,
1233    target: &BackupTarget,
1234    cfg: &StreamConfig<'_>,
1235) -> Result<StreamOutcome> {
1236    let current = seam.wal_state()?;
1237    // R858-B19: the salt is what actually names the WAL generation. Read it
1238    // before anything else touches the sink, so the restart decision below is
1239    // made against the live WAL rather than inferred from a sequence number
1240    // that a writer restart silently resets.
1241    let current_generation = WalGeneration {
1242        checkpoint_seq: current.checkpoint_seq,
1243        salt: read_wal_salt(seam, cfg.page_size, current.last_frame)?,
1244    };
1245    let persisted = read_watermark(&target.store, &target.watermark_key()).await?;
1246    // R574-T4: staleness measured at call start, against the *prior*
1247    // sidecar write — i.e. how long the sink had gone without durable
1248    // progress before this call ran. On the steady cadence the caller's
1249    // scheduler promises, this hovers at the invocation interval; a
1250    // skipped/late run pushes it past `rpo_target` and flips `breached`.
1251    let rpo = rpo_status(cfg.rpo_target, persisted.as_ref());
1252
1253    // R732-F2 (W245) / R736-T2 (W250): the two-level fence, checked before
1254    // anything is written. A sink stamped with a higher epoch than ours means
1255    // ownership moved within the cell while we were away; a sink stamped with
1256    // a higher pointer generation means ownership moved to a *different*
1257    // cell — the two raft groups don't share an epoch counter, so the epoch
1258    // check alone is blind to that move. Either means every byte we are about
1259    // to upload belongs to somebody else's stream. Bounce here and the call
1260    // is a pure read; bounce anywhere later and we have already interleaved
1261    // frames into the real owner's range.
1262    if let Some(p) = &persisted {
1263        if p.epoch > cfg.epoch || p.pointer_generation > cfg.pointer_generation {
1264            return Ok(StreamOutcome::Fenced {
1265                current_epoch: p.epoch,
1266                our_epoch: cfg.epoch,
1267                current_pointer_generation: p.pointer_generation,
1268                our_pointer_generation: cfg.pointer_generation,
1269            });
1270        }
1271    }
1272
1273    // Borrowed, not consumed: the conditional watermark advance at the end of
1274    // this call needs the object version this read observed (R732-T3).
1275    let prior = persisted.as_ref().map(|p| p.watermark);
1276
1277    // R858-B19: the generation the sidecar was written under. `None` for a
1278    // sidecar that predates the salt field — which reads as *unknown*, not as
1279    // "the same WAL", and so forces the restart branch exactly once. After that
1280    // one full re-upload the sink carries a salt and the stream self-heals.
1281    let prior_generation = persisted.as_ref().map(|p| p.generation);
1282
1283    // Decide the range to upload. Resuming at `last_frame + 1` is only sound
1284    // when the live WAL is PROVABLY the one the watermark was taken from: a
1285    // writer-process restart deletes the `-wal` file and the next writer starts
1286    // a fresh WAL at checkpoint-sequence 0 with a new salt, so an unchanged
1287    // sequence is not evidence of anything. Anything short of proof restarts.
1288    let (start_frame, restarted) = match prior_generation {
1289        Some(g) if g.is_provably_same_as(&current_generation) => {
1290            (prior.map(|p| p.last_frame).unwrap_or(0) + 1, false)
1291        }
1292        Some(_) => (1, true),
1293        None => (1, false),
1294    };
1295
1296    if current.last_frame < start_frame {
1297        return Ok(StreamOutcome::Empty { watermark: current, rpo });
1298    }
1299
1300    // Drain frames in ascending order through the bounded spill buffer +
1301    // backpressure policy (R574-F2). A partial upload leaves a prefix (the
1302    // manifest is written last, so a prefix without a manifest is invisible
1303    // to restore — the next tail just overwrites the same keys).
1304    let drain = drain_frames_with_backpressure(
1305        seam,
1306        target,
1307        cfg,
1308        cfg.epoch,
1309        current.checkpoint_seq,
1310        start_frame,
1311        current.last_frame,
1312    )
1313    .await?;
1314
1315    let Some(uploaded_through) = drain.uploaded_through else {
1316        // BackpressurePolicy::Shed dropped everything before any frame
1317        // persisted — no manifest/watermark write, so the next call
1318        // re-attempts this exact range from the unchanged prior watermark.
1319        return Ok(StreamOutcome::Shed {
1320            checkpoint_seq: current.checkpoint_seq,
1321            first_frame: start_frame,
1322            last_frame: current.last_frame,
1323            backpressure: drain.report,
1324            rpo,
1325        });
1326    };
1327
1328    // R858-B19: everything above sampled the generation ONCE, before the
1329    // drain. A WAL recreate between calls is what this ticket is about, but
1330    // nothing stops one landing *during* a call — and then the frames just
1331    // uploaded are a mix of two WALs, which is the same corruption arriving by
1332    // a narrower door. Re-read the salt and refuse if it moved.
1333    //
1334    // Refusing here is cheap and complete: the watermark has not advanced and
1335    // no manifest has been written, so the frames that landed are orphaned
1336    // under keys nothing references — invisible to restore, exactly like the
1337    // `Shed` path — and the next tail re-derives everything from the sidecar.
1338    let after = seam.wal_state()?;
1339    let after_salt = read_wal_salt(seam, cfg.page_size, after.last_frame)?;
1340    if after_salt != current_generation.salt {
1341        anyhow::bail!(
1342            "the source WAL was recreated while this tail was uploading ({} -> {}) — the frames \
1343             this call read span two WAL generations, so nothing is published and the sink is left \
1344             exactly as it was; the next tail will restart cleanly against the new WAL",
1345            current_generation.describe(),
1346            WalGeneration { checkpoint_seq: after.checkpoint_seq, salt: after_salt }.describe(),
1347        );
1348    }
1349
1350    // Claim the range with a conditional watermark advance, THEN publish the
1351    // generation manifest. Under a Shed policy that persisted a partial
1352    // prefix, both cover only `start_frame..=uploaded_through`, not the full
1353    // engine range.
1354    //
1355    // R732-T3 reordered these two. The manifest used to be written last, so
1356    // that a manifest never referenced a missing frame — that invariant is
1357    // untouched, because frames still precede both. What the old order could
1358    // not do is fence: a writer that lost the sidecar race had already
1359    // published its manifest, which would then sit in the chain at a regressed
1360    // epoch and make every future restore refuse. Publishing only after the
1361    // CAS means a fenced writer leaves no manifest at all.
1362    let cas = write_watermark(
1363        &target.store,
1364        &target.watermark_key(),
1365        Watermark { checkpoint_seq: current.checkpoint_seq, last_frame: uploaded_through },
1366        current_generation.salt,
1367        cfg.epoch,
1368        cfg.pointer_generation,
1369        persisted.as_ref().and_then(|p| p.version.as_ref()),
1370    )
1371    .await?;
1372    if cas == WatermarkCas::Contended {
1373        // Somebody replaced the sidecar under us. Re-read to learn who.
1374        let now = read_watermark(&target.store, &target.watermark_key()).await?;
1375        let current_epoch = now.as_ref().map_or(0, |p| p.epoch);
1376        let current_pointer_generation = now.as_ref().map_or(0, |p| p.pointer_generation);
1377        if current_epoch > cfg.epoch || current_pointer_generation > cfg.pointer_generation {
1378            // A newer owner won the race. Our frames are orphaned under our
1379            // own epoch prefix with no manifest naming them, so restore never
1380            // sees them — the sink is exactly what the winner left.
1381            return Ok(StreamOutcome::Fenced {
1382                current_epoch,
1383                our_epoch: cfg.epoch,
1384                current_pointer_generation,
1385                our_pointer_generation: cfg.pointer_generation,
1386            });
1387        }
1388        // Same or lower epoch AND generation: neither fence can tell these two
1389        // writers apart. That means two processes are streaming the same
1390        // tenant under the SAME tokens — a caller bug (a duplicate streamer,
1391        // or an ownership token handed out twice), and exactly the
1392        // condition that must not be papered over with a retry.
1393        anyhow::bail!(
1394            "watermark CAS for {} lost to a concurrent writer at epoch {current_epoch} (pointer generation {current_pointer_generation}) while we hold epoch {} (pointer generation {}) — two streamers share one fencing token",
1395            target.watermark_key(),
1396            cfg.epoch,
1397            cfg.pointer_generation,
1398        );
1399    }
1400
1401    let nanos = unix_nanos();
1402    let gen_key = target.generation_key(nanos);
1403    let manifest = format_generation_manifest(GenerationManifest {
1404        base_snapshot_key: cfg.base_snapshot_key,
1405        page_size: cfg.page_size,
1406        checkpoint_seq: current.checkpoint_seq,
1407        // R858-B19: stamp the generation this range actually came from, so
1408        // restore can refuse a chain that spans two WALs instead of splicing
1409        // them. Always `Some` on this path — a `None` salt means an empty WAL,
1410        // and an empty WAL took the `Empty` return above.
1411        salt: current_generation.salt,
1412        first_frame: start_frame,
1413        last_frame: uploaded_through,
1414        epoch: cfg.epoch,
1415        owner: cfg.owner,
1416        // R761-F2: exactly the batch objects the drain persisted. Restore
1417        // derives its keys from this list, so it is not a summary of the range
1418        // — it IS the range's index.
1419        frame_batches: &drain.frame_batches,
1420    });
1421    target
1422        .store
1423        .put(&gen_key, manifest.into_bytes().into())
1424        .await
1425        .with_context(|| format!("writing generation manifest {gen_key}"))?;
1426
1427    let frame_count = uploaded_through - start_frame + 1;
1428    let gen_key = gen_key.to_string();
1429    if restarted {
1430        Ok(StreamOutcome::Restarted {
1431            generation_key: gen_key,
1432            previous_generation: prior_generation.unwrap_or_default(),
1433            new_generation: current_generation,
1434            first_frame: start_frame,
1435            last_frame: uploaded_through,
1436            frame_count,
1437            backpressure: drain.report,
1438            rpo,
1439        })
1440    } else {
1441        Ok(StreamOutcome::Streamed {
1442            generation_key: gen_key,
1443            first_frame: start_frame,
1444            last_frame: uploaded_through,
1445            checkpoint_seq: current.checkpoint_seq,
1446            frame_count,
1447            backpressure: drain.report,
1448            rpo,
1449        })
1450    }
1451}
1452
1453/// Outcome of [`drain_frames_with_backpressure`]: the last frame_no
1454/// successfully persisted this call (`None` if every buffered frame was
1455/// shed before any of them landed) plus the backpressure activity report.
1456struct DrainOutcome {
1457    uploaded_through: Option<u64>,
1458    /// R761-F2: the `(first, last)` range of every batch object that actually
1459    /// landed, ascending and gap-free from the call's `start_frame` through
1460    /// `uploaded_through` (a batch is popped only on a successful upload, so a
1461    /// partial drain truncates this list rather than holing it). Copied into
1462    /// the generation manifest, which is what tells restore where to look.
1463    frame_batches: Vec<(u64, u64)>,
1464    report: BackpressureReport,
1465}
1466
1467/// Read WAL frames `start_frame..=last_frame` into a bounded spill buffer
1468/// (`cfg.backpressure.spill_buffer_frames`) and drain them to `target`'s
1469/// object store through [`put_with_backoff`] (R574-F2), one **batch object**
1470/// per buffer-full (R761-F2).
1471///
1472/// The buffer only refills up to its bound, so once full the loop must
1473/// resolve the buffered batch — either by a successful upload or by the
1474/// configured [`BackpressurePolicy`] deciding what "stuck" means:
1475/// `Block` retries the batch forever on a throttling error (no data loss,
1476/// but the call can run long); `Fail` bubbles the error once
1477/// `backoff.max_retries` is exhausted (nothing persists past what already
1478/// landed); `Shed` drops the whole buffered backlog (loudly, via the
1479/// returned report) and returns whatever prefix already persisted. Any
1480/// non-throttling error bubbles immediately regardless of policy — this
1481/// mechanism is specifically for R2 throttling, not general fault
1482/// tolerance.
1483///
1484/// R761-F2 made the upload unit the buffer's contents rather than its head
1485/// frame, which is why `spill_buffer_frames` also sizes the largest object
1486/// this sink will write: `spill_buffer_frames * (24 + page_size)` bytes, ~1 MB
1487/// at the 256-frame default and a 4 KB page. That coupling is deliberate — the
1488/// bound already promises a memory ceiling, and a second knob for batch size
1489/// would only ever be set to some fraction of it.
1490async fn drain_frames_with_backpressure<S: WalSeam>(
1491    seam: &S,
1492    target: &BackupTarget,
1493    cfg: &StreamConfig<'_>,
1494    epoch: u64,
1495    checkpoint_seq: u32,
1496    start_frame: u64,
1497    last_frame: u64,
1498) -> Result<DrainOutcome> {
1499    let bp = &cfg.backpressure;
1500    let frame_size = WAL_FRAME_HEADER_SIZE + cfg.page_size;
1501    let bound = bp.spill_buffer_frames.max(1);
1502
1503    let mut report = BackpressureReport { policy: bp.policy, ..Default::default() };
1504    let mut pending: VecDeque<(u64, Vec<u8>)> = VecDeque::new();
1505    let mut uploaded_through: Option<u64> = None;
1506    let mut frame_batches: Vec<(u64, u64)> = Vec::new();
1507    let mut next_to_read = start_frame;
1508    let mut buf = vec![0u8; frame_size];
1509
1510    loop {
1511        // Refill up to the bound while there's more WAL to read. Once full,
1512        // the buffered batch must be resolved before we accept more.
1513        while pending.len() < bound && next_to_read <= last_frame {
1514            seam.wal_get_frame(next_to_read, &mut buf)
1515                .with_context(|| format!("reading wal frame {next_to_read}"))?;
1516            pending.push_back((next_to_read, buf.clone()));
1517            next_to_read += 1;
1518            report.high_water_frames = report.high_water_frames.max(pending.len());
1519        }
1520        let Some(&(first, _)) = pending.front() else {
1521            break; // Fully drained: nothing buffered, nothing left to read.
1522        };
1523        // Non-empty (we just matched `front`), and the buffer is filled in
1524        // ascending order without gaps, so the back frame closes the range.
1525        let last = pending.back().map_or(first, |&(n, _)| n);
1526        let mut body = Vec::with_capacity(pending.len() * frame_size);
1527        for (_, bytes) in &pending {
1528            body.extend_from_slice(bytes);
1529        }
1530        let key = target.frame_batch_key(epoch, checkpoint_seq, first, last);
1531        match put_with_backoff(&target.store, &key, body.into(), bp, &mut report.throttle_retries)
1532            .await
1533        {
1534            Ok(_) => {
1535                pending.clear();
1536                uploaded_through = Some(last);
1537                frame_batches.push((first, last));
1538            }
1539            Err(e) => match bp.policy {
1540                BackpressurePolicy::Shed => {
1541                    report.frames_shed += pending.len() as u64;
1542                    return Ok(DrainOutcome { uploaded_through, frame_batches, report });
1543                }
1544                BackpressurePolicy::Block | BackpressurePolicy::Fail => {
1545                    return Err(e).with_context(|| {
1546                        format!("uploading wal frames {first}-{last} to {key}")
1547                    });
1548                }
1549            },
1550        }
1551    }
1552    Ok(DrainOutcome { uploaded_through, frame_batches, report })
1553}
1554
1555/// R858-B18 — how many times [`raw_consistent_copy_live`] will re-take a copy
1556/// whose [`SourceFingerprint`] moved underneath it before giving up loudly.
1557///
1558/// Bounded on purpose. An unbounded retry against a database under sustained
1559/// write pressure is an infinite loop that looks like a hang; a caller that
1560/// wants to keep trying should be the one deciding how long to keep trying, on
1561/// its own schedule. Four attempts is enough to ride out an isolated
1562/// checkpoint (headscale's database measures 94 KB, so an attempt is
1563/// sub-millisecond) and few enough that a genuinely hot database is reported as
1564/// hot within a few milliseconds instead of being ground at.
1565pub const COPY_VALIDATION_ATTEMPTS: u32 = 4;
1566
1567/// Take a raw, point-in-time-consistent byte image of the database at `db_path`
1568/// WITHOUT folding the WAL via a `TRUNCATE` checkpoint, and WITHOUT locking out
1569/// the process that owns it. The live-writer pair of
1570/// [`crate::dedup::raw_consistent_copy`], which a concurrent writer can make
1571/// return `busy`.
1572///
1573/// Algorithm — the read-only "copy main + replay WAL frames ourselves" path
1574/// flagged in the working doc and the R005-T1 spike, wrapped in R858-B18's
1575/// optimistic validation:
1576///
1577/// 1. Sample the source's [`SourceFingerprint`] — WAL salt + checkpoint
1578///    sequence, WAL length, main length, main change counter — straight off
1579///    disk, with no engine open.
1580/// 2. Open a fresh `turso_core::Connection` via [`CoreWalSeam::open_reader`],
1581///    which takes **no** whole-file lock (so the application keeps working) and
1582///    calls `wal_auto_actions_disable` so our seam can't auto-checkpoint or
1583///    restart the WAL header mid-read on our connection.
1584/// 3. Read the main DB file bytes from disk. With auto-actions disabled on our
1585///    connection the file cannot be folded by us; a concurrent writer in a
1586///    separate connection only ever extends the WAL (the main file is only
1587///    written by a checkpoint).
1588/// 4. Walk WAL frames `1..=max_frame`, replaying each page into the in-memory
1589///    image at offset `(page_no - 1) * page_size`. Track the last
1590///    `is_commit_frame` and the corresponding `db_size`. Any frames past the
1591///    last commit are uncommitted mid-transaction garbage — drop them
1592///    (crash-consistency, matching restore's `wal_insert_end(false)`).
1593/// 5. Truncate / grow the image to `db_size * page_size`.
1594/// 6. Re-sample the fingerprint. **Accept the image only if
1595///    [`SourceFingerprint::stable_across`] holds** — i.e. only if nothing that
1596///    can tear the copy moved. A foreign checkpoint inside our read window would
1597///    splice pre- and post-fold state; discard that image and retry from step 1,
1598///    up to [`COPY_VALIDATION_ATTEMPTS`] times, then fail. (A plain WAL append
1599///    is *not* movement for this purpose, and that distinction is what keeps
1600///    this usable against a database that is actually in use — see
1601///    `stable_across`.)
1602///
1603/// Step 6 is the entire correctness story against a foreign engine, and it is
1604/// why this function never returns an unvalidated image: a torn WAL replay
1605/// still produces a structurally valid SQLite file that passes
1606/// `PRAGMA integrity_check`, so "it parsed" proves nothing. See
1607/// [`SourceFingerprint`] for why validation rather than SQLite's real
1608/// `-shm` reader protocol.
1609///
1610/// Returned bytes are a self-contained vanilla-SQLite image (page-offset
1611/// stable, no `-wal` sidecar required), ready to feed
1612/// [`crate::dedup::snapshot_dedup`]'s content-addressed chunking under a
1613/// concurrent writer.
1614pub async fn raw_consistent_copy_live(db_path: &str, page_size: usize) -> Result<Vec<u8>> {
1615    let image = validated_against_source(db_path, "live copy", || async {
1616        let seam = CoreWalSeam::open_reader(db_path)
1617            .with_context(|| format!("opening read-only WAL seam on {db_path}"))?;
1618        let main_bytes = std::fs::read(db_path)
1619            .with_context(|| format!("reading main db file {db_path}"))?;
1620        replay_wal_onto_main(&seam, main_bytes, page_size)
1621        // Seam dropped at the end of this block, so nothing of ours holds the
1622        // file while the post-copy fingerprint is sampled.
1623    })
1624    .await?;
1625    anyhow::ensure!(
1626        image.starts_with(b"SQLite format 3\0"),
1627        "live consistent copy of {db_path} is not a SQLite database"
1628    );
1629    Ok(image)
1630}
1631
1632/// R858-B18 — run `take` against the live database at `db_path` and return its
1633/// result **only** if the source provably held still for the duration.
1634///
1635/// This is the optimistic-validation protocol both source-side tiers share:
1636/// tier 2's WAL-replay copy ([`raw_consistent_copy_live`]) and tier 1a's
1637/// `VACUUM INTO` (`crate::snapshot`). Both read a database a foreign engine may
1638/// be checkpointing underneath them, neither can take a lock that engine
1639/// respects, and both are therefore only correct if a fold inside their read
1640/// window is *detected*. `what` names the operation in the refusal message.
1641///
1642/// Each attempt samples a [`SourceFingerprint`] before and after, then
1643/// classifies the outcome on both axes — did it succeed, and did the source
1644/// move:
1645///
1646/// | | source held still | source moved |
1647/// |---|---|---|
1648/// | **`Ok`** | accept | discard, retry (it may splice pre-/post-fold state) |
1649/// | **`Err`** | return the error — nothing raced us, so it is real | retry (a symptom of the race) |
1650///
1651/// That bottom-right cell is why the fingerprint is load-bearing beyond
1652/// accept/reject: R858-B18's probe H measured the *common* shape of a caught
1653/// race not as a clean torn image but as `short read on WAL frame` — a foreign
1654/// checkpoint truncating the WAL while turso walked it. Without the fingerprint
1655/// there is no way to tell that transient apart from a genuinely corrupt
1656/// database, and the two want opposite handling.
1657///
1658/// After [`COPY_VALIDATION_ATTEMPTS`] attempts all of which saw movement, this
1659/// fails loudly and returns nothing.
1660pub(crate) async fn validated_against_source<T, F, Fut>(
1661    db_path: &str,
1662    what: &str,
1663    mut take: F,
1664) -> Result<T>
1665where
1666    F: FnMut() -> Fut,
1667    Fut: std::future::Future<Output = Result<T>>,
1668{
1669    let mut moved: Option<(SourceFingerprint, SourceFingerprint, Option<anyhow::Error>)> = None;
1670    for attempt in 1..=COPY_VALIDATION_ATTEMPTS {
1671        if attempt > 1 {
1672            // Back off between attempts: retrying instantly against a writer
1673            // mid-checkpoint just spends the whole budget inside one fold.
1674            tokio::time::sleep(std::time::Duration::from_millis(20 * u64::from(attempt - 1)))
1675                .await;
1676        }
1677        let before = SourceFingerprint::read(db_path)?;
1678        let attempted = take().await;
1679        let after = SourceFingerprint::read(db_path)?;
1680        let stable = before.stable_across(&after);
1681        match attempted {
1682            Ok(value) if stable => return Ok(value),
1683            Ok(_) => moved = Some((before, after, None)),
1684            Err(e) if !stable => moved = Some((before, after, Some(e))),
1685            Err(e) => {
1686                return Err(e).with_context(|| {
1687                    format!(
1688                        "{what} of {db_path} failed against a source that did NOT move during \
1689                         the attempt ({}), so this is not a concurrent-writer race",
1690                        before.describe()
1691                    )
1692                })
1693            }
1694        }
1695    }
1696    let (before, after, last_err) = moved.expect("COPY_VALIDATION_ATTEMPTS is non-zero");
1697    let because = match last_err {
1698        Some(e) => format!("the last attempt also failed mid-read ({e:#})"),
1699        None => "each attempt produced a result that could not be validated".to_string(),
1700    };
1701    anyhow::bail!(
1702        "refusing a {what} of {db_path}: the source moved under every one of \
1703         {COPY_VALIDATION_ATTEMPTS} attempts, so nothing could be validated as \
1704         point-in-time — {because}. The last attempt saw [{}] before and [{}] after, so a \
1705         foreign writer or checkpointer is active. Retry on your own schedule; NO result \
1706         is returned, because a torn read of a SQLite database still parses as valid SQLite.",
1707        before.describe(),
1708        after.describe(),
1709    )
1710}
1711
1712/// Pure replay of every committed WAL frame visible through `seam` onto
1713/// `main_bytes`. Split out from [`raw_consistent_copy_live`] so it can be
1714/// driven by a mock seam in unit tests; the live entry point layers disk I/O
1715/// and magic-byte validation on top.
1716///
1717/// Contract: the highest `is_commit_frame` in `1..=wal_state.last_frame`
1718/// defines both the post-replay page count and the cutoff for which frames
1719/// are applied. If no frame in that window is a commit, `main_bytes` is
1720/// returned unchanged.
1721pub(crate) fn replay_wal_onto_main<S: WalSeam>(
1722    seam: &S,
1723    mut main_bytes: Vec<u8>,
1724    page_size: usize,
1725) -> Result<Vec<u8>> {
1726    anyhow::ensure!(page_size > 0, "page_size must be non-zero");
1727    let watermark = seam.wal_state()?;
1728
1729    let frame_size = WAL_FRAME_HEADER_SIZE + page_size;
1730    let mut buf = vec![0u8; frame_size];
1731
1732    struct PendingFrame {
1733        page_no: u32,
1734        db_size: u32,
1735        page_bytes: Vec<u8>,
1736    }
1737    let mut frames: Vec<PendingFrame> = Vec::new();
1738    let mut last_commit_idx: Option<usize> = None;
1739    for frame_no in 1..=watermark.last_frame {
1740        let info = seam
1741            .wal_get_frame(frame_no, &mut buf)
1742            .with_context(|| format!("reading WAL frame {frame_no}"))?;
1743        anyhow::ensure!(
1744            info.page_no >= 1,
1745            "WAL frame {frame_no}: page_no must be >= 1"
1746        );
1747        frames.push(PendingFrame {
1748            page_no: info.page_no,
1749            db_size: info.db_size,
1750            page_bytes: buf[WAL_FRAME_HEADER_SIZE..].to_vec(),
1751        });
1752        if info.is_commit_frame() {
1753            last_commit_idx = Some(frames.len() - 1);
1754        }
1755    }
1756
1757    let Some(last_commit) = last_commit_idx else {
1758        // No committed frames in our view — main file alone is the image.
1759        // Any uncommitted suffix in the WAL is dropped by construction.
1760        return Ok(main_bytes);
1761    };
1762    let final_db_size = frames[last_commit].db_size as usize;
1763    let target_size = final_db_size
1764        .checked_mul(page_size)
1765        .context("db_size * page_size overflow")?;
1766
1767    // Grow image to hold the final committed image AND any page slot we'll
1768    // touch in the commit prefix (a frame may write a page above db_size
1769    // mid-grow; the final truncate cuts that back to db_size).
1770    let max_off_needed: usize = frames[..=last_commit]
1771        .iter()
1772        .map(|f| f.page_no as usize * page_size)
1773        .max()
1774        .unwrap_or(0);
1775    let need = target_size.max(max_off_needed);
1776    if main_bytes.len() < need {
1777        main_bytes.resize(need, 0);
1778    }
1779    for f in &frames[..=last_commit] {
1780        let off = (f.page_no as usize - 1) * page_size;
1781        main_bytes[off..off + page_size].copy_from_slice(&f.page_bytes);
1782    }
1783    main_bytes.truncate(target_size);
1784
1785    Ok(main_bytes)
1786}
1787
1788/// Summary of a [`restore_latest_stream`] call.
1789#[derive(Debug, Clone, PartialEq, Eq)]
1790pub struct RestoreOutcome {
1791    /// The tier-1a snapshot key every generation manifest referenced (must agree).
1792    pub base_snapshot_key: String,
1793    /// Sole `checkpoint_seq_no` across the replayed manifests (v1 refuses to
1794    /// span a WAL restart — see [`validate_generation_chain`]).
1795    pub checkpoint_seq: u32,
1796    /// Number of generation manifests replayed (≥ 1).
1797    pub generation_count: usize,
1798    /// Total frames inserted across all generations (`last_frame - 0`, since
1799    /// the chain is required to start at frame 1 and be gap-free).
1800    pub frames_replayed: u64,
1801    /// The last frame position written into the destination WAL.
1802    pub last_frame: u64,
1803    /// R732-F2: the highest fencing epoch contributing to this restore (`0`
1804    /// for a chain written before fencing existed). Reported so an operator
1805    /// restoring after an ownership transfer can see which owner's data they
1806    /// actually got.
1807    pub epoch: u64,
1808}
1809
1810/// A consistency-checked sequence of generation manifests, ready to drive a
1811/// replay. The fields are the single base / page size / checkpoint sequence
1812/// shared by every manifest in the chain.
1813#[derive(Debug, Clone, PartialEq, Eq)]
1814pub(crate) struct ValidatedChain {
1815    pub base_snapshot_key: String,
1816    pub page_size: usize,
1817    /// R858-B19: the one WAL generation every manifest in the chain agrees on.
1818    /// `salt` is `Some` for any chain of two or more manifests — a chain that
1819    /// could not prove a single generation never gets this far.
1820    pub generation: WalGeneration,
1821    pub total_frames: u64,
1822    /// R732-F2: the highest fencing epoch in the chain — i.e. the most recent
1823    /// owner that contributed frames. Unlike the other fields this is a
1824    /// *maximum*, not a shared constant: ownership legitimately moves
1825    /// mid-chain, so a chain may span epochs as long as they never go
1826    /// backwards.
1827    pub epoch: u64,
1828}
1829
1830/// Validate that a sorted list of generation manifests forms a single,
1831/// replayable chain: same base snapshot, same page size, single checkpoint
1832/// sequence, frames starting at 1 and contiguous across manifests.
1833///
1834/// V1 refuses to span a WAL restart (multiple `checkpoint_seq` values). A
1835/// restart implies the source engine folded the WAL into main between
1836/// generations — replaying the post-restart frames onto our pre-restart base
1837/// would skip that fold and corrupt the result. The remediation is a fresh
1838/// tier-1a snapshot, not heroics in restore.
1839///
1840/// # R858-B19 — refuse what cannot be proven
1841///
1842/// The `checkpoint_seq` test above was the *only* generation check here, and it
1843/// is blind to the fold that matters: a writer-process restart deletes the
1844/// `-wal` file and the next writer starts a fresh WAL back at sequence `0`, so
1845/// two unrelated WALs both report `0` and a spliced chain sailed through. That
1846/// produced a restore that reported SUCCESS while writing a stale-but-plausible
1847/// image (probe B: 9 rows against a 12-row source, `integrity_check ok`) or one
1848/// upstream sqlite3 calls malformed (probe F). **A silent wrong image is the
1849/// specific failure this function now exists to make impossible.**
1850///
1851/// So the salt is checked too, and — the part that matters — a chain that
1852/// cannot be *shown* to come from one WAL is refused rather than replayed:
1853///
1854/// - Two or more manifests, any of them lacking a salt (written before `v4`):
1855///   REFUSE. The splice is exactly what a pre-R858-B19 writer produced, and
1856///   nothing in those manifests records which WAL each range came from.
1857/// - Two or more manifests with disagreeing salts: REFUSE, naming both.
1858/// - A single manifest: accepted with whatever salt it has, including none.
1859///   One generation is not a splice; there is nothing to prove.
1860pub(crate) fn validate_generation_chain(
1861    manifests: &[OwnedGenerationManifest],
1862) -> Result<ValidatedChain> {
1863    let first = manifests
1864        .first()
1865        .context("validate_generation_chain: empty manifest list")?;
1866    let mut expected_next_frame: u64 = 1;
1867    let mut chain_epoch: u64 = 0;
1868    let multi = manifests.len() > 1;
1869    for (i, m) in manifests.iter().enumerate() {
1870        // R732-F2 (W245): generations are ordered by write time, so a chain
1871        // whose epoch goes BACKWARDS says a stale owner wrote after the sink
1872        // had already moved on — precisely the split-brain the fence exists to
1873        // stop, caught here on the read side too. Non-decreasing is fine and
1874        // expected: ownership transfers mid-stream and the new owner keeps
1875        // appending frames to the same contiguous range.
1876        if m.epoch < chain_epoch {
1877            anyhow::bail!(
1878                "generation #{i} was written at epoch {} but an earlier generation in the chain is at epoch {} — a fenced (stale) owner wrote to this sink, restore refuses rather than replay interleaved frames",
1879                m.epoch,
1880                chain_epoch,
1881            );
1882        }
1883        chain_epoch = m.epoch;
1884        if m.base_snapshot_key != first.base_snapshot_key {
1885            anyhow::bail!(
1886                "generation #{i} references base {:?}, expected {:?} — chain spans bases, restore needs a fresh tier-1a snapshot",
1887                m.base_snapshot_key,
1888                first.base_snapshot_key,
1889            );
1890        }
1891        if m.page_size != first.page_size {
1892            anyhow::bail!(
1893                "generation #{i} page_size {} differs from chain page_size {} — corrupt manifest or mixed sinks",
1894                m.page_size,
1895                first.page_size,
1896            );
1897        }
1898        if m.checkpoint_seq != first.checkpoint_seq {
1899            anyhow::bail!(
1900                "generation #{i} checkpoint_seq {} differs from chain checkpoint_seq {} — WAL restart between generations, restore needs a fresh tier-1a snapshot",
1901                m.checkpoint_seq,
1902                first.checkpoint_seq,
1903            );
1904        }
1905        // R858-B19: the check `checkpoint_seq` cannot make. Only enforced on a
1906        // multi-manifest chain — a lone generation has nothing to be spliced
1907        // to, so refusing it would strand every pre-v4 single-generation backup
1908        // for no safety gain.
1909        if multi {
1910            let Some(s) = m.salt else {
1911                anyhow::bail!(
1912                    "generation #{i} carries no WAL salt (written by a pre-R858-B19 writer) and this chain spans {} generations — which WAL each range came from was never recorded, so a chain spliced across a WAL recreate is indistinguishable from a good one. Restore refuses rather than replay a plausible wrong image; take a fresh tier-1a snapshot",
1913                    manifests.len(),
1914                );
1915            };
1916            if let Some(fs) = first.salt {
1917                if s != fs {
1918                    anyhow::bail!(
1919                        "generation #{i} was written under WAL salt {s} but the chain starts at salt {fs} (both at checkpoint_seq {}) — the source WAL was RECREATED mid-stream, so these frames belong to two different WALs and replaying them as one would corrupt the image. Restore needs a fresh tier-1a snapshot",
1920                        first.checkpoint_seq,
1921                    );
1922                }
1923            }
1924        }
1925        if m.first_frame != expected_next_frame {
1926            anyhow::bail!(
1927                "generation #{i} starts at frame {} but the previous generation ended at frame {} — gap in stream",
1928                m.first_frame,
1929                expected_next_frame.saturating_sub(1),
1930            );
1931        }
1932        if m.last_frame < m.first_frame {
1933            anyhow::bail!(
1934                "generation #{i} has last_frame {} < first_frame {} — corrupt manifest",
1935                m.last_frame,
1936                m.first_frame,
1937            );
1938        }
1939        expected_next_frame = m.last_frame + 1;
1940    }
1941    Ok(ValidatedChain {
1942        base_snapshot_key: first.base_snapshot_key.clone(),
1943        page_size: first.page_size,
1944        generation: WalGeneration {
1945            checkpoint_seq: first.checkpoint_seq,
1946            salt: first.salt,
1947        },
1948        total_frames: expected_next_frame - 1,
1949        epoch: chain_epoch,
1950    })
1951}
1952
1953/// Fetch one generation's frames and hand each to `on_frame` in ascending
1954/// frame order, skipping anything before `from_frame`. Returns how many frames
1955/// were delivered.
1956///
1957/// R761-F2: the one read path that resolves a manifest to frame bytes,
1958/// batched or not — `BackupTarget::frame_objects_of` decides which layout
1959/// the generation used, and a batch object is split here at
1960/// `24 + page_size` boundaries. Both restore's replay and
1961/// [`crate::puller::WalPuller`] call it, so a layout the writer can produce
1962/// can never be readable by one and not the other.
1963pub(crate) async fn for_each_frame_in_generation<F>(
1964    target: &BackupTarget,
1965    m: &OwnedGenerationManifest,
1966    from_frame: u64,
1967    mut on_frame: F,
1968) -> Result<u64>
1969where
1970    F: FnMut(u64, &[u8]) -> Result<()>,
1971{
1972    let frame_size = WAL_FRAME_HEADER_SIZE + m.page_size;
1973    let mut delivered: u64 = 0;
1974    for (key, first, last) in target.frame_objects_of(m) {
1975        if last < from_frame {
1976            continue; // wholly behind the caller's cursor — don't pay for the GET
1977        }
1978        let bytes = target
1979            .store
1980            .get(&key)
1981            .await
1982            .with_context(|| format!("fetching frame object {key}"))?
1983            .bytes()
1984            .await
1985            .with_context(|| format!("reading frame object body {key}"))?;
1986        let frames_in_object = (last - first + 1) as usize;
1987        let want = frames_in_object * frame_size;
1988        if bytes.len() != want {
1989            anyhow::bail!(
1990                "frame object {key} is {} bytes, expected {want} ({frames_in_object} frame(s) x [{WAL_FRAME_HEADER_SIZE} header + {} page])",
1991                bytes.len(),
1992                m.page_size,
1993            );
1994        }
1995        for (i, frame_no) in (first..=last).enumerate() {
1996            if frame_no < from_frame {
1997                continue;
1998            }
1999            let at = i * frame_size;
2000            on_frame(frame_no, &bytes[at..at + frame_size])?;
2001            delivered += 1;
2002        }
2003    }
2004    Ok(delivered)
2005}
2006
2007/// Replay every frame named by `manifests` into `seam`, in (checkpoint_seq,
2008/// frame_no) order. Caller must have already called `wal_insert_begin` on the
2009/// seam; `wal_insert_end` is also the caller's responsibility (so a test or a
2010/// future fault-injection path can choose `force_commit`).
2011///
2012/// Each manifest's own `page_size` sizes its frames —
2013/// [`validate_generation_chain`] has already established they all agree, so
2014/// there is no separate chain-level page size to thread through.
2015async fn replay_frames_into<S: WalInsertSeam>(
2016    target: &BackupTarget,
2017    seam: &S,
2018    manifests: &[OwnedGenerationManifest],
2019) -> Result<u64> {
2020    let mut total: u64 = 0;
2021    for m in manifests {
2022        total += for_each_frame_in_generation(target, m, 0, |frame_no, bytes| {
2023            seam.wal_insert_frame(frame_no, bytes)
2024                .with_context(|| format!("inserting frame {frame_no}"))
2025        })
2026        .await?;
2027    }
2028    Ok(total)
2029}
2030
2031/// Restore a database from a streamed backup: download the base tier-1a
2032/// snapshot referenced by the generation manifests, then replay every uploaded
2033/// WAL frame onto it via the [`WalInsertSeam`] of a fresh `turso_core`
2034/// connection. The session ends with `force_commit = false` so any frames
2035/// captured after the last commit frame are dropped (crash-consistency).
2036///
2037/// Errors loudly when:
2038/// - There are no generation manifests under the prefix (caller should restore
2039///   via the tier-1a path instead).
2040/// - The chain spans more than one `checkpoint_seq` or more than one base
2041///   snapshot (restart / mixed bases — needs a fresh tier-1a snapshot).
2042/// - Frame ranges are not gap-free starting at 1.
2043/// - A referenced frame object is missing or the wrong byte length.
2044///
2045/// `dest_path`'s `-wal` and `-shm` sidecars are removed before the base is
2046/// written; any pre-existing turso connection on `dest_path` must be closed
2047/// by the caller.
2048///
2049/// A caller that has already listed the chain for its own reasons should call
2050/// [`restore_stream_from_manifests`] instead and skip the re-listing this one
2051/// does — see its docs for what that costs (R760-T18).
2052pub async fn restore_latest_stream(
2053    target: &BackupTarget,
2054    dest_path: &str,
2055) -> Result<RestoreOutcome> {
2056    let manifests = list_and_parse_generation_manifests(target).await?;
2057    restore_stream_from_manifests(target, dest_path, &manifests).await
2058}
2059
2060/// [`restore_latest_stream`] for a caller that has already listed and parsed
2061/// the chain — same replay, minus the listing.
2062///
2063/// R760-T18: a caller that must inspect the chain *before* deciding how to
2064/// restore otherwise pays for every manifest twice. roadcase's `SinkHydrator`
2065/// is the live example: it calls [`list_and_parse_generation_manifests`] to
2066/// ask whether the era has any frames at all (an era whose base is still the
2067/// whole story restores via `snapshot::restore_latest` instead), and then
2068/// `restore_latest_stream` re-listed and re-fetched the identical objects. A
2069/// cold start with `n` generations was measured at `3n+1` Class B operations
2070/// against a `2n+1` floor — one GET per manifest, one per generation's frame
2071/// batch, one for the base. Handing the parse straight in removes the `n`
2072/// duplicates.
2073///
2074/// `manifests` must be ascending by generation, which is the order
2075/// [`list_and_parse_generation_manifests`] returns them in; the chain is
2076/// validated here exactly as it is on the listing path, so a hand-assembled
2077/// list cannot smuggle past a check.
2078pub async fn restore_stream_from_manifests(
2079    target: &BackupTarget,
2080    dest_path: &str,
2081    manifests: &[OwnedGenerationManifest],
2082) -> Result<RestoreOutcome> {
2083    if manifests.is_empty() {
2084        anyhow::bail!(
2085            "no generation manifests under {} — restore tier-1a directly via snapshot::restore_latest",
2086            join_key(&target.prefix, "generations"),
2087        );
2088    }
2089    let chain = validate_generation_chain(manifests)?;
2090
2091    // Download the base snapshot and lay it down at dest_path. Strip any stale
2092    // WAL/-shm sidecars first — the base alone is the entire pre-replay image
2093    // (VACUUM INTO output has no WAL).
2094    let base_bytes = target
2095        .store
2096        .get(&ObjPath::from(chain.base_snapshot_key.clone()))
2097        .await
2098        .with_context(|| format!("downloading base snapshot {}", chain.base_snapshot_key))?
2099        .bytes()
2100        .await
2101        .with_context(|| format!("reading base snapshot body {}", chain.base_snapshot_key))?;
2102    for sfx in ["-wal", "-shm"] {
2103        let _ = std::fs::remove_file(format!("{dest_path}{sfx}"));
2104    }
2105    std::fs::write(dest_path, &base_bytes)
2106        .with_context(|| format!("writing restored base snapshot to {dest_path}"))?;
2107
2108    // Open a fresh seam on the laid-down base and replay. `CoreWalSeam::open`
2109    // disables auto-actions; `wal_insert_begin` further locks the session to
2110    // empty auto-actions for the txn, so restart can't race the replay.
2111    let seam = CoreWalSeam::open(dest_path)?;
2112    seam.wal_insert_begin()
2113        .context("starting WAL insert session on restore destination")?;
2114    let frames = match replay_frames_into(target, &seam, manifests).await {
2115        Ok(n) => n,
2116        Err(e) => {
2117            // Best-effort: roll the partial replay back so we never leave the
2118            // dest's WAL with an uncommitted suffix.
2119            let _ = seam.wal_insert_end(false);
2120            return Err(e);
2121        }
2122    };
2123    // force_commit=false: the engine drops any tail past the last commit
2124    // frame, which is exactly the crash-consistency story we want — a
2125    // mid-transaction tail captured by tail_frames gets truncated cleanly.
2126    seam.wal_insert_end(false)
2127        .context("closing WAL insert session on restore destination")?;
2128
2129    Ok(RestoreOutcome {
2130        base_snapshot_key: chain.base_snapshot_key,
2131        checkpoint_seq: chain.generation.checkpoint_seq,
2132        generation_count: manifests.len(),
2133        frames_replayed: frames,
2134        last_frame: chain.total_frames,
2135        epoch: chain.epoch,
2136    })
2137}
2138
2139/// Every generation manifest at a sink, in key (i.e. chronological) order.
2140///
2141/// **The GETs are serialized**, one manifest at a time, and so is the frame
2142/// replay in [`for_each_frame_in_generation`] — so a restore's wall clock is
2143/// `(2n+1) x RTT` plus transfer for `n` generations, not `max(RTT)`. That is
2144/// fine at the tens of generations a dedicated streamer accumulates between
2145/// checkpoints and is not fine at thousands; R760-T18 bounds it on the
2146/// *writer* side (roadcase caps generations per era) rather than by making
2147/// this concurrent, because concurrency here would mean a real
2148/// `futures`/`tokio::spawn` dependency in a crate that deliberately keeps
2149/// `futures_util` to `[dev-dependencies]`.
2150///
2151/// Public since R732-T5: a split-brain reconciliation oracle needs to ask what
2152/// a sink would actually replay — which owner wrote which frame range, under
2153/// which epoch — without running a full restore. Pair it with
2154/// [`validate_generation_chain`]'s public counterpart if you need the chain
2155/// checked rather than merely listed.
2156pub async fn list_and_parse_generation_manifests(
2157    target: &BackupTarget,
2158) -> Result<Vec<OwnedGenerationManifest>> {
2159    let prefix = join_key(&target.prefix, "generations");
2160    let listing = target
2161        .store
2162        .list_with_delimiter(Some(&prefix))
2163        .await
2164        .with_context(|| format!("listing generation manifests under {prefix}"))?;
2165    let mut keys: Vec<ObjPath> = listing.objects.into_iter().map(|o| o.location).collect();
2166    keys.sort();
2167    let mut manifests = Vec::with_capacity(keys.len());
2168    for key in &keys {
2169        let bytes = target
2170            .store
2171            .get(key)
2172            .await
2173            .with_context(|| format!("fetching generation manifest {key}"))?
2174            .bytes()
2175            .await
2176            .with_context(|| format!("reading generation manifest body {key}"))?;
2177        let m = parse_generation_manifest(&String::from_utf8_lossy(&bytes))
2178            .with_context(|| format!("parsing generation manifest {key}"))?;
2179        manifests.push(m);
2180    }
2181    Ok(manifests)
2182}
2183
2184/// What the sink's fence currently stands at — the two comparands
2185/// [`tail_frames`] checks a writer against, and nothing else.
2186///
2187/// R869: this is the off-fleet copy of a fencing token. It matters because
2188/// `epoch` is minted by a raft group, and a raft group can be *destroyed* —
2189/// wipe the raft dir, re-form, and the fresh cluster's first
2190/// `ClaimTenant` grants epoch 1 while this sidecar still says 5, so the
2191/// rebuilt cluster is fenced out of its own sink. Reading the fence back is how
2192/// a rebuild starts above the number its dead predecessor left here.
2193#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
2194pub struct FenceState {
2195    /// The fencing epoch of the writer that last advanced the sidecar. `0`
2196    /// means "unfenced" — either nothing has claimed this sink or it predates
2197    /// R732-F2 — and fences nobody.
2198    pub epoch: u64,
2199    /// The pointer generation of that same writer. `0` means unfenced, exactly
2200    /// as for `epoch`.
2201    pub pointer_generation: u64,
2202}
2203
2204/// Read the fence [`tail_frames`] would check a writer against, without writing
2205/// anything or needing a WAL.
2206///
2207/// `Ok(None)` means the sidecar does not exist: nothing has ever streamed to
2208/// this prefix, so **there is no fence here at all** and any writer is
2209/// accepted. Do not collapse that into `FenceState::default()` at a call site
2210/// that is deciding whether to fence — "unfenced because nobody has written"
2211/// and "unfenced because an old writer stamped 0" are the same *value* but a
2212/// caller that cares about the difference (a rebuild deciding whether a tenant
2213/// was ever live) needs the `Option`.
2214///
2215/// **A floor derived from `epoch` is a lower bound, not an upper one.** It is
2216/// the highest epoch anyone has *written under*, which is not the highest a
2217/// dead raft group *granted* — a tenant claimed twice while idle leaves a node
2218/// holding an epoch strictly above anything this sidecar ever saw. So a rebuild
2219/// that seeds from this number closes availability (its own writes are
2220/// accepted) and does **not** on its own fence a resurrected node; that takes
2221/// `pointer_generation`, which is minted off-fleet where a dead cluster cannot
2222/// reach it. `oss/yubaba/crates/yubaba/tests/raft_rebuild_fencing.rs` drives
2223/// both halves against this code.
2224pub async fn read_fence_state(target: &BackupTarget) -> Result<Option<FenceState>> {
2225    Ok(read_watermark(&target.store, &target.watermark_key())
2226        .await?
2227        .map(|p| FenceState {
2228            epoch: p.epoch,
2229            pointer_generation: p.pointer_generation,
2230        }))
2231}
2232
2233/// In-memory shape of a generation manifest. Format on disk:
2234///
2235/// ```text
2236/// TURSO-BACKUP STREAM v4
2237/// base_snapshot <key>
2238/// page_size <n>
2239/// checkpoint_seq <n>
2240/// wal_salt <salt1>-<salt2>
2241/// first_frame <n>
2242/// last_frame <n>
2243/// epoch <n>
2244/// owner <label>        (optional)
2245/// frame_batch <first>-<last>   (one per batch object, ascending, gap-free)
2246/// ```
2247///
2248/// Text, dependency-free (no serde), one field per line. Mirrors the
2249/// `dedup::Manifest` convention so the crate stays consistent.
2250///
2251/// R732-F2 added `epoch`/`owner` and moved the header to `v2`. R761-F2 added
2252/// the `frame_batch` list and moved it to `v3`. R858-B19 added `wal_salt` and
2253/// moved it to `v4`. Older manifests still parse — `v1` with `epoch = 0`,
2254/// `owner = None`; `v1`/`v2` with no batch list, which is what marks their
2255/// frames as living one-per-object; `v1`/`v2`/`v3` with `salt = None` — so
2256/// backups written before any of those changes stay *parseable*. No older
2257/// version is ever written any more.
2258///
2259/// R858-B19, and this is the one place a legacy manifest is not merely
2260/// second-class: a chain of two or more manifests where any of them lacks a
2261/// salt is **refused** by [`validate_generation_chain`], because a pre-R858-B19
2262/// writer could splice two WAL generations into a contiguous-looking chain and
2263/// nothing in the manifest records which WAL each range came from. A
2264/// single-manifest chain is still restored — one generation cannot be a splice.
2265#[derive(Debug, Clone, PartialEq, Eq)]
2266pub struct GenerationManifest<'a> {
2267    pub base_snapshot_key: &'a str,
2268    pub page_size: usize,
2269    pub checkpoint_seq: u32,
2270    /// R858-B19: the WAL header salt these frames were read under — the field
2271    /// that actually names the generation, since `checkpoint_seq` resets to 0
2272    /// across a writer restart. `None` only when re-formatting a pre-R858-B19
2273    /// manifest; [`tail_frames`] always has one.
2274    pub salt: Option<WalSalt>,
2275    pub first_frame: u64,
2276    pub last_frame: u64,
2277    /// R732-F2: the fencing epoch this generation was written under. `0` means
2278    /// it predates fencing (a `v1` manifest), which is also the key shape its
2279    /// frames live under — see [`BackupTarget::frame_key`].
2280    pub epoch: u64,
2281    /// R732-F2: opaque owner label, diagnostic only. See [`StreamConfig::owner`].
2282    pub owner: Option<&'a str>,
2283    /// R761-F2: the `(first, last)` frame range of each batch object holding
2284    /// this generation's frames, ascending and exactly tiling
2285    /// `first_frame..=last_frame`. Empty means the pre-batching layout (one
2286    /// object per frame); a writer emits it empty only when re-formatting a
2287    /// legacy manifest.
2288    pub frame_batches: &'a [(u64, u64)],
2289}
2290
2291/// Header of a manifest written by a fencing-aware writer (R732-F2). The
2292/// version is bumped rather than the `epoch` key just being added, because
2293/// [`parse_generation_manifest`] rejects unknown keys: a pre-R732 binary
2294/// reading a fenced manifest would otherwise fail with `unknown manifest key:
2295/// epoch`, which reads like corruption. Failing on the *header* says the real
2296/// thing — this backup was written by a newer writer.
2297const MANIFEST_HEADER_V2: &str = "TURSO-BACKUP STREAM v2";
2298/// Pre-fencing header. Still accepted on read (those backups must stay
2299/// restorable) and parses with `epoch = 0`, `owner = None`; never written.
2300const MANIFEST_HEADER_V1: &str = "TURSO-BACKUP STREAM v1";
2301/// R761-F2 — carries the `frame_batch` list. Bumped for the same reason
2302/// `v2` was: [`parse_generation_manifest`] rejects unknown keys, so a
2303/// pre-R761-F2 binary reading a batched manifest would fail with `unknown
2304/// manifest key: frame_batch`, which reads like corruption. Failing on the
2305/// header says the true thing — a newer writer wrote this sink. The refusal
2306/// matters more here than it did for fencing: an older reader that somehow
2307/// skipped the batch list would look for per-frame keys that do not exist.
2308const MANIFEST_HEADER_V3: &str = "TURSO-BACKUP STREAM v3";
2309/// R858-B19 — carries `wal_salt`, the field that makes a generation chain
2310/// checkable across a WAL recreate. Bumped for the same reason `v2` and `v3`
2311/// were: [`parse_generation_manifest`] rejects unknown keys, so an older binary
2312/// reading one of these would fail with `unknown manifest key: wal_salt`, which
2313/// reads like corruption. Failing on the header says the true thing.
2314const MANIFEST_HEADER_V4: &str = "TURSO-BACKUP STREAM v4";
2315
2316pub(crate) fn format_generation_manifest(m: GenerationManifest<'_>) -> String {
2317    // Each header promises the fields that version introduced, so stamp the
2318    // newest version whose promises this manifest can actually keep. The only
2319    // way to reach anything but v4 is re-formatting a legacy manifest (a test,
2320    // or a repair tool); tail_frames always has both a salt and batches.
2321    // A v4 manifest promises BOTH the salt and the batch list, so a legacy
2322    // shape missing either falls back rather than stamping a version whose
2323    // promises it cannot keep. `salt` without batches has no representation and
2324    // cannot occur: only tail_frames produces a salt, and it always batches.
2325    let salt = if m.frame_batches.is_empty() { None } else { m.salt };
2326    let header = if salt.is_some() {
2327        MANIFEST_HEADER_V4
2328    } else if m.frame_batches.is_empty() {
2329        MANIFEST_HEADER_V2
2330    } else {
2331        MANIFEST_HEADER_V3
2332    };
2333    let mut out = format!(
2334        "{}\nbase_snapshot {}\npage_size {}\ncheckpoint_seq {}\n",
2335        header, m.base_snapshot_key, m.page_size, m.checkpoint_seq,
2336    );
2337    if let Some(s) = salt {
2338        out.push_str(&format!("wal_salt {}-{}\n", s.salt1, s.salt2));
2339    }
2340    out.push_str(&format!(
2341        "first_frame {}\nlast_frame {}\nepoch {}\n",
2342        m.first_frame, m.last_frame, m.epoch,
2343    ));
2344    // Omitted rather than written empty: the parser splits on the first space,
2345    // so `owner ` with no value would round-trip to `Some("")`.
2346    if let Some(owner) = m.owner {
2347        out.push_str(&format!("owner {owner}\n"));
2348    }
2349    for (first, last) in m.frame_batches {
2350        out.push_str(&format!("frame_batch {first}-{last}\n"));
2351    }
2352    out
2353}
2354
2355/// Parse a generation manifest. Tolerant to trailing whitespace; rejects any
2356/// other shape (so a corrupt manifest fails loudly during restore, doesn't
2357/// silently degrade).
2358#[derive(Debug, Clone, PartialEq, Eq)]
2359pub struct OwnedGenerationManifest {
2360    pub base_snapshot_key: String,
2361    pub page_size: usize,
2362    pub checkpoint_seq: u32,
2363    /// R858-B19: `None` for a pre-`v4` manifest, which means the WAL
2364    /// generation behind these frames was never recorded and cannot be
2365    /// recovered. Unknown, not "the same as its neighbour" — see
2366    /// [`validate_generation_chain`].
2367    pub salt: Option<WalSalt>,
2368    pub first_frame: u64,
2369    pub last_frame: u64,
2370    /// R732-F2: `0` for a `v1` (pre-fencing) manifest.
2371    pub epoch: u64,
2372    /// R732-F2: diagnostic only; `None` when the writer did not label itself.
2373    pub owner: Option<String>,
2374    /// R761-F2: the batch objects covering `first_frame..=last_frame`,
2375    /// ascending and gap-free (enforced by [`parse_generation_manifest`]).
2376    /// **Empty means the pre-batching layout** — one object per frame — which
2377    /// is how a `v1`/`v2` manifest keeps restoring. Resolve it to keys with
2378    /// `BackupTarget::frame_objects_of` (crate-private) rather than branching
2379    /// at each call site.
2380    pub frame_batches: Vec<(u64, u64)>,
2381}
2382
2383pub fn parse_generation_manifest(text: &str) -> Result<OwnedGenerationManifest> {
2384    let mut lines = text.lines();
2385    let header = lines.next().context("empty generation manifest")?.trim();
2386    let (v2, v3, v4) = match header {
2387        MANIFEST_HEADER_V4 => (true, true, true),
2388        MANIFEST_HEADER_V3 => (true, true, false),
2389        MANIFEST_HEADER_V2 => (true, false, false),
2390        MANIFEST_HEADER_V1 => (false, false, false),
2391        other => anyhow::bail!("unexpected manifest header: {other:?}"),
2392    };
2393    let mut base_snapshot_key: Option<String> = None;
2394    let mut page_size: Option<usize> = None;
2395    let mut checkpoint_seq: Option<u32> = None;
2396    let mut salt: Option<WalSalt> = None;
2397    let mut first_frame: Option<u64> = None;
2398    let mut last_frame: Option<u64> = None;
2399    let mut epoch: Option<u64> = None;
2400    let mut owner: Option<String> = None;
2401    let mut frame_batches: Vec<(u64, u64)> = Vec::new();
2402    for line in lines {
2403        let line = line.trim();
2404        if line.is_empty() {
2405            continue;
2406        }
2407        let (k, v) = line
2408            .split_once(' ')
2409            .with_context(|| format!("malformed manifest line: {line:?}"))?;
2410        match k {
2411            "base_snapshot" => base_snapshot_key = Some(v.to_string()),
2412            "page_size" => page_size = Some(v.parse().context("page_size")?),
2413            "checkpoint_seq" => checkpoint_seq = Some(v.parse().context("checkpoint_seq")?),
2414            "wal_salt" => {
2415                let (s1, s2) = v
2416                    .split_once('-')
2417                    .with_context(|| format!("malformed wal_salt pair: {v:?}"))?;
2418                salt = Some(WalSalt {
2419                    salt1: s1.trim().parse().context("wal_salt salt1")?,
2420                    salt2: s2.trim().parse().context("wal_salt salt2")?,
2421                });
2422            }
2423            "first_frame" => first_frame = Some(v.parse().context("first_frame")?),
2424            "last_frame" => last_frame = Some(v.parse().context("last_frame")?),
2425            "epoch" => epoch = Some(v.parse().context("epoch")?),
2426            "owner" => owner = Some(v.to_string()),
2427            "frame_batch" => {
2428                let (first, last) = v
2429                    .split_once('-')
2430                    .with_context(|| format!("malformed frame_batch range: {v:?}"))?;
2431                frame_batches.push((
2432                    first.trim().parse().context("frame_batch first")?,
2433                    last.trim().parse().context("frame_batch last")?,
2434                ));
2435            }
2436            other => anyhow::bail!("unknown manifest key: {other}"),
2437        }
2438    }
2439    // A v2 manifest without an epoch is corrupt, not legacy — the writer that
2440    // stamped the v2 header always writes one. Defaulting it to 0 would
2441    // silently demote a fenced generation to unfenced, which is the one
2442    // direction this whole mechanism must never fail in.
2443    if v2 && epoch.is_none() {
2444        anyhow::bail!("{MANIFEST_HEADER_V2} manifest is missing `epoch`");
2445    }
2446    // R858-B19: same reasoning as the `epoch` check above. A v4 writer always
2447    // stamps the salt, so a v4 manifest without one is corrupt, not legacy —
2448    // and defaulting it to `None` would silently demote a checkable generation
2449    // to an unknown one, which is the direction this mechanism must never fail
2450    // in. The converse guard matters just as much: a pre-v4 header carrying a
2451    // `wal_salt` line is hand-edited or truncated, and trusting that salt would
2452    // let a forged line wave a spliced chain through.
2453    if v4 && salt.is_none() {
2454        anyhow::bail!("{MANIFEST_HEADER_V4} manifest is missing `wal_salt`");
2455    }
2456    if !v4 && salt.is_some() {
2457        anyhow::bail!(
2458            "manifest header {header:?} carries a `wal_salt` line — the WAL salt is a {MANIFEST_HEADER_V4} field, so this manifest is corrupt or hand-edited"
2459        );
2460    }
2461    let first_frame = first_frame.context("missing first_frame")?;
2462    let last_frame = last_frame.context("missing last_frame")?;
2463    // R761-F2: the batch list IS the frame index, so a v3 manifest that does
2464    // not tile its own range exactly would send restore looking for objects
2465    // that were never written — caught here, at parse, rather than as a 404
2466    // halfway through a replay.
2467    if v3 {
2468        anyhow::ensure!(
2469            !frame_batches.is_empty(),
2470            "{MANIFEST_HEADER_V3} manifest has no `frame_batch` lines — it cannot say where its frames are"
2471        );
2472        let mut expected = first_frame;
2473        for &(first, last) in &frame_batches {
2474            anyhow::ensure!(
2475                first == expected && last >= first,
2476                "frame_batch {first}-{last} does not continue the range at frame {expected}"
2477            );
2478            expected = last + 1;
2479        }
2480        anyhow::ensure!(
2481            expected == last_frame + 1,
2482            "frame_batch list covers frames {first_frame}..={} but the manifest claims {first_frame}..={last_frame}",
2483            expected - 1,
2484        );
2485    } else if !frame_batches.is_empty() {
2486        anyhow::bail!(
2487            "manifest header {header:?} carries `frame_batch` lines — batching is a {MANIFEST_HEADER_V3} feature, so this manifest is corrupt or hand-edited"
2488        );
2489    }
2490    Ok(OwnedGenerationManifest {
2491        base_snapshot_key: base_snapshot_key.context("missing base_snapshot")?,
2492        page_size: page_size.context("missing page_size")?,
2493        checkpoint_seq: checkpoint_seq.context("missing checkpoint_seq")?,
2494        salt,
2495        first_frame,
2496        last_frame,
2497        epoch: epoch.unwrap_or(0),
2498        owner,
2499        frame_batches,
2500    })
2501}
2502
2503/// A watermark as read back from the sidecar: the engine position plus the
2504/// wall-clock instant the sidecar was last written (R574-T4 — this is what
2505/// [`RpoStatus::watermark_age`] measures against). `written_at_nanos` is
2506/// `None` for sidecars written before the timestamp field existed.
2507#[derive(Debug, Clone, PartialEq, Eq)]
2508struct PersistedWatermark {
2509    watermark: Watermark,
2510    /// R858-B19: the WAL generation `watermark.last_frame` is a position
2511    /// *within*. `generation.salt == None` marks a sidecar written before the
2512    /// salt fields existed — unknown, therefore never provably equal to the
2513    /// live WAL, therefore a forced restart on the next tail. That is the
2514    /// deliberate choice: one redundant re-upload beats resuming into a WAL
2515    /// nobody can show is the same one.
2516    generation: WalGeneration,
2517    written_at_nanos: Option<u128>,
2518    /// R732-F2: the fencing epoch of the writer that last advanced this
2519    /// sidecar. `0` for a sidecar written before the field existed, which
2520    /// reads as "unfenced" and therefore fences nobody — the same
2521    /// behaviour-preserving default as [`StreamConfig::epoch`].
2522    ///
2523    /// This is the value [`tail_frames`] compares against, so it is the single
2524    /// piece of state the whole fence rests on. R732-T3 makes advancing it a
2525    /// compare-and-swap (R732-T3), so the *concurrent* hole is closed too:
2526    /// two writers racing the read-modify-write cannot both win, because the
2527    /// loser's conditional put fails against the version the winner replaced.
2528    epoch: u64,
2529    /// R736-T2: the pointer generation of the writer that last advanced this
2530    /// sidecar. `0` for a sidecar written before the field existed, which
2531    /// reads as "unfenced" and therefore fences nobody — the same
2532    /// behaviour-preserving default as [`StreamConfig::pointer_generation`].
2533    /// Checked alongside `epoch` in [`tail_frames`]; either being stale
2534    /// bounces the writer.
2535    pointer_generation: u64,
2536    /// R732-T3: the object version this record was read at, carried so the
2537    /// next advance can be conditional on it. `None` only when the sidecar
2538    /// does not exist yet, which selects [`PutMode::Create`] instead of
2539    /// [`PutMode::Update`] — "I believe nobody owns this sink" is as much a
2540    /// precondition as "I believe it is still at version V".
2541    version: Option<UpdateVersion>,
2542}
2543
2544/// Sidecar format:
2545/// `<checkpoint_seq> <last_frame> [<written_at_unix_nanos> [<epoch>
2546/// [<pointer_generation> [<salt1> <salt2>]]]]`.
2547///
2548/// The third field was added by R574-T4, the fourth by R732-F2, the fifth by
2549/// R736-T2, and the sixth and seventh by R858-B19; all are positional appends,
2550/// which is what keeps this readable in both directions. A shorter sidecar (an
2551/// older writer) still parses, with the missing tail defaulting to `None` /
2552/// `0`, and an older reader stops after its last known field and never sees the
2553/// newer ones. Fields are only ever appended for exactly that reason — do not
2554/// reorder them.
2555///
2556/// R858-B19: a missing salt pair parses to `WalGeneration { salt: None }`,
2557/// which is *unknown*, not *unchanged*. `tail_frames` restarts against it
2558/// rather than resuming — see [`PersistedWatermark::generation`]. The two
2559/// fields are read as a pair: a sidecar carrying only one of them is corrupt
2560/// (this writer emits both or neither) and is rejected rather than
2561/// half-trusted.
2562async fn read_watermark(
2563    store: &Arc<dyn ObjectStore>,
2564    key: &ObjPath,
2565) -> Result<Option<PersistedWatermark>> {
2566    match store.get(key).await {
2567        Ok(res) => {
2568            // Capture the version BEFORE consuming the body — `bytes()` takes
2569            // `res` by value. Both fields are kept because stores differ in
2570            // which one they honour for a conditional put (object_store's own
2571            // `UpdateVersion` docs say to preserve both).
2572            let version = UpdateVersion {
2573                e_tag: res.meta.e_tag.clone(),
2574                version: res.meta.version.clone(),
2575            };
2576            let bytes = res.bytes().await.context("reading watermark sidecar")?;
2577            let s = String::from_utf8_lossy(&bytes);
2578            let mut fields = s.split_whitespace();
2579            let seq = fields
2580                .next()
2581                .context("watermark sidecar must be '<checkpoint_seq> <last_frame> [<nanos>]'")?;
2582            let frame = fields
2583                .next()
2584                .context("watermark sidecar must be '<checkpoint_seq> <last_frame> [<nanos>]'")?;
2585            let written_at_nanos = fields
2586                .next()
2587                .map(|n| n.parse::<u128>().context("watermark written_at nanos"))
2588                .transpose()?;
2589            let epoch = fields
2590                .next()
2591                .map(|e| e.parse::<u64>().context("watermark epoch"))
2592                .transpose()?
2593                .unwrap_or(0);
2594            let pointer_generation = fields
2595                .next()
2596                .map(|g| g.parse::<u64>().context("watermark pointer_generation"))
2597                .transpose()?
2598                .unwrap_or(0);
2599            // R858-B19: both salt words or neither. A lone `salt1` is not a
2600            // half-known generation, it is a torn write — say so instead of
2601            // silently downgrading it to "unknown" and papering over it.
2602            let salt = match (fields.next(), fields.next()) {
2603                (Some(s1), Some(s2)) => Some(WalSalt {
2604                    salt1: s1.parse().context("watermark salt1")?,
2605                    salt2: s2.parse().context("watermark salt2")?,
2606                }),
2607                (None, _) => None,
2608                (Some(_), None) => anyhow::bail!(
2609                    "watermark sidecar carries salt1 but no salt2 — truncated or hand-edited; \
2610                     refusing rather than guessing which WAL generation it names"
2611                ),
2612            };
2613            let checkpoint_seq: u32 = seq.parse().context("watermark checkpoint_seq")?;
2614            Ok(Some(PersistedWatermark {
2615                watermark: Watermark {
2616                    checkpoint_seq,
2617                    last_frame: frame.parse().context("watermark last_frame")?,
2618                },
2619                generation: WalGeneration { checkpoint_seq, salt },
2620                written_at_nanos,
2621                epoch,
2622                pointer_generation,
2623                version: Some(version),
2624            }))
2625        }
2626        Err(object_store::Error::NotFound { .. }) => Ok(None),
2627        Err(e) => Err(e).context("fetching watermark sidecar"),
2628    }
2629}
2630
2631/// R732-T3: result of a conditional watermark advance.
2632#[derive(Debug, Clone, Copy, PartialEq, Eq)]
2633enum WatermarkCas {
2634    /// We held the version we read, and the sidecar now names us.
2635    Advanced,
2636    /// Somebody else replaced the sidecar between our read and our write.
2637    /// Says nothing about *who* — the caller re-reads to find out whether it
2638    /// was a newer owner (we are fenced) or a same-epoch racer (a bug).
2639    Contended,
2640}
2641
2642/// Advance the watermark sidecar **conditionally** on the version it was read
2643/// at (R732-T3 / W245).
2644///
2645/// This is the atomic half of the fence. The epoch check in [`tail_frames`]
2646/// stops a *sequential* stale owner — one that returns after a transfer and
2647/// reads a sidecar already stamped higher. It cannot stop a *concurrent* one:
2648/// two writers that both read the old sidecar in the same instant would both
2649/// pass that check and then both blindly overwrite, last write winning, which
2650/// is exactly the corruption W245 describes. Making the advance a
2651/// compare-and-swap removes the window — at most one of them holds the version
2652/// the other replaced.
2653///
2654/// No new CAS primitive was needed: `object_store` already models this as
2655/// [`PutMode::Update`] (→ [`object_store::Error::Precondition`]) and
2656/// [`PutMode::Create`] (→ `AlreadyExists`) for the first write, and R2 honours
2657/// the underlying `If-Match` / `If-None-Match`.
2658///
2659/// **Deployment caveat, not a code path:** `AmazonS3Builder` must be
2660/// configured for conditional puts against a store that supports them. If the
2661/// backend silently degrades to unconditional writes, this returns `Advanced`
2662/// unconditionally and the concurrent window reopens — the sequential fence in
2663/// `tail_frames` still holds, but the race does not.
2664async fn write_watermark(
2665    store: &Arc<dyn ObjectStore>,
2666    key: &ObjPath,
2667    w: Watermark,
2668    salt: Option<WalSalt>,
2669    epoch: u64,
2670    pointer_generation: u64,
2671    expected: Option<&UpdateVersion>,
2672) -> Result<WatermarkCas> {
2673    let mut body = format!(
2674        "{} {} {} {} {}",
2675        w.checkpoint_seq,
2676        w.last_frame,
2677        unix_nanos(),
2678        epoch,
2679        pointer_generation
2680    );
2681    // R858-B19: omitted entirely rather than written as a sentinel, so an
2682    // unknown generation is one shape (absent) on both the read and write
2683    // sides, and no magic value can ever be mistaken for a real salt.
2684    if let Some(s) = salt {
2685        body.push_str(&format!(" {} {}", s.salt1, s.salt2));
2686    }
2687    body.push('\n');
2688    let opts = PutOptions {
2689        mode: match expected {
2690            Some(v) => PutMode::Update(v.clone()),
2691            // No sidecar when we read: assert that is *still* true, so two
2692            // writers bootstrapping the same fresh sink cannot both proceed.
2693            None => PutMode::Create,
2694        },
2695        ..Default::default()
2696    };
2697    match store.put_opts(key, body.into_bytes().into(), opts).await {
2698        Ok(_) => Ok(WatermarkCas::Advanced),
2699        Err(object_store::Error::Precondition { .. })
2700        | Err(object_store::Error::AlreadyExists { .. }) => Ok(WatermarkCas::Contended),
2701        Err(e) => Err(e).with_context(|| format!("writing watermark sidecar {key}")),
2702    }
2703}
2704
2705/// Which conditional-put mode a preflight probe found unenforced.
2706///
2707/// Both matter and they fail independently — a store can honour `If-None-Match`
2708/// (guarding the bootstrap put) while ignoring `If-Match` (guarding every
2709/// steady-state advance), or the reverse. Naming which one degraded is the
2710/// difference between an operator fixing a bucket setting in a minute and
2711/// bisecting a corruption in a week.
2712#[derive(Debug, Clone, Copy, PartialEq, Eq)]
2713pub enum PreflightStage {
2714    /// [`PutMode::Create`] / `If-None-Match` — the guard on two writers
2715    /// bootstrapping the same fresh sink.
2716    Create,
2717    /// [`PutMode::Update`] / `If-Match` — the guard on every subsequent
2718    /// watermark advance, and therefore the one the steady state rests on.
2719    Update,
2720}
2721
2722impl PreflightStage {
2723    pub fn as_str(&self) -> &'static str {
2724        match self {
2725            PreflightStage::Create => "PutMode::Create (If-None-Match)",
2726            PreflightStage::Update => "PutMode::Update (If-Match)",
2727        }
2728    }
2729}
2730
2731/// What [`probe_conditional_puts`] observed at a real sink.
2732#[derive(Debug, Clone, Copy, PartialEq, Eq)]
2733pub enum PreconditionSupport {
2734    /// Both conditional modes were enforced — a put that should have been
2735    /// rejected was rejected. The watermark CAS is real at this sink.
2736    Honoured,
2737    /// A put that *must* have failed its precondition succeeded instead, so
2738    /// this backend has silently degraded to unconditional writes.
2739    Degraded { stage: PreflightStage },
2740}
2741
2742impl PreconditionSupport {
2743    pub fn is_honoured(&self) -> bool {
2744        matches!(self, PreconditionSupport::Honoured)
2745    }
2746}
2747
2748/// R732-T4: prove, against the *real* configured sink, that conditional puts
2749/// are actually enforced — and hand the caller a hard answer so it can refuse
2750/// to start when they are not.
2751///
2752/// [`write_watermark`] carries a deployment caveat that no test can close:
2753/// if `AmazonS3Builder` is pointed at a store that does not honour
2754/// `If-Match` / `If-None-Match`, every conditional put silently succeeds, the
2755/// CAS degrades to last-write-wins, and the concurrent split-brain window
2756/// W245 exists to shut reopens — while every unit test stays green, because
2757/// the in-memory store used in tests does honour them. Support is a property
2758/// of configuration, not of code, so a runtime probe is the only guard that
2759/// can exist.
2760///
2761/// The probe writes a uniquely-keyed canary under `<prefix>/preflight/`,
2762/// exercises both modes with puts that MUST be rejected, and deletes it. It
2763/// touches no watermark, no manifest and no frame, so it is safe to run at
2764/// startup against a live sink another node owns — and it is deliberately
2765/// keyed per-process-per-nanosecond so two nodes probing at once cannot fail
2766/// each other.
2767///
2768/// `Ok(Degraded)` is the interesting return and is NOT an error: the probe
2769/// worked perfectly, and reported that the store is unsafe. An `Err` means the
2770/// probe could not reach a verdict at all (sink unreachable, credentials
2771/// wrong), which is also a refuse-to-start condition but a different one for
2772/// an operator to read.
2773pub async fn probe_conditional_puts(target: &BackupTarget) -> Result<PreconditionSupport> {
2774    let key = join_key(
2775        &target.prefix,
2776        &format!("preflight/conditional-put-{:020}-{}.canary", unix_nanos(), std::process::id()),
2777    );
2778    let result = probe_at_key(&target.store, &key).await;
2779    // Best-effort cleanup: a leaked canary is inert (nothing reads
2780    // `preflight/`), so a delete failure must not mask the verdict — which is
2781    // the whole reason the caller ran this.
2782    if let Err(e) = target.store.delete(&key).await {
2783        tracing_delete_failure(&key, &e);
2784    }
2785    result
2786}
2787
2788/// The probe body, split out so the canary is deleted on every path.
2789async fn probe_at_key(
2790    store: &Arc<dyn ObjectStore>,
2791    key: &ObjPath,
2792) -> Result<PreconditionSupport> {
2793    let create = PutOptions { mode: PutMode::Create, ..Default::default() };
2794
2795    // 1. Claim the key. This one is *expected* to succeed; if it doesn't, the
2796    //    probe cannot reach a verdict (and `AlreadyExists` here means the
2797    //    per-process-per-nanosecond key collided, which is a bug, not a store
2798    //    property — so it stays an error rather than a `Degraded` verdict).
2799    let first = store
2800        .put_opts(key, b"preflight-1".as_slice().into(), create.clone())
2801        .await
2802        .with_context(|| format!("preflight canary could not be created at {key}"))?;
2803    let v1 = UpdateVersion { e_tag: first.e_tag.clone(), version: first.version.clone() };
2804
2805    // 2. Create again over a key that now exists. A store honouring
2806    //    `If-None-Match` rejects this; one that succeeds has degraded.
2807    match store.put_opts(key, b"preflight-2".as_slice().into(), create).await {
2808        Err(object_store::Error::AlreadyExists { .. })
2809        | Err(object_store::Error::Precondition { .. }) => {}
2810        Ok(_) => return Ok(PreconditionSupport::Degraded { stage: PreflightStage::Create }),
2811        Err(e) => {
2812            return Err(e).with_context(|| format!("preflight Create probe failed at {key}"))
2813        }
2814    }
2815
2816    // 3. Advance the canary conditionally on the version we hold, to obtain a
2817    //    *superseded* version. Expected to succeed — it is the ordinary
2818    //    steady-state write `write_watermark` makes.
2819    store
2820        .put_opts(
2821            key,
2822            b"preflight-3".as_slice().into(),
2823            PutOptions { mode: PutMode::Update(v1.clone()), ..Default::default() },
2824        )
2825        .await
2826        .with_context(|| format!("preflight Update probe could not advance {key}"))?;
2827
2828    // 4. Update again on the now-stale version — exactly the shape of a fenced
2829    //    writer losing the watermark race. A store honouring `If-Match`
2830    //    rejects it.
2831    match store
2832        .put_opts(
2833            key,
2834            b"preflight-4".as_slice().into(),
2835            PutOptions { mode: PutMode::Update(v1), ..Default::default() },
2836        )
2837        .await
2838    {
2839        Err(object_store::Error::Precondition { .. })
2840        | Err(object_store::Error::AlreadyExists { .. }) => Ok(PreconditionSupport::Honoured),
2841        Ok(_) => Ok(PreconditionSupport::Degraded { stage: PreflightStage::Update }),
2842        Err(e) => Err(e).with_context(|| format!("preflight Update probe failed at {key}")),
2843    }
2844}
2845
2846/// turso-backup takes no logging dependency (it is a library consumed by
2847/// binaries that pick their own), so a failed canary cleanup goes to stderr
2848/// rather than through `tracing`.
2849fn tracing_delete_failure(key: &ObjPath, e: &object_store::Error) {
2850    eprintln!("turso-backup: preflight canary {key} could not be deleted: {e}");
2851}
2852
2853/// R574-T4: compute the watermark-staleness snapshot for this call. Pure so
2854/// the breach edge cases are unit-testable without staging a sidecar.
2855fn rpo_status(target: Option<Duration>, prior: Option<&PersistedWatermark>) -> RpoStatus {
2856    let watermark_age = prior
2857        .and_then(|p| p.written_at_nanos)
2858        .and_then(|written_at| unix_nanos().checked_sub(written_at))
2859        .map(|nanos| Duration::from_nanos(u64::try_from(nanos).unwrap_or(u64::MAX)));
2860    let breached = matches!((target, watermark_age), (Some(t), Some(age)) if age > t);
2861    RpoStatus { target, watermark_age, breached }
2862}
2863
2864fn join_key(prefix: &str, leaf: &str) -> ObjPath {
2865    let prefix = prefix.trim_matches('/');
2866    if prefix.is_empty() {
2867        ObjPath::from(leaf)
2868    } else {
2869        ObjPath::from(format!("{prefix}/{leaf}"))
2870    }
2871}
2872
2873fn unix_nanos() -> u128 {
2874    SystemTime::now()
2875        .duration_since(UNIX_EPOCH)
2876        .map(|d| d.as_nanos())
2877        .unwrap_or(0)
2878}
2879
2880/// Default [`StreamGcConfig::grace`] — 24 hours.
2881///
2882/// Conservative on purpose. The window has to cover the longest gap between an
2883/// object becoming unreachable-looking and a live reader or writer finishing
2884/// with it, and there are three such gaps, the largest of which is not bounded
2885/// by anything this crate controls:
2886///
2887/// 1. **A restore in flight.** [`restore_stream_from_manifests`] lists the
2888///    chain, then fetches the base, then walks every frame object one at a
2889///    time with the GETs serialized (see
2890///    [`list_and_parse_generation_manifests`] on why). A cold restore of a
2891///    large database over a slow link is minutes to hours, and it is reading
2892///    objects it named *before* the GC listed anything.
2893/// 2. **A tail mid-upload.** [`tail_frames`] uploads frame batches and only
2894///    then writes the generation manifest naming them, so between those two
2895///    points a perfectly live batch is referenced by no manifest and looks
2896///    exactly like an orphan.
2897/// 3. **A rebase mid-flight.** `crate::tail::rebase` publishes the new base
2898///    *before* deleting the old generation manifests, so in that window the
2899///    surviving manifests name the OLD base and the new one is reachable from
2900///    nothing. The grace window is the only thing that keeps a concurrent GC
2901///    from deleting the base a recovery just published.
2902///
2903/// 24 hours also swallows clock skew between the object store's
2904/// `last_modified` and this process's wall clock, which is what the age
2905/// comparison is made against. A day of retained garbage costs storage; an
2906/// hour too few costs a restore.
2907pub const DEFAULT_STREAM_GC_GRACE: Duration = Duration::from_secs(24 * 60 * 60);
2908
2909/// Knobs for [`gc_stream`].
2910///
2911/// [`Default`] is `{ grace: DEFAULT_STREAM_GC_GRACE, dry_run: true }` — the
2912/// posture for something a human points at a live bucket.
2913#[derive(Debug, Clone, PartialEq, Eq)]
2914pub struct StreamGcConfig {
2915    /// Never collect an object younger than this. See
2916    /// [`DEFAULT_STREAM_GC_GRACE`].
2917    pub grace: Duration,
2918    /// Report what would be collected and delete nothing. **Defaults to
2919    /// `true`.**
2920    pub dry_run: bool,
2921}
2922
2923impl Default for StreamGcConfig {
2924    fn default() -> Self {
2925        Self {
2926            grace: DEFAULT_STREAM_GC_GRACE,
2927            dry_run: true,
2928        }
2929    }
2930}
2931
2932/// What a [`gc_stream`] pass found — and, when `dry_run` was false, deleted.
2933///
2934/// The collected keys are listed rather than counted because the primary
2935/// consumer is a human reading a dry run before authorizing the real one.
2936#[derive(Debug, Clone, PartialEq, Eq, Default)]
2937pub struct StreamGcOutcome {
2938    /// Echo of [`StreamGcConfig::dry_run`]. `true` means nothing was deleted
2939    /// and `collected_*` is a proposal.
2940    pub dry_run: bool,
2941    /// Every snapshot key a restore could still reach, sorted. See
2942    /// [`gc_stream`]'s "Liveness" section for how this is derived.
2943    pub live_base_snapshot_keys: Vec<String>,
2944    /// Superseded base snapshots, sorted.
2945    pub collected_snapshots: Vec<String>,
2946    /// Orphaned frame objects (batch or legacy per-frame), sorted.
2947    pub collected_frame_objects: Vec<String>,
2948    /// Total bytes across `collected_snapshots` + `collected_frame_objects`.
2949    pub collected_bytes: u64,
2950    /// Snapshot objects kept — because they are live, or because the whole
2951    /// snapshot half was skipped (see `base_snapshots_skipped`).
2952    pub retained_snapshots: usize,
2953    /// `true` when the prefix held **no generation manifests**, so this is not
2954    /// provably a tier-2 sink and no base snapshot was collected.
2955    ///
2956    /// A `snapshots/` prefix with no chain over it is indistinguishable from a
2957    /// plain tier-1a sink, whose older snapshots are *history* rather than
2958    /// garbage — [`crate::snapshot::restore_latest`] takes the newest, but an
2959    /// operator restoring a point in time names an older key by hand. Pruning
2960    /// those is [`crate::dedup::gc_dedup`]'s `keep_n` decision to make, not
2961    /// this sweep's. A live tier-2 sink always has a chain, so in the steady
2962    /// state this is `false` and the superseded bases are collected; the one
2963    /// way to see it `true` on a real stream is a GC that lands inside
2964    /// `crate::tail::rebase`'s manifests-deleted-frames-not-yet-written window,
2965    /// where skipping one pass costs nothing.
2966    pub base_snapshots_skipped: bool,
2967    /// Frame objects kept because a generation manifest still names them.
2968    pub retained_frame_objects: usize,
2969    /// Objects that were unreachable but too young to touch — the grace window
2970    /// did its job. A number that never falls to zero across consecutive runs
2971    /// means the grace is longer than the churn interval, not that the sweep
2972    /// is broken.
2973    pub spared_by_grace: usize,
2974}
2975
2976/// Tier-2 GC: reclaim the objects a `crate::tail::rebase` orphans.
2977///
2978/// The tier-1b counterpart is [`crate::dedup::gc_dedup`], and this deliberately
2979/// mirrors its shape: an explicitly-invoked sweep that deletes only what
2980/// nothing can reach, never a step on the write or recovery path. `rebase`
2981/// leaves its garbage behind on purpose — deleting data as part of a recovery
2982/// path is how a recovery path becomes the outage — and this is where that
2983/// debt is settled.
2984///
2985/// ## Two orphan classes, not one
2986///
2987/// R850-T3 was filed against the frames alone. The frames are the **smaller**
2988/// term:
2989///
2990/// - **Orphaned frame objects.** Each `rebase` abandons
2991///   `frames/{old_checkpoint_seq}/…` (or `frames/{epoch}/{seq}/…`) once
2992///   `crate::tail::delete_generation_manifests` removes the manifests that
2993///   named them. Bounded per rebase by SQLite's autocheckpoint threshold —
2994///   ~1000 pages, i.e. a handful of batch objects.
2995/// - **Superseded base snapshots.** `rebase` calls
2996///   [`crate::snapshot::upload_base_snapshot`], which is the explicitly
2997///   *non*-deduplicating one-shot variant: it `put`s a new
2998///   `snapshots/snapshot-{nanos}.db` and deletes nothing. Every rebase
2999///   therefore leaves a complete permanent copy of the database behind, and
3000///   for any database bigger than a few megabytes that dwarfs the frames.
3001///   Found while implementing this ticket; covered here rather than filed,
3002///   because it is the same leak in the same function.
3003///
3004/// A sweep that reclaimed only the frames would leave the larger leak running.
3005///
3006/// ## Liveness
3007///
3008/// Nothing is deleted unless *no* restore path can reach it. The rule is the
3009/// exact complement of what a restore selects, and the two are pinned together
3010/// by `gc_liveness_is_the_complement_of_restore_selection` in this module's
3011/// tests so they cannot drift:
3012///
3013/// - **Frames.** Live iff some generation manifest names the object, resolved
3014///   through `BackupTarget::frame_objects_of` — the same function restore's
3015///   replay and [`crate::puller::WalPuller`] go through, so batch and legacy
3016///   per-frame layouts and both epoch key shapes are handled by construction
3017///   rather than by a second copy of the layout rules here.
3018/// - **Base snapshots.** Live iff *either* some generation manifest names it,
3019///   *or* it is the lexically-greatest key under `snapshots/`. Those are the
3020///   two selections a restore makes:
3021///   [`restore_stream_from_manifests`] takes the chain's `base_snapshot_key`,
3022///   and a prefix with no generations falls back to
3023///   [`crate::snapshot::restore_latest`], which takes the newest key (see
3024///   [`crate::hydrate`]'s `restore_subject`). Their union is kept.
3025///
3026/// The base rule uses *every* manifest's key rather than
3027/// `validate_generation_chain`'s single answer, and that is the safe
3028/// direction: a chain that spans a WAL restart does not validate at all, and a
3029/// GC that refuses to guess keeps both bases instead of deleting the one the
3030/// next rebase is about to adopt.
3031///
3032/// And when there is no chain at all, the snapshot half is skipped entirely
3033/// rather than falling back to "keep the newest, collect the rest" — see
3034/// [`StreamGcOutcome::base_snapshots_skipped`] for why that distinction is not
3035/// paranoia. The frame half still runs: a `frames/` prefix under a sink with no
3036/// generation manifests is unreachable by construction.
3037///
3038/// ## Grace window
3039///
3040/// An object younger than [`StreamGcConfig::grace`] is never collected, no
3041/// matter how unreachable it looks. [`DEFAULT_STREAM_GC_GRACE`] documents the
3042/// three live windows that depend on it.
3043///
3044/// ## Order
3045///
3046/// Frames first, then snapshots. Unlike [`crate::dedup::gc_dedup`] the order is
3047/// not load-bearing — this sweep deletes no manifests, so nothing retained ever
3048/// points at anything collected, at any point during the pass or after a crash
3049/// in the middle of one.
3050pub async fn gc_stream(target: &BackupTarget, cfg: &StreamGcConfig) -> Result<StreamGcOutcome> {
3051    let manifests = list_and_parse_generation_manifests(target).await?;
3052
3053    // Normalize through ObjPath so a manifest's textual key compares equal to
3054    // the same object's listed location regardless of slash padding.
3055    let mut live_bases: HashSet<String> = manifests
3056        .iter()
3057        .map(|m| ObjPath::from(m.base_snapshot_key.as_str()).to_string())
3058        .collect();
3059    if let Some(newest) = crate::snapshot::latest_snapshot_key(target).await? {
3060        live_bases.insert(ObjPath::from(newest).to_string());
3061    }
3062
3063    let mut live_frames: HashSet<String> = HashSet::new();
3064    for m in &manifests {
3065        for (key, _first, _last) in target.frame_objects_of(m) {
3066            live_frames.insert(key.to_string());
3067        }
3068    }
3069
3070    let mut out = StreamGcOutcome {
3071        dry_run: cfg.dry_run,
3072        base_snapshots_skipped: manifests.is_empty(),
3073        ..Default::default()
3074    };
3075    out.live_base_snapshot_keys = live_bases.iter().cloned().collect();
3076    out.live_base_snapshot_keys.sort();
3077
3078    let now_secs = SystemTime::now()
3079        .duration_since(UNIX_EPOCH)
3080        .map(|d| d.as_secs())
3081        .unwrap_or(0) as i64;
3082    let grace_secs = i64::try_from(cfg.grace.as_secs()).unwrap_or(i64::MAX);
3083    let too_young = |meta: &object_store::ObjectMeta| {
3084        now_secs.saturating_sub(meta.last_modified.timestamp()) < grace_secs
3085    };
3086
3087    for meta in list_recursive(&target.store, &join_key(&target.prefix, "frames")).await? {
3088        if live_frames.contains(&meta.location.to_string()) {
3089            out.retained_frame_objects += 1;
3090            continue;
3091        }
3092        if too_young(&meta) {
3093            out.spared_by_grace += 1;
3094            continue;
3095        }
3096        if !cfg.dry_run {
3097            delete_collected(target, &meta.location).await?;
3098        }
3099        out.collected_bytes += meta.size;
3100        out.collected_frame_objects.push(meta.location.to_string());
3101    }
3102
3103    let snapshots_prefix = join_key(&target.prefix, "snapshots");
3104    let listing = target
3105        .store
3106        .list_with_delimiter(Some(&snapshots_prefix))
3107        .await
3108        .with_context(|| format!("listing base snapshots under {snapshots_prefix}"))?;
3109    for meta in listing.objects {
3110        // No chain over this prefix means it is not provably a tier-2 sink, and
3111        // a tier-1a sink's older snapshots are history rather than garbage —
3112        // see `StreamGcOutcome::base_snapshots_skipped`.
3113        if out.base_snapshots_skipped || live_bases.contains(&meta.location.to_string()) {
3114            out.retained_snapshots += 1;
3115            continue;
3116        }
3117        if too_young(&meta) {
3118            out.spared_by_grace += 1;
3119            continue;
3120        }
3121        if !cfg.dry_run {
3122            delete_collected(target, &meta.location).await?;
3123        }
3124        out.collected_bytes += meta.size;
3125        out.collected_snapshots.push(meta.location.to_string());
3126    }
3127
3128    out.collected_frame_objects.sort();
3129    out.collected_snapshots.sort();
3130    Ok(out)
3131}
3132
3133/// Delete one collected object, treating "already gone" as success.
3134///
3135/// Two GC passes can legitimately overlap, and a retry after a partial one
3136/// lands here too — in both cases the object being absent is the outcome we
3137/// asked for, exactly as it is for `crate::tail::rebase`'s deletes.
3138async fn delete_collected(target: &BackupTarget, key: &ObjPath) -> Result<()> {
3139    match target.store.delete(key).await {
3140        Ok(()) => Ok(()),
3141        Err(object_store::Error::NotFound { .. }) => Ok(()),
3142        Err(e) => Err(anyhow::Error::new(e)).with_context(|| format!("collecting {key}")),
3143    }
3144}
3145
3146/// Every object at or below `root`, walked with `list_with_delimiter` rather
3147/// than the streaming `list`.
3148///
3149/// `ObjectStore::list` returns a `BoxStream`, which needs `futures_util` to
3150/// drain — and this crate deliberately keeps `futures_util` in
3151/// `[dev-dependencies]` (see `Cargo.toml`), so a streaming drain here would add
3152/// a real dependency to the published crate for one listing. A worklist over
3153/// `list_with_delimiter` costs one request per directory instead of one per
3154/// page, and the tier-2 frame tree is three levels deep at most
3155/// (`frames/{epoch}/{seq}/{object}`), so that is a handful of requests.
3156async fn list_recursive(
3157    store: &Arc<dyn ObjectStore>,
3158    root: &ObjPath,
3159) -> Result<Vec<object_store::ObjectMeta>> {
3160    let mut found = Vec::new();
3161    let mut pending = vec![root.clone()];
3162    while let Some(dir) = pending.pop() {
3163        let listing = store
3164            .list_with_delimiter(Some(&dir))
3165            .await
3166            .with_context(|| format!("listing {dir}"))?;
3167        found.extend(listing.objects);
3168        pending.extend(listing.common_prefixes);
3169    }
3170    Ok(found)
3171}
3172
3173#[cfg(test)]
3174mod tests {
3175    use super::*;
3176    use crate::backpressure::fault_injection::{Fault, FaultyStore};
3177    use crate::backpressure::BackoffConfig;
3178    use object_store::memory::InMemory;
3179    use std::cell::RefCell;
3180    use std::time::Duration;
3181
3182    /// In-memory WAL seam: a fixed page_size, a vector of frames the test
3183    /// appends to, and an advancing checkpoint_seq the test can bump.
3184    struct MockWal {
3185        page_size: usize,
3186        state: RefCell<MockState>,
3187    }
3188    struct MockState {
3189        checkpoint_seq: u32,
3190        /// R858-B19: the WAL header salt this mock's frames are stamped with.
3191        /// Modelled explicitly because the two fold regimes move it
3192        /// differently, and the whole bug was reading only the sequence:
3193        /// [`MockWal::restart`] moves both, [`MockWal::recreate`] moves only
3194        /// this one.
3195        salt: WalSalt,
3196        frames: Vec<MockFrame>,
3197        auto_actions_disabled: bool,
3198        /// R858-B19: re-roll `salt` once this many frame reads have happened,
3199        /// so a test can land a WAL recreate *inside* a single `tail_frames`
3200        /// call rather than only between two. See
3201        /// [`MockWal::recreate_after_reads`].
3202        swap_salt_after_reads: Option<u32>,
3203        reads: u32,
3204    }
3205    #[derive(Clone)]
3206    struct MockFrame {
3207        info: FrameInfo,
3208        page_bytes: Vec<u8>,
3209    }
3210
3211    impl MockWal {
3212        fn new(page_size: usize) -> Self {
3213            Self {
3214                page_size,
3215                state: RefCell::new(MockState {
3216                    checkpoint_seq: 0,
3217                    salt: WalSalt { salt1: 0xd492_ea8a, salt2: 0x7acf_42a3 },
3218                    frames: Vec::new(),
3219                    auto_actions_disabled: false,
3220                    swap_salt_after_reads: None,
3221                    reads: 0,
3222                }),
3223            }
3224        }
3225        fn append(&self, page_no: u32, db_size: u32, fill: u8) {
3226            let mut s = self.state.borrow_mut();
3227            s.frames.push(MockFrame {
3228                info: FrameInfo { page_no, db_size },
3229                page_bytes: vec![fill; self.page_size],
3230            });
3231        }
3232        /// Simulate an **in-process** WAL restart — a long-lived connection
3233        /// folding its own WAL at the autocheckpoint threshold. The WAL file is
3234        /// reused, so `checkpoint_seq` advances and `salt1` advances with it
3235        /// (measured `d492ea8a -> d492ea8b` alongside seq `0 -> 1`). This is
3236        /// the regime the pre-R858-B19 `checkpoint_seq` test already caught.
3237        fn restart(&self) {
3238            let mut s = self.state.borrow_mut();
3239            s.checkpoint_seq += 1;
3240            s.salt.salt1 = s.salt.salt1.wrapping_add(1);
3241            s.frames.clear();
3242        }
3243        /// R858-B19 — simulate a **writer-process restart**: the last
3244        /// connection closed, SQLite checkpointed and DELETED the `-wal` file,
3245        /// and the next writer created a fresh WAL. `checkpoint_seq` goes back
3246        /// to 0 (it never moves, from a watcher's point of view: `0 -> 0`) and
3247        /// the salt is fresh randomness, unrelated to the old one.
3248        ///
3249        /// This is the regime `checkpoint_seq` is blind to, and the reason the
3250        /// mock models a salt at all.
3251        fn recreate(&self, salt: WalSalt) {
3252            let mut s = self.state.borrow_mut();
3253            s.checkpoint_seq = 0;
3254            s.salt = salt;
3255            s.frames.clear();
3256        }
3257        /// R858-B19: arm a mid-call WAL recreate — the salt flips once `reads`
3258        /// frame reads have been served, which lands it inside the drain loop
3259        /// rather than between two `tail_frames` calls.
3260        fn recreate_after_reads(&self, reads: u32) {
3261            self.state.borrow_mut().swap_salt_after_reads = Some(reads);
3262        }
3263        fn frame_count(&self) -> u64 {
3264            self.state.borrow().frames.len() as u64
3265        }
3266    }
3267
3268    impl WalSeam for MockWal {
3269        fn wal_state(&self) -> Result<Watermark> {
3270            let s = self.state.borrow();
3271            Ok(Watermark {
3272                checkpoint_seq: s.checkpoint_seq,
3273                last_frame: s.frames.len() as u64,
3274            })
3275        }
3276        fn wal_get_frame(&self, frame_no: u64, buf: &mut [u8]) -> Result<FrameInfo> {
3277            let mut s = self.state.borrow_mut();
3278            s.reads += 1;
3279            if s.swap_salt_after_reads == Some(s.reads) {
3280                s.salt.salt1 = !s.salt.salt1;
3281            }
3282            let idx = frame_no
3283                .checked_sub(1)
3284                .context("frame_no must be >= 1")? as usize;
3285            let f = s
3286                .frames
3287                .get(idx)
3288                .with_context(|| format!("frame {frame_no} out of range"))?;
3289            // Synthesize a 24-byte header: big-endian page_no, db_size, the
3290            // WAL's salt pair (R858-B19 reads the generation out of exactly
3291            // these bytes), then zeros for the checksums (never validated).
3292            buf[0..4].copy_from_slice(&f.info.page_no.to_be_bytes());
3293            buf[4..8].copy_from_slice(&f.info.db_size.to_be_bytes());
3294            buf[8..12].copy_from_slice(&s.salt.salt1.to_be_bytes());
3295            buf[12..16].copy_from_slice(&s.salt.salt2.to_be_bytes());
3296            buf[16..WAL_FRAME_HEADER_SIZE].fill(0);
3297            buf[WAL_FRAME_HEADER_SIZE..].copy_from_slice(&f.page_bytes);
3298            Ok(f.info)
3299        }
3300        fn wal_auto_actions_disable(&self) {
3301            self.state.borrow_mut().auto_actions_disabled = true;
3302        }
3303    }
3304
3305    /// A store that accepts every put unconditionally — the exact failure the
3306    /// preflight probe exists to catch. It is not a contrived shape: it is
3307    /// what an S3-compatible backend without conditional-write support looks
3308    /// like from `object_store`'s side, and what `AmazonS3Builder` degrades to
3309    /// when pointed at one. Everything but `put_opts` delegates.
3310    #[derive(Debug)]
3311    struct UnconditionalStore {
3312        inner: Arc<dyn ObjectStore>,
3313    }
3314
3315    impl std::fmt::Display for UnconditionalStore {
3316        fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
3317            write!(f, "UnconditionalStore({})", self.inner)
3318        }
3319    }
3320
3321    #[async_trait::async_trait]
3322    impl ObjectStore for UnconditionalStore {
3323        async fn put_opts(
3324            &self,
3325            location: &ObjPath,
3326            payload: object_store::PutPayload,
3327            mut opts: PutOptions,
3328        ) -> object_store::Result<object_store::PutResult> {
3329            opts.mode = PutMode::Overwrite;
3330            self.inner.put_opts(location, payload, opts).await
3331        }
3332        async fn put_multipart_opts(
3333            &self,
3334            location: &ObjPath,
3335            opts: object_store::PutMultipartOptions,
3336        ) -> object_store::Result<Box<dyn object_store::MultipartUpload>> {
3337            self.inner.put_multipart_opts(location, opts).await
3338        }
3339        async fn get_opts(
3340            &self,
3341            location: &ObjPath,
3342            options: object_store::GetOptions,
3343        ) -> object_store::Result<object_store::GetResult> {
3344            self.inner.get_opts(location, options).await
3345        }
3346        fn delete_stream(
3347            &self,
3348            locations: futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>>,
3349        ) -> futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>> {
3350            self.inner.delete_stream(locations)
3351        }
3352        fn list(
3353            &self,
3354            prefix: Option<&ObjPath>,
3355        ) -> futures_util::stream::BoxStream<'static, object_store::Result<object_store::ObjectMeta>>
3356        {
3357            self.inner.list(prefix)
3358        }
3359        async fn list_with_delimiter(
3360            &self,
3361            prefix: Option<&ObjPath>,
3362        ) -> object_store::Result<object_store::ListResult> {
3363            self.inner.list_with_delimiter(prefix).await
3364        }
3365        async fn copy_opts(
3366            &self,
3367            from: &ObjPath,
3368            to: &ObjPath,
3369            options: object_store::CopyOptions,
3370        ) -> object_store::Result<()> {
3371            self.inner.copy_opts(from, to, options).await
3372        }
3373    }
3374
3375    /// A store that honours `If-None-Match` but not `If-Match`. The nastiest
3376    /// real-world shape, because the bootstrap put looks fine and only the
3377    /// steady-state advance — the one every tail after the first depends on —
3378    /// is unguarded.
3379    #[derive(Debug)]
3380    struct CreateOnlyStore {
3381        inner: Arc<dyn ObjectStore>,
3382    }
3383
3384    impl std::fmt::Display for CreateOnlyStore {
3385        fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
3386            write!(f, "CreateOnlyStore({})", self.inner)
3387        }
3388    }
3389
3390    #[async_trait::async_trait]
3391    impl ObjectStore for CreateOnlyStore {
3392        async fn put_opts(
3393            &self,
3394            location: &ObjPath,
3395            payload: object_store::PutPayload,
3396            mut opts: PutOptions,
3397        ) -> object_store::Result<object_store::PutResult> {
3398            if matches!(opts.mode, PutMode::Update(_)) {
3399                opts.mode = PutMode::Overwrite;
3400            }
3401            self.inner.put_opts(location, payload, opts).await
3402        }
3403        async fn put_multipart_opts(
3404            &self,
3405            location: &ObjPath,
3406            opts: object_store::PutMultipartOptions,
3407        ) -> object_store::Result<Box<dyn object_store::MultipartUpload>> {
3408            self.inner.put_multipart_opts(location, opts).await
3409        }
3410        async fn get_opts(
3411            &self,
3412            location: &ObjPath,
3413            options: object_store::GetOptions,
3414        ) -> object_store::Result<object_store::GetResult> {
3415            self.inner.get_opts(location, options).await
3416        }
3417        fn delete_stream(
3418            &self,
3419            locations: futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>>,
3420        ) -> futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>> {
3421            self.inner.delete_stream(locations)
3422        }
3423        fn list(
3424            &self,
3425            prefix: Option<&ObjPath>,
3426        ) -> futures_util::stream::BoxStream<'static, object_store::Result<object_store::ObjectMeta>>
3427        {
3428            self.inner.list(prefix)
3429        }
3430        async fn list_with_delimiter(
3431            &self,
3432            prefix: Option<&ObjPath>,
3433        ) -> object_store::Result<object_store::ListResult> {
3434            self.inner.list_with_delimiter(prefix).await
3435        }
3436        async fn copy_opts(
3437            &self,
3438            from: &ObjPath,
3439            to: &ObjPath,
3440            options: object_store::CopyOptions,
3441        ) -> object_store::Result<()> {
3442            self.inner.copy_opts(from, to, options).await
3443        }
3444    }
3445
3446    /// The happy path: a store that honours both modes passes, and — the part
3447    /// that matters for running this at startup against a live sink — leaves
3448    /// nothing behind.
3449    #[tokio::test]
3450    async fn the_preflight_probe_passes_on_a_conditional_store_and_leaves_no_trace() {
3451        let target = fresh_target();
3452        assert_eq!(
3453            probe_conditional_puts(&target).await.unwrap(),
3454            PreconditionSupport::Honoured
3455        );
3456
3457        let leftovers = objects_under(&target).await;
3458        assert!(
3459            leftovers.is_empty(),
3460            "the probe must clean up its canary — found {leftovers:?}"
3461        );
3462    }
3463
3464    /// The whole point: a backend that silently ignores preconditions is
3465    /// *reported*, not tolerated. Without this the watermark CAS degrades to
3466    /// last-write-wins and every other test in this file still passes.
3467    #[tokio::test]
3468    async fn the_preflight_probe_catches_a_store_that_ignores_preconditions() {
3469        let target = BackupTarget {
3470            store: Arc::new(UnconditionalStore { inner: Arc::new(InMemory::new()) }),
3471            prefix: "backups".into(),
3472        };
3473        assert_eq!(
3474            probe_conditional_puts(&target).await.unwrap(),
3475            PreconditionSupport::Degraded { stage: PreflightStage::Create },
3476            "an unconditional store fails at the first guard it meets"
3477        );
3478    }
3479
3480    /// Half-degraded stores are the ones that actually ship. `If-None-Match`
3481    /// works, so bootstrapping looks healthy; `If-Match` does not, so every
3482    /// steady-state advance is unguarded. The probe must name `Update`
3483    /// specifically — "conditional puts are broken" would send an operator to
3484    /// the wrong setting.
3485    #[tokio::test]
3486    async fn the_preflight_probe_names_update_when_only_if_match_is_ignored() {
3487        let target = BackupTarget {
3488            store: Arc::new(CreateOnlyStore { inner: Arc::new(InMemory::new()) }),
3489            prefix: "backups".into(),
3490        };
3491        assert_eq!(
3492            probe_conditional_puts(&target).await.unwrap(),
3493            PreconditionSupport::Degraded { stage: PreflightStage::Update }
3494        );
3495    }
3496
3497    /// The canary is keyed per process per nanosecond, so two nodes probing
3498    /// the same sink at once each get a real verdict instead of failing each
3499    /// other. A startup probe that flaked under concurrency would be turned
3500    /// off within a week.
3501    #[tokio::test]
3502    async fn concurrent_preflight_probes_do_not_collide() {
3503        let target = fresh_target();
3504        let (a, b) = tokio::join!(
3505            probe_conditional_puts(&target),
3506            probe_conditional_puts(&target)
3507        );
3508        assert_eq!(a.unwrap(), PreconditionSupport::Honoured);
3509        assert_eq!(b.unwrap(), PreconditionSupport::Honoured);
3510        assert!(objects_under(&target).await.is_empty());
3511    }
3512
3513    async fn objects_under(target: &BackupTarget) -> Vec<String> {
3514        use futures_util::StreamExt;
3515        target
3516            .store
3517            .list(None)
3518            .map(|m| m.unwrap().location.to_string())
3519            .collect()
3520            .await
3521    }
3522
3523    fn fresh_target() -> BackupTarget {
3524        BackupTarget {
3525            store: Arc::new(InMemory::new()),
3526            prefix: "backups".into(),
3527        }
3528    }
3529
3530    fn cfg() -> StreamConfig<'static> {
3531        StreamConfig {
3532            base_snapshot_key: "backups/snapshots/snapshot-00000000000000000001.db",
3533            page_size: 4096,
3534            backpressure: BackpressureConfig::default(),
3535            rpo_target: None,
3536            epoch: 0,
3537            owner: None,
3538            pointer_generation: 0,
3539        }
3540    }
3541
3542    /// Fast-ticking backoff for backpressure tests — a few ms, never the
3543    /// production defaults, so retry-heavy tests stay fast.
3544    fn fast_backoff() -> BackoffConfig {
3545        BackoffConfig {
3546            initial_delay: Duration::from_millis(1),
3547            max_delay: Duration::from_millis(4),
3548            multiplier: 2.0,
3549            max_retries: 3,
3550        }
3551    }
3552
3553    fn bp_cfg(policy: BackpressurePolicy, spill_buffer_frames: usize) -> StreamConfig<'static> {
3554        StreamConfig {
3555            base_snapshot_key: "backups/snapshots/snapshot-00000000000000000001.db",
3556            page_size: 4096,
3557            backpressure: BackpressureConfig {
3558                spill_buffer_frames,
3559                policy,
3560                backoff: fast_backoff(),
3561            },
3562            rpo_target: None,
3563            epoch: 0,
3564            owner: None,
3565            pointer_generation: 0,
3566        }
3567    }
3568
3569    fn faulty_target(faults: impl IntoIterator<Item = Fault>) -> BackupTarget {
3570        BackupTarget {
3571            store: Arc::new(FaultyStore::new(Arc::new(InMemory::new()), faults)),
3572            prefix: "backups".into(),
3573        }
3574    }
3575
3576    /// Five WAL frames (1..=5), the last a commit — a small, deterministic
3577    /// range for the backpressure tests below.
3578    fn five_frame_seam() -> MockWal {
3579        let seam = MockWal::new(4096);
3580        for i in 1..=4u32 {
3581            seam.append(i, 0, i as u8);
3582        }
3583        seam.append(5, 5, 5); // commit
3584        seam
3585    }
3586
3587    // --- R574-F2: explicit R2 backpressure ---------------------------------
3588
3589    /// A single 429 triggers exactly one retry, then the upload succeeds —
3590    /// the whole range still lands and the report shows the retry.
3591    #[tokio::test]
3592    async fn throttled_429_retries_then_succeeds() {
3593        let seam = five_frame_seam();
3594        let target = faulty_target([Fault::TooManyRequests]);
3595        let cfg = bp_cfg(BackpressurePolicy::Fail, 8);
3596        let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3597        match out {
3598            StreamOutcome::Streamed { frame_count, backpressure, .. } => {
3599                assert_eq!(frame_count, 5);
3600                assert_eq!(backpressure.throttle_retries, 1);
3601                assert_eq!(backpressure.frames_shed, 0);
3602            }
3603            other => panic!("expected Streamed, got {other:?}"),
3604        }
3605    }
3606
3607    /// A 503 is classified the same as a 429 and also triggers backoff then
3608    /// retry — two 503s in a row cost exactly two retries.
3609    #[tokio::test]
3610    async fn throttled_503_retries_then_succeeds() {
3611        let seam = five_frame_seam();
3612        let target = faulty_target([Fault::ServiceUnavailable, Fault::ServiceUnavailable]);
3613        let cfg = bp_cfg(BackpressurePolicy::Fail, 8);
3614        let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3615        match out {
3616            StreamOutcome::Streamed { frame_count, backpressure, .. } => {
3617                assert_eq!(frame_count, 5);
3618                assert_eq!(backpressure.throttle_retries, 2);
3619            }
3620            other => panic!("expected Streamed, got {other:?}"),
3621        }
3622    }
3623
3624    /// Buffer at bound + `Fail`: sustained throttling exhausts backoff and
3625    /// the whole call errors — nothing is persisted.
3626    #[tokio::test]
3627    async fn buffer_at_bound_fail_policy_errors_without_persisting() {
3628        let seam = five_frame_seam();
3629        let faults = std::iter::repeat_n(Fault::TooManyRequests, 50);
3630        let target = faulty_target(faults);
3631        let cfg = bp_cfg(BackpressurePolicy::Fail, 2);
3632        let err = tail_frames(&seam, &target, &cfg).await.unwrap_err();
3633        assert!(format!("{err}").contains("uploading wal frame"), "err was {err}");
3634        assert!(
3635            read_watermark(&target.store, &target.watermark_key())
3636                .await
3637                .unwrap()
3638                .is_none(),
3639            "Fail must not persist a watermark when it gives up"
3640        );
3641    }
3642
3643    /// Buffer at bound + `Shed`: the backlog is dropped with a loud report
3644    /// instead of erroring; nothing persists, so the next call would
3645    /// re-attempt the same range.
3646    #[tokio::test]
3647    async fn buffer_at_bound_shed_policy_drops_backlog_without_error() {
3648        let seam = five_frame_seam();
3649        let faults = std::iter::repeat_n(Fault::TooManyRequests, 50);
3650        let target = faulty_target(faults);
3651        let cfg = bp_cfg(BackpressurePolicy::Shed, 2);
3652        let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3653        match out {
3654            StreamOutcome::Shed { first_frame, last_frame, backpressure, .. } => {
3655                assert_eq!(first_frame, 1);
3656                assert_eq!(last_frame, 5);
3657                assert_eq!(backpressure.high_water_frames, 2, "capped at the bound");
3658                assert!(backpressure.frames_shed >= 1, "must report the dropped backlog");
3659            }
3660            other => panic!("expected Shed, got {other:?}"),
3661        }
3662        assert!(
3663            read_watermark(&target.store, &target.watermark_key())
3664                .await
3665                .unwrap()
3666                .is_none(),
3667            "Shed must not persist a watermark for a fully-dropped batch"
3668        );
3669    }
3670
3671    /// Buffer at bound + `Block`: retries never give up on a throttling
3672    /// error; once the store recovers, the full range still lands.
3673    #[tokio::test]
3674    async fn buffer_at_bound_block_policy_eventually_drains() {
3675        let seam = five_frame_seam();
3676        // More failures than fast_backoff's max_retries would tolerate under
3677        // Fail/Shed — Block must push through them anyway.
3678        let faults = std::iter::repeat_n(Fault::TooManyRequests, 4);
3679        let target = faulty_target(faults);
3680        let cfg = bp_cfg(BackpressurePolicy::Block, 2);
3681        let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3682        match out {
3683            StreamOutcome::Streamed { frame_count, backpressure, .. } => {
3684                assert_eq!(frame_count, 5);
3685                assert_eq!(backpressure.high_water_frames, 2);
3686                assert_eq!(backpressure.throttle_retries, 4);
3687                assert_eq!(backpressure.frames_shed, 0);
3688            }
3689            other => panic!("expected Streamed, got {other:?}"),
3690        }
3691    }
3692
3693    /// R761-F2: a call with more frames than the spill buffer holds splits
3694    /// into one batch object per buffer-full, the manifest indexes every one
3695    /// of them, and replay puts the frames back in order. The bound is the
3696    /// only thing sizing an object, so this is also the assertion that the
3697    /// largest object this sink writes stays bounded.
3698    #[tokio::test]
3699    async fn a_call_larger_than_the_spill_buffer_splits_into_several_batch_objects() {
3700        let seam = MockWal::new(4096);
3701        for i in 1..=5u32 {
3702            seam.append(i, i, i as u8);
3703        }
3704        let target = fresh_target();
3705        let cfg = bp_cfg(BackpressurePolicy::Fail, 2);
3706        let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3707        assert!(matches!(out, StreamOutcome::Streamed { frame_count: 5, .. }));
3708
3709        let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
3710        assert_eq!(
3711            manifests[0].frame_batches,
3712            vec![(1, 2), (3, 4), (5, 5)],
3713            "bound 2 over 5 frames is two full batches and a remainder"
3714        );
3715        let frame_size = WAL_FRAME_HEADER_SIZE + 4096;
3716        for (first, last) in [(1u64, 2u64), (3, 4), (5, 5)] {
3717            let bytes = target
3718                .store
3719                .get(&target.frame_batch_key(0, 0, first, last))
3720                .await
3721                .unwrap()
3722                .bytes()
3723                .await
3724                .unwrap();
3725            assert_eq!(bytes.len() as u64, (last - first + 1) * frame_size as u64);
3726        }
3727
3728        // And it all comes back, in order, through the normal read path.
3729        let insert = MockInsertSeam::new();
3730        assert_eq!(replay_frames_into(&target, &insert, &manifests).await.unwrap(), 5);
3731        let frames: Vec<u64> = insert
3732            .events()
3733            .into_iter()
3734            .filter_map(|e| match e {
3735                MockInsertEvent::Frame { frame_no, .. } => Some(frame_no),
3736                _ => None,
3737            })
3738            .collect();
3739        assert_eq!(frames, vec![1, 2, 3, 4, 5]);
3740    }
3741
3742    /// R761-F2 + R574-F2: a drain that lands one batch and then gets stuck
3743    /// under `Shed` publishes ONLY the batch that landed. The manifest's index
3744    /// is what restore follows, so a batch list claiming a shed object would
3745    /// be a 404 mid-replay — worse than the frames simply not being there.
3746    #[tokio::test]
3747    async fn a_partially_shed_drain_indexes_only_the_batches_that_landed() {
3748        let seam = five_frame_seam();
3749        // First batch through; the next one throttled until Shed gives up.
3750        // Exactly enough faults to exhaust one put's retry budget and no more,
3751        // so the watermark + manifest writes that follow the shed still land
3752        // (they go through the same store).
3753        let faults = std::iter::once(Fault::Pass).chain(std::iter::repeat_n(
3754            Fault::TooManyRequests,
3755            fast_backoff().max_retries as usize + 1,
3756        ));
3757        let target = faulty_target(faults);
3758        let cfg = bp_cfg(BackpressurePolicy::Shed, 2);
3759        let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3760        match out {
3761            StreamOutcome::Streamed { first_frame, last_frame, frame_count, backpressure, .. } => {
3762                assert_eq!((first_frame, last_frame, frame_count), (1, 2, 2));
3763                assert_eq!(backpressure.frames_shed, 2, "frames 3-4 were buffered and dropped");
3764            }
3765            other => panic!("expected a Streamed prefix, got {other:?}"),
3766        }
3767        let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
3768        assert_eq!(manifests.len(), 1);
3769        assert_eq!(manifests[0].last_frame, 2);
3770        assert_eq!(manifests[0].frame_batches, vec![(1, 2)]);
3771        // Every object the manifest names actually exists — the property that
3772        // makes the published prefix restorable.
3773        let insert = MockInsertSeam::new();
3774        assert_eq!(replay_frames_into(&target, &insert, &manifests).await.unwrap(), 2);
3775    }
3776
3777    /// The high-water mark reports the peak spill-buffer occupancy for the
3778    /// call, capped at the configured bound even when more frames remain.
3779    #[tokio::test]
3780    async fn high_water_reports_peak_buffered_frames() {
3781        let seam = five_frame_seam();
3782        let target = faulty_target([]); // no faults — pure high-water measurement
3783        let cfg = bp_cfg(BackpressurePolicy::Fail, 3);
3784        let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3785        match out {
3786            StreamOutcome::Streamed { backpressure, .. } => {
3787                assert_eq!(backpressure.high_water_frames, 3, "capped at the configured bound");
3788            }
3789            other => panic!("expected Streamed, got {other:?}"),
3790        }
3791    }
3792
3793    // --- R574-T4: explicit RPO knob --------------------------------------
3794
3795    /// First tail ever: no sidecar, so age is unknown and a configured
3796    /// target cannot be breached (there is nothing to measure against).
3797    #[tokio::test]
3798    async fn rpo_first_tail_has_unknown_age_and_no_breach() {
3799        let seam = five_frame_seam();
3800        let target = fresh_target();
3801        let mut cfg = cfg();
3802        cfg.rpo_target = Some(Duration::from_secs(1));
3803        let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3804        match out {
3805            StreamOutcome::Streamed { rpo, .. } => {
3806                assert_eq!(rpo.target, Some(Duration::from_secs(1)));
3807                assert_eq!(rpo.watermark_age, None);
3808                assert!(!rpo.breached);
3809            }
3810            other => panic!("expected Streamed, got {other:?}"),
3811        }
3812    }
3813
3814    /// Second tail past the target: the sidecar's stamped write instant is
3815    /// older than `rpo_target`, so the outcome flags a breach — the
3816    /// staleness emission the orchestrator alerts on.
3817    #[tokio::test]
3818    async fn rpo_stale_watermark_past_target_reports_breach() {
3819        let seam = MockWal::new(4096);
3820        seam.append(1, 1, 0xAA);
3821        let target = fresh_target();
3822        let mut cfg = cfg();
3823        // Zero target: any measurable gap between the two tails is a breach.
3824        cfg.rpo_target = Some(Duration::ZERO);
3825        let _ = tail_frames(&seam, &target, &cfg).await.unwrap();
3826        tokio::time::sleep(Duration::from_millis(5)).await;
3827        seam.append(2, 2, 0xBB);
3828        let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3829        match out {
3830            StreamOutcome::Streamed { rpo, .. } => {
3831                let age = rpo.watermark_age.expect("age known after first sidecar write");
3832                assert!(age >= Duration::from_millis(5), "age was {age:?}");
3833                assert!(rpo.breached, "zero target must flag any nonzero age");
3834            }
3835            other => panic!("expected Streamed, got {other:?}"),
3836        }
3837    }
3838
3839    /// A generous target with a prompt second tail: age is reported but the
3840    /// bound holds — and with no target at all, `breached` is always false.
3841    #[tokio::test]
3842    async fn rpo_within_target_and_no_target_do_not_breach() {
3843        let seam = MockWal::new(4096);
3844        seam.append(1, 1, 0xAA);
3845        let target = fresh_target();
3846        let mut with_target = cfg();
3847        with_target.rpo_target = Some(Duration::from_secs(3600));
3848        let _ = tail_frames(&seam, &target, &with_target).await.unwrap();
3849
3850        seam.append(2, 2, 0xBB);
3851        let out = tail_frames(&seam, &target, &with_target).await.unwrap();
3852        match out {
3853            StreamOutcome::Streamed { rpo, .. } => {
3854                assert!(rpo.watermark_age.is_some());
3855                assert!(!rpo.breached, "an hour budget can't be blown in-process");
3856            }
3857            other => panic!("expected Streamed, got {other:?}"),
3858        }
3859
3860        // Same staged sidecar, target removed: age still reported, never
3861        // breached (observe-before-you-pick-a-number mode).
3862        seam.append(3, 3, 0xCC);
3863        let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
3864        match out {
3865            StreamOutcome::Streamed { rpo, .. } => {
3866                assert_eq!(rpo.target, None);
3867                assert!(rpo.watermark_age.is_some());
3868                assert!(!rpo.breached);
3869            }
3870            other => panic!("expected Streamed, got {other:?}"),
3871        }
3872    }
3873
3874    /// A pre-T4 two-field sidecar still parses (written_at unknown), and the
3875    /// next write upgrades it to the stamped three-field format.
3876    #[tokio::test]
3877    async fn rpo_legacy_two_field_sidecar_parses_and_upgrades() {
3878        let target = fresh_target();
3879        target
3880            .store
3881            .put(&target.watermark_key(), b"0 1\n".to_vec().into())
3882            .await
3883            .unwrap();
3884        let legacy = read_watermark(&target.store, &target.watermark_key())
3885            .await
3886            .unwrap()
3887            .unwrap();
3888        assert_eq!(legacy.watermark, Watermark { checkpoint_seq: 0, last_frame: 1 });
3889        assert_eq!(legacy.written_at_nanos, None);
3890
3891        // A tail against the legacy sidecar reports unknown age (not a
3892        // breach), uploads the frames, and re-stamps the sidecar.
3893        //
3894        // R858-B19 CHANGED THE OUTCOME HERE, deliberately: this used to assert
3895        // `Streamed` with `first_frame == 2` ("resumes after the legacy
3896        // watermark"). A two-field sidecar records no salt, so nothing says the
3897        // WAL it names is the WAL in front of us — and resuming at frame 2 on
3898        // that basis is the exact inference that spliced two generations. The
3899        // sidecar's RPO semantics (an *unknown* age is not a breach) are what
3900        // this test is about and they are untouched.
3901        let seam = MockWal::new(4096);
3902        seam.append(1, 0, 0xAA);
3903        seam.append(2, 2, 0xBB);
3904        let mut cfg = cfg();
3905        cfg.rpo_target = Some(Duration::ZERO);
3906        let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3907        match out {
3908            StreamOutcome::Restarted { rpo, first_frame, previous_generation, .. } => {
3909                assert_eq!(first_frame, 1, "an unverifiable watermark is re-uploaded, not resumed");
3910                assert_eq!(previous_generation.salt, None);
3911                assert_eq!(rpo.watermark_age, None);
3912                assert!(!rpo.breached, "unknown age is not a breach even at zero target");
3913            }
3914            other => panic!("expected Restarted, got {other:?}"),
3915        }
3916        let upgraded = read_watermark(&target.store, &target.watermark_key())
3917            .await
3918            .unwrap()
3919            .unwrap();
3920        assert!(upgraded.written_at_nanos.is_some(), "rewrite stamps the timestamp");
3921    }
3922
3923    /// Pure rpo_status edge cases that don't need a staged store.
3924    #[test]
3925    fn rpo_status_truth_table() {
3926        let stamped = |nanos_ago: u128| PersistedWatermark {
3927            watermark: Watermark { checkpoint_seq: 0, last_frame: 1 },
3928            generation: WalGeneration::default(),
3929            written_at_nanos: Some(unix_nanos().saturating_sub(nanos_ago)),
3930            epoch: 0,
3931            pointer_generation: 0,
3932            version: None,
3933        };
3934        // No prior at all.
3935        let s = rpo_status(Some(Duration::from_secs(1)), None);
3936        assert_eq!((s.watermark_age, s.breached), (None, false));
3937        // Prior without a stamp (legacy sidecar).
3938        let legacy = PersistedWatermark {
3939            watermark: Watermark::default(),
3940            generation: WalGeneration::default(),
3941            written_at_nanos: None,
3942            epoch: 0,
3943            pointer_generation: 0,
3944            version: None,
3945        };
3946        let s = rpo_status(Some(Duration::ZERO), Some(&legacy));
3947        assert_eq!((s.watermark_age, s.breached), (None, false));
3948        // Old stamp vs tight target: breached.
3949        let s = rpo_status(Some(Duration::from_millis(1)), Some(&stamped(5_000_000_000)));
3950        assert!(s.watermark_age.unwrap() >= Duration::from_secs(4));
3951        assert!(s.breached);
3952        // Old stamp, no target: age known, never breached.
3953        let s = rpo_status(None, Some(&stamped(5_000_000_000)));
3954        assert!(s.watermark_age.is_some());
3955        assert!(!s.breached);
3956    }
3957
3958    /// R761-T1: the shipped cadence default is an arithmetic consequence of
3959    /// the 2026-08-13 `tail_sweep_harness` table, so the arithmetic is a
3960    /// test rather than a claim in a comment nobody re-checks. If a future
3961    /// harness run moves the measured points, this test is where the
3962    /// mismatch surfaces — re-derive the default, do not relax the test.
3963    #[test]
3964    fn default_tail_cadence_matches_the_measured_write_op_curve() {
3965        // Every uploading tail_frames call writes exactly two fixed objects
3966        // beyond the frames: one generation manifest, one watermark CAS.
3967        const FIXED_OBJECTS_PER_TAIL: f64 = 2.0;
3968        // Measured, flat across the whole sweep — a property of the schema
3969        // and transaction shape, not of the cadence.
3970        const MEASURED_FRAMES_PER_WRITE: f64 = 2.03;
3971        let puts_per_write =
3972            |w: f64| MEASURED_FRAMES_PER_WRITE + FIXED_OBJECTS_PER_TAIL / w;
3973
3974        // The model reproduces all four measured rows.
3975        for (writes_per_tail, measured) in [(1.0, 4.03), (5.0, 2.43), (25.0, 2.11), (100.0, 2.05)] {
3976            let modelled = puts_per_write(writes_per_tail);
3977            assert!(
3978                (modelled - measured).abs() < 0.005,
3979                "writes_per_tail={writes_per_tail}: model {modelled:.3} vs measured {measured:.3}"
3980            );
3981        }
3982
3983        // The default is tailed at half the stated bound, so one missed tick
3984        // still lands inside the promise.
3985        assert_eq!(DEFAULT_RPO_TARGET, DEFAULT_TAIL_INTERVAL * 2);
3986
3987        // At the burst rates the default was chosen against (~0.1-1 write/s
3988        // during an active session), 60s lands in the flat part of the curve:
3989        // the fixed-object term is under a fifth of the frame floor even at
3990        // the slow end, where 15s would still be paying 1.33.
3991        let slow_burst_writes_per_tail = 0.1 * DEFAULT_TAIL_INTERVAL.as_secs_f64();
3992        let fixed_term = FIXED_OBJECTS_PER_TAIL / slow_burst_writes_per_tail;
3993        assert!(
3994            fixed_term < MEASURED_FRAMES_PER_WRITE / 5.0,
3995            "fixed-object term {fixed_term:.3} at the slow-burst end is no longer small \
3996             relative to the {MEASURED_FRAMES_PER_WRITE} frame floor"
3997        );
3998        assert!(
3999            puts_per_write(slow_burst_writes_per_tail) < 2.4,
4000            "the default's worst modelled case should sit below the measured \
4001             writes_per_tail=5 point (2.43)"
4002        );
4003
4004        // R761-F2 removed the frames/write term: a tail call whose frames fit
4005        // one batch writes the batch, the manifest and the watermark, full
4006        // stop. Same curve shape, no floor.
4007        let batched_puts_per_write = |w: f64| (1.0 + FIXED_OBJECTS_PER_TAIL) / w;
4008        for (writes_per_tail, pre_batching) in [(1.0, 4.03), (5.0, 2.43), (25.0, 2.11), (100.0, 2.05)]
4009        {
4010            let cut = 1.0 - batched_puts_per_write(writes_per_tail) / pre_batching;
4011            let want = match writes_per_tail as u32 {
4012                1 => 0.256,
4013                5 => 0.753,
4014                25 => 0.943,
4015                _ => 0.985,
4016            };
4017            assert!(
4018                (cut - want).abs() < 0.005,
4019                "writes_per_tail={writes_per_tail}: batching cuts {:.1}%, expected {:.1}%",
4020                cut * 100.0,
4021                want * 100.0,
4022            );
4023        }
4024        // Which is why the cadence default did not move: what is left to win
4025        // past 60s is now a hundredth of a PUT per write at the fast-burst end
4026        // of the same band, against an RPO window that would grow 5x.
4027        assert!(
4028            batched_puts_per_write(1.0 * DEFAULT_TAIL_INTERVAL.as_secs_f64())
4029                - batched_puts_per_write(5.0 * DEFAULT_TAIL_INTERVAL.as_secs_f64())
4030                < 0.05,
4031            "batching should have flattened the cadence lever at the fast-burst end"
4032        );
4033    }
4034
4035    /// R782: `StreamOutcome::rpo()` is the accessor `tenant-streamer` pushes
4036    /// through — every variant that carries an `RpoStatus` gives it back, and
4037    /// `Fenced` (which deliberately carries none, per its own doc) gives
4038    /// `None` rather than a default/synthesized one.
4039    #[test]
4040    fn stream_outcome_rpo_accessor_covers_every_variant() {
4041        let rpo = RpoStatus { target: None, watermark_age: Some(Duration::from_secs(1)), breached: false };
4042        assert_eq!(StreamOutcome::Empty { watermark: Watermark::default(), rpo }.rpo(), Some(&rpo));
4043        assert_eq!(
4044            StreamOutcome::Streamed {
4045                generation_key: String::new(),
4046                first_frame: 1,
4047                last_frame: 1,
4048                checkpoint_seq: 0,
4049                frame_count: 1,
4050                backpressure: Default::default(),
4051                rpo,
4052            }
4053            .rpo(),
4054            Some(&rpo)
4055        );
4056        assert_eq!(
4057            StreamOutcome::Restarted {
4058                generation_key: String::new(),
4059                previous_generation: WalGeneration::default(),
4060                new_generation: WalGeneration {
4061                    checkpoint_seq: 1,
4062                    salt: Some(WalSalt { salt1: 1, salt2: 2 }),
4063                },
4064                first_frame: 1,
4065                last_frame: 1,
4066                frame_count: 1,
4067                backpressure: Default::default(),
4068                rpo,
4069            }
4070            .rpo(),
4071            Some(&rpo)
4072        );
4073        assert_eq!(
4074            StreamOutcome::Shed {
4075                checkpoint_seq: 0,
4076                first_frame: 1,
4077                last_frame: 1,
4078                backpressure: Default::default(),
4079                rpo,
4080            }
4081            .rpo(),
4082            Some(&rpo)
4083        );
4084        assert_eq!(
4085            StreamOutcome::Fenced {
4086                current_epoch: 2,
4087                our_epoch: 1,
4088                current_pointer_generation: 0,
4089                our_pointer_generation: 0,
4090            }
4091            .rpo(),
4092            None
4093        );
4094    }
4095
4096    /// First tail with no frames yet — Empty, no manifest, no watermark.
4097    #[tokio::test]
4098    async fn empty_when_no_frames() {
4099        let seam = MockWal::new(4096);
4100        let target = fresh_target();
4101        let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4102        match out {
4103            StreamOutcome::Empty { watermark, rpo } => {
4104                assert_eq!(watermark, Watermark::default());
4105                // No sidecar has ever been written — age is unknowable.
4106                assert_eq!(rpo, RpoStatus::default());
4107            }
4108            other => panic!("expected Empty, got {other:?}"),
4109        }
4110        // No manifest, no watermark sidecar.
4111        assert!(
4112            read_watermark(&target.store, &target.watermark_key())
4113                .await
4114                .unwrap()
4115                .is_none()
4116        );
4117    }
4118
4119    /// First tail with 3 frames — Streamed, manifest written, watermark
4120    /// recorded, every frame object retrievable at the right key.
4121    #[tokio::test]
4122    async fn streams_initial_frames_and_records_watermark() {
4123        let seam = MockWal::new(4096);
4124        seam.append(1, 0, 0xAA);
4125        seam.append(2, 0, 0xBB);
4126        seam.append(3, 3, 0xCC); // commit frame: db_size = 3 pages
4127
4128        let target = fresh_target();
4129        let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4130        let gen_key = match out {
4131            StreamOutcome::Streamed {
4132                generation_key,
4133                first_frame,
4134                last_frame,
4135                checkpoint_seq,
4136                frame_count,
4137                ..
4138            } => {
4139                assert_eq!(first_frame, 1);
4140                assert_eq!(last_frame, 3);
4141                assert_eq!(checkpoint_seq, 0);
4142                assert_eq!(frame_count, 3);
4143                assert!(generation_key.starts_with("backups/generations/gen-"));
4144                generation_key
4145            }
4146            other => panic!("expected Streamed, got {other:?}"),
4147        };
4148
4149        // R761-F2: ONE batch object holds all three frames, at the ranged key,
4150        // each frame at its own offset inside it.
4151        let frame_size = WAL_FRAME_HEADER_SIZE + 4096;
4152        let bytes = target
4153            .store
4154            .get(&target.frame_batch_key(0, 0, 1, 3))
4155            .await
4156            .unwrap()
4157            .bytes()
4158            .await
4159            .unwrap();
4160        assert_eq!(bytes.len(), 3 * frame_size);
4161        for frame_no in 1..=3u64 {
4162            let at = (frame_no as usize - 1) * frame_size;
4163            let page_no = u32::from_be_bytes(bytes[at..at + 4].try_into().unwrap());
4164            assert_eq!(page_no, frame_no as u32);
4165        }
4166        // And nothing was written per-frame.
4167        assert!(
4168            target.store.get(&target.frame_key(0, 0, 1)).await.is_err(),
4169            "the pre-R761-F2 per-frame key must not be written any more"
4170        );
4171
4172        // Manifest is parseable and round-trips.
4173        let m_bytes = target
4174            .store
4175            .get(&ObjPath::from(gen_key.clone()))
4176            .await
4177            .unwrap()
4178            .bytes()
4179            .await
4180            .unwrap();
4181        let parsed = parse_generation_manifest(&String::from_utf8_lossy(&m_bytes)).unwrap();
4182        assert_eq!(parsed.first_frame, 1);
4183        assert_eq!(parsed.last_frame, 3);
4184        assert_eq!(parsed.checkpoint_seq, 0);
4185        assert_eq!(parsed.page_size, 4096);
4186        assert_eq!(parsed.base_snapshot_key, cfg().base_snapshot_key);
4187        assert_eq!(parsed.frame_batches, vec![(1, 3)], "R761-F2: one batch, indexed");
4188
4189        // Watermark sidecar matches.
4190        let wm = read_watermark(&target.store, &target.watermark_key())
4191            .await
4192            .unwrap()
4193            .unwrap();
4194        assert_eq!(wm.watermark, Watermark { checkpoint_seq: 0, last_frame: 3 });
4195        assert!(wm.written_at_nanos.is_some(), "R574-T4: writes stamp the sidecar");
4196
4197        // The seam was told to take WAL ownership? Not directly via
4198        // tail_frames — callers do that at construction (CoreWalSeam::open).
4199        // Verify the mock invariant separately.
4200        assert!(!seam.state.borrow().auto_actions_disabled);
4201    }
4202
4203    /// Second tail with no new frames returns Empty without uploading.
4204    #[tokio::test]
4205    async fn second_tail_with_no_new_frames_is_empty() {
4206        let seam = MockWal::new(4096);
4207        seam.append(1, 1, 0xAA);
4208        let target = fresh_target();
4209        let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4210
4211        let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4212        match out {
4213            StreamOutcome::Empty { watermark, rpo } => {
4214                assert_eq!(watermark, Watermark { checkpoint_seq: 0, last_frame: 1 });
4215                // A sidecar exists from the first tail, so age is known even
4216                // with no rpo_target configured — and no target means no breach.
4217                assert!(rpo.watermark_age.is_some());
4218                assert!(!rpo.breached);
4219            }
4220            other => panic!("expected Empty, got {other:?}"),
4221        }
4222    }
4223
4224    /// Second tail with new frames uploads only the new ones and the
4225    /// generation manifest covers the incremental range. Previous frames
4226    /// remain at their prior keys (idempotent put on same key).
4227    #[tokio::test]
4228    async fn second_tail_uploads_only_new_frames() {
4229        let seam = MockWal::new(4096);
4230        seam.append(1, 0, 0x11);
4231        seam.append(2, 2, 0x22); // commit
4232        let target = fresh_target();
4233        let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4234
4235        // Append more.
4236        seam.append(3, 0, 0x33);
4237        seam.append(4, 4, 0x44); // commit
4238
4239        let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4240        match out {
4241            StreamOutcome::Streamed {
4242                first_frame,
4243                last_frame,
4244                frame_count,
4245                ..
4246            } => {
4247                assert_eq!(first_frame, 3);
4248                assert_eq!(last_frame, 4);
4249                assert_eq!(frame_count, 2);
4250            }
4251            other => panic!("expected Streamed, got {other:?}"),
4252        }
4253        assert_eq!(seam.frame_count(), 4);
4254        // Frames 1..=4 all retrievable under checkpoint_seq=0, as one batch
4255        // object per tail call — the first call's object is untouched by the
4256        // second (disjoint ranges, so no overwrite and no re-upload).
4257        for (first, last) in [(1u64, 2u64), (3, 4)] {
4258            target
4259                .store
4260                .get(&target.frame_batch_key(0, 0, first, last))
4261                .await
4262                .unwrap();
4263        }
4264    }
4265
4266    /// WAL restart bumps checkpoint_seq; the next tail uploads under the new
4267    /// sequence and reports Restarted.
4268    #[tokio::test]
4269    async fn wal_restart_emits_restarted_under_new_sequence() {
4270        let seam = MockWal::new(4096);
4271        seam.append(1, 1, 0xAA);
4272        seam.append(2, 2, 0xBB);
4273        let target = fresh_target();
4274        let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4275
4276        // Simulate the engine restarting the WAL (e.g. after a checkpoint).
4277        seam.restart();
4278        seam.append(1, 1, 0xCC);
4279
4280        let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4281        match out {
4282            StreamOutcome::Restarted {
4283                previous_generation,
4284                new_generation,
4285                first_frame,
4286                last_frame,
4287                frame_count,
4288                ..
4289            } => {
4290                assert_eq!(previous_generation.checkpoint_seq, 0);
4291                assert_eq!(new_generation.checkpoint_seq, 1);
4292                // R858-B19: an in-process restart moves the salt too, so this
4293                // regime is now caught twice over.
4294                assert_ne!(previous_generation.salt, new_generation.salt);
4295                assert_eq!(first_frame, 1);
4296                assert_eq!(last_frame, 1);
4297                assert_eq!(frame_count, 1);
4298            }
4299            other => panic!("expected Restarted, got {other:?}"),
4300        }
4301
4302        // Old seq=0 frames still in place; new seq=1 frame under its own
4303        // sequence — same key shape, different directory.
4304        target
4305            .store
4306            .get(&target.frame_batch_key(0, 0, 1, 2))
4307            .await
4308            .unwrap();
4309        target
4310            .store
4311            .get(&target.frame_batch_key(0, 1, 1, 1))
4312            .await
4313            .unwrap();
4314    }
4315
4316    /// R858-B19, the property this whole ticket turns on: **a WAL whose salt
4317    /// changed is a different WAL, even when `checkpoint_seq` did not move.**
4318    ///
4319    /// This is the writer-process-restart regime. The last connection closed,
4320    /// SQLite checkpointed and deleted the `-wal`, and the next writer built a
4321    /// fresh WAL back at checkpoint-sequence 0. Both tails therefore see
4322    /// `checkpoint_seq == 0`, which is precisely why the old
4323    /// `p.checkpoint_seq == current.checkpoint_seq` test resumed at
4324    /// `last_frame + 1` and spliced frames from two unrelated WALs into one
4325    /// generation chain — a restore that reported success and produced a stale
4326    /// or malformed image, with nothing raised anywhere.
4327    ///
4328    /// The probe (`examples/foreign_checkpoint_probe.rs -- b f`) drives the
4329    /// same fold end-to-end against the real system `sqlite3`. This test exists
4330    /// so the property is pinned here too: the probe needs an upstream sqlite3
4331    /// binary and half a minute, and a property this load-bearing should fail
4332    /// in `cargo test` when someone re-derives "the sequence is enough".
4333    #[tokio::test]
4334    async fn a_salt_change_is_a_restart_even_when_checkpoint_seq_is_unchanged() {
4335        let seam = MockWal::new(4096);
4336        seam.append(1, 1, 0xAA);
4337        seam.append(2, 2, 0xBB);
4338        seam.append(3, 3, 0xCC);
4339        let target = fresh_target();
4340        let first = tail_frames(&seam, &target, &cfg()).await.unwrap();
4341        assert!(matches!(first, StreamOutcome::Streamed { .. }), "got {first:?}");
4342
4343        // The foreign writer restarted: brand-new WAL, unrelated salt, and a
4344        // checkpoint-sequence that reads 0 both before and after.
4345        seam.recreate(WalSalt { salt1: 0x9ca8_9e29, salt2: 0x2be5_66fc });
4346        seam.append(1, 1, 0xDD);
4347        seam.append(2, 2, 0xEE);
4348        assert_eq!(
4349            seam.wal_state().unwrap().checkpoint_seq,
4350            0,
4351            "the premise: the sequence did NOT move across the recreate"
4352        );
4353
4354        let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4355        let StreamOutcome::Restarted {
4356            previous_generation,
4357            new_generation,
4358            first_frame,
4359            last_frame,
4360            ..
4361        } = &out
4362        else {
4363            panic!(
4364                "a recreated WAL must report Restarted — got {out:?}. If this is Streamed with \
4365                 first_frame 4, the salt check has been removed and two WAL generations are being \
4366                 spliced into one chain again (R858-B19)"
4367            );
4368        };
4369        assert_eq!(previous_generation.checkpoint_seq, new_generation.checkpoint_seq);
4370        assert_ne!(
4371            previous_generation.salt, new_generation.salt,
4372            "the salt is the only thing that moved, and it is what must be noticed"
4373        );
4374        assert_eq!((*first_frame, *last_frame), (1, 2), "re-upload from the top of the new WAL");
4375
4376        // And the chain that results is refused rather than restored: two
4377        // generations both starting at frame 1 is not a stream, and restore
4378        // must say so instead of producing a plausible-looking wrong image.
4379        let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
4380        assert_eq!(manifests.len(), 2);
4381        let err = validate_generation_chain(&manifests).unwrap_err().to_string();
4382        assert!(
4383            err.contains("WAL") || err.contains("gap in stream"),
4384            "restore must refuse the post-recreate chain loudly, got: {err}"
4385        );
4386    }
4387
4388    /// R858-B19 — the narrower door: a WAL recreate that lands **inside** one
4389    /// `tail_frames` call rather than between two. The pre-drain sample says
4390    /// generation A, the frames that got read are a mix of A and B, and no
4391    /// comparison of two watermarks can see it. The call must publish nothing.
4392    #[tokio::test]
4393    async fn a_wal_recreated_mid_drain_publishes_nothing() {
4394        let seam = MockWal::new(4096);
4395        for n in 1..=3u32 {
4396            seam.append(n, n, n as u8);
4397        }
4398        let target = fresh_target();
4399        // Read #1 is the pre-drain salt probe; reads #2..#4 are the drain. Flip
4400        // the WAL underneath us on the second frame of the drain.
4401        seam.recreate_after_reads(3);
4402
4403        let err = tail_frames(&seam, &target, &cfg()).await.unwrap_err().to_string();
4404        assert!(err.contains("recreated while this tail was uploading"), "got: {err}");
4405
4406        // The sink is untouched: no watermark to resume from, no manifest to
4407        // restore. Whatever frame objects landed are orphaned and unreferenced.
4408        assert!(read_watermark(&target.store, &target.watermark_key()).await.unwrap().is_none());
4409        assert!(list_and_parse_generation_manifests(&target).await.unwrap().is_empty());
4410    }
4411
4412    /// R858-B19 — the sidecar half of the same property. A watermark persisted
4413    /// by a pre-salt writer says nothing about which WAL it names, so the next
4414    /// tail must restart rather than resume from `last_frame + 1`. Unknown is
4415    /// not "unchanged".
4416    #[tokio::test]
4417    async fn a_saltless_legacy_watermark_forces_a_restart_rather_than_a_resume() {
4418        let target = fresh_target();
4419        // Exactly what a pre-R858-B19 writer left behind: five positional
4420        // fields, no salt pair.
4421        target
4422            .store
4423            .put(&target.watermark_key(), format!("0 2 {} 0 0\n", unix_nanos()).into_bytes().into())
4424            .await
4425            .unwrap();
4426
4427        let seam = MockWal::new(4096);
4428        seam.append(1, 1, 0xAA);
4429        seam.append(2, 2, 0xBB);
4430        seam.append(3, 3, 0xCC);
4431
4432        let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4433        let StreamOutcome::Restarted { previous_generation, first_frame, .. } = &out else {
4434            panic!(
4435                "an unverifiable watermark must restart, not resume — got {out:?}. Resuming here \
4436                 trusts a position in a WAL nobody can show is the same one (R858-B19)"
4437            );
4438        };
4439        assert_eq!(previous_generation.salt, None, "the prior generation is unknown, not equal");
4440        assert_eq!(*first_frame, 1, "re-upload everything rather than trust the old position");
4441
4442        // Self-healing: the sidecar now carries a salt, so the very next tail
4443        // resumes normally. One redundant re-upload, not a permanent restart
4444        // loop.
4445        seam.append(4, 4, 0xDD);
4446        let next = tail_frames(&seam, &target, &cfg()).await.unwrap();
4447        let StreamOutcome::Streamed { first_frame, .. } = &next else {
4448            panic!("the salt is recorded now, so this must resume — got {next:?}");
4449        };
4450        assert_eq!(*first_frame, 4);
4451    }
4452
4453    /// R858-B19 — the restore-side refusal, on the shape a pre-`v4` writer
4454    /// could actually leave in a bucket: a chain that looks perfectly
4455    /// contiguous but whose manifests never recorded which WAL they came from.
4456    /// `checkpoint_seq` agrees, the frame ranges tile, and it is still
4457    /// unrestorable, because that is exactly what a splice looks like.
4458    #[test]
4459    fn validate_chain_refuses_a_multi_generation_chain_with_no_recorded_salt() {
4460        let err = validate_generation_chain(&[
4461            mk_manifest_salted("base.db", 4096, 0, None, 1, 5),
4462            mk_manifest_salted("base.db", 4096, 0, None, 6, 9),
4463        ])
4464        .unwrap_err()
4465        .to_string();
4466        assert!(err.contains("no WAL salt"), "got: {err}");
4467
4468        // One generation is not a splice — a legacy single-manifest backup
4469        // stays restorable, because there is nothing here to prove.
4470        let chain =
4471            validate_generation_chain(&[mk_manifest_salted("base.db", 4096, 0, None, 1, 5)])
4472                .unwrap();
4473        assert_eq!(chain.generation.salt, None);
4474        assert_eq!(chain.total_frames, 5);
4475    }
4476
4477    /// R858-B19 — and the same refusal when the salts are recorded and
4478    /// *disagree*: a contiguous frame range across two different WALs. This is
4479    /// the case `checkpoint_seq` can never catch, since a recreate resets it to
4480    /// the value it already had.
4481    #[test]
4482    fn validate_chain_refuses_a_chain_whose_salt_changes_mid_stream() {
4483        let other = WalSalt { salt1: 0xea10_2175, salt2: 0xb4ff_221f };
4484        let err = validate_generation_chain(&[
4485            mk_manifest_salted("base.db", 4096, 0, Some(CHAIN_SALT), 1, 30),
4486            mk_manifest_salted("base.db", 4096, 0, Some(other), 31, 57),
4487        ])
4488        .unwrap_err()
4489        .to_string();
4490        assert!(err.contains("RECREATED mid-stream"), "got: {err}");
4491    }
4492
4493    /// R858-B19 — manifest `v4` carries the salt, and the version guards move
4494    /// with it: a `v4` header without a `wal_salt` is corrupt (not legacy), and
4495    /// a pre-`v4` header *with* one is forged (not a newer writer).
4496    #[test]
4497    fn manifest_v4_round_trips_the_wal_salt_and_guards_both_directions() {
4498        let batches = [(12u64, 34u64)];
4499        let salt = WalSalt { salt1: 0xb83c_03f5, salt2: 0x0000_0001 };
4500        let text = format_generation_manifest(GenerationManifest {
4501            base_snapshot_key: "backups/snapshots/snapshot-1.db",
4502            page_size: 4096,
4503            checkpoint_seq: 7,
4504            salt: Some(salt),
4505            first_frame: 12,
4506            last_frame: 34,
4507            epoch: 9,
4508            owner: Some("node-3"),
4509            frame_batches: &batches,
4510        });
4511        assert!(text.starts_with(MANIFEST_HEADER_V4), "got: {text}");
4512        let parsed = parse_generation_manifest(&text).unwrap();
4513        assert_eq!(parsed.salt, Some(salt));
4514        assert_eq!(parsed.checkpoint_seq, 7);
4515        assert_eq!(parsed.frame_batches, batches);
4516
4517        let no_salt = text.replace(&format!("wal_salt {}-{}\n", salt.salt1, salt.salt2), "");
4518        let err = parse_generation_manifest(&no_salt).unwrap_err().to_string();
4519        assert!(err.contains("missing `wal_salt`"), "got: {err}");
4520
4521        let forged = text.replace(MANIFEST_HEADER_V4, MANIFEST_HEADER_V3);
4522        let err = parse_generation_manifest(&forged).unwrap_err().to_string();
4523        assert!(err.contains("corrupt or hand-edited"), "got: {err}");
4524    }
4525
4526    /// R858-B19 — the sidecar's salt pair is read as a pair. Half a pair is a
4527    /// torn write, and downgrading it to "unknown" would hide that.
4528    #[tokio::test]
4529    async fn watermark_sidecar_round_trips_the_salt_and_rejects_half_a_pair() {
4530        let target = fresh_target();
4531        let key = target.watermark_key();
4532        let salt = WalSalt { salt1: 0xd492_ea8a, salt2: 0x7acf_42a3 };
4533        write_watermark(
4534            &target.store,
4535            &key,
4536            Watermark { checkpoint_seq: 2, last_frame: 11 },
4537            Some(salt),
4538            6,
4539            3,
4540            None,
4541        )
4542        .await
4543        .unwrap();
4544        let read = read_watermark(&target.store, &key).await.unwrap().unwrap();
4545        assert_eq!(read.generation, WalGeneration { checkpoint_seq: 2, salt: Some(salt) });
4546
4547        target
4548            .store
4549            .put(&key, format!("2 11 {} 6 3 {}\n", unix_nanos(), salt.salt1).into_bytes().into())
4550            .await
4551            .unwrap();
4552        let err = read_watermark(&target.store, &key).await.unwrap_err().to_string();
4553        assert!(err.contains("salt1 but no salt2"), "got: {err}");
4554    }
4555
4556    /// Captured frame insert: a mock [`WalInsertSeam`] records every call so
4557    /// tests can assert ordering and content without a real turso connection.
4558    struct MockInsertSeam {
4559        log: RefCell<Vec<MockInsertEvent>>,
4560    }
4561    #[derive(Debug, Clone, PartialEq, Eq)]
4562    enum MockInsertEvent {
4563        Begin,
4564        Frame { frame_no: u64, page_no: u32, db_size: u32 },
4565        End { force_commit: bool },
4566    }
4567    impl MockInsertSeam {
4568        fn new() -> Self {
4569            Self { log: RefCell::new(Vec::new()) }
4570        }
4571        fn events(&self) -> Vec<MockInsertEvent> {
4572            self.log.borrow().clone()
4573        }
4574    }
4575    impl WalInsertSeam for MockInsertSeam {
4576        fn wal_insert_begin(&self) -> Result<()> {
4577            self.log.borrow_mut().push(MockInsertEvent::Begin);
4578            Ok(())
4579        }
4580        fn wal_insert_frame(&self, frame_no: u64, frame: &[u8]) -> Result<()> {
4581            let page_no = u32::from_be_bytes(frame[0..4].try_into().unwrap());
4582            let db_size = u32::from_be_bytes(frame[4..8].try_into().unwrap());
4583            self.log.borrow_mut().push(MockInsertEvent::Frame { frame_no, page_no, db_size });
4584            Ok(())
4585        }
4586        fn wal_insert_end(&self, force_commit: bool) -> Result<()> {
4587            self.log.borrow_mut().push(MockInsertEvent::End { force_commit });
4588            Ok(())
4589        }
4590    }
4591
4592    fn mk_manifest(
4593        base: &str,
4594        page_size: usize,
4595        checkpoint_seq: u32,
4596        first_frame: u64,
4597        last_frame: u64,
4598    ) -> OwnedGenerationManifest {
4599        mk_manifest_salted(
4600            base,
4601            page_size,
4602            checkpoint_seq,
4603            Some(CHAIN_SALT),
4604            first_frame,
4605            last_frame,
4606        )
4607    }
4608
4609    /// R858-B19: the salt every `mk_manifest` fixture shares, so a chain built
4610    /// from them is one generation unless a test deliberately says otherwise.
4611    const CHAIN_SALT: WalSalt = WalSalt { salt1: 0x1109_ca5e, salt2: 0x7acf_42a3 };
4612
4613    /// R858-B19: `mk_manifest` with the generation salt spelled out — `None`
4614    /// reproduces a manifest written by a pre-`v4` writer.
4615    fn mk_manifest_salted(
4616        base: &str,
4617        page_size: usize,
4618        checkpoint_seq: u32,
4619        salt: Option<WalSalt>,
4620        first_frame: u64,
4621        last_frame: u64,
4622    ) -> OwnedGenerationManifest {
4623        OwnedGenerationManifest {
4624            base_snapshot_key: base.to_string(),
4625            page_size,
4626            checkpoint_seq,
4627            salt,
4628            first_frame,
4629            last_frame,
4630            epoch: 0,
4631            owner: None,
4632            // Chain validation is about frame ranges, not object layout, so
4633            // these fixtures stay on the pre-R761-F2 shape — which also keeps
4634            // them exercising the legacy read path.
4635            frame_batches: Vec::new(),
4636        }
4637    }
4638
4639    /// validate_generation_chain accepts a single well-formed manifest.
4640    #[test]
4641    fn validate_chain_accepts_single_generation() {
4642        let chain = validate_generation_chain(&[mk_manifest("base.db", 4096, 0, 1, 5)]).unwrap();
4643        assert_eq!(chain.base_snapshot_key, "base.db");
4644        assert_eq!(chain.page_size, 4096);
4645        assert_eq!(chain.generation.checkpoint_seq, 0);
4646        assert_eq!(chain.total_frames, 5);
4647    }
4648
4649    /// validate_generation_chain accepts a contiguous multi-generation chain
4650    /// and sums the frame count across generations.
4651    #[test]
4652    fn validate_chain_accepts_contiguous_multi_generation() {
4653        let chain = validate_generation_chain(&[
4654            mk_manifest("base.db", 4096, 3, 1, 5),
4655            mk_manifest("base.db", 4096, 3, 6, 9),
4656            mk_manifest("base.db", 4096, 3, 10, 12),
4657        ])
4658        .unwrap();
4659        assert_eq!(chain.generation.checkpoint_seq, 3);
4660        assert_eq!(chain.total_frames, 12);
4661    }
4662
4663    /// Empty chain is the "no generations under prefix" signal.
4664    #[test]
4665    fn validate_chain_rejects_empty() {
4666        assert!(validate_generation_chain(&[]).is_err());
4667    }
4668
4669    /// A gap between manifest ranges is corruption — fail loudly.
4670    #[test]
4671    fn validate_chain_rejects_gap() {
4672        let err = validate_generation_chain(&[
4673            mk_manifest("base.db", 4096, 0, 1, 5),
4674            mk_manifest("base.db", 4096, 0, 7, 9), // missing frame 6
4675        ])
4676        .unwrap_err();
4677        assert!(format!("{err}").contains("gap in stream"), "err was {err}");
4678    }
4679
4680    /// A chain that does not start at frame 1 means the base+chain don't match
4681    /// (some early frames were never uploaded, or the chain was truncated by a
4682    /// retention sweep). Refuse.
4683    #[test]
4684    fn validate_chain_rejects_non_one_start() {
4685        let err = validate_generation_chain(&[mk_manifest("base.db", 4096, 0, 5, 9)]).unwrap_err();
4686        assert!(format!("{err}").contains("gap in stream"), "err was {err}");
4687    }
4688
4689    /// A WAL restart between generations means the source folded the WAL into
4690    /// main; the post-restart frames do not replay onto a pre-restart base.
4691    /// Refuse and direct the caller at a fresh tier-1a snapshot.
4692    #[test]
4693    fn validate_chain_rejects_restart() {
4694        let err = validate_generation_chain(&[
4695            mk_manifest("base.db", 4096, 0, 1, 5),
4696            mk_manifest("base.db", 4096, 1, 1, 3), // restart, new seq
4697        ])
4698        .unwrap_err();
4699        let s = format!("{err}");
4700        assert!(s.contains("WAL restart"), "err was {s}");
4701        assert!(s.contains("fresh tier-1a snapshot"), "err was {s}");
4702    }
4703
4704    /// Generations referencing different bases mean the chain mixed sinks /
4705    /// the base was rotated mid-stream. Refuse.
4706    #[test]
4707    fn validate_chain_rejects_base_mismatch() {
4708        let err = validate_generation_chain(&[
4709            mk_manifest("base-A.db", 4096, 0, 1, 5),
4710            mk_manifest("base-B.db", 4096, 0, 6, 9),
4711        ])
4712        .unwrap_err();
4713        assert!(format!("{err}").contains("chain spans bases"), "err was {err}");
4714    }
4715
4716    /// Page-size mismatch across manifests is a corrupt-manifest signal.
4717    #[test]
4718    fn validate_chain_rejects_page_size_mismatch() {
4719        let err = validate_generation_chain(&[
4720            mk_manifest("base.db", 4096, 0, 1, 5),
4721            mk_manifest("base.db", 8192, 0, 6, 9),
4722        ])
4723        .unwrap_err();
4724        assert!(format!("{err}").contains("page_size"), "err was {err}");
4725    }
4726
4727    /// replay_frames_into walks every manifest's range in (seq, frame_no) order
4728    /// and pushes frames into the insert seam. The mock seam records the call
4729    /// log so we can verify ordering and frame content.
4730    #[tokio::test]
4731    async fn replay_walks_manifests_in_frame_order() {
4732        // Stage two generations' worth of frames into the object store via
4733        // tail_frames, then point replay at the resulting manifests.
4734        let seam = MockWal::new(4096);
4735        seam.append(1, 0, 0xAA);
4736        seam.append(2, 2, 0xBB); // commit
4737        let target = fresh_target();
4738        let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4739        seam.append(3, 0, 0xCC);
4740        seam.append(4, 4, 0xDD); // commit
4741        let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4742
4743        let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
4744        let insert = MockInsertSeam::new();
4745        let total = replay_frames_into(&target, &insert, &manifests)
4746            .await
4747            .unwrap();
4748        assert_eq!(total, 4);
4749
4750        let events = insert.events();
4751        assert_eq!(events.len(), 4, "begin/end are the caller's responsibility");
4752        for (i, ev) in events.iter().enumerate() {
4753            let want_frame_no = (i + 1) as u64;
4754            let want_page = want_frame_no as u32;
4755            let want_db_size = if want_frame_no.is_multiple_of(2) { want_frame_no as u32 } else { 0 };
4756            assert_eq!(
4757                ev,
4758                &MockInsertEvent::Frame {
4759                    frame_no: want_frame_no,
4760                    page_no: want_page,
4761                    db_size: want_db_size,
4762                }
4763            );
4764        }
4765    }
4766
4767    /// R761-F2 read compatibility, and the reason `frame_batches` being empty
4768    /// has to MEAN something rather than merely be absent: a sink written by a
4769    /// pre-batching writer — per-frame objects, a v2 manifest — still replays,
4770    /// and a chain that straddles the change replays as one stream. Written
4771    /// here by hand at the object level, because the writer that produced this
4772    /// shape no longer exists to produce it.
4773    #[tokio::test]
4774    async fn a_chain_straddling_the_batching_change_replays_as_one_stream() {
4775        let target = fresh_target();
4776        let frame_size = WAL_FRAME_HEADER_SIZE + 4096;
4777        let frame = |frame_no: u64, page_no: u32, db_size: u32| {
4778            let mut b = vec![0u8; frame_size];
4779            b[0..4].copy_from_slice(&page_no.to_be_bytes());
4780            b[4..8].copy_from_slice(&db_size.to_be_bytes());
4781            b[WAL_FRAME_HEADER_SIZE..].fill(frame_no as u8);
4782            b
4783        };
4784
4785        // Generation 1: the old layout — one object per frame, v2 manifest
4786        // with no batch list.
4787        for frame_no in 1..=2u64 {
4788            target
4789                .store
4790                .put(
4791                    &target.frame_key(0, 0, frame_no),
4792                    frame(frame_no, frame_no as u32, frame_no as u32).into(),
4793                )
4794                .await
4795                .unwrap();
4796        }
4797        let legacy = format_generation_manifest(GenerationManifest {
4798            base_snapshot_key: "b.db",
4799            page_size: 4096,
4800            checkpoint_seq: 0,
4801            salt: None,
4802            first_frame: 1,
4803            last_frame: 2,
4804            epoch: 0,
4805            owner: None,
4806            frame_batches: &[],
4807        });
4808        assert!(legacy.starts_with("TURSO-BACKUP STREAM v2\n"), "{legacy}");
4809        target
4810            .store
4811            .put(&target.generation_key(1), legacy.into_bytes().into())
4812            .await
4813            .unwrap();
4814
4815        // Generation 2: the new layout, batched, appended to the same chain.
4816        let mut batch = frame(3, 3, 0);
4817        batch.extend_from_slice(&frame(4, 4, 4));
4818        target
4819            .store
4820            .put(&target.frame_batch_key(0, 0, 3, 4), batch.into())
4821            .await
4822            .unwrap();
4823        let batched = format_generation_manifest(GenerationManifest {
4824            base_snapshot_key: "b.db",
4825            page_size: 4096,
4826            checkpoint_seq: 0,
4827            salt: None,
4828            first_frame: 3,
4829            last_frame: 4,
4830            epoch: 0,
4831            owner: None,
4832            frame_batches: &[(3, 4)],
4833        });
4834        target
4835            .store
4836            .put(&target.generation_key(2), batched.into_bytes().into())
4837            .await
4838            .unwrap();
4839
4840        let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
4841        assert!(manifests[0].frame_batches.is_empty(), "gen 1 is the legacy layout");
4842        assert_eq!(manifests[1].frame_batches, vec![(3, 4)]);
4843
4844        // R858-B19 CHANGED THIS ASSERTION, deliberately. It used to read
4845        // `validate_generation_chain(&manifests).expect("layout is not a chain
4846        // property")`. Layout still is not a chain property — that claim is
4847        // what the replay assertion below tests, and it still holds. What
4848        // changed is PROVENANCE: this fixture is a two-generation chain in
4849        // which neither manifest records which WAL its frames came from,
4850        // because both predate the `wal_salt` field. That is byte-for-byte the
4851        // shape a WAL recreate produced (probe F: contiguous ranges, one
4852        // checkpoint_seq, two different WALs), so restore can no longer tell
4853        // this chain from a spliced one and must refuse rather than hand back a
4854        // plausible wrong image. A chain this old needs a fresh tier-1a
4855        // snapshot; one generation of it would still restore.
4856        let err = validate_generation_chain(&manifests).unwrap_err().to_string();
4857        assert!(err.contains("no WAL salt"), "got: {err}");
4858
4859        let insert = MockInsertSeam::new();
4860        assert_eq!(replay_frames_into(&target, &insert, &manifests).await.unwrap(), 4);
4861        assert_eq!(
4862            insert.events(),
4863            (1..=4u64)
4864                .map(|n| MockInsertEvent::Frame {
4865                    frame_no: n,
4866                    page_no: n as u32,
4867                    db_size: if n == 3 { 0 } else { n as u32 },
4868                })
4869                .collect::<Vec<_>>(),
4870            "both layouts deliver the same frames in the same order"
4871        );
4872    }
4873
4874    /// restore_latest_stream errors loudly when there are no manifests under
4875    /// the prefix — the caller should be using tier-1a restore instead.
4876    #[tokio::test]
4877    async fn restore_errors_when_no_generations() {
4878        let target = fresh_target();
4879        let dest = std::env::temp_dir().join(format!(
4880            "turso-backup-restore-empty-{}-{}.db",
4881            std::process::id(),
4882            unix_nanos()
4883        ));
4884        let err = restore_latest_stream(&target, dest.to_str().unwrap())
4885            .await
4886            .unwrap_err();
4887        assert!(format!("{err}").contains("no generation manifests"), "err was {err}");
4888        // The dest file is never written when validation fails up front.
4889        assert!(!dest.exists());
4890    }
4891
4892    /// A frame uploaded with a wrong byte length corrupts the chain; restore's
4893    /// replay path catches it before touching the destination's WAL.
4894    #[tokio::test]
4895    async fn replay_rejects_wrong_size_frame() {
4896        let seam = MockWal::new(4096);
4897        seam.append(1, 1, 0xAA);
4898        let target = fresh_target();
4899        let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4900
4901        // Overwrite the one uploaded frame object with garbage of the wrong
4902        // length.
4903        target
4904            .store
4905            .put(&target.frame_batch_key(0, 0, 1, 1), b"too short".to_vec().into())
4906            .await
4907            .unwrap();
4908
4909        let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
4910        let insert = MockInsertSeam::new();
4911        let err = replay_frames_into(&target, &insert, &manifests)
4912            .await
4913            .unwrap_err();
4914        assert!(format!("{err}").contains("expected"), "err was {err}");
4915    }
4916
4917    /// Live-DB end-to-end fixture: a temp db path that cleans itself up
4918    /// (including `-wal` / `-shm` sidecars).
4919    struct TempDb(std::path::PathBuf);
4920    impl TempDb {
4921        fn new(tag: &str) -> Self {
4922            TempDb(std::env::temp_dir().join(format!(
4923                "turso-backup-stream-{tag}-{}-{}.db",
4924                std::process::id(),
4925                unix_nanos(),
4926            )))
4927        }
4928        fn path(&self) -> &str {
4929            self.0.to_str().unwrap()
4930        }
4931    }
4932    impl Drop for TempDb {
4933        fn drop(&mut self) {
4934            for sfx in ["", "-wal", "-shm"] {
4935                let _ = std::fs::remove_file(format!("{}{sfx}", self.0.display()));
4936            }
4937        }
4938    }
4939
4940    async fn seed_rows(path: &str, start: i64, count: i64) {
4941        let db = turso::Builder::new_local(path).build().await.unwrap();
4942        let conn = db.connect().unwrap();
4943        conn.execute("CREATE TABLE IF NOT EXISTS t (id INTEGER PRIMARY KEY, v TEXT)", ())
4944            .await
4945            .unwrap();
4946        conn.execute("BEGIN", ()).await.unwrap();
4947        for i in start..start + count {
4948            conn.execute("INSERT INTO t (id, v) VALUES (?, ?)", (i, format!("v{i}")))
4949                .await
4950                .unwrap();
4951        }
4952        conn.execute("COMMIT", ()).await.unwrap();
4953    }
4954
4955    async fn checkpoint_truncate(path: &str) {
4956        let db = turso::Builder::new_local(path).build().await.unwrap();
4957        let conn = db.connect().unwrap();
4958        let mut rows = conn.query("PRAGMA wal_checkpoint(TRUNCATE)", ()).await.unwrap();
4959        while rows.next().await.unwrap().is_some() {}
4960    }
4961
4962    async fn count_rows(path: &str) -> i64 {
4963        let db = turso::Builder::new_local(path).build().await.unwrap();
4964        let conn = db.connect().unwrap();
4965        let mut r = conn.query("SELECT COUNT(*) FROM t", ()).await.unwrap();
4966        let row = r.next().await.unwrap().unwrap();
4967        row.get::<i64>(0).unwrap()
4968    }
4969
4970    /// End-to-end against a real turso DB and the real `CoreWalSeam`:
4971    /// seed → snapshot → checkpoint (resets WAL, bumps checkpoint_seq) →
4972    /// append more rows (creates frames under the new seq) → tail → restore →
4973    /// row count on the dest matches the post-append source.
4974    ///
4975    /// This is the live-DB ping-pong R005-F2's handoff named: it exercises
4976    /// CoreWalSeam's `wal_state` / `wal_get_frame` AND the new
4977    /// `wal_insert_begin` / `wal_insert_frame` / `wal_insert_end` impls, plus
4978    /// the full manifest discovery + replay pipeline.
4979    #[tokio::test]
4980    async fn live_db_seed_snapshot_tail_restore_round_trips() {
4981        let src = TempDb::new("src");
4982        let dest = TempDb::new("dest");
4983
4984        // Seed 50 rows. These end up in the base snapshot via VACUUM INTO.
4985        seed_rows(src.path(), 0, 50).await;
4986
4987        let target = BackupTarget {
4988            store: Arc::new(InMemory::new()),
4989            prefix: "backups".into(),
4990        };
4991        let base_key = match crate::snapshot::snapshot_and_upload(src.path(), &target)
4992            .await
4993            .unwrap()
4994        {
4995            crate::snapshot::SnapshotOutcome::Uploaded { key, .. } => key,
4996            other => panic!("expected Uploaded base snapshot, got {other:?}"),
4997        };
4998
4999        // Checkpoint to fold prior WAL into main and reset the WAL header —
5000        // any subsequent writes land in fresh frames under a new
5001        // checkpoint_seq. This mirrors how a real orchestrator would hand off
5002        // from tier-1a (full snapshot) into tier-2 streaming.
5003        checkpoint_truncate(src.path()).await;
5004
5005        // Append 25 more rows — these become the WAL frames we stream.
5006        seed_rows(src.path(), 1000, 25).await;
5007
5008        // Tail the new frames into the sink via the real CoreWalSeam.
5009        {
5010            let seam = CoreWalSeam::open(src.path()).unwrap();
5011            let cfg = StreamConfig {
5012                base_snapshot_key: &base_key,
5013                page_size: 4096,
5014                backpressure: BackpressureConfig::default(),
5015                rpo_target: None,
5016                epoch: 0,
5017                owner: None,
5018                pointer_generation: 0,
5019            };
5020            let outcome = tail_frames(&seam, &target, &cfg).await.unwrap();
5021            match outcome {
5022                StreamOutcome::Streamed { frame_count, .. } => {
5023                    assert!(frame_count > 0, "expected at least one frame uploaded");
5024                }
5025                StreamOutcome::Restarted { frame_count, .. } => {
5026                    // Acceptable when this is the first tail after a checkpoint:
5027                    // there's a prior implicit watermark via the WAL-restart bit.
5028                    assert!(frame_count > 0);
5029                }
5030                other => panic!("expected Streamed/Restarted, got {other:?}"),
5031            }
5032        } // drop seam (and its turso_core connection) before restore opens its own.
5033
5034        // Restore into a fresh dest. Replay must produce the post-append state.
5035        let outcome = restore_latest_stream(&target, dest.path()).await.unwrap();
5036        assert_eq!(outcome.base_snapshot_key, base_key);
5037        assert!(outcome.frames_replayed > 0);
5038        assert!(outcome.generation_count >= 1);
5039
5040        // The restored DB is openable via vanilla turso and has the right rows.
5041        // (snapshot::restore_latest's tests already cover the vanilla-sqlite3
5042        // exit ramp; here the contract is "turso reopens" since wal_insert_*
5043        // is the only consumer side that knows the engine's frame layout.)
5044        let restored_rows = count_rows(dest.path()).await;
5045        assert_eq!(restored_rows, 75, "expected 50 (base) + 25 (replayed) = 75");
5046    }
5047
5048    /// Crash-consistency: the last frames captured by `tail_frames` are
5049    /// uncommitted (mid-transaction, `db_size == 0`). Restore must drop that
5050    /// suffix and end up at the previous commit's state — `wal_insert_end`
5051    /// with `force_commit = false` is the engine knob that does this.
5052    ///
5053    /// We can't easily force turso to commit-then-leak-a-partial-write in a
5054    /// unit test, so we drive the seam directly: a single `MockInsertSeam`-
5055    /// recorded restore would show the begin/end protocol, but to verify the
5056    /// engine actually truncates we need the real `CoreWalSeam`. The
5057    /// `manifest_with_uncommitted_tail_rolls_back_via_insert_end` test below
5058    /// stages frame bytes in the store by hand and aims them at a real DB.
5059    #[tokio::test]
5060    async fn manifest_with_uncommitted_tail_rolls_back_via_insert_end() {
5061        // Seed two rows via real turso so we have committed page 1 + page 2.
5062        // Then capture the WAL frames from the source. Then stage a *fake*
5063        // extra frame whose db_size = 0 (mid-transaction) in the sink, and
5064        // append it to the manifest range. Restore should drop that fake
5065        // frame and leave the dest at the 2-row state.
5066        let src = TempDb::new("ccsrc");
5067        let dest = TempDb::new("ccdest");
5068        seed_rows(src.path(), 0, 2).await;
5069        let target = BackupTarget {
5070            store: Arc::new(InMemory::new()),
5071            prefix: "backups".into(),
5072        };
5073        let base_key = match crate::snapshot::snapshot_and_upload(src.path(), &target)
5074            .await
5075            .unwrap()
5076        {
5077            crate::snapshot::SnapshotOutcome::Uploaded { key, .. } => key,
5078            other => panic!("expected Uploaded, got {other:?}"),
5079        };
5080        checkpoint_truncate(src.path()).await;
5081        seed_rows(src.path(), 100, 1).await; // one more row → at least one commit frame.
5082
5083        let (checkpoint_seq, last_committed_frame, frame_size) = {
5084            let seam = CoreWalSeam::open(src.path()).unwrap();
5085            let cfg = StreamConfig {
5086                base_snapshot_key: &base_key,
5087                page_size: 4096,
5088                backpressure: BackpressureConfig::default(),
5089                rpo_target: None,
5090                epoch: 0,
5091                owner: None,
5092                pointer_generation: 0,
5093            };
5094            let _ = tail_frames(&seam, &target, &cfg).await.unwrap();
5095            let w = seam.wal_state().unwrap();
5096            (w.checkpoint_seq, w.last_frame, WAL_FRAME_HEADER_SIZE + 4096)
5097        };
5098
5099        // Manually upload one extra "uncommitted" frame: copy the last
5100        // committed frame's bytes but zero db_size in the header. This
5101        // simulates tail_frames having captured a mid-transaction tail.
5102        // R761-F2: the frames live in one batch object, so take the last
5103        // frame's slice out of it rather than fetching a per-frame key.
5104        let batch_key = target.frame_batch_key(0, checkpoint_seq, 1, last_committed_frame);
5105        let batch = target
5106            .store
5107            .get(&batch_key)
5108            .await
5109            .unwrap()
5110            .bytes()
5111            .await
5112            .unwrap();
5113        let at = (last_committed_frame as usize - 1) * frame_size;
5114        let mut tail_bytes = batch[at..at + frame_size].to_vec();
5115        // Zero the big-endian db_size at offset 4..8 to mark this as non-commit.
5116        tail_bytes[4..8].copy_from_slice(&0u32.to_be_bytes());
5117        // Repad to ensure exact length — paranoia.
5118        assert_eq!(tail_bytes.len(), frame_size);
5119        let phantom_frame_no = last_committed_frame + 1;
5120        target
5121            .store
5122            .put(
5123                &target.frame_batch_key(0, checkpoint_seq, phantom_frame_no, phantom_frame_no),
5124                tail_bytes.into(),
5125            )
5126            .await
5127            .unwrap();
5128
5129        // Extend the latest generation manifest to claim the phantom frame.
5130        let mut keys: Vec<_> = target
5131            .store
5132            .list_with_delimiter(Some(&join_key(&target.prefix, "generations")))
5133            .await
5134            .unwrap()
5135            .objects
5136            .into_iter()
5137            .map(|o| o.location)
5138            .collect();
5139        keys.sort();
5140        let last_manifest_key = keys.last().unwrap().clone();
5141        let bytes = target
5142            .store
5143            .get(&last_manifest_key)
5144            .await
5145            .unwrap()
5146            .bytes()
5147            .await
5148            .unwrap();
5149        let mut m = parse_generation_manifest(&String::from_utf8_lossy(&bytes)).unwrap();
5150        m.last_frame = phantom_frame_no;
5151        // The phantom frame went up as its own single-frame batch object, so
5152        // the manifest's index has to name it too — the range alone is no
5153        // longer enough to find a frame (R761-F2).
5154        m.frame_batches.push((phantom_frame_no, phantom_frame_no));
5155        let new_text = format_generation_manifest(GenerationManifest {
5156            base_snapshot_key: &m.base_snapshot_key,
5157            page_size: m.page_size,
5158            checkpoint_seq: m.checkpoint_seq,
5159            salt: None,
5160            first_frame: m.first_frame,
5161            last_frame: m.last_frame,
5162            epoch: m.epoch,
5163            owner: m.owner.as_deref(),
5164            frame_batches: &m.frame_batches,
5165        });
5166        target
5167            .store
5168            .put(&last_manifest_key, new_text.into_bytes().into())
5169            .await
5170            .unwrap();
5171
5172        // Restore. The phantom uncommitted frame must be dropped by
5173        // wal_insert_end(false); the row count reflects the last *commit*.
5174        let _ = restore_latest_stream(&target, dest.path()).await.unwrap();
5175        let restored = count_rows(dest.path()).await;
5176        assert_eq!(restored, 3, "expected 2 (base) + 1 (committed) — phantom rolled back");
5177    }
5178
5179    /// `replay_wal_onto_main` with an empty WAL returns the main bytes
5180    /// untouched — the main file alone is the committed image.
5181    #[test]
5182    fn replay_empty_wal_returns_main_unchanged() {
5183        let seam = MockWal::new(4096);
5184        let mut main = vec![0u8; 4096 * 3];
5185        main[0..16].copy_from_slice(b"SQLite format 3\0");
5186        let img = replay_wal_onto_main(&seam, main.clone(), 4096).unwrap();
5187        assert_eq!(img, main);
5188    }
5189
5190    /// A single commit frame updates its page slot and the final image is
5191    /// sized to the commit's `db_size`. Unrelated pages in `main` are left
5192    /// in place.
5193    #[test]
5194    fn replay_single_commit_frame_applies_page() {
5195        let seam = MockWal::new(4096);
5196        // Frame 1: page 2, commit at db_size = 3 pages.
5197        seam.append(2, 3, 0xCC);
5198        let mut main = vec![0u8; 4096 * 3];
5199        main[0..16].copy_from_slice(b"SQLite format 3\0");
5200
5201        let img = replay_wal_onto_main(&seam, main, 4096).unwrap();
5202        assert_eq!(img.len(), 4096 * 3);
5203        // Page 1 untouched (still has the magic + zeros).
5204        assert!(img.starts_with(b"SQLite format 3\0"));
5205        // Page 2 = 0xCC fill from the WAL frame.
5206        assert!(img[4096..4096 * 2].iter().all(|&b| b == 0xCC));
5207        // Page 3 untouched.
5208        assert!(img[4096 * 2..].iter().all(|&b| b == 0));
5209    }
5210
5211    /// Uncommitted frames past the last commit are dropped and the image is
5212    /// truncated to the last commit's `db_size` — the crash-consistency story
5213    /// matching restore's `wal_insert_end(false)`.
5214    #[test]
5215    fn replay_drops_uncommitted_tail() {
5216        let seam = MockWal::new(4096);
5217        seam.append(1, 0, 0x11); // mid-txn
5218        seam.append(2, 2, 0x22); // commit at db_size = 2
5219        seam.append(3, 0, 0x33); // uncommitted (dropped)
5220        seam.append(4, 0, 0x44); // uncommitted (dropped)
5221
5222        let main = vec![0u8; 4096 * 4]; // pre-grown so a buggy replay would keep junk
5223        let img = replay_wal_onto_main(&seam, main, 4096).unwrap();
5224        assert_eq!(img.len(), 4096 * 2, "image must be truncated to db_size=2");
5225        assert!(img[0..4096].iter().all(|&b| b == 0x11));
5226        assert!(img[4096..].iter().all(|&b| b == 0x22));
5227    }
5228
5229    /// Multi-commit chain: the full commit prefix is applied and the image is
5230    /// grown to the final commit's `db_size`. Intermediate uncommitted frames
5231    /// between commits are also applied (they are part of the committed
5232    /// suffix once a later frame in the same batch becomes a commit).
5233    #[test]
5234    fn replay_applies_full_commit_prefix_and_grows_image() {
5235        let seam = MockWal::new(4096);
5236        seam.append(1, 0, 0xAA);
5237        seam.append(2, 2, 0xBB); // first commit
5238        seam.append(3, 0, 0xCC);
5239        seam.append(4, 4, 0xDD); // second commit
5240
5241        let main = vec![0u8; 4096 * 2];
5242        let img = replay_wal_onto_main(&seam, main, 4096).unwrap();
5243        assert_eq!(img.len(), 4096 * 4, "image grown to db_size=4");
5244        assert!(img[0..4096].iter().all(|&b| b == 0xAA));
5245        assert!(img[4096..4096 * 2].iter().all(|&b| b == 0xBB));
5246        assert!(img[4096 * 2..4096 * 3].iter().all(|&b| b == 0xCC));
5247        assert!(img[4096 * 3..].iter().all(|&b| b == 0xDD));
5248    }
5249
5250    /// WAL has frames but none are committed (writer mid-transaction at the
5251    /// instant we sampled). The main file alone is the image; the uncommitted
5252    /// suffix is dropped wholesale.
5253    #[test]
5254    fn replay_returns_main_when_only_uncommitted_frames_present() {
5255        let seam = MockWal::new(4096);
5256        seam.append(1, 0, 0x11);
5257        seam.append(2, 0, 0x22);
5258        let mut main = vec![0u8; 4096 * 3];
5259        main[0..16].copy_from_slice(b"SQLite format 3\0");
5260        let img = replay_wal_onto_main(&seam, main.clone(), 4096).unwrap();
5261        assert_eq!(img, main);
5262    }
5263
5264    /// End-to-end against a real turso DB and `CoreWalSeam`: seed, fold the
5265    /// first batch into main via TRUNCATE, then seed more rows so the WAL has
5266    /// uncheckpointed committed frames. `raw_consistent_copy_live` must
5267    /// reproduce the full row count without itself calling TRUNCATE — that is
5268    /// the live-writer contract a TRUNCATE-busy concurrent writer would
5269    /// otherwise block.
5270    #[tokio::test]
5271    async fn live_db_consistent_copy_without_truncate_round_trips() {
5272        let src = TempDb::new("live-cc-src");
5273        seed_rows(src.path(), 0, 25).await;
5274        checkpoint_truncate(src.path()).await; // first batch into main, WAL reset
5275        seed_rows(src.path(), 1000, 10).await; // second batch lives in WAL
5276
5277        // No TRUNCATE here — this is the live-writer path.
5278        let image = raw_consistent_copy_live(src.path(), 4096).await.unwrap();
5279        assert!(image.starts_with(b"SQLite format 3\0"));
5280
5281        // Write the image to a fresh path (no -wal sidecar — the replay
5282        // folded WAL in) and re-open to confirm all 35 rows are present.
5283        let restored = TempDb::new("live-cc-restored");
5284        std::fs::write(restored.path(), &image).unwrap();
5285        let rows = count_rows(restored.path()).await;
5286        assert_eq!(
5287            rows, 35,
5288            "expected 25 (checkpointed) + 10 (replayed from WAL)"
5289        );
5290    }
5291
5292    // ── R858-B18: reading a source nobody will lock for us ────────────────
5293    //
5294    // The property under test throughout: the copy is either a validated
5295    // point in time or a refusal, never a plausible wrong image.
5296
5297    /// The two salt readers must agree. [`read_wal_salt`] pulls it out of frame
5298    /// 1 through the [`WalSeam`] trait (the only path a `from_conn` seam has);
5299    /// [`WalFileHeader::read`] pulls it out of the 32-byte `-wal` header with no
5300    /// engine at all. They are separate code paths against separate byte
5301    /// offsets, and R858-B18 depends on them naming the same generation — if
5302    /// they could disagree, validation would compare a salt the streamer never
5303    /// recorded.
5304    #[tokio::test]
5305    async fn wal_file_header_salt_matches_the_salt_read_through_the_seam() {
5306        let src = TempDb::new("b18-salt-agree");
5307        seed_rows(src.path(), 0, 5).await;
5308
5309        let seam = CoreWalSeam::open_reader(src.path()).unwrap();
5310        let state = seam.wal_state().unwrap();
5311        assert!(state.last_frame > 0, "seed should leave frames in the WAL");
5312        let via_seam = read_wal_salt(&seam, 4096, state.last_frame).unwrap().unwrap();
5313        drop(seam);
5314
5315        let via_file = WalFileHeader::read(src.path()).unwrap().unwrap();
5316        assert_eq!(via_file.salt, via_seam, "frame 1 carries the WAL header's salt verbatim");
5317        assert_eq!(
5318            via_file.checkpoint_seq, state.checkpoint_seq,
5319            "and the on-disk sequence is the one the engine reports"
5320        );
5321    }
5322
5323    /// A `-wal` shorter than one header names no generation — an honest
5324    /// unknown, not an error (same posture as `WalGeneration { salt: None }`).
5325    #[test]
5326    fn wal_file_header_parse_needs_a_whole_header() {
5327        assert!(WalFileHeader::parse(&[0u8; WalFileHeader::SIZE - 1]).is_none());
5328        let mut hdr = [0u8; WalFileHeader::SIZE];
5329        hdr[8..12].copy_from_slice(&4096u32.to_be_bytes());
5330        hdr[12..16].copy_from_slice(&7u32.to_be_bytes());
5331        hdr[16..20].copy_from_slice(&0xb83c_03f5u32.to_be_bytes());
5332        hdr[20..24].copy_from_slice(&0x7acf_42a3u32.to_be_bytes());
5333        let parsed = WalFileHeader::parse(&hdr).unwrap();
5334        assert_eq!(parsed.page_size, 4096);
5335        assert_eq!(parsed.checkpoint_seq, 7);
5336        assert_eq!(parsed.salt, WalSalt { salt1: 0xb83c_03f5, salt2: 0x7acf_42a3 });
5337    }
5338
5339    /// The central asymmetry of [`SourceFingerprint::stable_across`]: a plain
5340    /// append is NOT movement (frames `1..=max_frame` are immutable within a
5341    /// generation, so our image is merely an earlier point in time), while a
5342    /// checkpoint IS (it rewrites the main file our copy already read).
5343    ///
5344    /// Getting this backwards is not a small error in either direction: treat
5345    /// an append as movement and every copy of a database that is actually in
5346    /// use is refused; treat a checkpoint as harmless and we hand back spliced
5347    /// state that passes `integrity_check`.
5348    #[tokio::test]
5349    async fn fingerprint_ignores_an_append_and_catches_a_checkpoint() {
5350        let src = TempDb::new("b18-fingerprint");
5351        seed_rows(src.path(), 0, 5).await;
5352
5353        let before = SourceFingerprint::read(src.path()).unwrap();
5354        seed_rows(src.path(), 100, 5).await;
5355        let appended = SourceFingerprint::read(src.path()).unwrap();
5356        assert!(
5357            appended.wal_len > before.wal_len,
5358            "the append must actually have grown the WAL, or this proves nothing \
5359             (before {}B, after {}B)",
5360            before.wal_len,
5361            appended.wal_len
5362        );
5363        assert!(
5364            before.stable_across(&appended),
5365            "an append is not movement: {} -> {}",
5366            before.describe(),
5367            appended.describe()
5368        );
5369
5370        checkpoint_truncate(src.path()).await;
5371        let folded = SourceFingerprint::read(src.path()).unwrap();
5372        assert!(
5373            !before.stable_across(&folded),
5374            "a checkpoint IS movement and must be caught: {} -> {}",
5375            before.describe(),
5376            folded.describe()
5377        );
5378    }
5379
5380    /// The protocol accepts when the source holds still, and the accepted value
5381    /// is the one the attempt produced.
5382    #[tokio::test]
5383    async fn validation_accepts_a_quiescent_source() {
5384        let src = TempDb::new("b18-quiescent");
5385        seed_rows(src.path(), 0, 3).await;
5386        let got = validated_against_source(src.path(), "test read", || async { Ok(41 + 1) })
5387            .await
5388            .unwrap();
5389        assert_eq!(got, 42);
5390    }
5391
5392    /// A source that moves under every attempt yields a REFUSAL, not a value —
5393    /// and the message names both samples so an operator can see what moved.
5394    #[tokio::test]
5395    async fn validation_refuses_when_the_source_moves_under_every_attempt() {
5396        let src = TempDb::new("b18-moving");
5397        seed_rows(src.path(), 0, 3).await;
5398        let path = src.path().to_string();
5399        let attempts = std::cell::Cell::new(0u32);
5400        let err = validated_against_source(src.path(), "test read", || {
5401            // Grow the main file inside the attempt window, which is exactly
5402            // the shape of a foreign checkpoint folding pages into it.
5403            attempts.set(attempts.get() + 1);
5404            let path = path.clone();
5405            async move {
5406                let mut f = std::fs::OpenOptions::new().append(true).open(&path)?;
5407                std::io::Write::write_all(&mut f, &[0u8; 4096])?;
5408                Ok(())
5409            }
5410        })
5411        .await
5412        .unwrap_err();
5413        assert_eq!(
5414            attempts.get(),
5415            COPY_VALIDATION_ATTEMPTS,
5416            "every attempt in the budget must be spent before refusing"
5417        );
5418        let msg = format!("{err:#}");
5419        assert!(msg.contains("refusing a test read"), "{msg}");
5420        assert!(msg.contains("before and"), "message must name both samples: {msg}");
5421    }
5422
5423    /// An attempt that fails against a source that did NOT move is a real
5424    /// error, not a race — it is returned as itself on the first attempt rather
5425    /// than retried and then reported as concurrency.
5426    #[tokio::test]
5427    async fn validation_surfaces_a_real_error_without_burning_the_budget() {
5428        let src = TempDb::new("b18-real-error");
5429        seed_rows(src.path(), 0, 3).await;
5430        let attempts = std::cell::Cell::new(0u32);
5431        let err = validated_against_source(src.path(), "test read", || {
5432            attempts.set(attempts.get() + 1);
5433            async { Err::<(), _>(anyhow::anyhow!("page 3 checksum mismatch")) }
5434        })
5435        .await
5436        .unwrap_err();
5437        assert_eq!(attempts.get(), 1, "a stable source means retrying cannot help");
5438        let msg = format!("{err:#}");
5439        assert!(msg.contains("did NOT move"), "{msg}");
5440        assert!(msg.contains("page 3 checksum mismatch"), "{msg}");
5441    }
5442
5443    /// The reader open is the one the backup path uses, and it must produce the
5444    /// same watermark the writable open does. (Cross-process non-exclusivity —
5445    /// the point of the flag — is measured in `examples/foreign_checkpoint_probe.rs`
5446    /// probes G2r/G3c, since it needs a second process.)
5447    #[tokio::test]
5448    async fn read_only_seam_reports_the_same_watermark_as_the_writable_one() {
5449        let src = TempDb::new("b18-reader-watermark");
5450        seed_rows(src.path(), 0, 4).await;
5451        let writable = CoreWalSeam::open(src.path()).unwrap().wal_state().unwrap();
5452        let reader = CoreWalSeam::open_reader(src.path()).unwrap().wal_state().unwrap();
5453        assert_eq!(writable, reader);
5454    }
5455
5456    /// A second call to `raw_consistent_copy_live` after more writes captures
5457    /// the new state — the primitive is callable repeatedly without per-call
5458    /// setup, mirroring how a snapshot loop would drive it.
5459    #[tokio::test]
5460    async fn live_db_consistent_copy_reflects_new_writes() {
5461        let src = TempDb::new("live-cc-incr-src");
5462        seed_rows(src.path(), 0, 5).await;
5463        let img1 = raw_consistent_copy_live(src.path(), 4096).await.unwrap();
5464        let r1 = TempDb::new("live-cc-incr-r1");
5465        std::fs::write(r1.path(), &img1).unwrap();
5466        assert_eq!(count_rows(r1.path()).await, 5);
5467
5468        seed_rows(src.path(), 100, 7).await;
5469        let img2 = raw_consistent_copy_live(src.path(), 4096).await.unwrap();
5470        let r2 = TempDb::new("live-cc-incr-r2");
5471        std::fs::write(r2.path(), &img2).unwrap();
5472        assert_eq!(count_rows(r2.path()).await, 12);
5473    }
5474
5475    // ── R732-F2 (W245): tenant fencing epochs on the R2 write path ────────
5476    //
5477    // The property under test throughout: a writer holding a stale fencing
5478    // token is REJECTED, not merely unlucky. Every test below names the
5479    // split-brain it rules out.
5480
5481    fn cfg_at_epoch(epoch: u64, owner: Option<&'static str>) -> StreamConfig<'static> {
5482        StreamConfig { epoch, owner, ..cfg() }
5483    }
5484
5485    /// R736-T2: like `cfg_at_epoch`, but also names the cross-cell pointer
5486    /// generation this writer believes it holds.
5487    fn cfg_at(epoch: u64, pointer_generation: u64, owner: Option<&'static str>) -> StreamConfig<'static> {
5488        StreamConfig { epoch, pointer_generation, owner, ..cfg() }
5489    }
5490
5491    /// THE canonical F2 test, and the reason the whole relay exists: two
5492    /// owners tailing the same tenant. The one at the older epoch bounces and
5493    /// writes nothing; the sink is byte-for-byte what the newer owner left.
5494    #[tokio::test]
5495    async fn a_stale_writer_is_fenced_and_writes_zero_frames() {
5496        let target = fresh_target();
5497
5498        // The real owner (epoch 2) streams three frames.
5499        let winner = MockWal::new(4096);
5500        for i in 1..=3u32 {
5501            winner.append(i, i, i as u8);
5502        }
5503        let out = tail_frames(&winner, &target, &cfg_at_epoch(2, Some("node-2")))
5504            .await
5505            .unwrap();
5506        assert!(matches!(out, StreamOutcome::Streamed { frame_count: 3, .. }));
5507
5508        let manifests_before = list_and_parse_generation_manifests(&target).await.unwrap();
5509        let watermark_before = read_watermark(&target.store, &target.watermark_key())
5510            .await
5511            .unwrap()
5512            .unwrap();
5513
5514        // The partitioned old owner (epoch 1) wakes up with its own frames and
5515        // tails the same sink, unaware it has been transferred away.
5516        let loser = MockWal::new(4096);
5517        for i in 1..=9u32 {
5518            loser.append(i, i, 0xff);
5519        }
5520        let out = tail_frames(&loser, &target, &cfg_at_epoch(1, Some("node-1")))
5521            .await
5522            .unwrap();
5523        assert_eq!(
5524            out,
5525            StreamOutcome::Fenced {
5526                current_epoch: 2,
5527                our_epoch: 1,
5528                current_pointer_generation: 0,
5529                our_pointer_generation: 0,
5530            },
5531            "the stale owner must be told it lost, with both epochs"
5532        );
5533
5534        // Zero side effects: no frames under its own epoch prefix, no extra
5535        // generation, and the watermark still names the winner.
5536        // Nothing at all under the loser's epoch prefix — asserted by listing
5537        // rather than by probing one key, so it holds whatever object layout
5538        // the writer would have used (R761-F2).
5539        let loser_prefix = "backups/frames/00000000000000000001/";
5540        assert!(
5541            !objects_under(&target)
5542                .await
5543                .iter()
5544                .any(|k| k.starts_with(loser_prefix)),
5545            "a fenced writer must not upload a single frame"
5546        );
5547        let manifests_after = list_and_parse_generation_manifests(&target).await.unwrap();
5548        assert_eq!(
5549            manifests_after, manifests_before,
5550            "a fenced writer must not write a generation manifest"
5551        );
5552        let watermark_after = read_watermark(&target.store, &target.watermark_key())
5553            .await
5554            .unwrap()
5555            .unwrap();
5556        assert_eq!(watermark_after.epoch, 2);
5557        assert_eq!(watermark_after.watermark, watermark_before.watermark);
5558    }
5559
5560    /// An UNFENCED (epoch 0) legacy streamer is fenced by a sink that has been
5561    /// claimed. This is the migration case: a node still running the old
5562    /// single-writer configuration must not be allowed to scribble over a
5563    /// tenant that yubaba has since handed to someone else.
5564    #[tokio::test]
5565    async fn an_unfenced_writer_is_fenced_by_a_claimed_sink() {
5566        let target = fresh_target();
5567        let owner = MockWal::new(4096);
5568        owner.append(1, 1, 1);
5569        tail_frames(&owner, &target, &cfg_at_epoch(1, None)).await.unwrap();
5570
5571        let legacy = MockWal::new(4096);
5572        legacy.append(1, 1, 2);
5573        assert_eq!(
5574            tail_frames(&legacy, &target, &cfg_at_epoch(0, None)).await.unwrap(),
5575            StreamOutcome::Fenced {
5576                current_epoch: 1,
5577                our_epoch: 0,
5578                current_pointer_generation: 0,
5579                our_pointer_generation: 0,
5580            }
5581        );
5582    }
5583
5584    // ── R736-T2 (W250): the second, cross-cell fence ───────────────────────
5585    //
5586    // The epoch alone is a *local* raft counter — it cannot see a tenant
5587    // that moved to a different cell's independent raft group. These tests
5588    // prove the pointer generation catches exactly the case the epoch can't:
5589    // a stale cell that is current on its own epoch.
5590
5591    /// THE canonical T2 test: a writer whose epoch is perfectly current for
5592    /// its own (now-stale) cell still bounces, because the global pointer
5593    /// says ownership moved elsewhere. Proves the epoch is necessary but not
5594    /// sufficient — this is the gap W250 exists to close.
5595    #[tokio::test]
5596    async fn a_stale_pointer_generation_fences_even_at_a_current_epoch() {
5597        let target = fresh_target();
5598
5599        // The new cell streams at generation 2, epoch 1 (its own local raft
5600        // is fresh — it just took ownership).
5601        let winner = MockWal::new(4096);
5602        for i in 1..=3u32 {
5603            winner.append(i, i, i as u8);
5604        }
5605        let out = tail_frames(&winner, &target, &cfg_at(1, 2, Some("cell-b/node-1")))
5606            .await
5607            .unwrap();
5608        assert!(matches!(out, StreamOutcome::Streamed { frame_count: 3, .. }));
5609
5610        // The old cell's writer wakes up unaware of the move. Its own local
5611        // epoch (1) is perfectly current for its own raft group — nothing
5612        // local told it to step down — but its pointer generation (1) is
5613        // behind the sink's (2).
5614        let loser = MockWal::new(4096);
5615        for i in 1..=9u32 {
5616            loser.append(i, i, 0xff);
5617        }
5618        let out = tail_frames(&loser, &target, &cfg_at(1, 1, Some("cell-a/node-1")))
5619            .await
5620            .unwrap();
5621        assert_eq!(
5622            out,
5623            StreamOutcome::Fenced {
5624                current_epoch: 1,
5625                our_epoch: 1,
5626                current_pointer_generation: 2,
5627                our_pointer_generation: 1,
5628            },
5629            "an equal, non-stale epoch must not mask a stale pointer generation"
5630        );
5631
5632        // Zero side effects, exactly like the epoch-only fence.
5633        let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
5634        assert_eq!(manifests.len(), 1, "only the winner's generation was written");
5635    }
5636
5637    /// The watermark sidecar carries the pointer generation as a fifth
5638    /// positional field, and a four-field sidecar (pre-R736-T2 writer) reads
5639    /// back generation 0 — unfenced, exactly like a pre-R732 sidecar reads
5640    /// back epoch 0.
5641    #[tokio::test]
5642    async fn watermark_sidecar_round_trips_the_pointer_generation() {
5643        let target = fresh_target();
5644        let key = target.watermark_key();
5645        write_watermark(
5646            &target.store,
5647            &key,
5648            Watermark { checkpoint_seq: 2, last_frame: 11 },
5649            None,
5650            6,
5651            3,
5652            None,
5653        )
5654        .await
5655        .unwrap();
5656        let read = read_watermark(&target.store, &key).await.unwrap().unwrap();
5657        assert_eq!((read.epoch, read.pointer_generation), (6, 3));
5658
5659        // A four-field sidecar (epoch, no generation) — the R732-F2 shape.
5660        target
5661            .store
5662            .put(&key, b"2 11 12345 6\n".to_vec().into())
5663            .await
5664            .unwrap();
5665        let legacy = read_watermark(&target.store, &key).await.unwrap().unwrap();
5666        assert_eq!(legacy.epoch, 6);
5667        assert_eq!(legacy.pointer_generation, 0, "a pre-R736-T2 sidecar fences nobody on generation");
5668    }
5669
5670    /// R869: `read_fence_state` reports exactly the two comparands
5671    /// [`tail_frames`] checks — including the pre-fencing sidecar shapes, which
5672    /// must read back as unfenced rather than erroring, since a rebuild will
5673    /// meet them on any sink written before R732-F2.
5674    #[tokio::test]
5675    async fn read_fence_state_reports_what_tail_frames_would_check() {
5676        let target = fresh_target();
5677        assert_eq!(
5678            read_fence_state(&target).await.unwrap(),
5679            None,
5680            "no sidecar means no fence at all, which is not the same as a zero fence"
5681        );
5682
5683        write_watermark(
5684            &target.store,
5685            &target.watermark_key(),
5686            Watermark {
5687                checkpoint_seq: 2,
5688                last_frame: 11,
5689            },
5690            None,
5691            6,
5692            3,
5693            None,
5694        )
5695        .await
5696        .unwrap();
5697        assert_eq!(
5698            read_fence_state(&target).await.unwrap(),
5699            Some(FenceState {
5700                epoch: 6,
5701                pointer_generation: 3
5702            })
5703        );
5704
5705        // A two-field sidecar — the original pre-fencing shape. Both comparands
5706        // read back 0, so a rebuild sees "unfenced" rather than a parse error.
5707        target
5708            .store
5709            .put(&target.watermark_key(), b"2 11\n".to_vec().into())
5710            .await
5711            .unwrap();
5712        assert_eq!(
5713            read_fence_state(&target).await.unwrap(),
5714            Some(FenceState::default())
5715        );
5716    }
5717
5718    /// The property a rebuild's epoch floor rests on, stated as a test rather
5719    /// than as a comment: seeding one above `read_fence_state().epoch` is
5720    /// exactly enough to stop being fenced, and one *below* it is not.
5721    #[tokio::test]
5722    async fn a_floor_taken_from_read_fence_state_is_what_unfences_a_rebuild() {
5723        let target = fresh_target();
5724        let seam = MockWal::new(4096);
5725        for i in 1..=3u32 {
5726            seam.append(i, i, i as u8);
5727        }
5728        tail_frames(&seam, &target, &cfg_at_epoch(5, Some("dead-fleet")))
5729            .await
5730            .unwrap();
5731
5732        let fence = read_fence_state(&target).await.unwrap().unwrap();
5733        assert_eq!(fence.epoch, 5);
5734
5735        // A rebuilt cluster that restarted its epochs at 1 is refused.
5736        seam.append(4, 4, 4);
5737        let out = tail_frames(&seam, &target, &cfg_at_epoch(1, Some("rebuilt")))
5738            .await
5739            .unwrap();
5740        assert!(matches!(out, StreamOutcome::Fenced { .. }), "got {out:?}");
5741
5742        // Seeded from the fence, it is not.
5743        let out = tail_frames(
5744            &seam,
5745            &target,
5746            &cfg_at_epoch(fence.epoch + 1, Some("rebuilt")),
5747        )
5748        .await
5749        .unwrap();
5750        assert!(matches!(out, StreamOutcome::Streamed { .. }), "got {out:?}");
5751    }
5752
5753    /// The same owner resuming at the same epoch is NOT fenced — fencing is
5754    /// strictly "someone newer exists", not "someone else wrote here". An
5755    /// owner that restarts under an unchanged token must keep streaming, or
5756    /// every process restart would wedge the tenant.
5757    #[tokio::test]
5758    async fn an_equal_epoch_writer_resumes_normally() {
5759        let target = fresh_target();
5760        let seam = MockWal::new(4096);
5761        seam.append(1, 1, 1);
5762        tail_frames(&seam, &target, &cfg_at_epoch(3, None)).await.unwrap();
5763        seam.append(2, 2, 2);
5764        let out = tail_frames(&seam, &target, &cfg_at_epoch(3, None)).await.unwrap();
5765        match out {
5766            StreamOutcome::Streamed { first_frame, last_frame, .. } => {
5767                assert_eq!((first_frame, last_frame), (2, 2), "resumes after the watermark");
5768            }
5769            other => panic!("expected Streamed, got {other:?}"),
5770        }
5771    }
5772
5773    /// Frame keys are namespaced by epoch, and epoch 0 keeps the pre-fencing
5774    /// two-level layout so existing backups stay addressable.
5775    #[test]
5776    fn frame_keys_are_namespaced_by_epoch_with_zero_keeping_the_legacy_layout() {
5777        let target = fresh_target();
5778        assert_eq!(
5779            target.frame_key(0, 7, 42).to_string(),
5780            "backups/frames/0000000007/00000000000000000042",
5781            "epoch 0 must keep the original key shape"
5782        );
5783        assert_eq!(
5784            target.frame_key(5, 7, 42).to_string(),
5785            "backups/frames/00000000000000000005/0000000007/00000000000000000042"
5786        );
5787        assert_ne!(target.frame_key(5, 7, 42), target.frame_key(6, 7, 42));
5788    }
5789
5790    /// R761-F2: a batch key carries its frame range, keeps the epoch
5791    /// namespacing and the zero-padding (so lexical order is still frame
5792    /// order), and cannot be confused with a pre-batching per-frame key even
5793    /// when the batch holds exactly one frame.
5794    #[test]
5795    fn batch_keys_carry_the_range_and_never_collide_with_a_per_frame_key() {
5796        let target = fresh_target();
5797        assert_eq!(
5798            target.frame_batch_key(0, 7, 42, 99).to_string(),
5799            "backups/frames/0000000007/00000000000000000042-00000000000000000099",
5800            "epoch 0 keeps the two-level layout for batches too"
5801        );
5802        assert_eq!(
5803            target.frame_batch_key(5, 7, 42, 99).to_string(),
5804            "backups/frames/00000000000000000005/0000000007/00000000000000000042-00000000000000000099"
5805        );
5806        assert_ne!(
5807            target.frame_batch_key(0, 7, 42, 42),
5808            target.frame_key(0, 7, 42),
5809            "a one-frame batch is still a batch — the two layouts must stay distinguishable"
5810        );
5811        // Lexical order matches frame order for the batches of one stream,
5812        // which never overlap.
5813        assert!(
5814            target.frame_batch_key(0, 7, 1, 8) < target.frame_batch_key(0, 7, 9, 16),
5815            "zero-padding must keep batches lexically ordered by first frame"
5816        );
5817    }
5818
5819    /// A takeover mid-stream leaves both owners' frames intact under their own
5820    /// prefixes — the key namespacing is the backstop behind the epoch check.
5821    #[tokio::test]
5822    async fn a_takeover_writes_under_its_own_epoch_prefix_without_disturbing_the_old_one() {
5823        let target = fresh_target();
5824        let seam = MockWal::new(4096);
5825        seam.append(1, 1, 0xaa);
5826        tail_frames(&seam, &target, &cfg_at_epoch(1, None)).await.unwrap();
5827        seam.append(2, 2, 0xbb);
5828        tail_frames(&seam, &target, &cfg_at_epoch(2, None)).await.unwrap();
5829
5830        let first = target.store.get(&target.frame_batch_key(1, 0, 1, 1)).await.unwrap();
5831        assert_eq!(first.bytes().await.unwrap().len(), WAL_FRAME_HEADER_SIZE + 4096);
5832        target
5833            .store
5834            .get(&target.frame_batch_key(2, 0, 2, 2))
5835            .await
5836            .expect("the new owner's frame lives under its own epoch");
5837        assert!(
5838            target.store.get(&target.frame_batch_key(1, 0, 2, 2)).await.is_err(),
5839            "the new owner must not write into the old owner's prefix"
5840        );
5841    }
5842
5843    /// Restore refuses a chain whose epoch goes backwards: a generation
5844    /// written by an owner that had already been fenced. Replaying it would
5845    /// interleave a stale owner's frames into the live stream.
5846    #[test]
5847    fn validate_chain_refuses_an_epoch_regression() {
5848        let mut newer = mk_manifest("base.db", 4096, 0, 1, 5);
5849        newer.epoch = 4;
5850        let mut stale = mk_manifest("base.db", 4096, 0, 6, 9);
5851        stale.epoch = 3;
5852        let err = validate_generation_chain(&[newer, stale]).unwrap_err();
5853        let msg = format!("{err}");
5854        assert!(msg.contains("epoch 3"), "err was {msg}");
5855        assert!(msg.contains("fenced"), "err must name the cause: {msg}");
5856    }
5857
5858    /// …but a chain that spans an ownership TRANSFER is fine. Epochs may
5859    /// advance mid-stream; only regression is corruption.
5860    #[test]
5861    fn validate_chain_accepts_a_transfer_mid_chain() {
5862        let mut first = mk_manifest("base.db", 4096, 0, 1, 5);
5863        first.epoch = 3;
5864        let mut second = mk_manifest("base.db", 4096, 0, 6, 9);
5865        second.epoch = 4;
5866        let chain = validate_generation_chain(&[first, second]).unwrap();
5867        assert_eq!(chain.total_frames, 9);
5868        assert_eq!(chain.epoch, 4, "the chain reports the most recent owner");
5869    }
5870
5871    /// A pre-fencing chain still validates and reports epoch 0.
5872    #[test]
5873    fn validate_chain_of_pre_fencing_manifests_reports_epoch_zero() {
5874        let chain = validate_generation_chain(&[mk_manifest("base.db", 4096, 0, 1, 5)]).unwrap();
5875        assert_eq!(chain.epoch, 0);
5876    }
5877
5878    /// Manifest v2 round-trips the epoch and the owner label.
5879    #[test]
5880    fn manifest_v2_round_trips_epoch_and_owner() {
5881        let text = format_generation_manifest(GenerationManifest {
5882            base_snapshot_key: "backups/snapshots/snapshot-1.db",
5883            page_size: 4096,
5884            checkpoint_seq: 7,
5885            salt: None,
5886            first_frame: 12,
5887            last_frame: 34,
5888            epoch: 9,
5889            owner: Some("node-3"),
5890            frame_batches: &[],
5891        });
5892        assert!(text.starts_with("TURSO-BACKUP STREAM v2\n"), "{text}");
5893        let parsed = parse_generation_manifest(&text).unwrap();
5894        assert_eq!(parsed.epoch, 9);
5895        assert_eq!(parsed.owner.as_deref(), Some("node-3"));
5896
5897        // No owner label → the key is omitted entirely, not written empty.
5898        let text = format_generation_manifest(GenerationManifest {
5899            base_snapshot_key: "b.db",
5900            page_size: 4096,
5901            checkpoint_seq: 0,
5902            salt: None,
5903            first_frame: 1,
5904            last_frame: 1,
5905            epoch: 1,
5906            owner: None,
5907            frame_batches: &[],
5908        });
5909        assert!(!text.contains("owner"), "{text}");
5910        assert_eq!(parse_generation_manifest(&text).unwrap().owner, None);
5911    }
5912
5913    /// R761-F2: a batch list round-trips, and its presence is what moves the
5914    /// header to v3 — the version and the layout are one fact, so a reader can
5915    /// never see a v3 header without an index or a v2 header with one.
5916    #[test]
5917    fn manifest_v3_round_trips_the_frame_batch_list() {
5918        let batches = [(12u64, 20u64), (21, 34)];
5919        let text = format_generation_manifest(GenerationManifest {
5920            base_snapshot_key: "backups/snapshots/snapshot-1.db",
5921            page_size: 4096,
5922            checkpoint_seq: 7,
5923            salt: None,
5924            first_frame: 12,
5925            last_frame: 34,
5926            epoch: 9,
5927            owner: Some("node-3"),
5928            frame_batches: &batches,
5929        });
5930        assert!(text.starts_with("TURSO-BACKUP STREAM v3\n"), "{text}");
5931        let parsed = parse_generation_manifest(&text).unwrap();
5932        assert_eq!(parsed.frame_batches, batches.to_vec());
5933        assert_eq!(parsed.first_frame, 12);
5934        assert_eq!(parsed.last_frame, 34);
5935        assert_eq!(parsed.owner.as_deref(), Some("node-3"));
5936    }
5937
5938    /// A v3 batch list that does not exactly tile the manifest's own frame
5939    /// range is corruption, and it has to fail at parse: the list IS the frame
5940    /// index, so a hole in it becomes a 404 halfway through a replay — after
5941    /// the destination has already been overwritten with the base snapshot.
5942    #[test]
5943    fn a_v3_manifest_whose_batches_do_not_tile_its_range_is_rejected() {
5944        let head = "TURSO-BACKUP STREAM v3\nbase_snapshot b.db\npage_size 4096\ncheckpoint_seq 0\nepoch 0\n";
5945        for (batches, want, why) in [
5946            ("frame_batch 1-3\nframe_batch 5-9\n", "does not continue", "gap"),
5947            ("frame_batch 1-3\n", "but the manifest claims", "short"),
5948            ("frame_batch 2-9\n", "does not continue", "wrong start"),
5949            ("frame_batch 1-4\nframe_batch 4-9\n", "does not continue", "overlap"),
5950            ("", "no `frame_batch` lines", "missing index"),
5951        ] {
5952            let text = format!("{head}first_frame 1\nlast_frame 9\n{batches}");
5953            let err = parse_generation_manifest(&text).unwrap_err();
5954            assert!(
5955                format!("{err}").contains(want),
5956                "{why}: expected {want:?}, err was {err}"
5957            );
5958        }
5959    }
5960
5961    /// The other direction: a `frame_batch` line under a v1/v2 header is
5962    /// corrupt too. Silently honouring it would let a hand-edited manifest
5963    /// claim a layout its header says it does not have.
5964    #[test]
5965    fn a_pre_v3_manifest_carrying_a_batch_line_is_rejected() {
5966        let bad = "TURSO-BACKUP STREAM v2\nbase_snapshot b.db\npage_size 4096\ncheckpoint_seq 0\nfirst_frame 1\nlast_frame 3\nepoch 0\nframe_batch 1-3\n";
5967        let err = parse_generation_manifest(bad).unwrap_err();
5968        assert!(format!("{err}").contains("corrupt or hand-edited"), "err was {err}");
5969    }
5970
5971    /// A v1 manifest — one written before fencing existed — still parses, as
5972    /// epoch 0. Those backups have to stay restorable.
5973    #[test]
5974    fn a_v1_manifest_parses_as_epoch_zero() {
5975        let legacy = "TURSO-BACKUP STREAM v1\nbase_snapshot b.db\npage_size 4096\ncheckpoint_seq 0\nfirst_frame 1\nlast_frame 3\n";
5976        let parsed = parse_generation_manifest(legacy).unwrap();
5977        assert_eq!(parsed.epoch, 0);
5978        assert_eq!(parsed.owner, None);
5979        assert_eq!(parsed.last_frame, 3);
5980    }
5981
5982    /// A v2 manifest missing its epoch is corrupt, not legacy. Defaulting it
5983    /// to 0 would silently demote a fenced generation to unfenced — the one
5984    /// direction this mechanism must never fail in.
5985    #[test]
5986    fn a_v2_manifest_without_an_epoch_is_rejected() {
5987        let bad = "TURSO-BACKUP STREAM v2\nbase_snapshot b.db\npage_size 4096\ncheckpoint_seq 0\nfirst_frame 1\nlast_frame 3\n";
5988        let err = parse_generation_manifest(bad).unwrap_err();
5989        assert!(format!("{err}").contains("missing `epoch`"), "err was {err}");
5990    }
5991
5992    /// The watermark sidecar carries the epoch as a fourth positional field,
5993    /// and a three-field sidecar (pre-R732 writer) reads back as epoch 0.
5994    #[tokio::test]
5995    async fn watermark_sidecar_round_trips_the_epoch() {
5996        let target = fresh_target();
5997        let key = target.watermark_key();
5998        write_watermark(
5999            &target.store,
6000            &key,
6001            Watermark { checkpoint_seq: 2, last_frame: 11 },
6002            None,
6003            6,
6004            0,
6005            None,
6006        )
6007        .await
6008        .unwrap();
6009        let read = read_watermark(&target.store, &key).await.unwrap().unwrap();
6010        assert_eq!(read.epoch, 6);
6011        assert_eq!(read.watermark.last_frame, 11);
6012        assert!(read.written_at_nanos.is_some());
6013
6014        target
6015            .store
6016            .put(&key, b"2 11 12345\n".to_vec().into())
6017            .await
6018            .unwrap();
6019        let legacy = read_watermark(&target.store, &key).await.unwrap().unwrap();
6020        assert_eq!(legacy.epoch, 0, "a pre-R732 sidecar fences nobody");
6021        assert_eq!(legacy.written_at_nanos, Some(12345));
6022    }
6023
6024    /// End-to-end on a real DB: an ownership transfer happens mid-stream and
6025    /// the restore is clean — every frame from both owners replays, and the
6026    /// outcome reports the winning epoch. This is the "the N+1 writer wins;
6027    /// restore is clean" half of the ticket's ask, run against turso_core
6028    /// rather than a mock.
6029    #[tokio::test]
6030    async fn a_transfer_mid_stream_restores_cleanly_and_reports_the_new_epoch() {
6031        let src = TempDb::new("src-epoch");
6032        let dest = TempDb::new("dest-epoch");
6033        seed_rows(src.path(), 0, 50).await;
6034
6035        let target = fresh_target();
6036        let base_key = match crate::snapshot::snapshot_and_upload(src.path(), &target)
6037            .await
6038            .unwrap()
6039        {
6040            crate::snapshot::SnapshotOutcome::Uploaded { key, .. } => key,
6041            other => panic!("expected Uploaded base snapshot, got {other:?}"),
6042        };
6043        checkpoint_truncate(src.path()).await;
6044
6045        // Owner A (epoch 1) streams the first batch of writes.
6046        seed_rows(src.path(), 1000, 10).await;
6047        {
6048            let seam = CoreWalSeam::open(src.path()).unwrap();
6049            let cfg = StreamConfig {
6050                base_snapshot_key: &base_key,
6051                epoch: 1,
6052                owner: Some("node-a"),
6053                ..cfg()
6054            };
6055            let out = tail_frames(&seam, &target, &cfg).await.unwrap();
6056            assert!(
6057                matches!(
6058                    out,
6059                    StreamOutcome::Streamed { .. } | StreamOutcome::Restarted { .. }
6060                ),
6061                "owner A should have streamed, got {out:?}"
6062            );
6063        }
6064
6065        // Ownership transfers. Owner B (epoch 2) picks up where A stopped.
6066        seed_rows(src.path(), 2000, 15).await;
6067        {
6068            let seam = CoreWalSeam::open(src.path()).unwrap();
6069            let cfg = StreamConfig {
6070                base_snapshot_key: &base_key,
6071                epoch: 2,
6072                owner: Some("node-b"),
6073                ..cfg()
6074            };
6075            let out = tail_frames(&seam, &target, &cfg).await.unwrap();
6076            assert!(
6077                matches!(
6078                    out,
6079                    StreamOutcome::Streamed { .. } | StreamOutcome::Restarted { .. }
6080                ),
6081                "owner B should have streamed, got {out:?}"
6082            );
6083        }
6084
6085        let outcome = restore_latest_stream(&target, dest.path()).await.unwrap();
6086        assert_eq!(outcome.base_snapshot_key, base_key);
6087        assert_eq!(outcome.epoch, 2, "restore reports the most recent owner");
6088        assert_eq!(
6089            count_rows(dest.path()).await,
6090            75,
6091            "50 (base) + 10 (owner A) + 15 (owner B)"
6092        );
6093    }
6094
6095    // ── R732-T3 (W245): the watermark advance is a compare-and-swap ───────
6096
6097    /// Bootstrapping a fresh sink uses `PutMode::Create`, so two writers
6098    /// racing to claim a brand-new tenant cannot both succeed. Without this
6099    /// the very first write — the one with no prior version to swap on —
6100    /// would be the one unguarded moment in the whole protocol.
6101    #[tokio::test]
6102    async fn a_second_bootstrap_of_a_fresh_sink_is_contended() {
6103        let target = fresh_target();
6104        let key = target.watermark_key();
6105        let w = Watermark { checkpoint_seq: 0, last_frame: 1 };
6106        assert_eq!(
6107            write_watermark(&target.store, &key, w, None, 1, 0, None).await.unwrap(),
6108            WatermarkCas::Advanced
6109        );
6110        assert_eq!(
6111            write_watermark(&target.store, &key, w, None, 1, 0, None).await.unwrap(),
6112            WatermarkCas::Contended,
6113            "the sink already exists — Create must not silently overwrite it"
6114        );
6115    }
6116
6117    /// An advance is conditional on the version actually read. A writer
6118    /// holding a version somebody else has already replaced loses.
6119    #[tokio::test]
6120    async fn an_advance_on_a_replaced_version_is_contended() {
6121        let target = fresh_target();
6122        let key = target.watermark_key();
6123        write_watermark(
6124            &target.store,
6125            &key,
6126            Watermark { checkpoint_seq: 0, last_frame: 1 },
6127            None,
6128            1,
6129            0,
6130            None,
6131        )
6132        .await
6133        .unwrap();
6134
6135        // Our writer reads the sidecar and holds onto that version…
6136        let stale = read_watermark(&target.store, &key).await.unwrap().unwrap();
6137        // …while somebody else advances it out from under us.
6138        write_watermark(
6139            &target.store,
6140            &key,
6141            Watermark { checkpoint_seq: 0, last_frame: 5 },
6142            None,
6143            2,
6144            0,
6145            read_watermark(&target.store, &key)
6146                .await
6147                .unwrap()
6148                .unwrap()
6149                .version
6150                .as_ref(),
6151        )
6152        .await
6153        .unwrap();
6154
6155        assert_eq!(
6156            write_watermark(
6157                &target.store,
6158                &key,
6159                Watermark { checkpoint_seq: 0, last_frame: 2 },
6160                None,
6161                1,
6162                0,
6163                stale.version.as_ref(),
6164            )
6165            .await
6166            .unwrap(),
6167            WatermarkCas::Contended,
6168            "the stale version must not be allowed to overwrite the newer one"
6169        );
6170        // And the sink still names the winner, not us.
6171        let now = read_watermark(&target.store, &key).await.unwrap().unwrap();
6172        assert_eq!((now.epoch, now.watermark.last_frame), (2, 5));
6173    }
6174
6175    /// Re-reading before advancing works: the point is the version, not the
6176    /// identity of the writer.
6177    #[tokio::test]
6178    async fn an_advance_on_the_current_version_succeeds() {
6179        let target = fresh_target();
6180        let key = target.watermark_key();
6181        write_watermark(
6182            &target.store,
6183            &key,
6184            Watermark { checkpoint_seq: 0, last_frame: 1 },
6185            None,
6186            1,
6187            0,
6188            None,
6189        )
6190        .await
6191        .unwrap();
6192        let cur = read_watermark(&target.store, &key).await.unwrap().unwrap();
6193        assert_eq!(
6194            write_watermark(
6195                &target.store,
6196                &key,
6197                Watermark { checkpoint_seq: 0, last_frame: 7 },
6198                None,
6199                1,
6200                0,
6201                cur.version.as_ref(),
6202            )
6203            .await
6204            .unwrap(),
6205            WatermarkCas::Advanced
6206        );
6207    }
6208
6209    /// A store that lets a test slip a competing writer in between our read
6210    /// of the watermark and our conditional write of it — the interleaving
6211    /// that a sequential epoch check cannot catch and the CAS must.
6212    ///
6213    /// On the first `put_opts` aimed at the watermark key it writes a rival
6214    /// sidecar (at `rival_epoch`) straight through to the inner store, then
6215    /// forwards our conditional put, which now finds a version it does not
6216    /// hold.
6217    struct RacingStore {
6218        inner: Arc<dyn ObjectStore>,
6219        watermark: ObjPath,
6220        rival_epoch: u64,
6221        /// R736-T2: the rival's pointer generation, so the same interleaving
6222        /// can be exercised for the cross-cell fence too.
6223        rival_pointer_generation: u64,
6224        fired: std::sync::atomic::AtomicBool,
6225    }
6226
6227    impl std::fmt::Display for RacingStore {
6228        fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
6229            write!(f, "RacingStore({})", self.inner)
6230        }
6231    }
6232    impl std::fmt::Debug for RacingStore {
6233        fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
6234            write!(f, "RacingStore({:?})", self.inner)
6235        }
6236    }
6237
6238    #[async_trait::async_trait]
6239    impl ObjectStore for RacingStore {
6240        async fn put_opts(
6241            &self,
6242            location: &ObjPath,
6243            payload: object_store::PutPayload,
6244            opts: PutOptions,
6245        ) -> object_store::Result<object_store::PutResult> {
6246            if location == &self.watermark
6247                && !self
6248                    .fired
6249                    .swap(true, std::sync::atomic::Ordering::SeqCst)
6250            {
6251                let rival =
6252                    format!("0 99 1 {} {}\n", self.rival_epoch, self.rival_pointer_generation);
6253                self.inner
6254                    .put(location, rival.into_bytes().into())
6255                    .await?;
6256            }
6257            self.inner.put_opts(location, payload, opts).await
6258        }
6259
6260        async fn put_multipart_opts(
6261            &self,
6262            location: &ObjPath,
6263            opts: object_store::PutMultipartOptions,
6264        ) -> object_store::Result<Box<dyn object_store::MultipartUpload>> {
6265            self.inner.put_multipart_opts(location, opts).await
6266        }
6267
6268        async fn get_opts(
6269            &self,
6270            location: &ObjPath,
6271            options: object_store::GetOptions,
6272        ) -> object_store::Result<object_store::GetResult> {
6273            self.inner.get_opts(location, options).await
6274        }
6275
6276        fn delete_stream(
6277            &self,
6278            locations: futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>>,
6279        ) -> futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>> {
6280            self.inner.delete_stream(locations)
6281        }
6282
6283        fn list(
6284            &self,
6285            prefix: Option<&ObjPath>,
6286        ) -> futures_util::stream::BoxStream<'static, object_store::Result<object_store::ObjectMeta>>
6287        {
6288            self.inner.list(prefix)
6289        }
6290
6291        async fn list_with_delimiter(
6292            &self,
6293            prefix: Option<&ObjPath>,
6294        ) -> object_store::Result<object_store::ListResult> {
6295            self.inner.list_with_delimiter(prefix).await
6296        }
6297
6298        async fn copy_opts(
6299            &self,
6300            from: &ObjPath,
6301            to: &ObjPath,
6302            options: object_store::CopyOptions,
6303        ) -> object_store::Result<()> {
6304            self.inner.copy_opts(from, to, options).await
6305        }
6306    }
6307
6308    fn racing_target(rival_epoch: u64) -> BackupTarget {
6309        racing_target_at(rival_epoch, 0)
6310    }
6311
6312    /// R736-T2: like `racing_target`, but also names the rival's pointer
6313    /// generation.
6314    fn racing_target_at(rival_epoch: u64, rival_pointer_generation: u64) -> BackupTarget {
6315        let inner: Arc<dyn ObjectStore> = Arc::new(InMemory::new());
6316        let prefix = "backups".to_string();
6317        let watermark = join_key(&prefix, "latest.stream-watermark");
6318        BackupTarget {
6319            store: Arc::new(RacingStore {
6320                inner,
6321                watermark,
6322                rival_epoch,
6323                rival_pointer_generation,
6324                fired: std::sync::atomic::AtomicBool::new(false),
6325            }),
6326            prefix,
6327        }
6328    }
6329
6330    /// The race the sequential check cannot see: a newer owner claims the
6331    /// sink *after* we read it and *before* we write it. The CAS catches it,
6332    /// and we report Fenced without publishing a manifest — which is why the
6333    /// manifest write had to move after the watermark advance.
6334    #[tokio::test]
6335    async fn losing_the_watermark_race_to_a_newer_owner_fences_us_before_we_publish() {
6336        let target = racing_target(9);
6337        let seam = MockWal::new(4096);
6338        for i in 1..=2u32 {
6339            seam.append(i, i, i as u8);
6340        }
6341        let out = tail_frames(&seam, &target, &cfg_at_epoch(4, Some("node-loser")))
6342            .await
6343            .unwrap();
6344        assert_eq!(
6345            out,
6346            StreamOutcome::Fenced {
6347                current_epoch: 9,
6348                our_epoch: 4,
6349                current_pointer_generation: 0,
6350                our_pointer_generation: 0,
6351            }
6352        );
6353
6354        // No manifest published — the chain stays clean, so restore is not
6355        // poisoned by a regressed generation from a writer that lost.
6356        assert!(
6357            list_and_parse_generation_manifests(&target)
6358                .await
6359                .unwrap()
6360                .is_empty(),
6361            "a writer that loses the CAS must publish nothing"
6362        );
6363        // The rival's sidecar survived untouched.
6364        let now = read_watermark(&target.store, &target.watermark_key())
6365            .await
6366            .unwrap()
6367            .unwrap();
6368        assert_eq!(now.epoch, 9);
6369    }
6370
6371    /// R736-T2: the same race, but the rival is a different cell claiming the
6372    /// tenant via a pointer CAS — our epoch is unchanged (we were never told
6373    /// to step down locally) but the generation moved under us mid-write.
6374    #[tokio::test]
6375    async fn losing_the_watermark_race_to_a_cross_cell_move_fences_us_before_we_publish() {
6376        let target = racing_target_at(4, 2);
6377        let seam = MockWal::new(4096);
6378        for i in 1..=2u32 {
6379            seam.append(i, i, i as u8);
6380        }
6381        let out = tail_frames(&seam, &target, &cfg_at(4, 1, Some("cell-a/node-loser")))
6382            .await
6383            .unwrap();
6384        assert_eq!(
6385            out,
6386            StreamOutcome::Fenced {
6387                current_epoch: 4,
6388                our_epoch: 4,
6389                current_pointer_generation: 2,
6390                our_pointer_generation: 1,
6391            },
6392            "an unchanged epoch must not mask a generation that moved under us"
6393        );
6394        assert!(
6395            list_and_parse_generation_manifests(&target)
6396                .await
6397                .unwrap()
6398                .is_empty(),
6399            "a writer that loses the CAS on generation must publish nothing"
6400        );
6401    }
6402
6403    /// Losing the race to a writer at our own epoch is NOT a fencing event —
6404    /// it means two streamers were handed the same token. That is a caller
6405    /// bug and must surface as a loud error, never be retried into success.
6406    #[tokio::test]
6407    async fn losing_the_race_to_an_equal_epoch_writer_is_a_loud_error() {
6408        let target = racing_target(4);
6409        let seam = MockWal::new(4096);
6410        seam.append(1, 1, 1);
6411        let err = tail_frames(&seam, &target, &cfg_at_epoch(4, None))
6412            .await
6413            .unwrap_err();
6414        let msg = format!("{err}");
6415        assert!(
6416            msg.contains("two streamers share one fencing token"),
6417            "err was {msg}"
6418        );
6419    }
6420
6421    /// Generation manifest round-trip and rejection of garbage.
6422    #[test]
6423    fn manifest_round_trip_and_rejection() {
6424        let text = format_generation_manifest(GenerationManifest {
6425            base_snapshot_key: "backups/snapshots/snapshot-1.db",
6426            page_size: 4096,
6427            checkpoint_seq: 7,
6428            salt: None,
6429            first_frame: 12,
6430            last_frame: 34,
6431            epoch: 0,
6432            owner: None,
6433            frame_batches: &[],
6434        });
6435        let parsed = parse_generation_manifest(&text).unwrap();
6436        assert_eq!(parsed.base_snapshot_key, "backups/snapshots/snapshot-1.db");
6437        assert_eq!(parsed.page_size, 4096);
6438        assert_eq!(parsed.checkpoint_seq, 7);
6439        assert_eq!(parsed.first_frame, 12);
6440        assert_eq!(parsed.last_frame, 34);
6441
6442        assert!(parse_generation_manifest("not a manifest").is_err());
6443        assert!(parse_generation_manifest("TURSO-BACKUP STREAM v1\nunknown 1").is_err());
6444        assert!(parse_generation_manifest("TURSO-BACKUP STREAM v1\nbase_snapshot k").is_err()); // missing fields
6445    }
6446
6447    // ---- R850-T3: tier-2 GC (gc_stream) --------------------------------
6448
6449    /// A tier-2 sink built object-by-object on an `InMemory` store.
6450    ///
6451    /// Hand-placed rather than produced by running `tail_frames`, because every
6452    /// one of these tests is about a sink in a state a *live* tailer never
6453    /// leaves behind — a superseded base, a generation's frames with no
6454    /// manifest. Driving the writer could not produce them without also
6455    /// producing the rebase that is the thing under test.
6456    struct GcSink {
6457        target: BackupTarget,
6458    }
6459
6460    impl GcSink {
6461        fn new() -> Self {
6462            GcSink {
6463                target: BackupTarget {
6464                    store: Arc::new(InMemory::new()),
6465                    prefix: "sink".to_string(),
6466                },
6467            }
6468        }
6469
6470        /// A base snapshot at `nanos`. Body is not a real database — nothing in
6471        /// the GC path parses it.
6472        async fn put_snapshot(&self, nanos: u128) -> String {
6473            let key = join_key(
6474                &self.target.prefix,
6475                &format!("snapshots/snapshot-{nanos:020}.db"),
6476            );
6477            self.target
6478                .store
6479                .put(&key, format!("base {nanos}").into_bytes().into())
6480                .await
6481                .unwrap();
6482            key.to_string()
6483        }
6484
6485        /// One batch object holding frames `first..=last` of generation
6486        /// `(epoch, seq)`, keyed exactly as `tail_frames` would key it.
6487        async fn put_frame_batch(&self, epoch: u64, seq: u32, first: u64, last: u64) -> String {
6488            let key = self.target.frame_batch_key(epoch, seq, first, last);
6489            self.target
6490                .store
6491                .put(&key, vec![0u8; 16].into())
6492                .await
6493                .unwrap();
6494            key.to_string()
6495        }
6496
6497        /// A generation manifest naming `base` and one batch `first..=last`.
6498        async fn put_manifest(
6499            &self,
6500            nanos: u128,
6501            base: &str,
6502            epoch: u64,
6503            seq: u32,
6504            first: u64,
6505            last: u64,
6506        ) {
6507            let batches = [(first, last)];
6508            let text = format_generation_manifest(GenerationManifest {
6509                base_snapshot_key: base,
6510                page_size: 4096,
6511                checkpoint_seq: seq,
6512                salt: Some(WalSalt {
6513                    salt1: 1,
6514                    salt2: 2,
6515                }),
6516                first_frame: first,
6517                last_frame: last,
6518                epoch,
6519                owner: None,
6520                frame_batches: &batches,
6521            });
6522            self.target
6523                .store
6524                .put(&self.target.generation_key(nanos), text.into_bytes().into())
6525                .await
6526                .unwrap();
6527        }
6528
6529        async fn exists(&self, key: &str) -> bool {
6530            self.target.store.get(&ObjPath::from(key)).await.is_ok()
6531        }
6532    }
6533
6534    /// Collect everything collectable: no grace, really delete.
6535    fn sweep_now() -> StreamGcConfig {
6536        StreamGcConfig {
6537            grace: Duration::ZERO,
6538            dry_run: false,
6539        }
6540    }
6541
6542    /// (a) The base a rebase superseded is reclaimed; the one a restore would
6543    /// pick is not.
6544    #[tokio::test]
6545    async fn gc_collects_a_superseded_base_snapshot_and_keeps_the_current_one() {
6546        let sink = GcSink::new();
6547        let old_base = sink.put_snapshot(1).await;
6548        let new_base = sink.put_snapshot(2).await;
6549        // The post-rebase steady state: the only chain names the new base.
6550        sink.put_manifest(10, &new_base, 0, 5, 1, 10).await;
6551        sink.put_frame_batch(0, 5, 1, 10).await;
6552
6553        let out = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6554
6555        assert_eq!(out.collected_snapshots, vec![old_base.clone()]);
6556        assert_eq!(out.retained_snapshots, 1);
6557        assert_eq!(out.live_base_snapshot_keys, vec![new_base.clone()]);
6558        assert!(!sink.exists(&old_base).await, "superseded base survived");
6559        assert!(sink.exists(&new_base).await, "live base was collected");
6560    }
6561
6562    /// The rebase-crash window: a new base is published before the old
6563    /// manifests are deleted, so the surviving chain names the OLD base. Both
6564    /// are live — the chain's pick and the newest — and the GC must refuse to
6565    /// choose between them.
6566    #[tokio::test]
6567    async fn gc_keeps_both_bases_while_a_rebase_is_half_landed() {
6568        let sink = GcSink::new();
6569        let old_base = sink.put_snapshot(1).await;
6570        let new_base = sink.put_snapshot(2).await;
6571        sink.put_manifest(10, &old_base, 0, 5, 1, 10).await;
6572
6573        let out = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6574
6575        assert!(out.collected_snapshots.is_empty());
6576        assert_eq!(out.retained_snapshots, 2);
6577        assert_eq!(out.live_base_snapshot_keys, vec![old_base, new_base]);
6578    }
6579
6580    /// (b) An orphaned generation's frames are reclaimed — under both key
6581    /// shapes — and a live generation's are not.
6582    #[tokio::test]
6583    async fn gc_collects_orphaned_frame_prefixes_and_keeps_the_live_generation() {
6584        let sink = GcSink::new();
6585        let base = sink.put_snapshot(2).await;
6586        sink.put_manifest(10, &base, 0, 5, 1, 10).await;
6587        let live = sink.put_frame_batch(0, 5, 1, 10).await;
6588        // Generation 4 was rebased away: its manifests are gone, its frames
6589        // are not. Once under the epoch-0 layout, once under the fenced one.
6590        let orphan_unfenced = sink.put_frame_batch(0, 4, 1, 6).await;
6591        let orphan_fenced = sink.put_frame_batch(7, 4, 1, 6).await;
6592
6593        let out = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6594
6595        let mut expected = vec![orphan_unfenced.clone(), orphan_fenced.clone()];
6596        expected.sort();
6597        assert_eq!(out.collected_frame_objects, expected);
6598        assert_eq!(out.retained_frame_objects, 1);
6599        assert_eq!(out.collected_bytes, 32, "two 16-byte batch objects");
6600        assert!(sink.exists(&live).await, "live generation's frames collected");
6601        assert!(!sink.exists(&orphan_unfenced).await);
6602        assert!(!sink.exists(&orphan_fenced).await);
6603    }
6604
6605    /// (c) Nothing inside the grace window is touched, however unreachable it
6606    /// looks. Everything an `InMemory` store holds was written moments ago, so
6607    /// a non-zero grace must spare the whole sweep.
6608    #[tokio::test]
6609    async fn gc_spares_everything_inside_the_grace_window() {
6610        let sink = GcSink::new();
6611        let stale_base = sink.put_snapshot(1).await;
6612        let base = sink.put_snapshot(2).await;
6613        sink.put_manifest(10, &base, 0, 5, 1, 10).await;
6614        sink.put_frame_batch(0, 5, 1, 10).await;
6615        let orphan = sink.put_frame_batch(0, 4, 1, 6).await;
6616
6617        let cfg = StreamGcConfig {
6618            grace: Duration::from_secs(3600),
6619            dry_run: false,
6620        };
6621        let out = gc_stream(&sink.target, &cfg).await.unwrap();
6622
6623        assert!(out.collected_frame_objects.is_empty());
6624        assert!(out.collected_snapshots.is_empty());
6625        assert_eq!(out.collected_bytes, 0);
6626        assert_eq!(out.spared_by_grace, 2, "the orphan batch and the stale base");
6627        assert!(sink.exists(&orphan).await);
6628        assert!(sink.exists(&stale_base).await);
6629    }
6630
6631    /// (d) A dry run reports exactly what the real sweep would collect, and
6632    /// deletes none of it.
6633    #[tokio::test]
6634    async fn gc_dry_run_reports_without_deleting() {
6635        let sink = GcSink::new();
6636        let stale_base = sink.put_snapshot(1).await;
6637        let base = sink.put_snapshot(2).await;
6638        sink.put_manifest(10, &base, 0, 5, 1, 10).await;
6639        sink.put_frame_batch(0, 5, 1, 10).await;
6640        let orphan = sink.put_frame_batch(0, 4, 1, 6).await;
6641
6642        let dry = gc_stream(
6643            &sink.target,
6644            &StreamGcConfig {
6645                grace: Duration::ZERO,
6646                dry_run: true,
6647            },
6648        )
6649        .await
6650        .unwrap();
6651
6652        assert!(dry.dry_run);
6653        assert_eq!(dry.collected_snapshots, vec![stale_base.clone()]);
6654        assert_eq!(dry.collected_frame_objects, vec![orphan.clone()]);
6655        assert!(dry.collected_bytes > 0);
6656        assert!(sink.exists(&stale_base).await, "dry run deleted a snapshot");
6657        assert!(sink.exists(&orphan).await, "dry run deleted a frame object");
6658
6659        // The proposal is what the real sweep then executes, key for key.
6660        let wet = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6661        assert!(!wet.dry_run);
6662        assert_eq!(wet.collected_snapshots, dry.collected_snapshots);
6663        assert_eq!(wet.collected_frame_objects, dry.collected_frame_objects);
6664        assert_eq!(wet.collected_bytes, dry.collected_bytes);
6665        assert!(!sink.exists(&stale_base).await);
6666        assert!(!sink.exists(&orphan).await);
6667    }
6668
6669    /// The pin `gc_stream`'s docs promise: its live set is the exact complement
6670    /// of what a restore selects, on BOTH restore branches. If either selection
6671    /// rule is ever changed without changing the other, this fails.
6672    #[tokio::test]
6673    async fn gc_liveness_is_the_complement_of_restore_selection() {
6674        // Branch 1 — no generations. `hydrate::restore_subject` falls back to
6675        // `snapshot::restore_latest`, which takes the lexically-greatest key.
6676        // That key is live; the older ones are not collected either, because a
6677        // chainless prefix is indistinguishable from a tier-1a sink whose older
6678        // snapshots are history (see `base_snapshots_skipped`).
6679        let sink = GcSink::new();
6680        let _older = sink.put_snapshot(1).await;
6681        let newest = sink.put_snapshot(2).await;
6682
6683        let restore_would_pick = crate::snapshot::latest_snapshot_key(&sink.target)
6684            .await
6685            .unwrap()
6686            .unwrap();
6687        assert_eq!(restore_would_pick, newest);
6688
6689        let out = gc_stream(
6690            &sink.target,
6691            &StreamGcConfig {
6692                grace: Duration::ZERO,
6693                dry_run: true,
6694            },
6695        )
6696        .await
6697        .unwrap();
6698        assert_eq!(out.live_base_snapshot_keys, vec![restore_would_pick]);
6699        assert!(out.base_snapshots_skipped);
6700        assert!(out.collected_snapshots.is_empty());
6701        assert_eq!(out.retained_snapshots, 2);
6702
6703        // Branch 2 — a validating chain. `restore_stream_from_manifests` takes
6704        // the chain's base, which here is the OLDER snapshot. The GC must keep
6705        // it, and keep the newest as well, since the fallback branch is still
6706        // reachable for any reader that finds no manifests.
6707        let chained = GcSink::new();
6708        let chain_base = chained.put_snapshot(1).await;
6709        let unreferenced_newer = chained.put_snapshot(2).await;
6710        chained.put_manifest(10, &chain_base, 0, 5, 1, 10).await;
6711        chained.put_frame_batch(0, 5, 1, 10).await;
6712
6713        let manifests = list_and_parse_generation_manifests(&chained.target)
6714            .await
6715            .unwrap();
6716        let chain = validate_generation_chain(&manifests).unwrap();
6717        assert_eq!(chain.base_snapshot_key, chain_base);
6718
6719        let out = gc_stream(
6720            &chained.target,
6721            &StreamGcConfig {
6722                grace: Duration::ZERO,
6723                dry_run: true,
6724            },
6725        )
6726        .await
6727        .unwrap();
6728        assert!(!out.base_snapshots_skipped, "a chain makes this a tier-2 sink");
6729        assert!(out.live_base_snapshot_keys.contains(&chain.base_snapshot_key));
6730        assert!(out.live_base_snapshot_keys.contains(&unreferenced_newer));
6731        assert!(out.collected_snapshots.is_empty());
6732
6733        // And the retained frame set is exactly what a replay would fetch.
6734        let replayed: Vec<String> = manifests
6735            .iter()
6736            .flat_map(|m| chained.target.frame_objects_of(m))
6737            .map(|(key, _, _)| key.to_string())
6738            .collect();
6739        assert_eq!(out.retained_frame_objects, replayed.len());
6740        assert!(out.collected_frame_objects.is_empty());
6741    }
6742
6743    /// A prefix with no chain keeps every snapshot — its older ones could be a
6744    /// tier-1a sink's history — but its `frames/` are unreachable by
6745    /// construction and still go.
6746    #[tokio::test]
6747    async fn gc_without_a_chain_keeps_snapshots_but_still_collects_frames() {
6748        let sink = GcSink::new();
6749        let older = sink.put_snapshot(1).await;
6750        let newest = sink.put_snapshot(2).await;
6751        let orphan = sink.put_frame_batch(0, 4, 1, 6).await;
6752
6753        let out = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6754
6755        assert!(out.base_snapshots_skipped);
6756        assert!(out.collected_snapshots.is_empty());
6757        assert_eq!(out.collected_frame_objects, vec![orphan.clone()]);
6758        assert!(sink.exists(&older).await);
6759        assert!(sink.exists(&newest).await);
6760        assert!(!sink.exists(&orphan).await);
6761    }
6762
6763    /// An empty prefix is a no-op, not an error — a GC pointed at a sink that
6764    /// has never been written to must not fail.
6765    #[tokio::test]
6766    async fn gc_on_an_empty_prefix_collects_nothing() {
6767        let sink = GcSink::new();
6768        let out = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6769        assert_eq!(
6770            out,
6771            StreamGcOutcome {
6772                base_snapshots_skipped: true,
6773                ..Default::default()
6774            }
6775        );
6776    }
6777}