turso_backup/stream.rs
1//! Tier 2 — WAL-frame streaming sink (R005-F2).
2//!
3//! Tee WAL frames from a live turso connection to an object store, anchored
4//! to a tier-1a base snapshot. Restore (R005-F3) replays the frames onto the
5//! snapshot. Near-zero RPO; engine-coupled but the public seam is small —
6//! see the spike findings in `.yah/docs/working/turso-s3-backup.md`.
7//!
8//! ## Object layout under `BackupTarget::prefix`
9//!
10//! ```text
11//! frames/{epoch:020}/{checkpoint_seq:010}/{first:020}-{last:020} batch of consecutive WAL frames
12//! frames/{checkpoint_seq:010}/{first:020}-{last:020} ditto, epoch 0 (unfenced)
13//! frames/{epoch:020}/{checkpoint_seq:010}/{frame_no:020} one raw frame, pre-R761-F2 layout
14//! frames/{checkpoint_seq:010}/{frame_no:020} ditto, epoch 0 (pre-fencing)
15//! generations/gen-{unix_nanos:020}.manifest one per `tail_frames` call that uploaded
16//! latest.stream-watermark text sidecar: "<checkpoint_seq> <last_frame> <written_at_nanos> <epoch> <pointer_generation> <salt1> <salt2>"
17//! ```
18//!
19//! R858-B19: the trailing salt pair is the WAL generation the `last_frame`
20//! position belongs to. `checkpoint_seq` alone does not identify a generation —
21//! a writer-process restart recreates the WAL back at sequence 0 — so the salt
22//! is what `tail_frames` compares and what a generation manifest stamps. See
23//! [`WalGeneration`].
24//!
25//! A frame object holds `n` consecutive frames, each `24 + page_size` bytes,
26//! concatenated in ascending frame order — so a batch is exactly the bytes the
27//! old per-frame objects held, glued together, and offset `i * frame_size`
28//! within it is frame `first + i`.
29//!
30//! The generation manifest names the base snapshot key, the page size, the
31//! frame range covered, which batch objects cover it (since R761-F2), and
32//! (since R732-F2) the fencing epoch and owner label. Object keys are
33//! zero-padded so lexical order matches chronological order (same convention as
34//! tier 1a snapshots / tier 1b manifests; clock-skew-immune).
35//!
36//! ## Frame batching (R761-F2, W248/W313 §9)
37//!
38//! One object per WAL frame made per-write Class A ops scale with *frames*,
39//! and a frame is one page — at 4 KB pages, every 4 KB of changed data was a
40//! billed op, which is a pathologically small object. A tail call's frames now
41//! go up as one ranged object per drain batch (bounded by
42//! [`BackpressureConfig::spill_buffer_frames`], the same bound that already
43//! caps how much this sink holds in memory), so an uploading call costs
44//! `ceil(frames / spill_buffer_frames) + 2` PUTs instead of `frames + 2` —
45//! typically **3, flat**, whatever the write volume in the interval.
46//!
47//! Reading stays compatible in the direction that matters: a manifest without
48//! a `frame_batch` list is a pre-R761-F2 generation and its frames are fetched
49//! one object each, so backups written before this change (and chains that
50//! straddle it) still restore. The reverse is a loud refusal by construction —
51//! the manifest header moved to `v3`, so an older binary reading a batched
52//! generation says "unexpected manifest header" rather than mis-reading it.
53//!
54//! ## Fencing (R732-F2 / R732-T3, W245; R736-T2, W250)
55//!
56//! [`StreamConfig::epoch`] is a per-tenant fencing token minted by yubaba's
57//! raft state machine. It is checked against the sidecar before any frame is
58//! uploaded, and the sidecar advance is a compare-and-swap on the version that
59//! check read — so a stale owner is *rejected* (with
60//! [`StreamOutcome::Fenced`]) both when it arrives late and when it races. An
61//! epoch of `0` means unfenced, which preserves single-writer behaviour but is
62//! still fenced *by* a claimed sink.
63//!
64//! That epoch is a **local** raft counter — it fences ownership moves within
65//! one cell, but two cells are independent raft groups that share no epoch
66//! counter, so it is blind to a tenant moving to a *different* cell.
67//! [`StreamConfig::pointer_generation`] is the second, cross-cell fence: the
68//! writer's belief about the global tenant→cell pointer's generation
69//! (`yah_tenant_pointer::PointerRecord`). It is checked alongside `epoch` at
70//! the same two points — up front against the sidecar, and again on a lost
71//! watermark CAS — so a stale generation bounces exactly like a stale epoch.
72//! A node is the real owner of a tenant only when **both** fences pass.
73//!
74//! ## Seam isolation
75//!
76//! Every call into `turso_core::Connection`'s `feature = "conn_raw_api"`
77//! surface goes through the [`WalSeam`] trait. If a future turso release
78//! renames or reshapes those calls, the delta is a single impl block — not a
79//! sed across the whole sink. The spike measured one breaking rename in three
80//! months (`wal_auto_checkpoint_disable` → `wal_auto_actions_disable`); the
81//! trait makes that a one-file fix.
82//!
83//! @yah:relay(R005, "Tier 2 — WAL-frame streaming (deferred)")
84//! @yah:at(2026-05-26T22:28:31Z)
85//! @yah:status(open)
86//! @yah:phase(P3)
87//! @yah:parent(Q002)
88//! @arch:see(.yah/docs/working/turso-s3-backup.md)
89//!
90//! @arch:see(.yah/docs/working/turso-s3-backup.md)
91//!
92//! @yah:ticket(R005-F2, "Frame-streaming sink: base snapshot + incremental frames + generation tracking")
93//! @yah:assignee(agent:claude)
94//! @yah:at(2026-05-26T22:30:09Z)
95//! @yah:status(review)
96//! @yah:phase(P3)
97//! @yah:parent(R005)
98//! @arch:see(.yah/docs/working/turso-s3-backup.md)
99//! @yah:handoff("Implemented in src/stream.rs (~590 LOC). Public API: WalSeam trait (3 methods: wal_state, wal_get_frame, wal_auto_actions_disable) + real CoreWalSeam impl over turso_core::Connection + StreamConfig{base_snapshot_key, page_size} + tail_frames(seam, target, cfg) -> StreamOutcome (Empty | Streamed | Restarted) + GenerationManifest format/parse + Watermark sidecar.")
100//! @yah:handoff("Object layout under prefix: frames/{checkpoint_seq:010}/{frame_no:020} for raw frames (24-byte header + page), generations/gen-{nanos:020}.manifest for per-call manifests, latest.stream-watermark for the (checkpoint_seq, last_frame) sidecar. Same zero-padded-key convention as tier 1a snapshots and tier 1b manifests — clock-skew-immune.")
101//! @yah:handoff("Tier-2 invariants from the R005-T1 spike are baked in: (checkpoint_seq, frame_no) compound key (not raw frame_no), so a WAL restart -> Restarted outcome under a new seq; page_size is a required cfg parameter (no 4096 hardcode like sync_server.rs); WalSeam isolates all turso_core::Connection calls (one-file delta if upstream renames again); CoreWalSeam::open() takes WAL ownership via wal_auto_actions_disable() at construction.")
102//! @yah:handoff("Cargo.toml: added turso_core = '0.6.1' with features = ['conn_raw_api'] as sibling dep to turso='0.6.1'. The friendly turso wrapper does NOT re-export the raw WAL API; pinned in lockstep — if either bumps, bump both.")
103//! @yah:handoff("Verified: 6 new stream unit tests (empty-on-empty-wal, initial-tail-records-watermark, second-tail-no-new-frames-is-Empty, second-tail-uploads-only-new, wal-restart-emits-Restarted-under-new-seq, manifest-roundtrip+rejection) using a mockable WalSeam. Full crate: 18/18 tests green. cargo clippy --all-targets -- --deny=warnings clean.")
104//! @yah:handoff("What's NOT verified at F2 level: live-DB ping-pong of CoreWalSeam (seed rows -> wal_state -> wal_get_frame -> assert is_commit_frame on the last frame). F3's restore path will be the natural end-to-end exercise. Optional sanity test could be added under F2 if you'd rather catch a CoreWalSeam regression here vs in F3.")
105//! @yah:handoff("Generation manifest text format: 'TURSO-BACKUP STREAM v1' header, then `base_snapshot <key>`, `page_size <n>`, `checkpoint_seq <n>`, `first_frame <n>`, `last_frame <n>`. Dependency-free, same convention as dedup::Manifest. parse_generation_manifest fails loudly on any other shape.")
106//! @yah:next("User: review/approve F2. If you want a live-DB CoreWalSeam sanity test before signoff, say so and I'll add it under F2; otherwise F3 picks it up naturally.")
107//! @yah:next("On approval: archive F2, claim R005-F3 (Restore via frame replay onto snapshot). F3 fetches latest gen-manifest, downloads referenced base_snapshot + frames in (checkpoint_seq, frame_no) order, replays into a writable DB via wal_insert_begin/wal_insert_frame/wal_insert_end (the same WalSeam trait, extended with the insert side).")
108//! @yah:next("Optional independent of F3: file a tiny upstream PR to re-export conn_raw_api from the `turso` wrapper crate so the sibling turso_core dep collapses to one.")
109//!
110//! @yah:ticket(R005-F3, "Restore via frame replay onto snapshot + restart/crash-consistency handling")
111//! @yah:assignee(agent:claude)
112//! @yah:at(2026-05-26T22:30:10Z)
113//! @yah:status(review)
114//! @yah:phase(P3)
115//! @yah:parent(R005)
116//! @arch:see(.yah/docs/working/turso-s3-backup.md)
117//! @yah:depends_on(R005-F2)
118//! @yah:handoff("Built restore_latest_stream(&BackupTarget, dest_path) + the WalInsertSeam trait extension on the existing WalSeam pattern. Public surface: WalInsertSeam{wal_insert_begin, wal_insert_frame, wal_insert_end} + impl for CoreWalSeam (wraps turso_core::Connection::wal_insert_*); RestoreOutcome{base_snapshot_key, checkpoint_seq, generation_count, frames_replayed, last_frame}; restore_latest_stream(target, dest_path)->RestoreOutcome.")
119//! @yah:handoff("Flow: list+sort all generations/*.manifest keys lexicographically → parse each → validate_generation_chain checks all share one base, one page_size, one checkpoint_seq, frames start at 1 and are contiguous → download base_snapshot to dest_path → CoreWalSeam::open(dest) → wal_insert_begin → replay_frames_into walks (seq, frame_no) in order, downloads frames/{seq:010}/{frame_no:020}, asserts byte length = 24+page_size, calls wal_insert_frame → wal_insert_end(force_commit=false). Crash-consistency story = the engine's own truncate-to-last-commit-frame on insert_end(false).")
120//! @yah:handoff("Restart handling = REFUSE: a chain spanning two checkpoint_seqs means the source engine folded WAL into main between generations, so the post-restart frames don't replay onto our pre-restart base. validate_generation_chain bails with 'WAL restart between generations, restore needs a fresh tier-1a snapshot'. Same refusal for cross-base chains and frame gaps. This is the right semantics — heroic restart-spanning replay would silently corrupt.")
121//! @yah:handoff("Verified with 13 new stream tests on top of F2's 6: validate_chain accepts single/contiguous-multi, rejects empty/gap/non-one-start/restart/base-mismatch/page_size-mismatch (7 tests); replay_walks_manifests_in_frame_order with MockInsertSeam; restore_errors_when_no_generations; replay_rejects_wrong_size_frame; live_db_seed_snapshot_tail_restore_round_trips (real turso + CoreWalSeam: seed 50 → snapshot → checkpoint → seed 25 more → tail → restore → assert dest has 75 rows); manifest_with_uncommitted_tail_rolls_back_via_insert_end (live: stage a phantom non-commit frame in the sink, extend the manifest, assert restore drops it and dest=3 rows = 2 base + 1 committed, NOT 4).")
122//! @yah:handoff("cargo test -p turso-backup = 31/31 green (up from 18); cargo clippy --all-targets -- --deny=warnings = clean. One clippy lint fixed in the new test code: manual_is_multiple_of (Rust 1.95 lint). The live tests use turso::Builder for writes/reads + CoreWalSeam for tailing — exclusive WAL lock means we drop the high-level conn before opening the low-level seam (matched by the snapshot tests' pattern).")
123//! @yah:handoff("Two known design decisions worth flagging to the reviewer: (1) replay starts at frame 1 in the dest's fresh WAL but the source's frames could overlap content already in the base snapshot (since VACUUM INTO point-in-time includes WAL state). turso's wal_insert_frame compares-and-returns-OK on identical content, so redundant frames are no-ops — clean orchestration via checkpoint-then-snapshot-then-stream avoids them entirely (which is what the live test does). (2) on replay error we attempt a best-effort wal_insert_end(false) before bubbling the error up, so we don't leave the dest's WAL with an uncommitted suffix half-open.")
124//! @yah:next("User: review/approve F3. Once green, archive F3 and the F2/F3-dependent state of R005 collapses to F4 only (concurrent-writer-safe raw copy).")
125//! @yah:verify("cargo test -p turso-backup")
126//! @yah:verify("cargo clippy --all-targets -- --deny=warnings")
127//!
128//! @yah:ticket(R005-F4, "Concurrent-writer-safe raw copy: read-only main + WAL-frame replay (no TRUNCATE-checkpoint dependency), for backing up under a live writer")
129//! @yah:assignee(agent:claude)
130//! @yah:at(2026-05-27T03:05:00Z)
131//! @yah:status(review)
132//! @yah:phase(P3)
133//! @yah:parent(R005)
134//! @arch:see(.yah/docs/working/turso-s3-backup.md)
135//! @yah:gotcha("R004's dedup::raw_consistent_copy assumes no live writer: it folds WAL->main via PRAGMA wal_checkpoint(TRUNCATE), which a concurrent writer can make return 'busy'. The live-writer alternative copies the main file under a read txn and replays WAL frames itself (wal_get_frame seam) — tier-2 territory, gated on R005-T1's seam assessment. Filed as an R004-T4 followup.")
136//! @yah:handoff("Implemented raw_consistent_copy_live(db_path, page_size) + pub(crate) replay_wal_onto_main inner. Algorithm: open CoreWalSeam (auto-actions disabled on our conn) -> wal_state for (cp_seq, max_frame) -> std::fs::read main -> walk frames 1..=max_frame collecting (page_no, db_size, page_bytes), locate the last is_commit_frame -> grow image to fit the largest page slot in the commit prefix -> overwrite each page at (page_no-1)*page_size -> truncate to db_size*page_size. Frames past the last commit are dropped wholesale (crash-consistency, matches restore's wal_insert_end(false)).")
137//! @yah:next("User: review/approve F4. Once green, archive F4 and R005 collapses to closed (R005-F2/F3/F4 all in review).")
138//! @yah:verify("cargo test -p turso-backup: 38/38 green (up from 31). New: 5 unit tests on replay_wal_onto_main with MockWal (empty-WAL/single-commit/uncommitted-tail-dropped/multi-commit-prefix-grows-image/only-uncommitted-frames-returns-main) + 2 live tests on raw_consistent_copy_live (live_db_consistent_copy_without_truncate_round_trips: seed -> checkpoint_truncate -> seed more uncheckpointed -> copy-live -> reopen = 35 rows; live_db_consistent_copy_reflects_new_writes: repeatable copy after subsequent writes).")
139//! @yah:verify("cargo clippy --all-targets -- --deny=warnings: clean.")
140//!
141//! ## R574-F2 — explicit R2 backpressure
142//!
143//! `tail_frames`'s upload loop drains frames through
144//! [`crate::backpressure::put_with_backoff`] against a bounded spill buffer
145//! (see [`drain_frames_with_backpressure`]) instead of putting each frame
146//! directly and bubbling any error raw. Full design in the `backpressure`
147//! module doc; ticket tracked in the W248 relay, not in-source.
148//! @arch:see(.yah/docs/working/W248-wal-streamer-hardening.md)
149//!
150//! ## R574-T4 — explicit RPO knob
151//!
152//! Cadence was entirely implicit in the caller's `tail_frames` invocation
153//! interval (doc §10). `StreamConfig::rpo_target` states that interval as a
154//! number; every `tail_frames` call reports [`RpoStatus`] (age of the last
155//! durably-persisted watermark, and whether that age has drifted past the
156//! target) on its [`StreamOutcome`] so an orchestrator can alert without
157//! reimplementing the bookkeeping. This crate still does not schedule
158//! anything itself — cadence/retention policy stays caller-driven per the
159//! doc's "Not in scope" — `rpo_target` documents the contract the caller's
160//! own scheduler is expected to uphold, and `RpoStatus` is the receipt.
161//! @arch:see(.yah/docs/working/W248-wal-streamer-hardening.md)
162//!
163//! ## R761-T1 — a measured default tail cadence
164//!
165//! [`DEFAULT_TAIL_INTERVAL`] (60 s) and [`DEFAULT_RPO_TARGET`] (120 s) are
166//! the cadence this crate recommends, derived from
167//! `examples/tail_sweep_harness.rs`'s 2026-08-13 sweep rather than chosen:
168//! every uploading `tail_frames` call writes two fixed objects (generation
169//! manifest + watermark CAS) on top of the frames, so per-write cost was
170//! `frames/write + 2/writes_per_tail` — 4.03 PUTs/write at one write per
171//! tail, 2.05 at a hundred. The constants' docs carry the full table and
172//! the reasoning for landing at 60 s instead of chasing the last 3 %. Still
173//! no scheduler in this crate; these are numbers for the caller's.
174//!
175//! R761-F2 then removed the `frames/write` term those numbers were floored
176//! by (see *Frame batching* above), so the cost is now `3/writes_per_tail`
177//! for a call whose frames fit one batch — the cadence lever and the layout
178//! lever compose, and the second is worth more the longer the interval.
179//! @arch:see(.yah/docs/working/W248-wal-streamer-hardening.md)
180//!
181//! ## R574-F3 — one-puller-per-box fan-out + warm applier
182//!
183//! [`crate::puller::WalPuller`] is the read side's counterpart to
184//! `tail_frames`'s write side: it pulls each newly-uploaded frame from R2
185//! exactly once per box and fans it out in-process to every attached warm
186//! applier (a [`WalInsertSeam`] with its page cache trimmed hard via
187//! [`CoreWalSeam::trim_page_cache_kb`]), so R2 read ops are O(boxes), not
188//! O(replicas). See the `puller` module doc for the full design and its v1
189//! scope cut (attach is cold-start-only; no mid-stream backlog replay).
190//! @arch:see(.yah/docs/working/W248-wal-streamer-hardening.md)
191
192use anyhow::{Context, Result};
193use object_store::path::Path as ObjPath;
194use object_store::{ObjectStore, ObjectStoreExt, PutMode, PutOptions, UpdateVersion};
195use std::collections::{HashSet, VecDeque};
196use std::sync::Arc;
197use std::time::{Duration, SystemTime, UNIX_EPOCH};
198
199use crate::backpressure::{put_with_backoff, BackpressureConfig, BackpressurePolicy, BackpressureReport};
200use crate::snapshot::BackupTarget;
201
202/// WAL frame layout: 24-byte frame header (page_no big-endian at 0..4,
203/// db_size big-endian at 4..8, salts/checksums in the remainder) followed by
204/// `page_size` bytes of page data. Constant per the SQLite WAL format; the
205/// page size is read from the base snapshot's header (page 1, byte 16, `u16`
206/// with the value `1` meaning 65536). `sync_server.rs` hardcodes 4096 — we
207/// don't, see [`StreamConfig::page_size`].
208pub const WAL_FRAME_HEADER_SIZE: usize = 24;
209
210/// Position in the WAL: a `(checkpoint_seq, last_frame)` pair. `max_frame`
211/// resets to 0 every time the WAL header restarts (`WalAutoActions::Restart`
212/// fires), and `checkpoint_seq_no` increments alongside it — so the pair is
213/// the right primary key for sink objects, not raw frame_no.
214///
215/// **R858-B19 — `checkpoint_seq` alone does NOT identify a WAL generation, and
216/// this doc used to claim it did ("increments monotonically across restarts").
217/// That claim is false and it cost a silent wrong restore.** It holds only for
218/// an *in-process* restart, where the same WAL file is reused. When the last
219/// connection to a SQLite database closes, the engine checkpoints and DELETES
220/// the `-wal` file; the next writer creates a fresh WAL back at
221/// checkpoint-sequence `0` with a brand-new random salt. Across that fold
222/// `checkpoint_seq` goes `0 -> 0` over two completely unrelated WALs (measured
223/// 2026-09-06 against system sqlite3 3.51.0 by
224/// `examples/foreign_checkpoint_probe.rs`, probes B and F).
225///
226/// Generation identity therefore lives in [`WalGeneration`], which pairs the
227/// sequence with the WAL header's salt. Never compare two `Watermark`s'
228/// `checkpoint_seq` to decide "same WAL" — that is exactly the inference this
229/// ticket exists to delete. `last_frame` is only meaningful *relative to a
230/// known generation*.
231#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
232pub struct Watermark {
233 pub checkpoint_seq: u32,
234 pub last_frame: u64,
235}
236
237/// R858-B19 — the WAL header's `(salt1, salt2)`, the two 32-bit values SQLite
238/// re-rolls on every WAL reset. Stored big-endian at bytes 16..24 of the
239/// 32-byte WAL header, and copied verbatim into bytes 8..16 of **every frame
240/// header** written under that header — which is where this crate reads it
241/// from, since `turso_core::WalState` exposes only `checkpoint_seq_no` and
242/// `max_frame`.
243///
244/// Why the salt and not the sequence: the salt moves in *both* fold regimes,
245/// and the sequence moves in only one.
246///
247/// - WAL recreated by a writer restart: fresh randomness (measured
248/// `b83c03f5 -> 9ca89e39`, unrelated), while `checkpoint_seq` resets `0 -> 0`.
249/// - In-process autocheckpoint: `salt1` increments in lockstep with the
250/// sequence (measured `d492ea8a -> ... -> d492ea92` alongside seq `0 -> 8`).
251#[derive(Debug, Clone, Copy, PartialEq, Eq)]
252pub struct WalSalt {
253 pub salt1: u32,
254 pub salt2: u32,
255}
256
257impl WalSalt {
258 /// Read the salt out of a WAL **frame** header (bytes 8..16 of the 24-byte
259 /// header, big-endian). Panics on a short slice — every caller here sizes
260 /// its buffer at `WAL_FRAME_HEADER_SIZE + page_size`.
261 pub(crate) fn from_frame_header(frame: &[u8]) -> Self {
262 let be = |o: usize| u32::from_be_bytes(frame[o..o + 4].try_into().unwrap());
263 Self { salt1: be(8), salt2: be(12) }
264 }
265
266 /// R858-B18 — read the salt out of the **32-byte `-wal` file header**
267 /// (bytes 16..24, big-endian), the copy SQLite writes once per WAL
268 /// generation and duplicates into every frame header.
269 ///
270 /// The two constructors differ only in offset, and they live together
271 /// deliberately: one type, one place, so the two ways this crate can reach
272 /// the same value can never disagree. [`from_frame_header`](Self::from_frame_header)
273 /// is the seam-only path ([`read_wal_salt`] — works through a
274 /// [`CoreWalSeam::from_conn`] that has no path); this one is the
275 /// path-only path ([`SourceFingerprint`] — works with no engine open at
276 /// all, which is what makes it usable as an independent check *on* a copy
277 /// the engine took).
278 ///
279 /// Panics on a slice shorter than 24 bytes; [`WalFileHeader::parse`] is the
280 /// length-checked entry point every caller here actually uses.
281 pub(crate) fn from_wal_file_header(hdr: &[u8]) -> Self {
282 let be = |o: usize| u32::from_be_bytes(hdr[o..o + 4].try_into().unwrap());
283 Self { salt1: be(16), salt2: be(20) }
284 }
285}
286
287impl std::fmt::Display for WalSalt {
288 fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
289 write!(f, "{:08x}/{:08x}", self.salt1, self.salt2)
290 }
291}
292
293/// R858-B19 — the identity of one WAL generation: the checkpoint sequence
294/// **and** the header salt that actually distinguishes it. This is what
295/// [`tail_frames`] compares across calls, what the watermark sidecar persists,
296/// and what a generation manifest stamps.
297///
298/// `salt: None` means **unknown generation**, and unknown is a value here, not
299/// a missing one: it arises from a sidecar or manifest written before this
300/// field existed, or from a WAL with no frames to read a salt out of. An
301/// unknown generation is never provably equal to anything — see
302/// [`WalGeneration::is_provably_same_as`] — so it forces a restart on the write
303/// side and a refusal on the restore side rather than a guess.
304#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
305pub struct WalGeneration {
306 pub checkpoint_seq: u32,
307 pub salt: Option<WalSalt>,
308}
309
310impl WalGeneration {
311 /// True only when both sides carry a **known** salt and every component
312 /// agrees — i.e. only when these are *provably* the same WAL.
313 ///
314 /// The asymmetry is the whole point. "Not provably the same" is treated as
315 /// "different", which costs a redundant re-upload (or a loud restore
316 /// refusal) in the worst case. The opposite default — "no evidence of a
317 /// change, so assume it is the same WAL" — is what spliced two WAL
318 /// generations into one chain and restored a plausible wrong image.
319 pub fn is_provably_same_as(&self, other: &WalGeneration) -> bool {
320 match (self.salt, other.salt) {
321 (Some(a), Some(b)) => a == b && self.checkpoint_seq == other.checkpoint_seq,
322 _ => false,
323 }
324 }
325
326 /// Render for an error message: `seq 4 salt 1109ca5e/7acf42a3`, or
327 /// `seq 4 salt <unknown>` when the salt was never recorded.
328 pub(crate) fn describe(&self) -> String {
329 match self.salt {
330 Some(s) => format!("seq {} salt {s}", self.checkpoint_seq),
331 None => format!("seq {} salt <unknown>", self.checkpoint_seq),
332 }
333 }
334}
335
336/// R858-B18 — the 32-byte `-wal` file header, read straight off disk with no
337/// engine open. Ground truth about which WAL generation is on disk and how far
338/// it has been written, independent of anything turso caches.
339///
340/// Only the three fields that move are kept. The rest of the header (magic,
341/// format version, the two header checksums) is either constant for a given
342/// build or a function of these; a change to any of it that did *not* move one
343/// of these three would not be a change this crate can act on.
344#[derive(Debug, Clone, Copy, PartialEq, Eq)]
345pub struct WalFileHeader {
346 /// Page size the WAL's frames carry (bytes 8..12).
347 pub page_size: u32,
348 /// Checkpoint sequence (bytes 12..16) — the field `tail_frames` used to key
349 /// restart detection on before [`WalGeneration`] paired it with the salt.
350 pub checkpoint_seq: u32,
351 /// The generation's salt (bytes 16..24).
352 pub salt: WalSalt,
353}
354
355impl WalFileHeader {
356 /// SQLite's WAL header is exactly this many bytes, ahead of frame 1.
357 pub const SIZE: usize = 32;
358
359 /// Parse a WAL header out of the first [`Self::SIZE`] bytes of a `-wal`
360 /// file. `None` for anything shorter — a truncated or freshly-created WAL
361 /// has no generation to name yet, which is an honest unknown rather than an
362 /// error (see [`WalGeneration`]'s `salt: None`).
363 pub fn parse(bytes: &[u8]) -> Option<Self> {
364 if bytes.len() < Self::SIZE {
365 return None;
366 }
367 let be = |o: usize| u32::from_be_bytes(bytes[o..o + 4].try_into().unwrap());
368 Some(Self {
369 page_size: be(8),
370 checkpoint_seq: be(12),
371 salt: WalSalt::from_wal_file_header(bytes),
372 })
373 }
374
375 /// Read `{db_path}-wal`'s header. `Ok(None)` when the sidecar is absent
376 /// (WAL folded and deleted, or journal_mode != WAL) or too short to parse.
377 pub fn read(db_path: &str) -> Result<Option<Self>> {
378 match std::fs::File::open(format!("{db_path}-wal")) {
379 Ok(mut f) => {
380 let mut buf = [0u8; Self::SIZE];
381 let mut filled = 0usize;
382 loop {
383 match std::io::Read::read(&mut f, &mut buf[filled..]) {
384 Ok(0) => break,
385 Ok(n) => filled += n,
386 Err(e) if e.kind() == std::io::ErrorKind::Interrupted => continue,
387 Err(e) => {
388 return Err(e).with_context(|| format!("reading {db_path}-wal header"))
389 }
390 }
391 if filled == Self::SIZE {
392 break;
393 }
394 }
395 Ok(Self::parse(&buf[..filled]))
396 }
397 Err(e) if e.kind() == std::io::ErrorKind::NotFound => Ok(None),
398 Err(e) => Err(e).with_context(|| format!("opening {db_path}-wal")),
399 }
400 }
401}
402
403/// R858-B18 — everything about a source database's **files** that must hold
404/// still for a copy of them to be a point-in-time image, sampled with no engine
405/// open and no lock taken.
406///
407/// This is the other half of reading a foreign-written database. A ReadOnly
408/// open ([`CoreWalSeam::open_reader`]) gets us in the door without stealing the
409/// whole-file lock, but it buys **no shared locking protocol** with upstream C
410/// SQLite: turso locks whole files with `fcntl`, C SQLite uses byte-range locks
411/// plus the `-shm` WAL index, and neither engine observes the other's. So a
412/// foreign checkpoint landing in the middle of our copy can fold WAL frames
413/// into the main file we have already half-read, and the result is a torn image
414/// that still passes `PRAGMA integrity_check`.
415///
416/// Rather than reimplement SQLite's reader protocol (registering a read-mark in
417/// the `-shm` WAL index — a research project and a permanent compatibility
418/// liability against an engine we do not control), [`raw_consistent_copy_live`]
419/// uses textbook **optimistic validation**: sample this before the copy, sample
420/// it again after, and accept the copy only if nothing moved. That converts
421/// "may silently read torn state" into "detects torn state and refuses", needs
422/// no cooperation from the foreign engine, and costs two stats and a 132-byte
423/// read per attempt.
424///
425/// ## Why these fields
426///
427/// - `wal` (salt + checkpoint_seq) moves on **every** WAL reset, in both fold
428/// regimes — fresh randomness on a writer restart, `salt1` incrementing on an
429/// in-process autocheckpoint (both measured; see [`WalSalt`]).
430/// - `main_len` moves when a checkpoint grows the main database.
431/// - `change_counter` (main header bytes 24..28) moves on every write to the
432/// main file — i.e. on every checkpoint — even one that leaves its length
433/// alone. It is meaningful here **only because the foreign writer is C
434/// SQLite**: turso does not maintain this field (it stays `1` in every
435/// journal mode, verified — see `snapshot.rs`'s two-gate rationale), which is
436/// exactly why the WAL salt carries the weight and this one is corroboration.
437/// - `wal_len` is recorded for the report but deliberately **not** part of the
438/// accept/reject test — see [`Self::stable_across`], which is the comparison
439/// to use. There is no `PartialEq` on this type on purpose: a bare `==` would
440/// silently include `wal_len` and refuse every copy taken while the
441/// application was merely writing.
442#[derive(Debug, Clone, Copy)]
443pub struct SourceFingerprint {
444 /// Length of the main database file.
445 pub main_len: u64,
446 /// The main header's change counter, or `None` when the file is too short
447 /// to carry a SQLite header at all (a database whose page 1 still lives
448 /// only in the WAL). Unknown-and-unknown compares equal, which is safe
449 /// because `main_len` participates in the same comparison.
450 pub change_counter: Option<u32>,
451 /// The `-wal` header, or `None` when there is no WAL sidecar.
452 pub wal: Option<WalFileHeader>,
453 /// Length of the `-wal` file (`0` when absent).
454 pub wal_len: u64,
455}
456
457impl SourceFingerprint {
458 /// Offset of the change counter in SQLite's 100-byte database header.
459 const CHANGE_COUNTER_OFFSET: usize = 24;
460
461 /// Sample the fingerprint of the database at `db_path`. Touches nothing:
462 /// two metadata calls plus a 28-byte and a 32-byte read.
463 pub fn read(db_path: &str) -> Result<Self> {
464 let mut hdr = [0u8; Self::CHANGE_COUNTER_OFFSET + 4];
465 let main_len = match std::fs::File::open(db_path) {
466 Ok(mut f) => {
467 let len = f
468 .metadata()
469 .with_context(|| format!("stat {db_path}"))?
470 .len();
471 if len as usize >= hdr.len() {
472 std::io::Read::read_exact(&mut f, &mut hdr)
473 .with_context(|| format!("reading {db_path} header"))?;
474 }
475 len
476 }
477 Err(e) => return Err(e).with_context(|| format!("opening source db {db_path}")),
478 };
479 let change_counter = (main_len as usize >= hdr.len()).then(|| {
480 u32::from_be_bytes(hdr[Self::CHANGE_COUNTER_OFFSET..].try_into().unwrap())
481 });
482 Ok(Self {
483 main_len,
484 change_counter,
485 wal: WalFileHeader::read(db_path)?,
486 wal_len: std::fs::metadata(format!("{db_path}-wal"))
487 .map(|m| m.len())
488 .unwrap_or(0),
489 })
490 }
491
492 /// True when nothing that can **tear** a copy moved between `self` (sampled
493 /// before) and `after` (sampled after). This is the accept test in
494 /// [`raw_consistent_copy_live`], and it is narrower than field equality on
495 /// purpose.
496 ///
497 /// ## What can tear the copy, and what cannot
498 ///
499 /// The copy reads the main file, then replays WAL frames `1..=max_frame`
500 /// captured when the seam opened. Against that algorithm:
501 ///
502 /// - **A checkpoint tears it.** It rewrites pages of the main file *and*
503 /// resets the WAL, so our already-read main bytes and our frame reads can
504 /// straddle the fold — replaying pre-fold frames over post-fold pages
505 /// rolls pages backwards. Caught: a checkpoint bumps `change_counter`
506 /// and/or `main_len`, and a WAL restart re-rolls the salt and the sequence
507 /// (measured in both regimes, `examples/foreign_checkpoint_probe.rs`
508 /// probes A and E).
509 /// - **A plain append does NOT tear it.** SQLite only ever appends frames
510 /// within a generation, and only a reset (which re-rolls `salt1`) lets it
511 /// overwrite an existing frame. So frames `1..=max_frame` are immutable
512 /// for as long as the salt holds, and a writer that commits during our
513 /// copy just means our image is a slightly earlier point in time — which
514 /// is what a point-in-time copy *is*.
515 ///
516 /// Which is why `wal_len` and the frame count are excluded. Including them
517 /// buys no additional safety and costs a refusal on every copy taken while
518 /// the application is writing at all — turning a working backup into one
519 /// that only succeeds against an idle database. R858-B18 measured that
520 /// difference rather than assuming it; see probe H.
521 pub fn stable_across(&self, after: &Self) -> bool {
522 self.main_len == after.main_len
523 && self.change_counter == after.change_counter
524 && self.wal == after.wal
525 }
526
527 /// One-line rendering for the refusal message, so a failure names what
528 /// actually moved instead of asserting "something did".
529 pub(crate) fn describe(&self) -> String {
530 let wal = match self.wal {
531 Some(h) => format!("seq {} salt {}", h.checkpoint_seq, h.salt),
532 None => "<no WAL>".to_string(),
533 };
534 format!(
535 "main {}B change_counter {} | wal {}B {wal}",
536 self.main_len,
537 self.change_counter.map_or("<unknown>".to_string(), |c| c.to_string()),
538 self.wal_len,
539 )
540 }
541}
542
543/// Subset of `turso_core::Connection`'s `feature = "conn_raw_api"` surface
544/// we use. Every WAL call into turso goes through this trait. A future
545/// signature shift becomes a one-impl delta.
546pub trait WalSeam {
547 /// Snapshot the WAL position (checkpoint seq + max_frame).
548 fn wal_state(&self) -> Result<Watermark>;
549
550 /// Fetch frame `frame_no` (1-based) into `buf` (must be
551 /// `WAL_FRAME_HEADER_SIZE + page_size` bytes). Returns the page number
552 /// the frame applies to and the post-frame DB size (in pages) — non-zero
553 /// `db_size` marks a commit frame.
554 fn wal_get_frame(&self, frame_no: u64, buf: &mut [u8]) -> Result<FrameInfo>;
555
556 /// Take WAL ownership: turn off both the auto-checkpoint and the
557 /// auto-WAL-restart so our watermark stays meaningful across calls. The
558 /// in-tree consumer (`cli/sync_server.rs`) makes the same move at startup.
559 fn wal_auto_actions_disable(&self);
560}
561
562/// Restore-side counterpart to [`WalSeam`]: the three `wal_insert_*` calls
563/// `cli/sync_server.rs` clients use to replay frames into a fresh DB. Same
564/// rationale — one impl block to update if upstream renames.
565///
566/// Ordering contract: `begin` → N × `insert_frame(monotonic frame_no)` → `end`.
567/// `end(false)` is the crash-safe default: the engine rolls back any suffix of
568/// frames written after the last commit frame (`db_size > 0`) in the session.
569pub trait WalInsertSeam {
570 /// Open a write transaction with auto-checkpoint/restart suppressed so our
571 /// monotonically-numbered inserts aren't reshuffled mid-session.
572 fn wal_insert_begin(&self) -> Result<()>;
573
574 /// Insert `frame` (a 24-byte header + `page_size` bytes of page data) at
575 /// position `frame_no` (1-based, must be exactly `prev + 1` — gaps error).
576 /// Identical content at an already-written position is a no-op (the engine
577 /// compares and returns OK).
578 fn wal_insert_frame(&self, frame_no: u64, frame: &[u8]) -> Result<()>;
579
580 /// Close the session. With `force_commit = false` (the restore default) the
581 /// engine drops any frames after the last commit frame — automatic
582 /// crash-consistency for a tail captured mid-transaction. `force_commit = true`
583 /// commits even an uncommitted suffix; not used by restore.
584 fn wal_insert_end(&self, force_commit: bool) -> Result<()>;
585}
586
587/// Frame-header projection (what we need for streaming + commit detection).
588/// Mirrors `turso_core::types::WalFrameInfo` without re-exporting the type.
589#[derive(Debug, Clone, Copy, PartialEq, Eq)]
590pub struct FrameInfo {
591 pub page_no: u32,
592 /// Number of pages in the DB after this frame's commit, or `0` if the
593 /// frame is mid-transaction.
594 pub db_size: u32,
595}
596
597impl FrameInfo {
598 pub fn is_commit_frame(&self) -> bool {
599 self.db_size > 0
600 }
601}
602
603/// Real WAL seam backed by a live `turso_core::Connection`.
604pub struct CoreWalSeam {
605 conn: Arc<turso_core::Connection>,
606}
607
608impl CoreWalSeam {
609 /// Open a **writable** connection at `path` and disable auto-checkpoint /
610 /// auto-restart so the caller owns WAL maintenance. Requires `turso_core`
611 /// with `features = ["conn_raw_api"]` (set in this crate's Cargo.toml).
612 ///
613 /// This takes turso's **whole-file exclusive `fcntl` lock**, so it is
614 /// mutually exclusive with any other process holding the database open —
615 /// including upstream C SQLite (measured: `examples/foreign_checkpoint_probe.rs`
616 /// probes D and G). That is correct for the restore/apply direction, which
617 /// owns the destination file outright, and wrong for reading a live source:
618 /// use [`Self::open_reader`] there.
619 pub fn open(path: &str) -> Result<Self> {
620 let io: Arc<dyn turso_core::IO> =
621 Arc::new(turso_core::PlatformIO::new().context("creating turso_core PlatformIO")?);
622 let db = turso_core::Database::open_file(io, path)
623 .with_context(|| format!("opening turso_core db {path}"))?;
624 let conn = db.connect().context("connecting to turso_core db")?;
625 // Take WAL ownership — same first move sync_server.rs makes.
626 conn.wal_auto_actions_disable();
627 Ok(Self { conn })
628 }
629
630 /// R858-B18 — open a **read-only** connection at `path`, taking **no**
631 /// whole-file lock. This is the constructor the backup direction wants: a
632 /// backup is a reader, and it must not lock out the application whose
633 /// database it is reading.
634 ///
635 /// `OpenFlags::ReadOnly` is what buys that, per handle and with no
636 /// process-wide effect. `turso_core-0.7.2/io/unix.rs:67-72` takes the
637 /// exclusive lock only when
638 /// `env::var(ENV_DISABLE_FILE_LOCK).is_err() && !flags.intersects(ReadOnly | NoLock)`;
639 /// `io_uring.rs:463` and `windows.rs:305` carry the identical condition, so
640 /// this is not a unix-only accident. The `LIMBO_DISABLE_FILE_LOCK=1` escape
641 /// hatch reaches the same no-lock state but does it for **every** open in
642 /// the process, including the writable ones — never use it here.
643 ///
644 /// ## Two things this does NOT give you
645 ///
646 /// 1. **No shared locking protocol.** Getting in without a lock is not
647 /// coordination: a foreign checkpoint can still land mid-read. That is
648 /// what [`SourceFingerprint`] validation is for, and why
649 /// [`raw_consistent_copy_live`] pairs the two rather than shipping this
650 /// flag alone — the flag alone converts "refuses to open" into "may
651 /// return a torn image", which is strictly worse.
652 /// 2. **No escape from turso's process-global registry.**
653 /// `Database::open_file_with_flags` consults `DATABASE_MANAGER`, keyed by
654 /// file id, *before* it looks at the flags (`lib.rs:940`), and hands back
655 /// an already-open `Database` with its own flags discarded. So in a
656 /// process that already holds a writable handle on this exact file, this
657 /// call returns that writable handle — same as `snapshot.rs`'s
658 /// `upload_base_snapshot` doc records. Cross-process (the headscale case)
659 /// is unaffected: the registry is per-process.
660 pub fn open_reader(path: &str) -> Result<Self> {
661 let io: Arc<dyn turso_core::IO> =
662 Arc::new(turso_core::PlatformIO::new().context("creating turso_core PlatformIO")?);
663 let db = turso_core::Database::open_file_with_flags(
664 io,
665 path,
666 turso_core::OpenFlags::ReadOnly,
667 turso_core::DatabaseOpts::new(),
668 None,
669 )
670 .with_context(|| format!("opening turso_core db {path} read-only"))?;
671 let conn = db.connect().context("connecting to turso_core db")?;
672 // Belt and braces: a read-only connection cannot checkpoint anyway, but
673 // the seam contract is that nothing we hold folds the WAL under us.
674 conn.wal_auto_actions_disable();
675 Ok(Self { conn })
676 }
677
678 /// Wrap an already-open connection. The caller is responsible for having
679 /// called `wal_auto_actions_disable()` on it.
680 pub fn from_conn(conn: Arc<turso_core::Connection>) -> Self {
681 Self { conn }
682 }
683}
684
685impl WalSeam for CoreWalSeam {
686 fn wal_state(&self) -> Result<Watermark> {
687 let s = self.conn.wal_state().context("turso_core wal_state")?;
688 Ok(Watermark {
689 checkpoint_seq: s.checkpoint_seq_no,
690 last_frame: s.max_frame,
691 })
692 }
693
694 fn wal_get_frame(&self, frame_no: u64, buf: &mut [u8]) -> Result<FrameInfo> {
695 let info = self
696 .conn
697 .wal_get_frame(frame_no, buf)
698 .with_context(|| format!("turso_core wal_get_frame({frame_no})"))?;
699 Ok(FrameInfo {
700 page_no: info.page_no,
701 db_size: info.db_size,
702 })
703 }
704
705 fn wal_auto_actions_disable(&self) {
706 self.conn.wal_auto_actions_disable();
707 }
708}
709
710impl WalInsertSeam for CoreWalSeam {
711 fn wal_insert_begin(&self) -> Result<()> {
712 self.conn
713 .wal_insert_begin()
714 .context("turso_core wal_insert_begin")
715 }
716
717 fn wal_insert_frame(&self, frame_no: u64, frame: &[u8]) -> Result<()> {
718 self.conn
719 .wal_insert_frame(frame_no, frame)
720 .with_context(|| format!("turso_core wal_insert_frame({frame_no})"))?;
721 Ok(())
722 }
723
724 fn wal_insert_end(&self, force_commit: bool) -> Result<()> {
725 self.conn
726 .wal_insert_end(force_commit)
727 .context("turso_core wal_insert_end")
728 }
729}
730
731impl CoreWalSeam {
732 /// Resize this connection's page cache toward `target_kb` kilobytes via
733 /// the standard `PRAGMA cache_size` surface (negative value = KB, per
734 /// SQLite's own convention — see `turso_core::translate::pragma`'s
735 /// `update_cache_size`) rather than reaching into `turso_core`'s
736 /// `Pager::change_page_cache_size` / `CacheResizeResult` directly: the
737 /// latter are technically reachable (`Connection::get_pager()` and
738 /// `Pager`/`Page`/`PageRef` are all `pub use`d at the crate root) but
739 /// `CacheResizeResult` itself is not re-exported, so calling it from
740 /// here would mean handling a value of an unnameable type. The pragma
741 /// path exercises the exact same resize logic through turso_core's own
742 /// public, documented SQL surface instead.
743 ///
744 /// R574-F3: a warm applier trims its cache hard immediately on attach —
745 /// it never serves reads, so cached pages are pure standing RSS cost
746 /// (see the R574-T1 measurement this sizes against).
747 pub(crate) fn trim_page_cache_kb(&self, target_kb: i64) -> Result<()> {
748 anyhow::ensure!(
749 target_kb > 0,
750 "trim_page_cache_kb: target_kb must be positive, got {target_kb}"
751 );
752 self.conn
753 .execute(format!("PRAGMA cache_size = -{target_kb}"))
754 .with_context(|| format!("PRAGMA cache_size = -{target_kb}"))
755 }
756}
757
758/// R761-T1: the tail cadence this crate recommends when a caller has no
759/// reason of its own to pick a different one — **60 s**, and the number is
760/// a measurement result, not a guess.
761///
762/// ## The measurement
763///
764/// `examples/tail_sweep_harness.rs`, run 2026-08-13 against a casual-app
765/// workload (single-row transactions into a table with a secondary index),
766/// sweeping `writes_per_tail` — how many application writes accumulate
767/// between two `tail_frames` calls:
768///
769/// | writes_per_tail | PUTs/write | frames/write |
770/// |----------------:|-----------:|-------------:|
771/// | 1 | 4.03 | 2.03 |
772/// | 5 | 2.43 | 2.03 |
773/// | 25 | 2.11 | 2.03 |
774/// | 100 | 2.05 | 2.03 |
775///
776/// Those four points are not four independent facts. Every `tail_frames`
777/// call that uploads anything writes exactly **two fixed objects** beyond
778/// the frames themselves — one generation manifest and one watermark CAS —
779/// so
780///
781/// ```text
782/// PUTs/write = frames/write + 2 / writes_per_tail
783/// ```
784///
785/// which reproduces all four measured rows to the last digit: `2.03 + 2/1`,
786/// `2.03 + 2/5`, `2.03 + 2/25`, `2.03 + 2/100`. frames/write is FLAT across
787/// the sweep — it is a property of the schema and transaction shape, and
788/// cadence cannot touch it. The entire lever this default pulls is the
789/// `2 / writes_per_tail` term.
790///
791/// ## Why 60 s and not longer
792///
793/// That term has sharply diminishing returns. Of the total reduction
794/// available (4.03 → 2.05), moving from `writes_per_tail` 1 → 5 captures
795/// 81 %, 1 → 25 captures 97 %, and everything from 25 → 100 is the last
796/// 3 %. So the target is the 25 region, not 100.
797///
798/// Converting that writes axis into a time interval needs a write RATE,
799/// and the honest one is the rate DURING an active session — W313 §3.1's
800/// "~10 writes/day" arrives clustered in bursts, not spread evenly, and an
801/// idle tenant costs nothing at any cadence (a tail call with no new frames
802/// uploads no objects at all, so a long interval only ever helps a tenant
803/// that is actively writing). Across the burst rates a casual app produces,
804/// ~0.1–1 write/s:
805///
806/// - 15 s (what a 30 s RPO bound derives today) spans 1.5–15 writes/tail
807/// → 3.36–2.16 PUTs/write. The slow-burst end is the worst case the
808/// ticket is named after, and it is nearly the full 4.03.
809/// - **60 s spans 6–60 writes/tail → 2.36–2.06 PUTs/write.**
810/// - 300 s spans 30–300 → 2.10–2.04: at most 0.26 PUTs/write better than
811/// 60 s, bought with a 5× wider data-loss window on exactly the tenants
812/// that are actively writing. Not a trade worth making for 3 % of a
813/// bill.
814///
815/// 60 s is where the curve has flattened but the exposure window is still
816/// something an operator can say out loud.
817///
818/// ## What it did not fix, and what did
819///
820/// The 2.03 frames/write floor was untouchable from here — it was one object
821/// per WAL frame, and cadence cannot amortize a per-frame cost. **R761-F2
822/// removed it** by uploading a tail call's frames as one ranged object, so
823/// the curve above is now
824///
825/// ```text
826/// PUTs/write = 3 / writes_per_tail
827/// ```
828///
829/// (one batch object + manifest + watermark), i.e. 3.00 / 0.60 / 0.12 / 0.03
830/// at the same four points — a 26 % cut at `writes_per_tail = 1` and 98.5 %
831/// at 100. **That does not move this default**, and the reasoning above is
832/// why rather than an accident: the 60 s choice was made against the shape
833/// of the curve, and what batching changes is its scale. Over the same
834/// 0.1–1 write/s burst band, going 60 s → 300 s now buys 0.50 → 0.10
835/// PUTs/write at the slow end and 0.05 → 0.01 at the fast one — against a
836/// pre-batching worst case of 3.36, an absolute difference small enough
837/// that RPO exposure is the only term still worth optimizing here. The
838/// measured points above are kept as the pre-batching baseline the
839/// reduction is stated against.
840///
841/// Re-run the harness and revisit both numbers whenever schema shape, page
842/// size, or turso's WAL behaviour changes:
843/// `cargo run -p turso-backup --example tail_sweep_harness`.
844///
845/// This crate still schedules nothing itself (R574-T4's "document the
846/// contract" choice stands) — the constant is the number a caller's
847/// scheduler should adopt absent a reason not to, and
848/// [`DEFAULT_RPO_TARGET`] is the bound that goes with it.
849pub const DEFAULT_TAIL_INTERVAL: Duration = Duration::from_secs(60);
850
851/// R761-T1: the RPO bound that goes with [`DEFAULT_TAIL_INTERVAL`] — twice
852/// it, because tailing *at* the bound makes ordinary scheduling jitter read
853/// as a breach, and tailing at half of it means one missed tick still lands
854/// inside the promise. See [`DEFAULT_TAIL_INTERVAL`] for why the cadence is
855/// 60 s; this is that number expressed as the promise rather than the
856/// mechanism, and it is what belongs in [`StreamConfig::rpo_target`].
857pub const DEFAULT_RPO_TARGET: Duration = Duration::from_secs(120);
858
859/// Configuration for a streaming session.
860pub struct StreamConfig<'a> {
861 /// Object-store key of the base tier-1a snapshot the frames replay onto.
862 /// Recorded in every generation manifest; restore re-fetches it.
863 pub base_snapshot_key: &'a str,
864 /// Page size of the base snapshot — read it from the snapshot header
865 /// (offset 16, `u16` big-endian; the on-disk value `1` means 65 536).
866 /// Required because the seam does not return it and `sync_server.rs`'s
867 /// 4 KB hardcode is the wrong default to inherit.
868 pub page_size: usize,
869 /// R574-F2: bounded spill buffer + overflow policy + 429/503 backoff
870 /// for the R2 upload side of the drain loop. `Default` is
871 /// behavior-preserving (`BackpressurePolicy::Fail` — errors bubble
872 /// immediately, same as before this field existed).
873 pub backpressure: BackpressureConfig,
874 /// R574-T4: the stated RPO bound — the caller's own scheduler is
875 /// expected to invoke `tail_frames` often enough that the watermark
876 /// never goes stale past this. `None` (the default) means no target is
877 /// asserted; [`RpoStatus::breached`] is always `false` in that case,
878 /// but [`RpoStatus::watermark_age`] is still reported so a caller can
879 /// observe the real gap before picking a number.
880 ///
881 /// R761-T1: [`DEFAULT_RPO_TARGET`] (120 s, tailed at
882 /// [`DEFAULT_TAIL_INTERVAL`] = 60 s) is the number to put here absent a
883 /// reason to pick another — its doc carries the measured write-op table
884 /// that justifies it. `None` stays the default so that a caller who has
885 /// not thought about RPO never gets a breach flag it did not ask for.
886 pub rpo_target: Option<Duration>,
887 /// R732-F2 (W245): this writer's **fencing token** — the per-tenant epoch
888 /// handed out by yubaba's raft state machine
889 /// (`YubabaState::tenant_fencing_token`). It is stamped into every frame
890 /// key, every generation manifest, and the watermark sidecar, and it is
891 /// checked before a single frame is uploaded: a writer whose epoch is
892 /// *lower* than the one already recorded at the sink is a stale owner and
893 /// bounces with [`StreamOutcome::Fenced`] rather than interleaving its
894 /// frames with the real owner's.
895 ///
896 /// **`0` means unfenced**, and is the behaviour-preserving default for a
897 /// single-writer deployment that has no ownership authority to ask. Note
898 /// that unfenced is not exempt: a `0` writer is still fenced by any sink
899 /// already stamped with a real epoch, which is exactly what should happen
900 /// when a tenant has been claimed and a legacy streamer is still running.
901 pub epoch: u64,
902 /// R732-F2: opaque label for *who* holds `epoch` — a yubaba node id, a
903 /// hostname, whatever the caller finds useful. Recorded in the manifest
904 /// and never interpreted here. Purely diagnostic: when you are staring at
905 /// a fenced stream at 3am, "which owner wrote generation 7" is the first
906 /// question, and the epoch alone does not answer it.
907 pub owner: Option<&'a str>,
908 /// R736-T2 (W250): the **cross-cell** fence — this writer's belief about
909 /// the global tenant→cell pointer's generation
910 /// (`yah_tenant_pointer::PointerRecord::generation`). `epoch` alone
911 /// fences *within* one raft group; it says nothing when ownership moves
912 /// to a different cell, because the two cells run independent raft
913 /// groups that share no epoch counter. Checked alongside `epoch` before
914 /// any frame is uploaded: a writer whose generation is *lower* than the
915 /// one already recorded at the sink is a stale cell and bounces with
916 /// [`StreamOutcome::Fenced`], exactly like a stale epoch.
917 ///
918 /// **`0` means unfenced** — the same behaviour-preserving default as
919 /// `epoch`, for a single-cell deployment that has no global pointer to
920 /// ask. This mirrors `yah_tenant_pointer`'s `FIRST_GENERATION = 1`: `0`
921 /// is what an unfenced writer defaults to and must never be mistakable
922 /// for a live cross-cell owner. The caller (yubaba's control plane)
923 /// resolves the pointer and hands the generation in as a plain `u64` —
924 /// this crate does not link `yah_tenant_pointer` to read it itself.
925 pub pointer_generation: u64,
926}
927
928/// What a [`tail_frames`] call did.
929#[derive(Debug, Clone, PartialEq, Eq)]
930pub enum StreamOutcome {
931 /// No new frames since the last tail — sink is already current.
932 Empty {
933 watermark: Watermark,
934 /// R574-T4: staleness of the last durably-persisted watermark —
935 /// still meaningful here, since "nothing new to stream" and "the
936 /// caller's scheduler stopped invoking us" look identical from the
937 /// engine's side and only this field tells them apart.
938 rpo: RpoStatus,
939 },
940 /// Uploaded a contiguous range of frames and wrote a generation manifest.
941 Streamed {
942 generation_key: String,
943 first_frame: u64,
944 last_frame: u64,
945 checkpoint_seq: u32,
946 frame_count: u64,
947 /// R574-F2: backpressure activity during this call (policy,
948 /// high-water spill-buffer occupancy, shed count, throttle retries).
949 backpressure: BackpressureReport,
950 /// R574-T4: see [`StreamOutcome::Empty::rpo`].
951 rpo: RpoStatus,
952 },
953 /// The live WAL is not provably the one the sidecar watermark was taken
954 /// from, so this call re-uploaded frames `1..N` from the top rather than
955 /// resuming.
956 ///
957 /// R858-B19 widened this from "`checkpoint_seq` advanced" to "the
958 /// [`WalGeneration`] did not prove itself unchanged", which is why both
959 /// fields are now generations rather than bare sequence numbers: the
960 /// motivating case is a writer restart where the sequence reads `0 -> 0`
961 /// and only the salt moved. `previous_generation.salt == None` names the
962 /// third case — a sidecar written before the salt existed, restarted
963 /// because it cannot be checked, not because it was seen to change.
964 Restarted {
965 generation_key: String,
966 previous_generation: WalGeneration,
967 new_generation: WalGeneration,
968 first_frame: u64,
969 last_frame: u64,
970 frame_count: u64,
971 /// R574-F2: see [`StreamOutcome::Streamed::backpressure`].
972 backpressure: BackpressureReport,
973 /// R574-T4: see [`StreamOutcome::Empty::rpo`].
974 rpo: RpoStatus,
975 },
976 /// R574-F2: `BackpressurePolicy::Shed` dropped every buffered frame in
977 /// this call before any of them persisted (R2 was throttling harder
978 /// than the spill buffer + backoff could absorb). No manifest/watermark
979 /// was written — the next `tail_frames` call re-attempts the same
980 /// range from the unchanged prior watermark. Distinct from `Empty`,
981 /// which means the engine itself had nothing new.
982 Shed {
983 checkpoint_seq: u32,
984 first_frame: u64,
985 last_frame: u64,
986 backpressure: BackpressureReport,
987 /// R574-T4: see [`StreamOutcome::Empty::rpo`].
988 rpo: RpoStatus,
989 },
990 /// R732-F2 (W245) / R736-T2 (W250): **this writer is a stale owner and
991 /// wrote nothing.** Either the sink's watermark is stamped with an epoch
992 /// higher than [`StreamConfig::epoch`] (ownership moved within the cell),
993 /// or with a pointer generation higher than
994 /// [`StreamConfig::pointer_generation`] (ownership moved to a different
995 /// cell) — the two-level fence bounces on either. Detected before the
996 /// first frame upload, so a fenced call is a pure read — no frames, no
997 /// manifest, no watermark write.
998 ///
999 /// This is the outcome the whole fencing design exists to produce. Without
1000 /// it a partitioned old master and a freshly-promoted new master both
1001 /// stream into the same prefix and silently corrupt each other; with it
1002 /// the loser finds out on its very next tail and can stop.
1003 ///
1004 /// Deliberately carries no [`RpoStatus`]: a fenced writer's view of
1005 /// watermark staleness is not its stream's RPO any more, and reporting one
1006 /// here would page the wrong operator about the wrong node.
1007 Fenced {
1008 /// The epoch recorded at the sink — the real owner's token.
1009 current_epoch: u64,
1010 /// The (lower) epoch this writer tried to stream under.
1011 our_epoch: u64,
1012 /// R736-T2: the pointer generation recorded at the sink.
1013 current_pointer_generation: u64,
1014 /// R736-T2: the (lower) generation this writer tried to stream under.
1015 our_pointer_generation: u64,
1016 },
1017}
1018
1019impl StreamOutcome {
1020 /// This call's RPO snapshot, or `None` for [`StreamOutcome::Fenced`] —
1021 /// see that variant's doc for why it deliberately carries none.
1022 ///
1023 /// R782: the accessor a caller (`tenant-streamer`'s tail loop) uses to
1024 /// push `watermark_age` onward without re-deriving this match on every
1025 /// call site that needs it.
1026 pub fn rpo(&self) -> Option<&RpoStatus> {
1027 match self {
1028 StreamOutcome::Empty { rpo, .. }
1029 | StreamOutcome::Streamed { rpo, .. }
1030 | StreamOutcome::Restarted { rpo, .. }
1031 | StreamOutcome::Shed { rpo, .. } => Some(rpo),
1032 StreamOutcome::Fenced { .. } => None,
1033 }
1034 }
1035
1036 /// This call's backpressure activity, or `None` for the two outcomes that
1037 /// never reached the drain loop ([`StreamOutcome::Empty`] had nothing to
1038 /// send, [`StreamOutcome::Fenced`] was refused before the first frame).
1039 ///
1040 /// R760-B8: the accessor a *multi-tenant* caller needs. `Streamed` is not
1041 /// the same thing as "the sink took everything" — under
1042 /// [`BackpressurePolicy::Shed`] a call that persisted a partial prefix and
1043 /// dropped the rest reports `Streamed` with a nonzero
1044 /// [`BackpressureReport::frames_shed`], and a caller that only matches the
1045 /// variant cannot tell that apart from a clean tail. roadcase's shard
1046 /// flusher reads this on every arm so a struggling cell is nameable from
1047 /// its own metrics rather than by bisecting tenants.
1048 pub fn backpressure(&self) -> Option<&BackpressureReport> {
1049 match self {
1050 StreamOutcome::Streamed { backpressure, .. }
1051 | StreamOutcome::Restarted { backpressure, .. }
1052 | StreamOutcome::Shed { backpressure, .. } => Some(backpressure),
1053 StreamOutcome::Empty { .. } | StreamOutcome::Fenced { .. } => None,
1054 }
1055 }
1056}
1057
1058/// R574-T4: watermark-staleness snapshot for RPO drift alerting, computed
1059/// fresh on every `tail_frames` call against [`StreamConfig::rpo_target`].
1060#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
1061pub struct RpoStatus {
1062 /// The stated bound from `StreamConfig::rpo_target`, echoed back for
1063 /// convenience (so a caller reading only the outcome still knows what
1064 /// was being enforced).
1065 pub target: Option<Duration>,
1066 /// Elapsed time since the watermark last durably advanced, measured at
1067 /// the start of this call. `None` when the sink has never persisted a
1068 /// watermark — there is nothing yet to measure staleness against.
1069 pub watermark_age: Option<Duration>,
1070 /// `true` when both `target` and `watermark_age` are set and
1071 /// `watermark_age > target` — the RPO bound is breached and an
1072 /// orchestrator should alert. Always `false` when no target is
1073 /// configured.
1074 pub breached: bool,
1075}
1076
1077impl BackupTarget {
1078 pub(crate) fn watermark_key(&self) -> ObjPath {
1079 join_key(&self.prefix, "latest.stream-watermark")
1080 }
1081
1082 /// R732-F2: frames are namespaced by the writer's fencing epoch, so two
1083 /// owners at different epochs cannot land on the same key even if they
1084 /// somehow both get as far as uploading. The epoch check in
1085 /// [`tail_frames`] is the guard; this key shape is the backstop that makes
1086 /// a guard failure recoverable (both owners' frames survive and the
1087 /// manifests say who wrote what) instead of a silent overwrite.
1088 ///
1089 /// Epoch `0` keeps the original two-level layout. That is not cosmetic:
1090 /// backups written before fencing existed are still restorable because
1091 /// their manifests say `epoch 0` and land back on this branch.
1092 ///
1093 /// R761-F2: this is the *read-side legacy* key shape now — nothing writes
1094 /// one-object-per-frame any more. See [`Self::frame_batch_key`].
1095 pub(crate) fn frame_key(&self, epoch: u64, checkpoint_seq: u32, frame_no: u64) -> ObjPath {
1096 let suffix = if epoch == 0 {
1097 format!("frames/{checkpoint_seq:010}/{frame_no:020}")
1098 } else {
1099 format!("frames/{epoch:020}/{checkpoint_seq:010}/{frame_no:020}")
1100 };
1101 join_key(&self.prefix, &suffix)
1102 }
1103
1104 /// R761-F2: key of a batch object holding frames `first..=last`
1105 /// concatenated. Same epoch namespacing and zero-padding as
1106 /// [`Self::frame_key`], so lexical order still matches frame order — one
1107 /// stream's batches never overlap, because each is a slice of a single
1108 /// monotonic drain.
1109 ///
1110 /// The range is in the key rather than only in the manifest so the sink
1111 /// stays self-describing: a `ls` of the prefix tells an operator exactly
1112 /// which frames are present, which is what made the per-frame layout easy
1113 /// to reason about and is worth keeping.
1114 ///
1115 /// Distinguishable from a legacy per-frame key by construction (that one
1116 /// has no `-`), so a stream that straddles the layout change can hold both
1117 /// shapes under the same directory without collision.
1118 pub(crate) fn frame_batch_key(
1119 &self,
1120 epoch: u64,
1121 checkpoint_seq: u32,
1122 first: u64,
1123 last: u64,
1124 ) -> ObjPath {
1125 let suffix = if epoch == 0 {
1126 format!("frames/{checkpoint_seq:010}/{first:020}-{last:020}")
1127 } else {
1128 format!("frames/{epoch:020}/{checkpoint_seq:010}/{first:020}-{last:020}")
1129 };
1130 join_key(&self.prefix, &suffix)
1131 }
1132
1133 /// Every object holding generation `m`'s frames, as `(key, first, last)`
1134 /// in ascending frame order.
1135 ///
1136 /// The single place that knows how a generation's frames are laid out:
1137 /// R761-F2 batches when the manifest carries a `frame_batch` list, the
1138 /// pre-R761-F2 one-object-per-frame layout when it does not. Restore's
1139 /// replay and [`crate::puller::WalPuller`] both go through here, so the
1140 /// two can never drift apart about where a frame lives.
1141 pub(crate) fn frame_objects_of(
1142 &self,
1143 m: &OwnedGenerationManifest,
1144 ) -> Vec<(ObjPath, u64, u64)> {
1145 if m.frame_batches.is_empty() {
1146 (m.first_frame..=m.last_frame)
1147 .map(|n| (self.frame_key(m.epoch, m.checkpoint_seq, n), n, n))
1148 .collect()
1149 } else {
1150 m.frame_batches
1151 .iter()
1152 .map(|&(first, last)| {
1153 (
1154 self.frame_batch_key(m.epoch, m.checkpoint_seq, first, last),
1155 first,
1156 last,
1157 )
1158 })
1159 .collect()
1160 }
1161 }
1162
1163 pub(crate) fn generation_key(&self, unix_nanos: u128) -> ObjPath {
1164 join_key(
1165 &self.prefix,
1166 &format!("generations/gen-{unix_nanos:020}.manifest"),
1167 )
1168 }
1169}
1170
1171/// R858-B19 — identify the WAL generation `seam` is currently reading, by
1172/// pulling the salt out of **frame 1's** header.
1173///
1174/// Frame 1 rather than the 32-byte WAL header because this goes through the
1175/// [`WalSeam`] trait, which is the crate's one boundary against `turso_core`:
1176/// `turso_core::WalState` reports only `checkpoint_seq_no` and `max_frame`, and
1177/// reading the `-wal` file directly would need a path that
1178/// [`CoreWalSeam::from_conn`] does not have. Every frame header carries a
1179/// verbatim copy of the WAL header's salt, so frame 1 answers the same question
1180/// through machinery every seam already implements — including the mocks, and
1181/// including roadcase's connection-backed seam.
1182///
1183/// `Ok(None)` means the WAL holds no frames, so there is no generation to name
1184/// yet. That is not a failure: it is the honest "unknown", and
1185/// [`WalGeneration::is_provably_same_as`] treats it as such.
1186///
1187/// Costs one page-sized read per [`tail_frames`] call — negligible next to the
1188/// frames that call is about to upload, and it is the only thing standing
1189/// between a WAL recreate and a silently spliced generation chain.
1190pub(crate) fn read_wal_salt<S: WalSeam>(
1191 seam: &S,
1192 page_size: usize,
1193 last_frame: u64,
1194) -> Result<Option<WalSalt>> {
1195 if last_frame == 0 {
1196 return Ok(None);
1197 }
1198 let mut buf = vec![0u8; WAL_FRAME_HEADER_SIZE + page_size];
1199 seam.wal_get_frame(1, &mut buf)
1200 .context("reading WAL frame 1 to identify the WAL generation (R858-B19)")?;
1201 Ok(Some(WalSalt::from_frame_header(&buf)))
1202}
1203
1204/// Tail new WAL frames from `seam` into `target`, anchored to a base
1205/// snapshot. Idempotent and resumable: on the second call only frames after
1206/// the recorded watermark are uploaded.
1207///
1208/// Ordering within a single call:
1209/// 1. Read [`Watermark`] from the seam (snapshot the current `(checkpoint_seq,
1210/// max_frame)`) and the current [`WalGeneration`] (that sequence plus the
1211/// WAL salt, via [`read_wal_salt`]).
1212/// 2. Read the prior watermark sidecar (if any).
1213/// 3. Unless the sidecar's generation is *provably* the same WAL as the live
1214/// one, treat this as a restart: upload frames `1..=max_frame`. R858-B19 —
1215/// the test is [`WalGeneration::is_provably_same_as`], not a `checkpoint_seq`
1216/// comparison, because a writer-process restart recreates the WAL at
1217/// sequence `0` and an unchanged sequence therefore proves nothing.
1218/// 4. Otherwise upload frames `prior.last_frame+1..=max_frame`.
1219/// 5. Advance the watermark sidecar with a compare-and-swap on the version
1220/// read in step 2, then write a generation manifest pointing at the base
1221/// snapshot + frame range + the batch objects covering it. Frames precede
1222/// both, so a manifest never references a missing frame; the manifest
1223/// follows the CAS, so a writer that loses the sidecar race publishes
1224/// nothing (R732-T3).
1225///
1226/// Returns [`StreamOutcome::Empty`] if there is nothing to do (max_frame
1227/// hasn't advanced and checkpoint_seq is unchanged), or
1228/// [`StreamOutcome::Fenced`] if this writer's [`StreamConfig::epoch`] has been
1229/// superseded — either observed up front in step 2 or discovered by losing the
1230/// CAS in step 5.
1231pub async fn tail_frames<S: WalSeam>(
1232 seam: &S,
1233 target: &BackupTarget,
1234 cfg: &StreamConfig<'_>,
1235) -> Result<StreamOutcome> {
1236 let current = seam.wal_state()?;
1237 // R858-B19: the salt is what actually names the WAL generation. Read it
1238 // before anything else touches the sink, so the restart decision below is
1239 // made against the live WAL rather than inferred from a sequence number
1240 // that a writer restart silently resets.
1241 let current_generation = WalGeneration {
1242 checkpoint_seq: current.checkpoint_seq,
1243 salt: read_wal_salt(seam, cfg.page_size, current.last_frame)?,
1244 };
1245 let persisted = read_watermark(&target.store, &target.watermark_key()).await?;
1246 // R574-T4: staleness measured at call start, against the *prior*
1247 // sidecar write — i.e. how long the sink had gone without durable
1248 // progress before this call ran. On the steady cadence the caller's
1249 // scheduler promises, this hovers at the invocation interval; a
1250 // skipped/late run pushes it past `rpo_target` and flips `breached`.
1251 let rpo = rpo_status(cfg.rpo_target, persisted.as_ref());
1252
1253 // R732-F2 (W245) / R736-T2 (W250): the two-level fence, checked before
1254 // anything is written. A sink stamped with a higher epoch than ours means
1255 // ownership moved within the cell while we were away; a sink stamped with
1256 // a higher pointer generation means ownership moved to a *different*
1257 // cell — the two raft groups don't share an epoch counter, so the epoch
1258 // check alone is blind to that move. Either means every byte we are about
1259 // to upload belongs to somebody else's stream. Bounce here and the call
1260 // is a pure read; bounce anywhere later and we have already interleaved
1261 // frames into the real owner's range.
1262 if let Some(p) = &persisted {
1263 if p.epoch > cfg.epoch || p.pointer_generation > cfg.pointer_generation {
1264 return Ok(StreamOutcome::Fenced {
1265 current_epoch: p.epoch,
1266 our_epoch: cfg.epoch,
1267 current_pointer_generation: p.pointer_generation,
1268 our_pointer_generation: cfg.pointer_generation,
1269 });
1270 }
1271 }
1272
1273 // Borrowed, not consumed: the conditional watermark advance at the end of
1274 // this call needs the object version this read observed (R732-T3).
1275 let prior = persisted.as_ref().map(|p| p.watermark);
1276
1277 // R858-B19: the generation the sidecar was written under. `None` for a
1278 // sidecar that predates the salt field — which reads as *unknown*, not as
1279 // "the same WAL", and so forces the restart branch exactly once. After that
1280 // one full re-upload the sink carries a salt and the stream self-heals.
1281 let prior_generation = persisted.as_ref().map(|p| p.generation);
1282
1283 // Decide the range to upload. Resuming at `last_frame + 1` is only sound
1284 // when the live WAL is PROVABLY the one the watermark was taken from: a
1285 // writer-process restart deletes the `-wal` file and the next writer starts
1286 // a fresh WAL at checkpoint-sequence 0 with a new salt, so an unchanged
1287 // sequence is not evidence of anything. Anything short of proof restarts.
1288 let (start_frame, restarted) = match prior_generation {
1289 Some(g) if g.is_provably_same_as(¤t_generation) => {
1290 (prior.map(|p| p.last_frame).unwrap_or(0) + 1, false)
1291 }
1292 Some(_) => (1, true),
1293 None => (1, false),
1294 };
1295
1296 if current.last_frame < start_frame {
1297 return Ok(StreamOutcome::Empty { watermark: current, rpo });
1298 }
1299
1300 // Drain frames in ascending order through the bounded spill buffer +
1301 // backpressure policy (R574-F2). A partial upload leaves a prefix (the
1302 // manifest is written last, so a prefix without a manifest is invisible
1303 // to restore — the next tail just overwrites the same keys).
1304 let drain = drain_frames_with_backpressure(
1305 seam,
1306 target,
1307 cfg,
1308 cfg.epoch,
1309 current.checkpoint_seq,
1310 start_frame,
1311 current.last_frame,
1312 )
1313 .await?;
1314
1315 let Some(uploaded_through) = drain.uploaded_through else {
1316 // BackpressurePolicy::Shed dropped everything before any frame
1317 // persisted — no manifest/watermark write, so the next call
1318 // re-attempts this exact range from the unchanged prior watermark.
1319 return Ok(StreamOutcome::Shed {
1320 checkpoint_seq: current.checkpoint_seq,
1321 first_frame: start_frame,
1322 last_frame: current.last_frame,
1323 backpressure: drain.report,
1324 rpo,
1325 });
1326 };
1327
1328 // R858-B19: everything above sampled the generation ONCE, before the
1329 // drain. A WAL recreate between calls is what this ticket is about, but
1330 // nothing stops one landing *during* a call — and then the frames just
1331 // uploaded are a mix of two WALs, which is the same corruption arriving by
1332 // a narrower door. Re-read the salt and refuse if it moved.
1333 //
1334 // Refusing here is cheap and complete: the watermark has not advanced and
1335 // no manifest has been written, so the frames that landed are orphaned
1336 // under keys nothing references — invisible to restore, exactly like the
1337 // `Shed` path — and the next tail re-derives everything from the sidecar.
1338 let after = seam.wal_state()?;
1339 let after_salt = read_wal_salt(seam, cfg.page_size, after.last_frame)?;
1340 if after_salt != current_generation.salt {
1341 anyhow::bail!(
1342 "the source WAL was recreated while this tail was uploading ({} -> {}) — the frames \
1343 this call read span two WAL generations, so nothing is published and the sink is left \
1344 exactly as it was; the next tail will restart cleanly against the new WAL",
1345 current_generation.describe(),
1346 WalGeneration { checkpoint_seq: after.checkpoint_seq, salt: after_salt }.describe(),
1347 );
1348 }
1349
1350 // Claim the range with a conditional watermark advance, THEN publish the
1351 // generation manifest. Under a Shed policy that persisted a partial
1352 // prefix, both cover only `start_frame..=uploaded_through`, not the full
1353 // engine range.
1354 //
1355 // R732-T3 reordered these two. The manifest used to be written last, so
1356 // that a manifest never referenced a missing frame — that invariant is
1357 // untouched, because frames still precede both. What the old order could
1358 // not do is fence: a writer that lost the sidecar race had already
1359 // published its manifest, which would then sit in the chain at a regressed
1360 // epoch and make every future restore refuse. Publishing only after the
1361 // CAS means a fenced writer leaves no manifest at all.
1362 let cas = write_watermark(
1363 &target.store,
1364 &target.watermark_key(),
1365 Watermark { checkpoint_seq: current.checkpoint_seq, last_frame: uploaded_through },
1366 current_generation.salt,
1367 cfg.epoch,
1368 cfg.pointer_generation,
1369 persisted.as_ref().and_then(|p| p.version.as_ref()),
1370 )
1371 .await?;
1372 if cas == WatermarkCas::Contended {
1373 // Somebody replaced the sidecar under us. Re-read to learn who.
1374 let now = read_watermark(&target.store, &target.watermark_key()).await?;
1375 let current_epoch = now.as_ref().map_or(0, |p| p.epoch);
1376 let current_pointer_generation = now.as_ref().map_or(0, |p| p.pointer_generation);
1377 if current_epoch > cfg.epoch || current_pointer_generation > cfg.pointer_generation {
1378 // A newer owner won the race. Our frames are orphaned under our
1379 // own epoch prefix with no manifest naming them, so restore never
1380 // sees them — the sink is exactly what the winner left.
1381 return Ok(StreamOutcome::Fenced {
1382 current_epoch,
1383 our_epoch: cfg.epoch,
1384 current_pointer_generation,
1385 our_pointer_generation: cfg.pointer_generation,
1386 });
1387 }
1388 // Same or lower epoch AND generation: neither fence can tell these two
1389 // writers apart. That means two processes are streaming the same
1390 // tenant under the SAME tokens — a caller bug (a duplicate streamer,
1391 // or an ownership token handed out twice), and exactly the
1392 // condition that must not be papered over with a retry.
1393 anyhow::bail!(
1394 "watermark CAS for {} lost to a concurrent writer at epoch {current_epoch} (pointer generation {current_pointer_generation}) while we hold epoch {} (pointer generation {}) — two streamers share one fencing token",
1395 target.watermark_key(),
1396 cfg.epoch,
1397 cfg.pointer_generation,
1398 );
1399 }
1400
1401 let nanos = unix_nanos();
1402 let gen_key = target.generation_key(nanos);
1403 let manifest = format_generation_manifest(GenerationManifest {
1404 base_snapshot_key: cfg.base_snapshot_key,
1405 page_size: cfg.page_size,
1406 checkpoint_seq: current.checkpoint_seq,
1407 // R858-B19: stamp the generation this range actually came from, so
1408 // restore can refuse a chain that spans two WALs instead of splicing
1409 // them. Always `Some` on this path — a `None` salt means an empty WAL,
1410 // and an empty WAL took the `Empty` return above.
1411 salt: current_generation.salt,
1412 first_frame: start_frame,
1413 last_frame: uploaded_through,
1414 epoch: cfg.epoch,
1415 owner: cfg.owner,
1416 // R761-F2: exactly the batch objects the drain persisted. Restore
1417 // derives its keys from this list, so it is not a summary of the range
1418 // — it IS the range's index.
1419 frame_batches: &drain.frame_batches,
1420 });
1421 target
1422 .store
1423 .put(&gen_key, manifest.into_bytes().into())
1424 .await
1425 .with_context(|| format!("writing generation manifest {gen_key}"))?;
1426
1427 let frame_count = uploaded_through - start_frame + 1;
1428 let gen_key = gen_key.to_string();
1429 if restarted {
1430 Ok(StreamOutcome::Restarted {
1431 generation_key: gen_key,
1432 previous_generation: prior_generation.unwrap_or_default(),
1433 new_generation: current_generation,
1434 first_frame: start_frame,
1435 last_frame: uploaded_through,
1436 frame_count,
1437 backpressure: drain.report,
1438 rpo,
1439 })
1440 } else {
1441 Ok(StreamOutcome::Streamed {
1442 generation_key: gen_key,
1443 first_frame: start_frame,
1444 last_frame: uploaded_through,
1445 checkpoint_seq: current.checkpoint_seq,
1446 frame_count,
1447 backpressure: drain.report,
1448 rpo,
1449 })
1450 }
1451}
1452
1453/// Outcome of [`drain_frames_with_backpressure`]: the last frame_no
1454/// successfully persisted this call (`None` if every buffered frame was
1455/// shed before any of them landed) plus the backpressure activity report.
1456struct DrainOutcome {
1457 uploaded_through: Option<u64>,
1458 /// R761-F2: the `(first, last)` range of every batch object that actually
1459 /// landed, ascending and gap-free from the call's `start_frame` through
1460 /// `uploaded_through` (a batch is popped only on a successful upload, so a
1461 /// partial drain truncates this list rather than holing it). Copied into
1462 /// the generation manifest, which is what tells restore where to look.
1463 frame_batches: Vec<(u64, u64)>,
1464 report: BackpressureReport,
1465}
1466
1467/// Read WAL frames `start_frame..=last_frame` into a bounded spill buffer
1468/// (`cfg.backpressure.spill_buffer_frames`) and drain them to `target`'s
1469/// object store through [`put_with_backoff`] (R574-F2), one **batch object**
1470/// per buffer-full (R761-F2).
1471///
1472/// The buffer only refills up to its bound, so once full the loop must
1473/// resolve the buffered batch — either by a successful upload or by the
1474/// configured [`BackpressurePolicy`] deciding what "stuck" means:
1475/// `Block` retries the batch forever on a throttling error (no data loss,
1476/// but the call can run long); `Fail` bubbles the error once
1477/// `backoff.max_retries` is exhausted (nothing persists past what already
1478/// landed); `Shed` drops the whole buffered backlog (loudly, via the
1479/// returned report) and returns whatever prefix already persisted. Any
1480/// non-throttling error bubbles immediately regardless of policy — this
1481/// mechanism is specifically for R2 throttling, not general fault
1482/// tolerance.
1483///
1484/// R761-F2 made the upload unit the buffer's contents rather than its head
1485/// frame, which is why `spill_buffer_frames` also sizes the largest object
1486/// this sink will write: `spill_buffer_frames * (24 + page_size)` bytes, ~1 MB
1487/// at the 256-frame default and a 4 KB page. That coupling is deliberate — the
1488/// bound already promises a memory ceiling, and a second knob for batch size
1489/// would only ever be set to some fraction of it.
1490async fn drain_frames_with_backpressure<S: WalSeam>(
1491 seam: &S,
1492 target: &BackupTarget,
1493 cfg: &StreamConfig<'_>,
1494 epoch: u64,
1495 checkpoint_seq: u32,
1496 start_frame: u64,
1497 last_frame: u64,
1498) -> Result<DrainOutcome> {
1499 let bp = &cfg.backpressure;
1500 let frame_size = WAL_FRAME_HEADER_SIZE + cfg.page_size;
1501 let bound = bp.spill_buffer_frames.max(1);
1502
1503 let mut report = BackpressureReport { policy: bp.policy, ..Default::default() };
1504 let mut pending: VecDeque<(u64, Vec<u8>)> = VecDeque::new();
1505 let mut uploaded_through: Option<u64> = None;
1506 let mut frame_batches: Vec<(u64, u64)> = Vec::new();
1507 let mut next_to_read = start_frame;
1508 let mut buf = vec![0u8; frame_size];
1509
1510 loop {
1511 // Refill up to the bound while there's more WAL to read. Once full,
1512 // the buffered batch must be resolved before we accept more.
1513 while pending.len() < bound && next_to_read <= last_frame {
1514 seam.wal_get_frame(next_to_read, &mut buf)
1515 .with_context(|| format!("reading wal frame {next_to_read}"))?;
1516 pending.push_back((next_to_read, buf.clone()));
1517 next_to_read += 1;
1518 report.high_water_frames = report.high_water_frames.max(pending.len());
1519 }
1520 let Some(&(first, _)) = pending.front() else {
1521 break; // Fully drained: nothing buffered, nothing left to read.
1522 };
1523 // Non-empty (we just matched `front`), and the buffer is filled in
1524 // ascending order without gaps, so the back frame closes the range.
1525 let last = pending.back().map_or(first, |&(n, _)| n);
1526 let mut body = Vec::with_capacity(pending.len() * frame_size);
1527 for (_, bytes) in &pending {
1528 body.extend_from_slice(bytes);
1529 }
1530 let key = target.frame_batch_key(epoch, checkpoint_seq, first, last);
1531 match put_with_backoff(&target.store, &key, body.into(), bp, &mut report.throttle_retries)
1532 .await
1533 {
1534 Ok(_) => {
1535 pending.clear();
1536 uploaded_through = Some(last);
1537 frame_batches.push((first, last));
1538 }
1539 Err(e) => match bp.policy {
1540 BackpressurePolicy::Shed => {
1541 report.frames_shed += pending.len() as u64;
1542 return Ok(DrainOutcome { uploaded_through, frame_batches, report });
1543 }
1544 BackpressurePolicy::Block | BackpressurePolicy::Fail => {
1545 return Err(e).with_context(|| {
1546 format!("uploading wal frames {first}-{last} to {key}")
1547 });
1548 }
1549 },
1550 }
1551 }
1552 Ok(DrainOutcome { uploaded_through, frame_batches, report })
1553}
1554
1555/// R858-B18 — how many times [`raw_consistent_copy_live`] will re-take a copy
1556/// whose [`SourceFingerprint`] moved underneath it before giving up loudly.
1557///
1558/// Bounded on purpose. An unbounded retry against a database under sustained
1559/// write pressure is an infinite loop that looks like a hang; a caller that
1560/// wants to keep trying should be the one deciding how long to keep trying, on
1561/// its own schedule. Four attempts is enough to ride out an isolated
1562/// checkpoint (headscale's database measures 94 KB, so an attempt is
1563/// sub-millisecond) and few enough that a genuinely hot database is reported as
1564/// hot within a few milliseconds instead of being ground at.
1565pub const COPY_VALIDATION_ATTEMPTS: u32 = 4;
1566
1567/// Take a raw, point-in-time-consistent byte image of the database at `db_path`
1568/// WITHOUT folding the WAL via a `TRUNCATE` checkpoint, and WITHOUT locking out
1569/// the process that owns it. The live-writer pair of
1570/// [`crate::dedup::raw_consistent_copy`], which a concurrent writer can make
1571/// return `busy`.
1572///
1573/// Algorithm — the read-only "copy main + replay WAL frames ourselves" path
1574/// flagged in the working doc and the R005-T1 spike, wrapped in R858-B18's
1575/// optimistic validation:
1576///
1577/// 1. Sample the source's [`SourceFingerprint`] — WAL salt + checkpoint
1578/// sequence, WAL length, main length, main change counter — straight off
1579/// disk, with no engine open.
1580/// 2. Open a fresh `turso_core::Connection` via [`CoreWalSeam::open_reader`],
1581/// which takes **no** whole-file lock (so the application keeps working) and
1582/// calls `wal_auto_actions_disable` so our seam can't auto-checkpoint or
1583/// restart the WAL header mid-read on our connection.
1584/// 3. Read the main DB file bytes from disk. With auto-actions disabled on our
1585/// connection the file cannot be folded by us; a concurrent writer in a
1586/// separate connection only ever extends the WAL (the main file is only
1587/// written by a checkpoint).
1588/// 4. Walk WAL frames `1..=max_frame`, replaying each page into the in-memory
1589/// image at offset `(page_no - 1) * page_size`. Track the last
1590/// `is_commit_frame` and the corresponding `db_size`. Any frames past the
1591/// last commit are uncommitted mid-transaction garbage — drop them
1592/// (crash-consistency, matching restore's `wal_insert_end(false)`).
1593/// 5. Truncate / grow the image to `db_size * page_size`.
1594/// 6. Re-sample the fingerprint. **Accept the image only if
1595/// [`SourceFingerprint::stable_across`] holds** — i.e. only if nothing that
1596/// can tear the copy moved. A foreign checkpoint inside our read window would
1597/// splice pre- and post-fold state; discard that image and retry from step 1,
1598/// up to [`COPY_VALIDATION_ATTEMPTS`] times, then fail. (A plain WAL append
1599/// is *not* movement for this purpose, and that distinction is what keeps
1600/// this usable against a database that is actually in use — see
1601/// `stable_across`.)
1602///
1603/// Step 6 is the entire correctness story against a foreign engine, and it is
1604/// why this function never returns an unvalidated image: a torn WAL replay
1605/// still produces a structurally valid SQLite file that passes
1606/// `PRAGMA integrity_check`, so "it parsed" proves nothing. See
1607/// [`SourceFingerprint`] for why validation rather than SQLite's real
1608/// `-shm` reader protocol.
1609///
1610/// Returned bytes are a self-contained vanilla-SQLite image (page-offset
1611/// stable, no `-wal` sidecar required), ready to feed
1612/// [`crate::dedup::snapshot_dedup`]'s content-addressed chunking under a
1613/// concurrent writer.
1614pub async fn raw_consistent_copy_live(db_path: &str, page_size: usize) -> Result<Vec<u8>> {
1615 let image = validated_against_source(db_path, "live copy", || async {
1616 let seam = CoreWalSeam::open_reader(db_path)
1617 .with_context(|| format!("opening read-only WAL seam on {db_path}"))?;
1618 let main_bytes = std::fs::read(db_path)
1619 .with_context(|| format!("reading main db file {db_path}"))?;
1620 replay_wal_onto_main(&seam, main_bytes, page_size)
1621 // Seam dropped at the end of this block, so nothing of ours holds the
1622 // file while the post-copy fingerprint is sampled.
1623 })
1624 .await?;
1625 anyhow::ensure!(
1626 image.starts_with(b"SQLite format 3\0"),
1627 "live consistent copy of {db_path} is not a SQLite database"
1628 );
1629 Ok(image)
1630}
1631
1632/// R858-B18 — run `take` against the live database at `db_path` and return its
1633/// result **only** if the source provably held still for the duration.
1634///
1635/// This is the optimistic-validation protocol both source-side tiers share:
1636/// tier 2's WAL-replay copy ([`raw_consistent_copy_live`]) and tier 1a's
1637/// `VACUUM INTO` (`crate::snapshot`). Both read a database a foreign engine may
1638/// be checkpointing underneath them, neither can take a lock that engine
1639/// respects, and both are therefore only correct if a fold inside their read
1640/// window is *detected*. `what` names the operation in the refusal message.
1641///
1642/// Each attempt samples a [`SourceFingerprint`] before and after, then
1643/// classifies the outcome on both axes — did it succeed, and did the source
1644/// move:
1645///
1646/// | | source held still | source moved |
1647/// |---|---|---|
1648/// | **`Ok`** | accept | discard, retry (it may splice pre-/post-fold state) |
1649/// | **`Err`** | return the error — nothing raced us, so it is real | retry (a symptom of the race) |
1650///
1651/// That bottom-right cell is why the fingerprint is load-bearing beyond
1652/// accept/reject: R858-B18's probe H measured the *common* shape of a caught
1653/// race not as a clean torn image but as `short read on WAL frame` — a foreign
1654/// checkpoint truncating the WAL while turso walked it. Without the fingerprint
1655/// there is no way to tell that transient apart from a genuinely corrupt
1656/// database, and the two want opposite handling.
1657///
1658/// After [`COPY_VALIDATION_ATTEMPTS`] attempts all of which saw movement, this
1659/// fails loudly and returns nothing.
1660pub(crate) async fn validated_against_source<T, F, Fut>(
1661 db_path: &str,
1662 what: &str,
1663 mut take: F,
1664) -> Result<T>
1665where
1666 F: FnMut() -> Fut,
1667 Fut: std::future::Future<Output = Result<T>>,
1668{
1669 let mut moved: Option<(SourceFingerprint, SourceFingerprint, Option<anyhow::Error>)> = None;
1670 for attempt in 1..=COPY_VALIDATION_ATTEMPTS {
1671 if attempt > 1 {
1672 // Back off between attempts: retrying instantly against a writer
1673 // mid-checkpoint just spends the whole budget inside one fold.
1674 tokio::time::sleep(std::time::Duration::from_millis(20 * u64::from(attempt - 1)))
1675 .await;
1676 }
1677 let before = SourceFingerprint::read(db_path)?;
1678 let attempted = take().await;
1679 let after = SourceFingerprint::read(db_path)?;
1680 let stable = before.stable_across(&after);
1681 match attempted {
1682 Ok(value) if stable => return Ok(value),
1683 Ok(_) => moved = Some((before, after, None)),
1684 Err(e) if !stable => moved = Some((before, after, Some(e))),
1685 Err(e) => {
1686 return Err(e).with_context(|| {
1687 format!(
1688 "{what} of {db_path} failed against a source that did NOT move during \
1689 the attempt ({}), so this is not a concurrent-writer race",
1690 before.describe()
1691 )
1692 })
1693 }
1694 }
1695 }
1696 let (before, after, last_err) = moved.expect("COPY_VALIDATION_ATTEMPTS is non-zero");
1697 let because = match last_err {
1698 Some(e) => format!("the last attempt also failed mid-read ({e:#})"),
1699 None => "each attempt produced a result that could not be validated".to_string(),
1700 };
1701 anyhow::bail!(
1702 "refusing a {what} of {db_path}: the source moved under every one of \
1703 {COPY_VALIDATION_ATTEMPTS} attempts, so nothing could be validated as \
1704 point-in-time — {because}. The last attempt saw [{}] before and [{}] after, so a \
1705 foreign writer or checkpointer is active. Retry on your own schedule; NO result \
1706 is returned, because a torn read of a SQLite database still parses as valid SQLite.",
1707 before.describe(),
1708 after.describe(),
1709 )
1710}
1711
1712/// Pure replay of every committed WAL frame visible through `seam` onto
1713/// `main_bytes`. Split out from [`raw_consistent_copy_live`] so it can be
1714/// driven by a mock seam in unit tests; the live entry point layers disk I/O
1715/// and magic-byte validation on top.
1716///
1717/// Contract: the highest `is_commit_frame` in `1..=wal_state.last_frame`
1718/// defines both the post-replay page count and the cutoff for which frames
1719/// are applied. If no frame in that window is a commit, `main_bytes` is
1720/// returned unchanged.
1721pub(crate) fn replay_wal_onto_main<S: WalSeam>(
1722 seam: &S,
1723 mut main_bytes: Vec<u8>,
1724 page_size: usize,
1725) -> Result<Vec<u8>> {
1726 anyhow::ensure!(page_size > 0, "page_size must be non-zero");
1727 let watermark = seam.wal_state()?;
1728
1729 let frame_size = WAL_FRAME_HEADER_SIZE + page_size;
1730 let mut buf = vec![0u8; frame_size];
1731
1732 struct PendingFrame {
1733 page_no: u32,
1734 db_size: u32,
1735 page_bytes: Vec<u8>,
1736 }
1737 let mut frames: Vec<PendingFrame> = Vec::new();
1738 let mut last_commit_idx: Option<usize> = None;
1739 for frame_no in 1..=watermark.last_frame {
1740 let info = seam
1741 .wal_get_frame(frame_no, &mut buf)
1742 .with_context(|| format!("reading WAL frame {frame_no}"))?;
1743 anyhow::ensure!(
1744 info.page_no >= 1,
1745 "WAL frame {frame_no}: page_no must be >= 1"
1746 );
1747 frames.push(PendingFrame {
1748 page_no: info.page_no,
1749 db_size: info.db_size,
1750 page_bytes: buf[WAL_FRAME_HEADER_SIZE..].to_vec(),
1751 });
1752 if info.is_commit_frame() {
1753 last_commit_idx = Some(frames.len() - 1);
1754 }
1755 }
1756
1757 let Some(last_commit) = last_commit_idx else {
1758 // No committed frames in our view — main file alone is the image.
1759 // Any uncommitted suffix in the WAL is dropped by construction.
1760 return Ok(main_bytes);
1761 };
1762 let final_db_size = frames[last_commit].db_size as usize;
1763 let target_size = final_db_size
1764 .checked_mul(page_size)
1765 .context("db_size * page_size overflow")?;
1766
1767 // Grow image to hold the final committed image AND any page slot we'll
1768 // touch in the commit prefix (a frame may write a page above db_size
1769 // mid-grow; the final truncate cuts that back to db_size).
1770 let max_off_needed: usize = frames[..=last_commit]
1771 .iter()
1772 .map(|f| f.page_no as usize * page_size)
1773 .max()
1774 .unwrap_or(0);
1775 let need = target_size.max(max_off_needed);
1776 if main_bytes.len() < need {
1777 main_bytes.resize(need, 0);
1778 }
1779 for f in &frames[..=last_commit] {
1780 let off = (f.page_no as usize - 1) * page_size;
1781 main_bytes[off..off + page_size].copy_from_slice(&f.page_bytes);
1782 }
1783 main_bytes.truncate(target_size);
1784
1785 Ok(main_bytes)
1786}
1787
1788/// Summary of a [`restore_latest_stream`] call.
1789#[derive(Debug, Clone, PartialEq, Eq)]
1790pub struct RestoreOutcome {
1791 /// The tier-1a snapshot key every generation manifest referenced (must agree).
1792 pub base_snapshot_key: String,
1793 /// Sole `checkpoint_seq_no` across the replayed manifests (v1 refuses to
1794 /// span a WAL restart — see [`validate_generation_chain`]).
1795 pub checkpoint_seq: u32,
1796 /// Number of generation manifests replayed (≥ 1).
1797 pub generation_count: usize,
1798 /// Total frames inserted across all generations (`last_frame - 0`, since
1799 /// the chain is required to start at frame 1 and be gap-free).
1800 pub frames_replayed: u64,
1801 /// The last frame position written into the destination WAL.
1802 pub last_frame: u64,
1803 /// R732-F2: the highest fencing epoch contributing to this restore (`0`
1804 /// for a chain written before fencing existed). Reported so an operator
1805 /// restoring after an ownership transfer can see which owner's data they
1806 /// actually got.
1807 pub epoch: u64,
1808}
1809
1810/// A consistency-checked sequence of generation manifests, ready to drive a
1811/// replay. The fields are the single base / page size / checkpoint sequence
1812/// shared by every manifest in the chain.
1813#[derive(Debug, Clone, PartialEq, Eq)]
1814pub(crate) struct ValidatedChain {
1815 pub base_snapshot_key: String,
1816 pub page_size: usize,
1817 /// R858-B19: the one WAL generation every manifest in the chain agrees on.
1818 /// `salt` is `Some` for any chain of two or more manifests — a chain that
1819 /// could not prove a single generation never gets this far.
1820 pub generation: WalGeneration,
1821 pub total_frames: u64,
1822 /// R732-F2: the highest fencing epoch in the chain — i.e. the most recent
1823 /// owner that contributed frames. Unlike the other fields this is a
1824 /// *maximum*, not a shared constant: ownership legitimately moves
1825 /// mid-chain, so a chain may span epochs as long as they never go
1826 /// backwards.
1827 pub epoch: u64,
1828}
1829
1830/// Validate that a sorted list of generation manifests forms a single,
1831/// replayable chain: same base snapshot, same page size, single checkpoint
1832/// sequence, frames starting at 1 and contiguous across manifests.
1833///
1834/// V1 refuses to span a WAL restart (multiple `checkpoint_seq` values). A
1835/// restart implies the source engine folded the WAL into main between
1836/// generations — replaying the post-restart frames onto our pre-restart base
1837/// would skip that fold and corrupt the result. The remediation is a fresh
1838/// tier-1a snapshot, not heroics in restore.
1839///
1840/// # R858-B19 — refuse what cannot be proven
1841///
1842/// The `checkpoint_seq` test above was the *only* generation check here, and it
1843/// is blind to the fold that matters: a writer-process restart deletes the
1844/// `-wal` file and the next writer starts a fresh WAL back at sequence `0`, so
1845/// two unrelated WALs both report `0` and a spliced chain sailed through. That
1846/// produced a restore that reported SUCCESS while writing a stale-but-plausible
1847/// image (probe B: 9 rows against a 12-row source, `integrity_check ok`) or one
1848/// upstream sqlite3 calls malformed (probe F). **A silent wrong image is the
1849/// specific failure this function now exists to make impossible.**
1850///
1851/// So the salt is checked too, and — the part that matters — a chain that
1852/// cannot be *shown* to come from one WAL is refused rather than replayed:
1853///
1854/// - Two or more manifests, any of them lacking a salt (written before `v4`):
1855/// REFUSE. The splice is exactly what a pre-R858-B19 writer produced, and
1856/// nothing in those manifests records which WAL each range came from.
1857/// - Two or more manifests with disagreeing salts: REFUSE, naming both.
1858/// - A single manifest: accepted with whatever salt it has, including none.
1859/// One generation is not a splice; there is nothing to prove.
1860pub(crate) fn validate_generation_chain(
1861 manifests: &[OwnedGenerationManifest],
1862) -> Result<ValidatedChain> {
1863 let first = manifests
1864 .first()
1865 .context("validate_generation_chain: empty manifest list")?;
1866 let mut expected_next_frame: u64 = 1;
1867 let mut chain_epoch: u64 = 0;
1868 let multi = manifests.len() > 1;
1869 for (i, m) in manifests.iter().enumerate() {
1870 // R732-F2 (W245): generations are ordered by write time, so a chain
1871 // whose epoch goes BACKWARDS says a stale owner wrote after the sink
1872 // had already moved on — precisely the split-brain the fence exists to
1873 // stop, caught here on the read side too. Non-decreasing is fine and
1874 // expected: ownership transfers mid-stream and the new owner keeps
1875 // appending frames to the same contiguous range.
1876 if m.epoch < chain_epoch {
1877 anyhow::bail!(
1878 "generation #{i} was written at epoch {} but an earlier generation in the chain is at epoch {} — a fenced (stale) owner wrote to this sink, restore refuses rather than replay interleaved frames",
1879 m.epoch,
1880 chain_epoch,
1881 );
1882 }
1883 chain_epoch = m.epoch;
1884 if m.base_snapshot_key != first.base_snapshot_key {
1885 anyhow::bail!(
1886 "generation #{i} references base {:?}, expected {:?} — chain spans bases, restore needs a fresh tier-1a snapshot",
1887 m.base_snapshot_key,
1888 first.base_snapshot_key,
1889 );
1890 }
1891 if m.page_size != first.page_size {
1892 anyhow::bail!(
1893 "generation #{i} page_size {} differs from chain page_size {} — corrupt manifest or mixed sinks",
1894 m.page_size,
1895 first.page_size,
1896 );
1897 }
1898 if m.checkpoint_seq != first.checkpoint_seq {
1899 anyhow::bail!(
1900 "generation #{i} checkpoint_seq {} differs from chain checkpoint_seq {} — WAL restart between generations, restore needs a fresh tier-1a snapshot",
1901 m.checkpoint_seq,
1902 first.checkpoint_seq,
1903 );
1904 }
1905 // R858-B19: the check `checkpoint_seq` cannot make. Only enforced on a
1906 // multi-manifest chain — a lone generation has nothing to be spliced
1907 // to, so refusing it would strand every pre-v4 single-generation backup
1908 // for no safety gain.
1909 if multi {
1910 let Some(s) = m.salt else {
1911 anyhow::bail!(
1912 "generation #{i} carries no WAL salt (written by a pre-R858-B19 writer) and this chain spans {} generations — which WAL each range came from was never recorded, so a chain spliced across a WAL recreate is indistinguishable from a good one. Restore refuses rather than replay a plausible wrong image; take a fresh tier-1a snapshot",
1913 manifests.len(),
1914 );
1915 };
1916 if let Some(fs) = first.salt {
1917 if s != fs {
1918 anyhow::bail!(
1919 "generation #{i} was written under WAL salt {s} but the chain starts at salt {fs} (both at checkpoint_seq {}) — the source WAL was RECREATED mid-stream, so these frames belong to two different WALs and replaying them as one would corrupt the image. Restore needs a fresh tier-1a snapshot",
1920 first.checkpoint_seq,
1921 );
1922 }
1923 }
1924 }
1925 if m.first_frame != expected_next_frame {
1926 anyhow::bail!(
1927 "generation #{i} starts at frame {} but the previous generation ended at frame {} — gap in stream",
1928 m.first_frame,
1929 expected_next_frame.saturating_sub(1),
1930 );
1931 }
1932 if m.last_frame < m.first_frame {
1933 anyhow::bail!(
1934 "generation #{i} has last_frame {} < first_frame {} — corrupt manifest",
1935 m.last_frame,
1936 m.first_frame,
1937 );
1938 }
1939 expected_next_frame = m.last_frame + 1;
1940 }
1941 Ok(ValidatedChain {
1942 base_snapshot_key: first.base_snapshot_key.clone(),
1943 page_size: first.page_size,
1944 generation: WalGeneration {
1945 checkpoint_seq: first.checkpoint_seq,
1946 salt: first.salt,
1947 },
1948 total_frames: expected_next_frame - 1,
1949 epoch: chain_epoch,
1950 })
1951}
1952
1953/// Fetch one generation's frames and hand each to `on_frame` in ascending
1954/// frame order, skipping anything before `from_frame`. Returns how many frames
1955/// were delivered.
1956///
1957/// R761-F2: the one read path that resolves a manifest to frame bytes,
1958/// batched or not — `BackupTarget::frame_objects_of` decides which layout
1959/// the generation used, and a batch object is split here at
1960/// `24 + page_size` boundaries. Both restore's replay and
1961/// [`crate::puller::WalPuller`] call it, so a layout the writer can produce
1962/// can never be readable by one and not the other.
1963pub(crate) async fn for_each_frame_in_generation<F>(
1964 target: &BackupTarget,
1965 m: &OwnedGenerationManifest,
1966 from_frame: u64,
1967 mut on_frame: F,
1968) -> Result<u64>
1969where
1970 F: FnMut(u64, &[u8]) -> Result<()>,
1971{
1972 let frame_size = WAL_FRAME_HEADER_SIZE + m.page_size;
1973 let mut delivered: u64 = 0;
1974 for (key, first, last) in target.frame_objects_of(m) {
1975 if last < from_frame {
1976 continue; // wholly behind the caller's cursor — don't pay for the GET
1977 }
1978 let bytes = target
1979 .store
1980 .get(&key)
1981 .await
1982 .with_context(|| format!("fetching frame object {key}"))?
1983 .bytes()
1984 .await
1985 .with_context(|| format!("reading frame object body {key}"))?;
1986 let frames_in_object = (last - first + 1) as usize;
1987 let want = frames_in_object * frame_size;
1988 if bytes.len() != want {
1989 anyhow::bail!(
1990 "frame object {key} is {} bytes, expected {want} ({frames_in_object} frame(s) x [{WAL_FRAME_HEADER_SIZE} header + {} page])",
1991 bytes.len(),
1992 m.page_size,
1993 );
1994 }
1995 for (i, frame_no) in (first..=last).enumerate() {
1996 if frame_no < from_frame {
1997 continue;
1998 }
1999 let at = i * frame_size;
2000 on_frame(frame_no, &bytes[at..at + frame_size])?;
2001 delivered += 1;
2002 }
2003 }
2004 Ok(delivered)
2005}
2006
2007/// Replay every frame named by `manifests` into `seam`, in (checkpoint_seq,
2008/// frame_no) order. Caller must have already called `wal_insert_begin` on the
2009/// seam; `wal_insert_end` is also the caller's responsibility (so a test or a
2010/// future fault-injection path can choose `force_commit`).
2011///
2012/// Each manifest's own `page_size` sizes its frames —
2013/// [`validate_generation_chain`] has already established they all agree, so
2014/// there is no separate chain-level page size to thread through.
2015async fn replay_frames_into<S: WalInsertSeam>(
2016 target: &BackupTarget,
2017 seam: &S,
2018 manifests: &[OwnedGenerationManifest],
2019) -> Result<u64> {
2020 let mut total: u64 = 0;
2021 for m in manifests {
2022 total += for_each_frame_in_generation(target, m, 0, |frame_no, bytes| {
2023 seam.wal_insert_frame(frame_no, bytes)
2024 .with_context(|| format!("inserting frame {frame_no}"))
2025 })
2026 .await?;
2027 }
2028 Ok(total)
2029}
2030
2031/// Restore a database from a streamed backup: download the base tier-1a
2032/// snapshot referenced by the generation manifests, then replay every uploaded
2033/// WAL frame onto it via the [`WalInsertSeam`] of a fresh `turso_core`
2034/// connection. The session ends with `force_commit = false` so any frames
2035/// captured after the last commit frame are dropped (crash-consistency).
2036///
2037/// Errors loudly when:
2038/// - There are no generation manifests under the prefix (caller should restore
2039/// via the tier-1a path instead).
2040/// - The chain spans more than one `checkpoint_seq` or more than one base
2041/// snapshot (restart / mixed bases — needs a fresh tier-1a snapshot).
2042/// - Frame ranges are not gap-free starting at 1.
2043/// - A referenced frame object is missing or the wrong byte length.
2044///
2045/// `dest_path`'s `-wal` and `-shm` sidecars are removed before the base is
2046/// written; any pre-existing turso connection on `dest_path` must be closed
2047/// by the caller.
2048///
2049/// A caller that has already listed the chain for its own reasons should call
2050/// [`restore_stream_from_manifests`] instead and skip the re-listing this one
2051/// does — see its docs for what that costs (R760-T18).
2052pub async fn restore_latest_stream(
2053 target: &BackupTarget,
2054 dest_path: &str,
2055) -> Result<RestoreOutcome> {
2056 let manifests = list_and_parse_generation_manifests(target).await?;
2057 restore_stream_from_manifests(target, dest_path, &manifests).await
2058}
2059
2060/// [`restore_latest_stream`] for a caller that has already listed and parsed
2061/// the chain — same replay, minus the listing.
2062///
2063/// R760-T18: a caller that must inspect the chain *before* deciding how to
2064/// restore otherwise pays for every manifest twice. roadcase's `SinkHydrator`
2065/// is the live example: it calls [`list_and_parse_generation_manifests`] to
2066/// ask whether the era has any frames at all (an era whose base is still the
2067/// whole story restores via `snapshot::restore_latest` instead), and then
2068/// `restore_latest_stream` re-listed and re-fetched the identical objects. A
2069/// cold start with `n` generations was measured at `3n+1` Class B operations
2070/// against a `2n+1` floor — one GET per manifest, one per generation's frame
2071/// batch, one for the base. Handing the parse straight in removes the `n`
2072/// duplicates.
2073///
2074/// `manifests` must be ascending by generation, which is the order
2075/// [`list_and_parse_generation_manifests`] returns them in; the chain is
2076/// validated here exactly as it is on the listing path, so a hand-assembled
2077/// list cannot smuggle past a check.
2078pub async fn restore_stream_from_manifests(
2079 target: &BackupTarget,
2080 dest_path: &str,
2081 manifests: &[OwnedGenerationManifest],
2082) -> Result<RestoreOutcome> {
2083 if manifests.is_empty() {
2084 anyhow::bail!(
2085 "no generation manifests under {} — restore tier-1a directly via snapshot::restore_latest",
2086 join_key(&target.prefix, "generations"),
2087 );
2088 }
2089 let chain = validate_generation_chain(manifests)?;
2090
2091 // Download the base snapshot and lay it down at dest_path. Strip any stale
2092 // WAL/-shm sidecars first — the base alone is the entire pre-replay image
2093 // (VACUUM INTO output has no WAL).
2094 let base_bytes = target
2095 .store
2096 .get(&ObjPath::from(chain.base_snapshot_key.clone()))
2097 .await
2098 .with_context(|| format!("downloading base snapshot {}", chain.base_snapshot_key))?
2099 .bytes()
2100 .await
2101 .with_context(|| format!("reading base snapshot body {}", chain.base_snapshot_key))?;
2102 for sfx in ["-wal", "-shm"] {
2103 let _ = std::fs::remove_file(format!("{dest_path}{sfx}"));
2104 }
2105 std::fs::write(dest_path, &base_bytes)
2106 .with_context(|| format!("writing restored base snapshot to {dest_path}"))?;
2107
2108 // Open a fresh seam on the laid-down base and replay. `CoreWalSeam::open`
2109 // disables auto-actions; `wal_insert_begin` further locks the session to
2110 // empty auto-actions for the txn, so restart can't race the replay.
2111 let seam = CoreWalSeam::open(dest_path)?;
2112 seam.wal_insert_begin()
2113 .context("starting WAL insert session on restore destination")?;
2114 let frames = match replay_frames_into(target, &seam, manifests).await {
2115 Ok(n) => n,
2116 Err(e) => {
2117 // Best-effort: roll the partial replay back so we never leave the
2118 // dest's WAL with an uncommitted suffix.
2119 let _ = seam.wal_insert_end(false);
2120 return Err(e);
2121 }
2122 };
2123 // force_commit=false: the engine drops any tail past the last commit
2124 // frame, which is exactly the crash-consistency story we want — a
2125 // mid-transaction tail captured by tail_frames gets truncated cleanly.
2126 seam.wal_insert_end(false)
2127 .context("closing WAL insert session on restore destination")?;
2128
2129 Ok(RestoreOutcome {
2130 base_snapshot_key: chain.base_snapshot_key,
2131 checkpoint_seq: chain.generation.checkpoint_seq,
2132 generation_count: manifests.len(),
2133 frames_replayed: frames,
2134 last_frame: chain.total_frames,
2135 epoch: chain.epoch,
2136 })
2137}
2138
2139/// Every generation manifest at a sink, in key (i.e. chronological) order.
2140///
2141/// **The GETs are serialized**, one manifest at a time, and so is the frame
2142/// replay in [`for_each_frame_in_generation`] — so a restore's wall clock is
2143/// `(2n+1) x RTT` plus transfer for `n` generations, not `max(RTT)`. That is
2144/// fine at the tens of generations a dedicated streamer accumulates between
2145/// checkpoints and is not fine at thousands; R760-T18 bounds it on the
2146/// *writer* side (roadcase caps generations per era) rather than by making
2147/// this concurrent, because concurrency here would mean a real
2148/// `futures`/`tokio::spawn` dependency in a crate that deliberately keeps
2149/// `futures_util` to `[dev-dependencies]`.
2150///
2151/// Public since R732-T5: a split-brain reconciliation oracle needs to ask what
2152/// a sink would actually replay — which owner wrote which frame range, under
2153/// which epoch — without running a full restore. Pair it with
2154/// [`validate_generation_chain`]'s public counterpart if you need the chain
2155/// checked rather than merely listed.
2156pub async fn list_and_parse_generation_manifests(
2157 target: &BackupTarget,
2158) -> Result<Vec<OwnedGenerationManifest>> {
2159 let prefix = join_key(&target.prefix, "generations");
2160 let listing = target
2161 .store
2162 .list_with_delimiter(Some(&prefix))
2163 .await
2164 .with_context(|| format!("listing generation manifests under {prefix}"))?;
2165 let mut keys: Vec<ObjPath> = listing.objects.into_iter().map(|o| o.location).collect();
2166 keys.sort();
2167 let mut manifests = Vec::with_capacity(keys.len());
2168 for key in &keys {
2169 let bytes = target
2170 .store
2171 .get(key)
2172 .await
2173 .with_context(|| format!("fetching generation manifest {key}"))?
2174 .bytes()
2175 .await
2176 .with_context(|| format!("reading generation manifest body {key}"))?;
2177 let m = parse_generation_manifest(&String::from_utf8_lossy(&bytes))
2178 .with_context(|| format!("parsing generation manifest {key}"))?;
2179 manifests.push(m);
2180 }
2181 Ok(manifests)
2182}
2183
2184/// What the sink's fence currently stands at — the two comparands
2185/// [`tail_frames`] checks a writer against, and nothing else.
2186///
2187/// R869: this is the off-fleet copy of a fencing token. It matters because
2188/// `epoch` is minted by a raft group, and a raft group can be *destroyed* —
2189/// wipe the raft dir, re-form, and the fresh cluster's first
2190/// `ClaimTenant` grants epoch 1 while this sidecar still says 5, so the
2191/// rebuilt cluster is fenced out of its own sink. Reading the fence back is how
2192/// a rebuild starts above the number its dead predecessor left here.
2193#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
2194pub struct FenceState {
2195 /// The fencing epoch of the writer that last advanced the sidecar. `0`
2196 /// means "unfenced" — either nothing has claimed this sink or it predates
2197 /// R732-F2 — and fences nobody.
2198 pub epoch: u64,
2199 /// The pointer generation of that same writer. `0` means unfenced, exactly
2200 /// as for `epoch`.
2201 pub pointer_generation: u64,
2202}
2203
2204/// Read the fence [`tail_frames`] would check a writer against, without writing
2205/// anything or needing a WAL.
2206///
2207/// `Ok(None)` means the sidecar does not exist: nothing has ever streamed to
2208/// this prefix, so **there is no fence here at all** and any writer is
2209/// accepted. Do not collapse that into `FenceState::default()` at a call site
2210/// that is deciding whether to fence — "unfenced because nobody has written"
2211/// and "unfenced because an old writer stamped 0" are the same *value* but a
2212/// caller that cares about the difference (a rebuild deciding whether a tenant
2213/// was ever live) needs the `Option`.
2214///
2215/// **A floor derived from `epoch` is a lower bound, not an upper one.** It is
2216/// the highest epoch anyone has *written under*, which is not the highest a
2217/// dead raft group *granted* — a tenant claimed twice while idle leaves a node
2218/// holding an epoch strictly above anything this sidecar ever saw. So a rebuild
2219/// that seeds from this number closes availability (its own writes are
2220/// accepted) and does **not** on its own fence a resurrected node; that takes
2221/// `pointer_generation`, which is minted off-fleet where a dead cluster cannot
2222/// reach it. `oss/yubaba/crates/yubaba/tests/raft_rebuild_fencing.rs` drives
2223/// both halves against this code.
2224pub async fn read_fence_state(target: &BackupTarget) -> Result<Option<FenceState>> {
2225 Ok(read_watermark(&target.store, &target.watermark_key())
2226 .await?
2227 .map(|p| FenceState {
2228 epoch: p.epoch,
2229 pointer_generation: p.pointer_generation,
2230 }))
2231}
2232
2233/// In-memory shape of a generation manifest. Format on disk:
2234///
2235/// ```text
2236/// TURSO-BACKUP STREAM v4
2237/// base_snapshot <key>
2238/// page_size <n>
2239/// checkpoint_seq <n>
2240/// wal_salt <salt1>-<salt2>
2241/// first_frame <n>
2242/// last_frame <n>
2243/// epoch <n>
2244/// owner <label> (optional)
2245/// frame_batch <first>-<last> (one per batch object, ascending, gap-free)
2246/// ```
2247///
2248/// Text, dependency-free (no serde), one field per line. Mirrors the
2249/// `dedup::Manifest` convention so the crate stays consistent.
2250///
2251/// R732-F2 added `epoch`/`owner` and moved the header to `v2`. R761-F2 added
2252/// the `frame_batch` list and moved it to `v3`. R858-B19 added `wal_salt` and
2253/// moved it to `v4`. Older manifests still parse — `v1` with `epoch = 0`,
2254/// `owner = None`; `v1`/`v2` with no batch list, which is what marks their
2255/// frames as living one-per-object; `v1`/`v2`/`v3` with `salt = None` — so
2256/// backups written before any of those changes stay *parseable*. No older
2257/// version is ever written any more.
2258///
2259/// R858-B19, and this is the one place a legacy manifest is not merely
2260/// second-class: a chain of two or more manifests where any of them lacks a
2261/// salt is **refused** by [`validate_generation_chain`], because a pre-R858-B19
2262/// writer could splice two WAL generations into a contiguous-looking chain and
2263/// nothing in the manifest records which WAL each range came from. A
2264/// single-manifest chain is still restored — one generation cannot be a splice.
2265#[derive(Debug, Clone, PartialEq, Eq)]
2266pub struct GenerationManifest<'a> {
2267 pub base_snapshot_key: &'a str,
2268 pub page_size: usize,
2269 pub checkpoint_seq: u32,
2270 /// R858-B19: the WAL header salt these frames were read under — the field
2271 /// that actually names the generation, since `checkpoint_seq` resets to 0
2272 /// across a writer restart. `None` only when re-formatting a pre-R858-B19
2273 /// manifest; [`tail_frames`] always has one.
2274 pub salt: Option<WalSalt>,
2275 pub first_frame: u64,
2276 pub last_frame: u64,
2277 /// R732-F2: the fencing epoch this generation was written under. `0` means
2278 /// it predates fencing (a `v1` manifest), which is also the key shape its
2279 /// frames live under — see [`BackupTarget::frame_key`].
2280 pub epoch: u64,
2281 /// R732-F2: opaque owner label, diagnostic only. See [`StreamConfig::owner`].
2282 pub owner: Option<&'a str>,
2283 /// R761-F2: the `(first, last)` frame range of each batch object holding
2284 /// this generation's frames, ascending and exactly tiling
2285 /// `first_frame..=last_frame`. Empty means the pre-batching layout (one
2286 /// object per frame); a writer emits it empty only when re-formatting a
2287 /// legacy manifest.
2288 pub frame_batches: &'a [(u64, u64)],
2289}
2290
2291/// Header of a manifest written by a fencing-aware writer (R732-F2). The
2292/// version is bumped rather than the `epoch` key just being added, because
2293/// [`parse_generation_manifest`] rejects unknown keys: a pre-R732 binary
2294/// reading a fenced manifest would otherwise fail with `unknown manifest key:
2295/// epoch`, which reads like corruption. Failing on the *header* says the real
2296/// thing — this backup was written by a newer writer.
2297const MANIFEST_HEADER_V2: &str = "TURSO-BACKUP STREAM v2";
2298/// Pre-fencing header. Still accepted on read (those backups must stay
2299/// restorable) and parses with `epoch = 0`, `owner = None`; never written.
2300const MANIFEST_HEADER_V1: &str = "TURSO-BACKUP STREAM v1";
2301/// R761-F2 — carries the `frame_batch` list. Bumped for the same reason
2302/// `v2` was: [`parse_generation_manifest`] rejects unknown keys, so a
2303/// pre-R761-F2 binary reading a batched manifest would fail with `unknown
2304/// manifest key: frame_batch`, which reads like corruption. Failing on the
2305/// header says the true thing — a newer writer wrote this sink. The refusal
2306/// matters more here than it did for fencing: an older reader that somehow
2307/// skipped the batch list would look for per-frame keys that do not exist.
2308const MANIFEST_HEADER_V3: &str = "TURSO-BACKUP STREAM v3";
2309/// R858-B19 — carries `wal_salt`, the field that makes a generation chain
2310/// checkable across a WAL recreate. Bumped for the same reason `v2` and `v3`
2311/// were: [`parse_generation_manifest`] rejects unknown keys, so an older binary
2312/// reading one of these would fail with `unknown manifest key: wal_salt`, which
2313/// reads like corruption. Failing on the header says the true thing.
2314const MANIFEST_HEADER_V4: &str = "TURSO-BACKUP STREAM v4";
2315
2316pub(crate) fn format_generation_manifest(m: GenerationManifest<'_>) -> String {
2317 // Each header promises the fields that version introduced, so stamp the
2318 // newest version whose promises this manifest can actually keep. The only
2319 // way to reach anything but v4 is re-formatting a legacy manifest (a test,
2320 // or a repair tool); tail_frames always has both a salt and batches.
2321 // A v4 manifest promises BOTH the salt and the batch list, so a legacy
2322 // shape missing either falls back rather than stamping a version whose
2323 // promises it cannot keep. `salt` without batches has no representation and
2324 // cannot occur: only tail_frames produces a salt, and it always batches.
2325 let salt = if m.frame_batches.is_empty() { None } else { m.salt };
2326 let header = if salt.is_some() {
2327 MANIFEST_HEADER_V4
2328 } else if m.frame_batches.is_empty() {
2329 MANIFEST_HEADER_V2
2330 } else {
2331 MANIFEST_HEADER_V3
2332 };
2333 let mut out = format!(
2334 "{}\nbase_snapshot {}\npage_size {}\ncheckpoint_seq {}\n",
2335 header, m.base_snapshot_key, m.page_size, m.checkpoint_seq,
2336 );
2337 if let Some(s) = salt {
2338 out.push_str(&format!("wal_salt {}-{}\n", s.salt1, s.salt2));
2339 }
2340 out.push_str(&format!(
2341 "first_frame {}\nlast_frame {}\nepoch {}\n",
2342 m.first_frame, m.last_frame, m.epoch,
2343 ));
2344 // Omitted rather than written empty: the parser splits on the first space,
2345 // so `owner ` with no value would round-trip to `Some("")`.
2346 if let Some(owner) = m.owner {
2347 out.push_str(&format!("owner {owner}\n"));
2348 }
2349 for (first, last) in m.frame_batches {
2350 out.push_str(&format!("frame_batch {first}-{last}\n"));
2351 }
2352 out
2353}
2354
2355/// Parse a generation manifest. Tolerant to trailing whitespace; rejects any
2356/// other shape (so a corrupt manifest fails loudly during restore, doesn't
2357/// silently degrade).
2358#[derive(Debug, Clone, PartialEq, Eq)]
2359pub struct OwnedGenerationManifest {
2360 pub base_snapshot_key: String,
2361 pub page_size: usize,
2362 pub checkpoint_seq: u32,
2363 /// R858-B19: `None` for a pre-`v4` manifest, which means the WAL
2364 /// generation behind these frames was never recorded and cannot be
2365 /// recovered. Unknown, not "the same as its neighbour" — see
2366 /// [`validate_generation_chain`].
2367 pub salt: Option<WalSalt>,
2368 pub first_frame: u64,
2369 pub last_frame: u64,
2370 /// R732-F2: `0` for a `v1` (pre-fencing) manifest.
2371 pub epoch: u64,
2372 /// R732-F2: diagnostic only; `None` when the writer did not label itself.
2373 pub owner: Option<String>,
2374 /// R761-F2: the batch objects covering `first_frame..=last_frame`,
2375 /// ascending and gap-free (enforced by [`parse_generation_manifest`]).
2376 /// **Empty means the pre-batching layout** — one object per frame — which
2377 /// is how a `v1`/`v2` manifest keeps restoring. Resolve it to keys with
2378 /// `BackupTarget::frame_objects_of` (crate-private) rather than branching
2379 /// at each call site.
2380 pub frame_batches: Vec<(u64, u64)>,
2381}
2382
2383pub fn parse_generation_manifest(text: &str) -> Result<OwnedGenerationManifest> {
2384 let mut lines = text.lines();
2385 let header = lines.next().context("empty generation manifest")?.trim();
2386 let (v2, v3, v4) = match header {
2387 MANIFEST_HEADER_V4 => (true, true, true),
2388 MANIFEST_HEADER_V3 => (true, true, false),
2389 MANIFEST_HEADER_V2 => (true, false, false),
2390 MANIFEST_HEADER_V1 => (false, false, false),
2391 other => anyhow::bail!("unexpected manifest header: {other:?}"),
2392 };
2393 let mut base_snapshot_key: Option<String> = None;
2394 let mut page_size: Option<usize> = None;
2395 let mut checkpoint_seq: Option<u32> = None;
2396 let mut salt: Option<WalSalt> = None;
2397 let mut first_frame: Option<u64> = None;
2398 let mut last_frame: Option<u64> = None;
2399 let mut epoch: Option<u64> = None;
2400 let mut owner: Option<String> = None;
2401 let mut frame_batches: Vec<(u64, u64)> = Vec::new();
2402 for line in lines {
2403 let line = line.trim();
2404 if line.is_empty() {
2405 continue;
2406 }
2407 let (k, v) = line
2408 .split_once(' ')
2409 .with_context(|| format!("malformed manifest line: {line:?}"))?;
2410 match k {
2411 "base_snapshot" => base_snapshot_key = Some(v.to_string()),
2412 "page_size" => page_size = Some(v.parse().context("page_size")?),
2413 "checkpoint_seq" => checkpoint_seq = Some(v.parse().context("checkpoint_seq")?),
2414 "wal_salt" => {
2415 let (s1, s2) = v
2416 .split_once('-')
2417 .with_context(|| format!("malformed wal_salt pair: {v:?}"))?;
2418 salt = Some(WalSalt {
2419 salt1: s1.trim().parse().context("wal_salt salt1")?,
2420 salt2: s2.trim().parse().context("wal_salt salt2")?,
2421 });
2422 }
2423 "first_frame" => first_frame = Some(v.parse().context("first_frame")?),
2424 "last_frame" => last_frame = Some(v.parse().context("last_frame")?),
2425 "epoch" => epoch = Some(v.parse().context("epoch")?),
2426 "owner" => owner = Some(v.to_string()),
2427 "frame_batch" => {
2428 let (first, last) = v
2429 .split_once('-')
2430 .with_context(|| format!("malformed frame_batch range: {v:?}"))?;
2431 frame_batches.push((
2432 first.trim().parse().context("frame_batch first")?,
2433 last.trim().parse().context("frame_batch last")?,
2434 ));
2435 }
2436 other => anyhow::bail!("unknown manifest key: {other}"),
2437 }
2438 }
2439 // A v2 manifest without an epoch is corrupt, not legacy — the writer that
2440 // stamped the v2 header always writes one. Defaulting it to 0 would
2441 // silently demote a fenced generation to unfenced, which is the one
2442 // direction this whole mechanism must never fail in.
2443 if v2 && epoch.is_none() {
2444 anyhow::bail!("{MANIFEST_HEADER_V2} manifest is missing `epoch`");
2445 }
2446 // R858-B19: same reasoning as the `epoch` check above. A v4 writer always
2447 // stamps the salt, so a v4 manifest without one is corrupt, not legacy —
2448 // and defaulting it to `None` would silently demote a checkable generation
2449 // to an unknown one, which is the direction this mechanism must never fail
2450 // in. The converse guard matters just as much: a pre-v4 header carrying a
2451 // `wal_salt` line is hand-edited or truncated, and trusting that salt would
2452 // let a forged line wave a spliced chain through.
2453 if v4 && salt.is_none() {
2454 anyhow::bail!("{MANIFEST_HEADER_V4} manifest is missing `wal_salt`");
2455 }
2456 if !v4 && salt.is_some() {
2457 anyhow::bail!(
2458 "manifest header {header:?} carries a `wal_salt` line — the WAL salt is a {MANIFEST_HEADER_V4} field, so this manifest is corrupt or hand-edited"
2459 );
2460 }
2461 let first_frame = first_frame.context("missing first_frame")?;
2462 let last_frame = last_frame.context("missing last_frame")?;
2463 // R761-F2: the batch list IS the frame index, so a v3 manifest that does
2464 // not tile its own range exactly would send restore looking for objects
2465 // that were never written — caught here, at parse, rather than as a 404
2466 // halfway through a replay.
2467 if v3 {
2468 anyhow::ensure!(
2469 !frame_batches.is_empty(),
2470 "{MANIFEST_HEADER_V3} manifest has no `frame_batch` lines — it cannot say where its frames are"
2471 );
2472 let mut expected = first_frame;
2473 for &(first, last) in &frame_batches {
2474 anyhow::ensure!(
2475 first == expected && last >= first,
2476 "frame_batch {first}-{last} does not continue the range at frame {expected}"
2477 );
2478 expected = last + 1;
2479 }
2480 anyhow::ensure!(
2481 expected == last_frame + 1,
2482 "frame_batch list covers frames {first_frame}..={} but the manifest claims {first_frame}..={last_frame}",
2483 expected - 1,
2484 );
2485 } else if !frame_batches.is_empty() {
2486 anyhow::bail!(
2487 "manifest header {header:?} carries `frame_batch` lines — batching is a {MANIFEST_HEADER_V3} feature, so this manifest is corrupt or hand-edited"
2488 );
2489 }
2490 Ok(OwnedGenerationManifest {
2491 base_snapshot_key: base_snapshot_key.context("missing base_snapshot")?,
2492 page_size: page_size.context("missing page_size")?,
2493 checkpoint_seq: checkpoint_seq.context("missing checkpoint_seq")?,
2494 salt,
2495 first_frame,
2496 last_frame,
2497 epoch: epoch.unwrap_or(0),
2498 owner,
2499 frame_batches,
2500 })
2501}
2502
2503/// A watermark as read back from the sidecar: the engine position plus the
2504/// wall-clock instant the sidecar was last written (R574-T4 — this is what
2505/// [`RpoStatus::watermark_age`] measures against). `written_at_nanos` is
2506/// `None` for sidecars written before the timestamp field existed.
2507#[derive(Debug, Clone, PartialEq, Eq)]
2508struct PersistedWatermark {
2509 watermark: Watermark,
2510 /// R858-B19: the WAL generation `watermark.last_frame` is a position
2511 /// *within*. `generation.salt == None` marks a sidecar written before the
2512 /// salt fields existed — unknown, therefore never provably equal to the
2513 /// live WAL, therefore a forced restart on the next tail. That is the
2514 /// deliberate choice: one redundant re-upload beats resuming into a WAL
2515 /// nobody can show is the same one.
2516 generation: WalGeneration,
2517 written_at_nanos: Option<u128>,
2518 /// R732-F2: the fencing epoch of the writer that last advanced this
2519 /// sidecar. `0` for a sidecar written before the field existed, which
2520 /// reads as "unfenced" and therefore fences nobody — the same
2521 /// behaviour-preserving default as [`StreamConfig::epoch`].
2522 ///
2523 /// This is the value [`tail_frames`] compares against, so it is the single
2524 /// piece of state the whole fence rests on. R732-T3 makes advancing it a
2525 /// compare-and-swap (R732-T3), so the *concurrent* hole is closed too:
2526 /// two writers racing the read-modify-write cannot both win, because the
2527 /// loser's conditional put fails against the version the winner replaced.
2528 epoch: u64,
2529 /// R736-T2: the pointer generation of the writer that last advanced this
2530 /// sidecar. `0` for a sidecar written before the field existed, which
2531 /// reads as "unfenced" and therefore fences nobody — the same
2532 /// behaviour-preserving default as [`StreamConfig::pointer_generation`].
2533 /// Checked alongside `epoch` in [`tail_frames`]; either being stale
2534 /// bounces the writer.
2535 pointer_generation: u64,
2536 /// R732-T3: the object version this record was read at, carried so the
2537 /// next advance can be conditional on it. `None` only when the sidecar
2538 /// does not exist yet, which selects [`PutMode::Create`] instead of
2539 /// [`PutMode::Update`] — "I believe nobody owns this sink" is as much a
2540 /// precondition as "I believe it is still at version V".
2541 version: Option<UpdateVersion>,
2542}
2543
2544/// Sidecar format:
2545/// `<checkpoint_seq> <last_frame> [<written_at_unix_nanos> [<epoch>
2546/// [<pointer_generation> [<salt1> <salt2>]]]]`.
2547///
2548/// The third field was added by R574-T4, the fourth by R732-F2, the fifth by
2549/// R736-T2, and the sixth and seventh by R858-B19; all are positional appends,
2550/// which is what keeps this readable in both directions. A shorter sidecar (an
2551/// older writer) still parses, with the missing tail defaulting to `None` /
2552/// `0`, and an older reader stops after its last known field and never sees the
2553/// newer ones. Fields are only ever appended for exactly that reason — do not
2554/// reorder them.
2555///
2556/// R858-B19: a missing salt pair parses to `WalGeneration { salt: None }`,
2557/// which is *unknown*, not *unchanged*. `tail_frames` restarts against it
2558/// rather than resuming — see [`PersistedWatermark::generation`]. The two
2559/// fields are read as a pair: a sidecar carrying only one of them is corrupt
2560/// (this writer emits both or neither) and is rejected rather than
2561/// half-trusted.
2562async fn read_watermark(
2563 store: &Arc<dyn ObjectStore>,
2564 key: &ObjPath,
2565) -> Result<Option<PersistedWatermark>> {
2566 match store.get(key).await {
2567 Ok(res) => {
2568 // Capture the version BEFORE consuming the body — `bytes()` takes
2569 // `res` by value. Both fields are kept because stores differ in
2570 // which one they honour for a conditional put (object_store's own
2571 // `UpdateVersion` docs say to preserve both).
2572 let version = UpdateVersion {
2573 e_tag: res.meta.e_tag.clone(),
2574 version: res.meta.version.clone(),
2575 };
2576 let bytes = res.bytes().await.context("reading watermark sidecar")?;
2577 let s = String::from_utf8_lossy(&bytes);
2578 let mut fields = s.split_whitespace();
2579 let seq = fields
2580 .next()
2581 .context("watermark sidecar must be '<checkpoint_seq> <last_frame> [<nanos>]'")?;
2582 let frame = fields
2583 .next()
2584 .context("watermark sidecar must be '<checkpoint_seq> <last_frame> [<nanos>]'")?;
2585 let written_at_nanos = fields
2586 .next()
2587 .map(|n| n.parse::<u128>().context("watermark written_at nanos"))
2588 .transpose()?;
2589 let epoch = fields
2590 .next()
2591 .map(|e| e.parse::<u64>().context("watermark epoch"))
2592 .transpose()?
2593 .unwrap_or(0);
2594 let pointer_generation = fields
2595 .next()
2596 .map(|g| g.parse::<u64>().context("watermark pointer_generation"))
2597 .transpose()?
2598 .unwrap_or(0);
2599 // R858-B19: both salt words or neither. A lone `salt1` is not a
2600 // half-known generation, it is a torn write — say so instead of
2601 // silently downgrading it to "unknown" and papering over it.
2602 let salt = match (fields.next(), fields.next()) {
2603 (Some(s1), Some(s2)) => Some(WalSalt {
2604 salt1: s1.parse().context("watermark salt1")?,
2605 salt2: s2.parse().context("watermark salt2")?,
2606 }),
2607 (None, _) => None,
2608 (Some(_), None) => anyhow::bail!(
2609 "watermark sidecar carries salt1 but no salt2 — truncated or hand-edited; \
2610 refusing rather than guessing which WAL generation it names"
2611 ),
2612 };
2613 let checkpoint_seq: u32 = seq.parse().context("watermark checkpoint_seq")?;
2614 Ok(Some(PersistedWatermark {
2615 watermark: Watermark {
2616 checkpoint_seq,
2617 last_frame: frame.parse().context("watermark last_frame")?,
2618 },
2619 generation: WalGeneration { checkpoint_seq, salt },
2620 written_at_nanos,
2621 epoch,
2622 pointer_generation,
2623 version: Some(version),
2624 }))
2625 }
2626 Err(object_store::Error::NotFound { .. }) => Ok(None),
2627 Err(e) => Err(e).context("fetching watermark sidecar"),
2628 }
2629}
2630
2631/// R732-T3: result of a conditional watermark advance.
2632#[derive(Debug, Clone, Copy, PartialEq, Eq)]
2633enum WatermarkCas {
2634 /// We held the version we read, and the sidecar now names us.
2635 Advanced,
2636 /// Somebody else replaced the sidecar between our read and our write.
2637 /// Says nothing about *who* — the caller re-reads to find out whether it
2638 /// was a newer owner (we are fenced) or a same-epoch racer (a bug).
2639 Contended,
2640}
2641
2642/// Advance the watermark sidecar **conditionally** on the version it was read
2643/// at (R732-T3 / W245).
2644///
2645/// This is the atomic half of the fence. The epoch check in [`tail_frames`]
2646/// stops a *sequential* stale owner — one that returns after a transfer and
2647/// reads a sidecar already stamped higher. It cannot stop a *concurrent* one:
2648/// two writers that both read the old sidecar in the same instant would both
2649/// pass that check and then both blindly overwrite, last write winning, which
2650/// is exactly the corruption W245 describes. Making the advance a
2651/// compare-and-swap removes the window — at most one of them holds the version
2652/// the other replaced.
2653///
2654/// No new CAS primitive was needed: `object_store` already models this as
2655/// [`PutMode::Update`] (→ [`object_store::Error::Precondition`]) and
2656/// [`PutMode::Create`] (→ `AlreadyExists`) for the first write, and R2 honours
2657/// the underlying `If-Match` / `If-None-Match`.
2658///
2659/// **Deployment caveat, not a code path:** `AmazonS3Builder` must be
2660/// configured for conditional puts against a store that supports them. If the
2661/// backend silently degrades to unconditional writes, this returns `Advanced`
2662/// unconditionally and the concurrent window reopens — the sequential fence in
2663/// `tail_frames` still holds, but the race does not.
2664async fn write_watermark(
2665 store: &Arc<dyn ObjectStore>,
2666 key: &ObjPath,
2667 w: Watermark,
2668 salt: Option<WalSalt>,
2669 epoch: u64,
2670 pointer_generation: u64,
2671 expected: Option<&UpdateVersion>,
2672) -> Result<WatermarkCas> {
2673 let mut body = format!(
2674 "{} {} {} {} {}",
2675 w.checkpoint_seq,
2676 w.last_frame,
2677 unix_nanos(),
2678 epoch,
2679 pointer_generation
2680 );
2681 // R858-B19: omitted entirely rather than written as a sentinel, so an
2682 // unknown generation is one shape (absent) on both the read and write
2683 // sides, and no magic value can ever be mistaken for a real salt.
2684 if let Some(s) = salt {
2685 body.push_str(&format!(" {} {}", s.salt1, s.salt2));
2686 }
2687 body.push('\n');
2688 let opts = PutOptions {
2689 mode: match expected {
2690 Some(v) => PutMode::Update(v.clone()),
2691 // No sidecar when we read: assert that is *still* true, so two
2692 // writers bootstrapping the same fresh sink cannot both proceed.
2693 None => PutMode::Create,
2694 },
2695 ..Default::default()
2696 };
2697 match store.put_opts(key, body.into_bytes().into(), opts).await {
2698 Ok(_) => Ok(WatermarkCas::Advanced),
2699 Err(object_store::Error::Precondition { .. })
2700 | Err(object_store::Error::AlreadyExists { .. }) => Ok(WatermarkCas::Contended),
2701 Err(e) => Err(e).with_context(|| format!("writing watermark sidecar {key}")),
2702 }
2703}
2704
2705/// Which conditional-put mode a preflight probe found unenforced.
2706///
2707/// Both matter and they fail independently — a store can honour `If-None-Match`
2708/// (guarding the bootstrap put) while ignoring `If-Match` (guarding every
2709/// steady-state advance), or the reverse. Naming which one degraded is the
2710/// difference between an operator fixing a bucket setting in a minute and
2711/// bisecting a corruption in a week.
2712#[derive(Debug, Clone, Copy, PartialEq, Eq)]
2713pub enum PreflightStage {
2714 /// [`PutMode::Create`] / `If-None-Match` — the guard on two writers
2715 /// bootstrapping the same fresh sink.
2716 Create,
2717 /// [`PutMode::Update`] / `If-Match` — the guard on every subsequent
2718 /// watermark advance, and therefore the one the steady state rests on.
2719 Update,
2720}
2721
2722impl PreflightStage {
2723 pub fn as_str(&self) -> &'static str {
2724 match self {
2725 PreflightStage::Create => "PutMode::Create (If-None-Match)",
2726 PreflightStage::Update => "PutMode::Update (If-Match)",
2727 }
2728 }
2729}
2730
2731/// What [`probe_conditional_puts`] observed at a real sink.
2732#[derive(Debug, Clone, Copy, PartialEq, Eq)]
2733pub enum PreconditionSupport {
2734 /// Both conditional modes were enforced — a put that should have been
2735 /// rejected was rejected. The watermark CAS is real at this sink.
2736 Honoured,
2737 /// A put that *must* have failed its precondition succeeded instead, so
2738 /// this backend has silently degraded to unconditional writes.
2739 Degraded { stage: PreflightStage },
2740}
2741
2742impl PreconditionSupport {
2743 pub fn is_honoured(&self) -> bool {
2744 matches!(self, PreconditionSupport::Honoured)
2745 }
2746}
2747
2748/// R732-T4: prove, against the *real* configured sink, that conditional puts
2749/// are actually enforced — and hand the caller a hard answer so it can refuse
2750/// to start when they are not.
2751///
2752/// [`write_watermark`] carries a deployment caveat that no test can close:
2753/// if `AmazonS3Builder` is pointed at a store that does not honour
2754/// `If-Match` / `If-None-Match`, every conditional put silently succeeds, the
2755/// CAS degrades to last-write-wins, and the concurrent split-brain window
2756/// W245 exists to shut reopens — while every unit test stays green, because
2757/// the in-memory store used in tests does honour them. Support is a property
2758/// of configuration, not of code, so a runtime probe is the only guard that
2759/// can exist.
2760///
2761/// The probe writes a uniquely-keyed canary under `<prefix>/preflight/`,
2762/// exercises both modes with puts that MUST be rejected, and deletes it. It
2763/// touches no watermark, no manifest and no frame, so it is safe to run at
2764/// startup against a live sink another node owns — and it is deliberately
2765/// keyed per-process-per-nanosecond so two nodes probing at once cannot fail
2766/// each other.
2767///
2768/// `Ok(Degraded)` is the interesting return and is NOT an error: the probe
2769/// worked perfectly, and reported that the store is unsafe. An `Err` means the
2770/// probe could not reach a verdict at all (sink unreachable, credentials
2771/// wrong), which is also a refuse-to-start condition but a different one for
2772/// an operator to read.
2773pub async fn probe_conditional_puts(target: &BackupTarget) -> Result<PreconditionSupport> {
2774 let key = join_key(
2775 &target.prefix,
2776 &format!("preflight/conditional-put-{:020}-{}.canary", unix_nanos(), std::process::id()),
2777 );
2778 let result = probe_at_key(&target.store, &key).await;
2779 // Best-effort cleanup: a leaked canary is inert (nothing reads
2780 // `preflight/`), so a delete failure must not mask the verdict — which is
2781 // the whole reason the caller ran this.
2782 if let Err(e) = target.store.delete(&key).await {
2783 tracing_delete_failure(&key, &e);
2784 }
2785 result
2786}
2787
2788/// The probe body, split out so the canary is deleted on every path.
2789async fn probe_at_key(
2790 store: &Arc<dyn ObjectStore>,
2791 key: &ObjPath,
2792) -> Result<PreconditionSupport> {
2793 let create = PutOptions { mode: PutMode::Create, ..Default::default() };
2794
2795 // 1. Claim the key. This one is *expected* to succeed; if it doesn't, the
2796 // probe cannot reach a verdict (and `AlreadyExists` here means the
2797 // per-process-per-nanosecond key collided, which is a bug, not a store
2798 // property — so it stays an error rather than a `Degraded` verdict).
2799 let first = store
2800 .put_opts(key, b"preflight-1".as_slice().into(), create.clone())
2801 .await
2802 .with_context(|| format!("preflight canary could not be created at {key}"))?;
2803 let v1 = UpdateVersion { e_tag: first.e_tag.clone(), version: first.version.clone() };
2804
2805 // 2. Create again over a key that now exists. A store honouring
2806 // `If-None-Match` rejects this; one that succeeds has degraded.
2807 match store.put_opts(key, b"preflight-2".as_slice().into(), create).await {
2808 Err(object_store::Error::AlreadyExists { .. })
2809 | Err(object_store::Error::Precondition { .. }) => {}
2810 Ok(_) => return Ok(PreconditionSupport::Degraded { stage: PreflightStage::Create }),
2811 Err(e) => {
2812 return Err(e).with_context(|| format!("preflight Create probe failed at {key}"))
2813 }
2814 }
2815
2816 // 3. Advance the canary conditionally on the version we hold, to obtain a
2817 // *superseded* version. Expected to succeed — it is the ordinary
2818 // steady-state write `write_watermark` makes.
2819 store
2820 .put_opts(
2821 key,
2822 b"preflight-3".as_slice().into(),
2823 PutOptions { mode: PutMode::Update(v1.clone()), ..Default::default() },
2824 )
2825 .await
2826 .with_context(|| format!("preflight Update probe could not advance {key}"))?;
2827
2828 // 4. Update again on the now-stale version — exactly the shape of a fenced
2829 // writer losing the watermark race. A store honouring `If-Match`
2830 // rejects it.
2831 match store
2832 .put_opts(
2833 key,
2834 b"preflight-4".as_slice().into(),
2835 PutOptions { mode: PutMode::Update(v1), ..Default::default() },
2836 )
2837 .await
2838 {
2839 Err(object_store::Error::Precondition { .. })
2840 | Err(object_store::Error::AlreadyExists { .. }) => Ok(PreconditionSupport::Honoured),
2841 Ok(_) => Ok(PreconditionSupport::Degraded { stage: PreflightStage::Update }),
2842 Err(e) => Err(e).with_context(|| format!("preflight Update probe failed at {key}")),
2843 }
2844}
2845
2846/// turso-backup takes no logging dependency (it is a library consumed by
2847/// binaries that pick their own), so a failed canary cleanup goes to stderr
2848/// rather than through `tracing`.
2849fn tracing_delete_failure(key: &ObjPath, e: &object_store::Error) {
2850 eprintln!("turso-backup: preflight canary {key} could not be deleted: {e}");
2851}
2852
2853/// R574-T4: compute the watermark-staleness snapshot for this call. Pure so
2854/// the breach edge cases are unit-testable without staging a sidecar.
2855fn rpo_status(target: Option<Duration>, prior: Option<&PersistedWatermark>) -> RpoStatus {
2856 let watermark_age = prior
2857 .and_then(|p| p.written_at_nanos)
2858 .and_then(|written_at| unix_nanos().checked_sub(written_at))
2859 .map(|nanos| Duration::from_nanos(u64::try_from(nanos).unwrap_or(u64::MAX)));
2860 let breached = matches!((target, watermark_age), (Some(t), Some(age)) if age > t);
2861 RpoStatus { target, watermark_age, breached }
2862}
2863
2864fn join_key(prefix: &str, leaf: &str) -> ObjPath {
2865 let prefix = prefix.trim_matches('/');
2866 if prefix.is_empty() {
2867 ObjPath::from(leaf)
2868 } else {
2869 ObjPath::from(format!("{prefix}/{leaf}"))
2870 }
2871}
2872
2873fn unix_nanos() -> u128 {
2874 SystemTime::now()
2875 .duration_since(UNIX_EPOCH)
2876 .map(|d| d.as_nanos())
2877 .unwrap_or(0)
2878}
2879
2880/// Default [`StreamGcConfig::grace`] — 24 hours.
2881///
2882/// Conservative on purpose. The window has to cover the longest gap between an
2883/// object becoming unreachable-looking and a live reader or writer finishing
2884/// with it, and there are three such gaps, the largest of which is not bounded
2885/// by anything this crate controls:
2886///
2887/// 1. **A restore in flight.** [`restore_stream_from_manifests`] lists the
2888/// chain, then fetches the base, then walks every frame object one at a
2889/// time with the GETs serialized (see
2890/// [`list_and_parse_generation_manifests`] on why). A cold restore of a
2891/// large database over a slow link is minutes to hours, and it is reading
2892/// objects it named *before* the GC listed anything.
2893/// 2. **A tail mid-upload.** [`tail_frames`] uploads frame batches and only
2894/// then writes the generation manifest naming them, so between those two
2895/// points a perfectly live batch is referenced by no manifest and looks
2896/// exactly like an orphan.
2897/// 3. **A rebase mid-flight.** `crate::tail::rebase` publishes the new base
2898/// *before* deleting the old generation manifests, so in that window the
2899/// surviving manifests name the OLD base and the new one is reachable from
2900/// nothing. The grace window is the only thing that keeps a concurrent GC
2901/// from deleting the base a recovery just published.
2902///
2903/// 24 hours also swallows clock skew between the object store's
2904/// `last_modified` and this process's wall clock, which is what the age
2905/// comparison is made against. A day of retained garbage costs storage; an
2906/// hour too few costs a restore.
2907pub const DEFAULT_STREAM_GC_GRACE: Duration = Duration::from_secs(24 * 60 * 60);
2908
2909/// Knobs for [`gc_stream`].
2910///
2911/// [`Default`] is `{ grace: DEFAULT_STREAM_GC_GRACE, dry_run: true }` — the
2912/// posture for something a human points at a live bucket.
2913#[derive(Debug, Clone, PartialEq, Eq)]
2914pub struct StreamGcConfig {
2915 /// Never collect an object younger than this. See
2916 /// [`DEFAULT_STREAM_GC_GRACE`].
2917 pub grace: Duration,
2918 /// Report what would be collected and delete nothing. **Defaults to
2919 /// `true`.**
2920 pub dry_run: bool,
2921}
2922
2923impl Default for StreamGcConfig {
2924 fn default() -> Self {
2925 Self {
2926 grace: DEFAULT_STREAM_GC_GRACE,
2927 dry_run: true,
2928 }
2929 }
2930}
2931
2932/// What a [`gc_stream`] pass found — and, when `dry_run` was false, deleted.
2933///
2934/// The collected keys are listed rather than counted because the primary
2935/// consumer is a human reading a dry run before authorizing the real one.
2936#[derive(Debug, Clone, PartialEq, Eq, Default)]
2937pub struct StreamGcOutcome {
2938 /// Echo of [`StreamGcConfig::dry_run`]. `true` means nothing was deleted
2939 /// and `collected_*` is a proposal.
2940 pub dry_run: bool,
2941 /// Every snapshot key a restore could still reach, sorted. See
2942 /// [`gc_stream`]'s "Liveness" section for how this is derived.
2943 pub live_base_snapshot_keys: Vec<String>,
2944 /// Superseded base snapshots, sorted.
2945 pub collected_snapshots: Vec<String>,
2946 /// Orphaned frame objects (batch or legacy per-frame), sorted.
2947 pub collected_frame_objects: Vec<String>,
2948 /// Total bytes across `collected_snapshots` + `collected_frame_objects`.
2949 pub collected_bytes: u64,
2950 /// Snapshot objects kept — because they are live, or because the whole
2951 /// snapshot half was skipped (see `base_snapshots_skipped`).
2952 pub retained_snapshots: usize,
2953 /// `true` when the prefix held **no generation manifests**, so this is not
2954 /// provably a tier-2 sink and no base snapshot was collected.
2955 ///
2956 /// A `snapshots/` prefix with no chain over it is indistinguishable from a
2957 /// plain tier-1a sink, whose older snapshots are *history* rather than
2958 /// garbage — [`crate::snapshot::restore_latest`] takes the newest, but an
2959 /// operator restoring a point in time names an older key by hand. Pruning
2960 /// those is [`crate::dedup::gc_dedup`]'s `keep_n` decision to make, not
2961 /// this sweep's. A live tier-2 sink always has a chain, so in the steady
2962 /// state this is `false` and the superseded bases are collected; the one
2963 /// way to see it `true` on a real stream is a GC that lands inside
2964 /// `crate::tail::rebase`'s manifests-deleted-frames-not-yet-written window,
2965 /// where skipping one pass costs nothing.
2966 pub base_snapshots_skipped: bool,
2967 /// Frame objects kept because a generation manifest still names them.
2968 pub retained_frame_objects: usize,
2969 /// Objects that were unreachable but too young to touch — the grace window
2970 /// did its job. A number that never falls to zero across consecutive runs
2971 /// means the grace is longer than the churn interval, not that the sweep
2972 /// is broken.
2973 pub spared_by_grace: usize,
2974}
2975
2976/// Tier-2 GC: reclaim the objects a `crate::tail::rebase` orphans.
2977///
2978/// The tier-1b counterpart is [`crate::dedup::gc_dedup`], and this deliberately
2979/// mirrors its shape: an explicitly-invoked sweep that deletes only what
2980/// nothing can reach, never a step on the write or recovery path. `rebase`
2981/// leaves its garbage behind on purpose — deleting data as part of a recovery
2982/// path is how a recovery path becomes the outage — and this is where that
2983/// debt is settled.
2984///
2985/// ## Two orphan classes, not one
2986///
2987/// R850-T3 was filed against the frames alone. The frames are the **smaller**
2988/// term:
2989///
2990/// - **Orphaned frame objects.** Each `rebase` abandons
2991/// `frames/{old_checkpoint_seq}/…` (or `frames/{epoch}/{seq}/…`) once
2992/// `crate::tail::delete_generation_manifests` removes the manifests that
2993/// named them. Bounded per rebase by SQLite's autocheckpoint threshold —
2994/// ~1000 pages, i.e. a handful of batch objects.
2995/// - **Superseded base snapshots.** `rebase` calls
2996/// [`crate::snapshot::upload_base_snapshot`], which is the explicitly
2997/// *non*-deduplicating one-shot variant: it `put`s a new
2998/// `snapshots/snapshot-{nanos}.db` and deletes nothing. Every rebase
2999/// therefore leaves a complete permanent copy of the database behind, and
3000/// for any database bigger than a few megabytes that dwarfs the frames.
3001/// Found while implementing this ticket; covered here rather than filed,
3002/// because it is the same leak in the same function.
3003///
3004/// A sweep that reclaimed only the frames would leave the larger leak running.
3005///
3006/// ## Liveness
3007///
3008/// Nothing is deleted unless *no* restore path can reach it. The rule is the
3009/// exact complement of what a restore selects, and the two are pinned together
3010/// by `gc_liveness_is_the_complement_of_restore_selection` in this module's
3011/// tests so they cannot drift:
3012///
3013/// - **Frames.** Live iff some generation manifest names the object, resolved
3014/// through `BackupTarget::frame_objects_of` — the same function restore's
3015/// replay and [`crate::puller::WalPuller`] go through, so batch and legacy
3016/// per-frame layouts and both epoch key shapes are handled by construction
3017/// rather than by a second copy of the layout rules here.
3018/// - **Base snapshots.** Live iff *either* some generation manifest names it,
3019/// *or* it is the lexically-greatest key under `snapshots/`. Those are the
3020/// two selections a restore makes:
3021/// [`restore_stream_from_manifests`] takes the chain's `base_snapshot_key`,
3022/// and a prefix with no generations falls back to
3023/// [`crate::snapshot::restore_latest`], which takes the newest key (see
3024/// [`crate::hydrate`]'s `restore_subject`). Their union is kept.
3025///
3026/// The base rule uses *every* manifest's key rather than
3027/// `validate_generation_chain`'s single answer, and that is the safe
3028/// direction: a chain that spans a WAL restart does not validate at all, and a
3029/// GC that refuses to guess keeps both bases instead of deleting the one the
3030/// next rebase is about to adopt.
3031///
3032/// And when there is no chain at all, the snapshot half is skipped entirely
3033/// rather than falling back to "keep the newest, collect the rest" — see
3034/// [`StreamGcOutcome::base_snapshots_skipped`] for why that distinction is not
3035/// paranoia. The frame half still runs: a `frames/` prefix under a sink with no
3036/// generation manifests is unreachable by construction.
3037///
3038/// ## Grace window
3039///
3040/// An object younger than [`StreamGcConfig::grace`] is never collected, no
3041/// matter how unreachable it looks. [`DEFAULT_STREAM_GC_GRACE`] documents the
3042/// three live windows that depend on it.
3043///
3044/// ## Order
3045///
3046/// Frames first, then snapshots. Unlike [`crate::dedup::gc_dedup`] the order is
3047/// not load-bearing — this sweep deletes no manifests, so nothing retained ever
3048/// points at anything collected, at any point during the pass or after a crash
3049/// in the middle of one.
3050pub async fn gc_stream(target: &BackupTarget, cfg: &StreamGcConfig) -> Result<StreamGcOutcome> {
3051 let manifests = list_and_parse_generation_manifests(target).await?;
3052
3053 // Normalize through ObjPath so a manifest's textual key compares equal to
3054 // the same object's listed location regardless of slash padding.
3055 let mut live_bases: HashSet<String> = manifests
3056 .iter()
3057 .map(|m| ObjPath::from(m.base_snapshot_key.as_str()).to_string())
3058 .collect();
3059 if let Some(newest) = crate::snapshot::latest_snapshot_key(target).await? {
3060 live_bases.insert(ObjPath::from(newest).to_string());
3061 }
3062
3063 let mut live_frames: HashSet<String> = HashSet::new();
3064 for m in &manifests {
3065 for (key, _first, _last) in target.frame_objects_of(m) {
3066 live_frames.insert(key.to_string());
3067 }
3068 }
3069
3070 let mut out = StreamGcOutcome {
3071 dry_run: cfg.dry_run,
3072 base_snapshots_skipped: manifests.is_empty(),
3073 ..Default::default()
3074 };
3075 out.live_base_snapshot_keys = live_bases.iter().cloned().collect();
3076 out.live_base_snapshot_keys.sort();
3077
3078 let now_secs = SystemTime::now()
3079 .duration_since(UNIX_EPOCH)
3080 .map(|d| d.as_secs())
3081 .unwrap_or(0) as i64;
3082 let grace_secs = i64::try_from(cfg.grace.as_secs()).unwrap_or(i64::MAX);
3083 let too_young = |meta: &object_store::ObjectMeta| {
3084 now_secs.saturating_sub(meta.last_modified.timestamp()) < grace_secs
3085 };
3086
3087 for meta in list_recursive(&target.store, &join_key(&target.prefix, "frames")).await? {
3088 if live_frames.contains(&meta.location.to_string()) {
3089 out.retained_frame_objects += 1;
3090 continue;
3091 }
3092 if too_young(&meta) {
3093 out.spared_by_grace += 1;
3094 continue;
3095 }
3096 if !cfg.dry_run {
3097 delete_collected(target, &meta.location).await?;
3098 }
3099 out.collected_bytes += meta.size;
3100 out.collected_frame_objects.push(meta.location.to_string());
3101 }
3102
3103 let snapshots_prefix = join_key(&target.prefix, "snapshots");
3104 let listing = target
3105 .store
3106 .list_with_delimiter(Some(&snapshots_prefix))
3107 .await
3108 .with_context(|| format!("listing base snapshots under {snapshots_prefix}"))?;
3109 for meta in listing.objects {
3110 // No chain over this prefix means it is not provably a tier-2 sink, and
3111 // a tier-1a sink's older snapshots are history rather than garbage —
3112 // see `StreamGcOutcome::base_snapshots_skipped`.
3113 if out.base_snapshots_skipped || live_bases.contains(&meta.location.to_string()) {
3114 out.retained_snapshots += 1;
3115 continue;
3116 }
3117 if too_young(&meta) {
3118 out.spared_by_grace += 1;
3119 continue;
3120 }
3121 if !cfg.dry_run {
3122 delete_collected(target, &meta.location).await?;
3123 }
3124 out.collected_bytes += meta.size;
3125 out.collected_snapshots.push(meta.location.to_string());
3126 }
3127
3128 out.collected_frame_objects.sort();
3129 out.collected_snapshots.sort();
3130 Ok(out)
3131}
3132
3133/// Delete one collected object, treating "already gone" as success.
3134///
3135/// Two GC passes can legitimately overlap, and a retry after a partial one
3136/// lands here too — in both cases the object being absent is the outcome we
3137/// asked for, exactly as it is for `crate::tail::rebase`'s deletes.
3138async fn delete_collected(target: &BackupTarget, key: &ObjPath) -> Result<()> {
3139 match target.store.delete(key).await {
3140 Ok(()) => Ok(()),
3141 Err(object_store::Error::NotFound { .. }) => Ok(()),
3142 Err(e) => Err(anyhow::Error::new(e)).with_context(|| format!("collecting {key}")),
3143 }
3144}
3145
3146/// Every object at or below `root`, walked with `list_with_delimiter` rather
3147/// than the streaming `list`.
3148///
3149/// `ObjectStore::list` returns a `BoxStream`, which needs `futures_util` to
3150/// drain — and this crate deliberately keeps `futures_util` in
3151/// `[dev-dependencies]` (see `Cargo.toml`), so a streaming drain here would add
3152/// a real dependency to the published crate for one listing. A worklist over
3153/// `list_with_delimiter` costs one request per directory instead of one per
3154/// page, and the tier-2 frame tree is three levels deep at most
3155/// (`frames/{epoch}/{seq}/{object}`), so that is a handful of requests.
3156async fn list_recursive(
3157 store: &Arc<dyn ObjectStore>,
3158 root: &ObjPath,
3159) -> Result<Vec<object_store::ObjectMeta>> {
3160 let mut found = Vec::new();
3161 let mut pending = vec![root.clone()];
3162 while let Some(dir) = pending.pop() {
3163 let listing = store
3164 .list_with_delimiter(Some(&dir))
3165 .await
3166 .with_context(|| format!("listing {dir}"))?;
3167 found.extend(listing.objects);
3168 pending.extend(listing.common_prefixes);
3169 }
3170 Ok(found)
3171}
3172
3173#[cfg(test)]
3174mod tests {
3175 use super::*;
3176 use crate::backpressure::fault_injection::{Fault, FaultyStore};
3177 use crate::backpressure::BackoffConfig;
3178 use object_store::memory::InMemory;
3179 use std::cell::RefCell;
3180 use std::time::Duration;
3181
3182 /// In-memory WAL seam: a fixed page_size, a vector of frames the test
3183 /// appends to, and an advancing checkpoint_seq the test can bump.
3184 struct MockWal {
3185 page_size: usize,
3186 state: RefCell<MockState>,
3187 }
3188 struct MockState {
3189 checkpoint_seq: u32,
3190 /// R858-B19: the WAL header salt this mock's frames are stamped with.
3191 /// Modelled explicitly because the two fold regimes move it
3192 /// differently, and the whole bug was reading only the sequence:
3193 /// [`MockWal::restart`] moves both, [`MockWal::recreate`] moves only
3194 /// this one.
3195 salt: WalSalt,
3196 frames: Vec<MockFrame>,
3197 auto_actions_disabled: bool,
3198 /// R858-B19: re-roll `salt` once this many frame reads have happened,
3199 /// so a test can land a WAL recreate *inside* a single `tail_frames`
3200 /// call rather than only between two. See
3201 /// [`MockWal::recreate_after_reads`].
3202 swap_salt_after_reads: Option<u32>,
3203 reads: u32,
3204 }
3205 #[derive(Clone)]
3206 struct MockFrame {
3207 info: FrameInfo,
3208 page_bytes: Vec<u8>,
3209 }
3210
3211 impl MockWal {
3212 fn new(page_size: usize) -> Self {
3213 Self {
3214 page_size,
3215 state: RefCell::new(MockState {
3216 checkpoint_seq: 0,
3217 salt: WalSalt { salt1: 0xd492_ea8a, salt2: 0x7acf_42a3 },
3218 frames: Vec::new(),
3219 auto_actions_disabled: false,
3220 swap_salt_after_reads: None,
3221 reads: 0,
3222 }),
3223 }
3224 }
3225 fn append(&self, page_no: u32, db_size: u32, fill: u8) {
3226 let mut s = self.state.borrow_mut();
3227 s.frames.push(MockFrame {
3228 info: FrameInfo { page_no, db_size },
3229 page_bytes: vec![fill; self.page_size],
3230 });
3231 }
3232 /// Simulate an **in-process** WAL restart — a long-lived connection
3233 /// folding its own WAL at the autocheckpoint threshold. The WAL file is
3234 /// reused, so `checkpoint_seq` advances and `salt1` advances with it
3235 /// (measured `d492ea8a -> d492ea8b` alongside seq `0 -> 1`). This is
3236 /// the regime the pre-R858-B19 `checkpoint_seq` test already caught.
3237 fn restart(&self) {
3238 let mut s = self.state.borrow_mut();
3239 s.checkpoint_seq += 1;
3240 s.salt.salt1 = s.salt.salt1.wrapping_add(1);
3241 s.frames.clear();
3242 }
3243 /// R858-B19 — simulate a **writer-process restart**: the last
3244 /// connection closed, SQLite checkpointed and DELETED the `-wal` file,
3245 /// and the next writer created a fresh WAL. `checkpoint_seq` goes back
3246 /// to 0 (it never moves, from a watcher's point of view: `0 -> 0`) and
3247 /// the salt is fresh randomness, unrelated to the old one.
3248 ///
3249 /// This is the regime `checkpoint_seq` is blind to, and the reason the
3250 /// mock models a salt at all.
3251 fn recreate(&self, salt: WalSalt) {
3252 let mut s = self.state.borrow_mut();
3253 s.checkpoint_seq = 0;
3254 s.salt = salt;
3255 s.frames.clear();
3256 }
3257 /// R858-B19: arm a mid-call WAL recreate — the salt flips once `reads`
3258 /// frame reads have been served, which lands it inside the drain loop
3259 /// rather than between two `tail_frames` calls.
3260 fn recreate_after_reads(&self, reads: u32) {
3261 self.state.borrow_mut().swap_salt_after_reads = Some(reads);
3262 }
3263 fn frame_count(&self) -> u64 {
3264 self.state.borrow().frames.len() as u64
3265 }
3266 }
3267
3268 impl WalSeam for MockWal {
3269 fn wal_state(&self) -> Result<Watermark> {
3270 let s = self.state.borrow();
3271 Ok(Watermark {
3272 checkpoint_seq: s.checkpoint_seq,
3273 last_frame: s.frames.len() as u64,
3274 })
3275 }
3276 fn wal_get_frame(&self, frame_no: u64, buf: &mut [u8]) -> Result<FrameInfo> {
3277 let mut s = self.state.borrow_mut();
3278 s.reads += 1;
3279 if s.swap_salt_after_reads == Some(s.reads) {
3280 s.salt.salt1 = !s.salt.salt1;
3281 }
3282 let idx = frame_no
3283 .checked_sub(1)
3284 .context("frame_no must be >= 1")? as usize;
3285 let f = s
3286 .frames
3287 .get(idx)
3288 .with_context(|| format!("frame {frame_no} out of range"))?;
3289 // Synthesize a 24-byte header: big-endian page_no, db_size, the
3290 // WAL's salt pair (R858-B19 reads the generation out of exactly
3291 // these bytes), then zeros for the checksums (never validated).
3292 buf[0..4].copy_from_slice(&f.info.page_no.to_be_bytes());
3293 buf[4..8].copy_from_slice(&f.info.db_size.to_be_bytes());
3294 buf[8..12].copy_from_slice(&s.salt.salt1.to_be_bytes());
3295 buf[12..16].copy_from_slice(&s.salt.salt2.to_be_bytes());
3296 buf[16..WAL_FRAME_HEADER_SIZE].fill(0);
3297 buf[WAL_FRAME_HEADER_SIZE..].copy_from_slice(&f.page_bytes);
3298 Ok(f.info)
3299 }
3300 fn wal_auto_actions_disable(&self) {
3301 self.state.borrow_mut().auto_actions_disabled = true;
3302 }
3303 }
3304
3305 /// A store that accepts every put unconditionally — the exact failure the
3306 /// preflight probe exists to catch. It is not a contrived shape: it is
3307 /// what an S3-compatible backend without conditional-write support looks
3308 /// like from `object_store`'s side, and what `AmazonS3Builder` degrades to
3309 /// when pointed at one. Everything but `put_opts` delegates.
3310 #[derive(Debug)]
3311 struct UnconditionalStore {
3312 inner: Arc<dyn ObjectStore>,
3313 }
3314
3315 impl std::fmt::Display for UnconditionalStore {
3316 fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
3317 write!(f, "UnconditionalStore({})", self.inner)
3318 }
3319 }
3320
3321 #[async_trait::async_trait]
3322 impl ObjectStore for UnconditionalStore {
3323 async fn put_opts(
3324 &self,
3325 location: &ObjPath,
3326 payload: object_store::PutPayload,
3327 mut opts: PutOptions,
3328 ) -> object_store::Result<object_store::PutResult> {
3329 opts.mode = PutMode::Overwrite;
3330 self.inner.put_opts(location, payload, opts).await
3331 }
3332 async fn put_multipart_opts(
3333 &self,
3334 location: &ObjPath,
3335 opts: object_store::PutMultipartOptions,
3336 ) -> object_store::Result<Box<dyn object_store::MultipartUpload>> {
3337 self.inner.put_multipart_opts(location, opts).await
3338 }
3339 async fn get_opts(
3340 &self,
3341 location: &ObjPath,
3342 options: object_store::GetOptions,
3343 ) -> object_store::Result<object_store::GetResult> {
3344 self.inner.get_opts(location, options).await
3345 }
3346 fn delete_stream(
3347 &self,
3348 locations: futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>>,
3349 ) -> futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>> {
3350 self.inner.delete_stream(locations)
3351 }
3352 fn list(
3353 &self,
3354 prefix: Option<&ObjPath>,
3355 ) -> futures_util::stream::BoxStream<'static, object_store::Result<object_store::ObjectMeta>>
3356 {
3357 self.inner.list(prefix)
3358 }
3359 async fn list_with_delimiter(
3360 &self,
3361 prefix: Option<&ObjPath>,
3362 ) -> object_store::Result<object_store::ListResult> {
3363 self.inner.list_with_delimiter(prefix).await
3364 }
3365 async fn copy_opts(
3366 &self,
3367 from: &ObjPath,
3368 to: &ObjPath,
3369 options: object_store::CopyOptions,
3370 ) -> object_store::Result<()> {
3371 self.inner.copy_opts(from, to, options).await
3372 }
3373 }
3374
3375 /// A store that honours `If-None-Match` but not `If-Match`. The nastiest
3376 /// real-world shape, because the bootstrap put looks fine and only the
3377 /// steady-state advance — the one every tail after the first depends on —
3378 /// is unguarded.
3379 #[derive(Debug)]
3380 struct CreateOnlyStore {
3381 inner: Arc<dyn ObjectStore>,
3382 }
3383
3384 impl std::fmt::Display for CreateOnlyStore {
3385 fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
3386 write!(f, "CreateOnlyStore({})", self.inner)
3387 }
3388 }
3389
3390 #[async_trait::async_trait]
3391 impl ObjectStore for CreateOnlyStore {
3392 async fn put_opts(
3393 &self,
3394 location: &ObjPath,
3395 payload: object_store::PutPayload,
3396 mut opts: PutOptions,
3397 ) -> object_store::Result<object_store::PutResult> {
3398 if matches!(opts.mode, PutMode::Update(_)) {
3399 opts.mode = PutMode::Overwrite;
3400 }
3401 self.inner.put_opts(location, payload, opts).await
3402 }
3403 async fn put_multipart_opts(
3404 &self,
3405 location: &ObjPath,
3406 opts: object_store::PutMultipartOptions,
3407 ) -> object_store::Result<Box<dyn object_store::MultipartUpload>> {
3408 self.inner.put_multipart_opts(location, opts).await
3409 }
3410 async fn get_opts(
3411 &self,
3412 location: &ObjPath,
3413 options: object_store::GetOptions,
3414 ) -> object_store::Result<object_store::GetResult> {
3415 self.inner.get_opts(location, options).await
3416 }
3417 fn delete_stream(
3418 &self,
3419 locations: futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>>,
3420 ) -> futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>> {
3421 self.inner.delete_stream(locations)
3422 }
3423 fn list(
3424 &self,
3425 prefix: Option<&ObjPath>,
3426 ) -> futures_util::stream::BoxStream<'static, object_store::Result<object_store::ObjectMeta>>
3427 {
3428 self.inner.list(prefix)
3429 }
3430 async fn list_with_delimiter(
3431 &self,
3432 prefix: Option<&ObjPath>,
3433 ) -> object_store::Result<object_store::ListResult> {
3434 self.inner.list_with_delimiter(prefix).await
3435 }
3436 async fn copy_opts(
3437 &self,
3438 from: &ObjPath,
3439 to: &ObjPath,
3440 options: object_store::CopyOptions,
3441 ) -> object_store::Result<()> {
3442 self.inner.copy_opts(from, to, options).await
3443 }
3444 }
3445
3446 /// The happy path: a store that honours both modes passes, and — the part
3447 /// that matters for running this at startup against a live sink — leaves
3448 /// nothing behind.
3449 #[tokio::test]
3450 async fn the_preflight_probe_passes_on_a_conditional_store_and_leaves_no_trace() {
3451 let target = fresh_target();
3452 assert_eq!(
3453 probe_conditional_puts(&target).await.unwrap(),
3454 PreconditionSupport::Honoured
3455 );
3456
3457 let leftovers = objects_under(&target).await;
3458 assert!(
3459 leftovers.is_empty(),
3460 "the probe must clean up its canary — found {leftovers:?}"
3461 );
3462 }
3463
3464 /// The whole point: a backend that silently ignores preconditions is
3465 /// *reported*, not tolerated. Without this the watermark CAS degrades to
3466 /// last-write-wins and every other test in this file still passes.
3467 #[tokio::test]
3468 async fn the_preflight_probe_catches_a_store_that_ignores_preconditions() {
3469 let target = BackupTarget {
3470 store: Arc::new(UnconditionalStore { inner: Arc::new(InMemory::new()) }),
3471 prefix: "backups".into(),
3472 };
3473 assert_eq!(
3474 probe_conditional_puts(&target).await.unwrap(),
3475 PreconditionSupport::Degraded { stage: PreflightStage::Create },
3476 "an unconditional store fails at the first guard it meets"
3477 );
3478 }
3479
3480 /// Half-degraded stores are the ones that actually ship. `If-None-Match`
3481 /// works, so bootstrapping looks healthy; `If-Match` does not, so every
3482 /// steady-state advance is unguarded. The probe must name `Update`
3483 /// specifically — "conditional puts are broken" would send an operator to
3484 /// the wrong setting.
3485 #[tokio::test]
3486 async fn the_preflight_probe_names_update_when_only_if_match_is_ignored() {
3487 let target = BackupTarget {
3488 store: Arc::new(CreateOnlyStore { inner: Arc::new(InMemory::new()) }),
3489 prefix: "backups".into(),
3490 };
3491 assert_eq!(
3492 probe_conditional_puts(&target).await.unwrap(),
3493 PreconditionSupport::Degraded { stage: PreflightStage::Update }
3494 );
3495 }
3496
3497 /// The canary is keyed per process per nanosecond, so two nodes probing
3498 /// the same sink at once each get a real verdict instead of failing each
3499 /// other. A startup probe that flaked under concurrency would be turned
3500 /// off within a week.
3501 #[tokio::test]
3502 async fn concurrent_preflight_probes_do_not_collide() {
3503 let target = fresh_target();
3504 let (a, b) = tokio::join!(
3505 probe_conditional_puts(&target),
3506 probe_conditional_puts(&target)
3507 );
3508 assert_eq!(a.unwrap(), PreconditionSupport::Honoured);
3509 assert_eq!(b.unwrap(), PreconditionSupport::Honoured);
3510 assert!(objects_under(&target).await.is_empty());
3511 }
3512
3513 async fn objects_under(target: &BackupTarget) -> Vec<String> {
3514 use futures_util::StreamExt;
3515 target
3516 .store
3517 .list(None)
3518 .map(|m| m.unwrap().location.to_string())
3519 .collect()
3520 .await
3521 }
3522
3523 fn fresh_target() -> BackupTarget {
3524 BackupTarget {
3525 store: Arc::new(InMemory::new()),
3526 prefix: "backups".into(),
3527 }
3528 }
3529
3530 fn cfg() -> StreamConfig<'static> {
3531 StreamConfig {
3532 base_snapshot_key: "backups/snapshots/snapshot-00000000000000000001.db",
3533 page_size: 4096,
3534 backpressure: BackpressureConfig::default(),
3535 rpo_target: None,
3536 epoch: 0,
3537 owner: None,
3538 pointer_generation: 0,
3539 }
3540 }
3541
3542 /// Fast-ticking backoff for backpressure tests — a few ms, never the
3543 /// production defaults, so retry-heavy tests stay fast.
3544 fn fast_backoff() -> BackoffConfig {
3545 BackoffConfig {
3546 initial_delay: Duration::from_millis(1),
3547 max_delay: Duration::from_millis(4),
3548 multiplier: 2.0,
3549 max_retries: 3,
3550 }
3551 }
3552
3553 fn bp_cfg(policy: BackpressurePolicy, spill_buffer_frames: usize) -> StreamConfig<'static> {
3554 StreamConfig {
3555 base_snapshot_key: "backups/snapshots/snapshot-00000000000000000001.db",
3556 page_size: 4096,
3557 backpressure: BackpressureConfig {
3558 spill_buffer_frames,
3559 policy,
3560 backoff: fast_backoff(),
3561 },
3562 rpo_target: None,
3563 epoch: 0,
3564 owner: None,
3565 pointer_generation: 0,
3566 }
3567 }
3568
3569 fn faulty_target(faults: impl IntoIterator<Item = Fault>) -> BackupTarget {
3570 BackupTarget {
3571 store: Arc::new(FaultyStore::new(Arc::new(InMemory::new()), faults)),
3572 prefix: "backups".into(),
3573 }
3574 }
3575
3576 /// Five WAL frames (1..=5), the last a commit — a small, deterministic
3577 /// range for the backpressure tests below.
3578 fn five_frame_seam() -> MockWal {
3579 let seam = MockWal::new(4096);
3580 for i in 1..=4u32 {
3581 seam.append(i, 0, i as u8);
3582 }
3583 seam.append(5, 5, 5); // commit
3584 seam
3585 }
3586
3587 // --- R574-F2: explicit R2 backpressure ---------------------------------
3588
3589 /// A single 429 triggers exactly one retry, then the upload succeeds —
3590 /// the whole range still lands and the report shows the retry.
3591 #[tokio::test]
3592 async fn throttled_429_retries_then_succeeds() {
3593 let seam = five_frame_seam();
3594 let target = faulty_target([Fault::TooManyRequests]);
3595 let cfg = bp_cfg(BackpressurePolicy::Fail, 8);
3596 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3597 match out {
3598 StreamOutcome::Streamed { frame_count, backpressure, .. } => {
3599 assert_eq!(frame_count, 5);
3600 assert_eq!(backpressure.throttle_retries, 1);
3601 assert_eq!(backpressure.frames_shed, 0);
3602 }
3603 other => panic!("expected Streamed, got {other:?}"),
3604 }
3605 }
3606
3607 /// A 503 is classified the same as a 429 and also triggers backoff then
3608 /// retry — two 503s in a row cost exactly two retries.
3609 #[tokio::test]
3610 async fn throttled_503_retries_then_succeeds() {
3611 let seam = five_frame_seam();
3612 let target = faulty_target([Fault::ServiceUnavailable, Fault::ServiceUnavailable]);
3613 let cfg = bp_cfg(BackpressurePolicy::Fail, 8);
3614 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3615 match out {
3616 StreamOutcome::Streamed { frame_count, backpressure, .. } => {
3617 assert_eq!(frame_count, 5);
3618 assert_eq!(backpressure.throttle_retries, 2);
3619 }
3620 other => panic!("expected Streamed, got {other:?}"),
3621 }
3622 }
3623
3624 /// Buffer at bound + `Fail`: sustained throttling exhausts backoff and
3625 /// the whole call errors — nothing is persisted.
3626 #[tokio::test]
3627 async fn buffer_at_bound_fail_policy_errors_without_persisting() {
3628 let seam = five_frame_seam();
3629 let faults = std::iter::repeat_n(Fault::TooManyRequests, 50);
3630 let target = faulty_target(faults);
3631 let cfg = bp_cfg(BackpressurePolicy::Fail, 2);
3632 let err = tail_frames(&seam, &target, &cfg).await.unwrap_err();
3633 assert!(format!("{err}").contains("uploading wal frame"), "err was {err}");
3634 assert!(
3635 read_watermark(&target.store, &target.watermark_key())
3636 .await
3637 .unwrap()
3638 .is_none(),
3639 "Fail must not persist a watermark when it gives up"
3640 );
3641 }
3642
3643 /// Buffer at bound + `Shed`: the backlog is dropped with a loud report
3644 /// instead of erroring; nothing persists, so the next call would
3645 /// re-attempt the same range.
3646 #[tokio::test]
3647 async fn buffer_at_bound_shed_policy_drops_backlog_without_error() {
3648 let seam = five_frame_seam();
3649 let faults = std::iter::repeat_n(Fault::TooManyRequests, 50);
3650 let target = faulty_target(faults);
3651 let cfg = bp_cfg(BackpressurePolicy::Shed, 2);
3652 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3653 match out {
3654 StreamOutcome::Shed { first_frame, last_frame, backpressure, .. } => {
3655 assert_eq!(first_frame, 1);
3656 assert_eq!(last_frame, 5);
3657 assert_eq!(backpressure.high_water_frames, 2, "capped at the bound");
3658 assert!(backpressure.frames_shed >= 1, "must report the dropped backlog");
3659 }
3660 other => panic!("expected Shed, got {other:?}"),
3661 }
3662 assert!(
3663 read_watermark(&target.store, &target.watermark_key())
3664 .await
3665 .unwrap()
3666 .is_none(),
3667 "Shed must not persist a watermark for a fully-dropped batch"
3668 );
3669 }
3670
3671 /// Buffer at bound + `Block`: retries never give up on a throttling
3672 /// error; once the store recovers, the full range still lands.
3673 #[tokio::test]
3674 async fn buffer_at_bound_block_policy_eventually_drains() {
3675 let seam = five_frame_seam();
3676 // More failures than fast_backoff's max_retries would tolerate under
3677 // Fail/Shed — Block must push through them anyway.
3678 let faults = std::iter::repeat_n(Fault::TooManyRequests, 4);
3679 let target = faulty_target(faults);
3680 let cfg = bp_cfg(BackpressurePolicy::Block, 2);
3681 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3682 match out {
3683 StreamOutcome::Streamed { frame_count, backpressure, .. } => {
3684 assert_eq!(frame_count, 5);
3685 assert_eq!(backpressure.high_water_frames, 2);
3686 assert_eq!(backpressure.throttle_retries, 4);
3687 assert_eq!(backpressure.frames_shed, 0);
3688 }
3689 other => panic!("expected Streamed, got {other:?}"),
3690 }
3691 }
3692
3693 /// R761-F2: a call with more frames than the spill buffer holds splits
3694 /// into one batch object per buffer-full, the manifest indexes every one
3695 /// of them, and replay puts the frames back in order. The bound is the
3696 /// only thing sizing an object, so this is also the assertion that the
3697 /// largest object this sink writes stays bounded.
3698 #[tokio::test]
3699 async fn a_call_larger_than_the_spill_buffer_splits_into_several_batch_objects() {
3700 let seam = MockWal::new(4096);
3701 for i in 1..=5u32 {
3702 seam.append(i, i, i as u8);
3703 }
3704 let target = fresh_target();
3705 let cfg = bp_cfg(BackpressurePolicy::Fail, 2);
3706 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3707 assert!(matches!(out, StreamOutcome::Streamed { frame_count: 5, .. }));
3708
3709 let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
3710 assert_eq!(
3711 manifests[0].frame_batches,
3712 vec![(1, 2), (3, 4), (5, 5)],
3713 "bound 2 over 5 frames is two full batches and a remainder"
3714 );
3715 let frame_size = WAL_FRAME_HEADER_SIZE + 4096;
3716 for (first, last) in [(1u64, 2u64), (3, 4), (5, 5)] {
3717 let bytes = target
3718 .store
3719 .get(&target.frame_batch_key(0, 0, first, last))
3720 .await
3721 .unwrap()
3722 .bytes()
3723 .await
3724 .unwrap();
3725 assert_eq!(bytes.len() as u64, (last - first + 1) * frame_size as u64);
3726 }
3727
3728 // And it all comes back, in order, through the normal read path.
3729 let insert = MockInsertSeam::new();
3730 assert_eq!(replay_frames_into(&target, &insert, &manifests).await.unwrap(), 5);
3731 let frames: Vec<u64> = insert
3732 .events()
3733 .into_iter()
3734 .filter_map(|e| match e {
3735 MockInsertEvent::Frame { frame_no, .. } => Some(frame_no),
3736 _ => None,
3737 })
3738 .collect();
3739 assert_eq!(frames, vec![1, 2, 3, 4, 5]);
3740 }
3741
3742 /// R761-F2 + R574-F2: a drain that lands one batch and then gets stuck
3743 /// under `Shed` publishes ONLY the batch that landed. The manifest's index
3744 /// is what restore follows, so a batch list claiming a shed object would
3745 /// be a 404 mid-replay — worse than the frames simply not being there.
3746 #[tokio::test]
3747 async fn a_partially_shed_drain_indexes_only_the_batches_that_landed() {
3748 let seam = five_frame_seam();
3749 // First batch through; the next one throttled until Shed gives up.
3750 // Exactly enough faults to exhaust one put's retry budget and no more,
3751 // so the watermark + manifest writes that follow the shed still land
3752 // (they go through the same store).
3753 let faults = std::iter::once(Fault::Pass).chain(std::iter::repeat_n(
3754 Fault::TooManyRequests,
3755 fast_backoff().max_retries as usize + 1,
3756 ));
3757 let target = faulty_target(faults);
3758 let cfg = bp_cfg(BackpressurePolicy::Shed, 2);
3759 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3760 match out {
3761 StreamOutcome::Streamed { first_frame, last_frame, frame_count, backpressure, .. } => {
3762 assert_eq!((first_frame, last_frame, frame_count), (1, 2, 2));
3763 assert_eq!(backpressure.frames_shed, 2, "frames 3-4 were buffered and dropped");
3764 }
3765 other => panic!("expected a Streamed prefix, got {other:?}"),
3766 }
3767 let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
3768 assert_eq!(manifests.len(), 1);
3769 assert_eq!(manifests[0].last_frame, 2);
3770 assert_eq!(manifests[0].frame_batches, vec![(1, 2)]);
3771 // Every object the manifest names actually exists — the property that
3772 // makes the published prefix restorable.
3773 let insert = MockInsertSeam::new();
3774 assert_eq!(replay_frames_into(&target, &insert, &manifests).await.unwrap(), 2);
3775 }
3776
3777 /// The high-water mark reports the peak spill-buffer occupancy for the
3778 /// call, capped at the configured bound even when more frames remain.
3779 #[tokio::test]
3780 async fn high_water_reports_peak_buffered_frames() {
3781 let seam = five_frame_seam();
3782 let target = faulty_target([]); // no faults — pure high-water measurement
3783 let cfg = bp_cfg(BackpressurePolicy::Fail, 3);
3784 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3785 match out {
3786 StreamOutcome::Streamed { backpressure, .. } => {
3787 assert_eq!(backpressure.high_water_frames, 3, "capped at the configured bound");
3788 }
3789 other => panic!("expected Streamed, got {other:?}"),
3790 }
3791 }
3792
3793 // --- R574-T4: explicit RPO knob --------------------------------------
3794
3795 /// First tail ever: no sidecar, so age is unknown and a configured
3796 /// target cannot be breached (there is nothing to measure against).
3797 #[tokio::test]
3798 async fn rpo_first_tail_has_unknown_age_and_no_breach() {
3799 let seam = five_frame_seam();
3800 let target = fresh_target();
3801 let mut cfg = cfg();
3802 cfg.rpo_target = Some(Duration::from_secs(1));
3803 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3804 match out {
3805 StreamOutcome::Streamed { rpo, .. } => {
3806 assert_eq!(rpo.target, Some(Duration::from_secs(1)));
3807 assert_eq!(rpo.watermark_age, None);
3808 assert!(!rpo.breached);
3809 }
3810 other => panic!("expected Streamed, got {other:?}"),
3811 }
3812 }
3813
3814 /// Second tail past the target: the sidecar's stamped write instant is
3815 /// older than `rpo_target`, so the outcome flags a breach — the
3816 /// staleness emission the orchestrator alerts on.
3817 #[tokio::test]
3818 async fn rpo_stale_watermark_past_target_reports_breach() {
3819 let seam = MockWal::new(4096);
3820 seam.append(1, 1, 0xAA);
3821 let target = fresh_target();
3822 let mut cfg = cfg();
3823 // Zero target: any measurable gap between the two tails is a breach.
3824 cfg.rpo_target = Some(Duration::ZERO);
3825 let _ = tail_frames(&seam, &target, &cfg).await.unwrap();
3826 tokio::time::sleep(Duration::from_millis(5)).await;
3827 seam.append(2, 2, 0xBB);
3828 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3829 match out {
3830 StreamOutcome::Streamed { rpo, .. } => {
3831 let age = rpo.watermark_age.expect("age known after first sidecar write");
3832 assert!(age >= Duration::from_millis(5), "age was {age:?}");
3833 assert!(rpo.breached, "zero target must flag any nonzero age");
3834 }
3835 other => panic!("expected Streamed, got {other:?}"),
3836 }
3837 }
3838
3839 /// A generous target with a prompt second tail: age is reported but the
3840 /// bound holds — and with no target at all, `breached` is always false.
3841 #[tokio::test]
3842 async fn rpo_within_target_and_no_target_do_not_breach() {
3843 let seam = MockWal::new(4096);
3844 seam.append(1, 1, 0xAA);
3845 let target = fresh_target();
3846 let mut with_target = cfg();
3847 with_target.rpo_target = Some(Duration::from_secs(3600));
3848 let _ = tail_frames(&seam, &target, &with_target).await.unwrap();
3849
3850 seam.append(2, 2, 0xBB);
3851 let out = tail_frames(&seam, &target, &with_target).await.unwrap();
3852 match out {
3853 StreamOutcome::Streamed { rpo, .. } => {
3854 assert!(rpo.watermark_age.is_some());
3855 assert!(!rpo.breached, "an hour budget can't be blown in-process");
3856 }
3857 other => panic!("expected Streamed, got {other:?}"),
3858 }
3859
3860 // Same staged sidecar, target removed: age still reported, never
3861 // breached (observe-before-you-pick-a-number mode).
3862 seam.append(3, 3, 0xCC);
3863 let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
3864 match out {
3865 StreamOutcome::Streamed { rpo, .. } => {
3866 assert_eq!(rpo.target, None);
3867 assert!(rpo.watermark_age.is_some());
3868 assert!(!rpo.breached);
3869 }
3870 other => panic!("expected Streamed, got {other:?}"),
3871 }
3872 }
3873
3874 /// A pre-T4 two-field sidecar still parses (written_at unknown), and the
3875 /// next write upgrades it to the stamped three-field format.
3876 #[tokio::test]
3877 async fn rpo_legacy_two_field_sidecar_parses_and_upgrades() {
3878 let target = fresh_target();
3879 target
3880 .store
3881 .put(&target.watermark_key(), b"0 1\n".to_vec().into())
3882 .await
3883 .unwrap();
3884 let legacy = read_watermark(&target.store, &target.watermark_key())
3885 .await
3886 .unwrap()
3887 .unwrap();
3888 assert_eq!(legacy.watermark, Watermark { checkpoint_seq: 0, last_frame: 1 });
3889 assert_eq!(legacy.written_at_nanos, None);
3890
3891 // A tail against the legacy sidecar reports unknown age (not a
3892 // breach), uploads the frames, and re-stamps the sidecar.
3893 //
3894 // R858-B19 CHANGED THE OUTCOME HERE, deliberately: this used to assert
3895 // `Streamed` with `first_frame == 2` ("resumes after the legacy
3896 // watermark"). A two-field sidecar records no salt, so nothing says the
3897 // WAL it names is the WAL in front of us — and resuming at frame 2 on
3898 // that basis is the exact inference that spliced two generations. The
3899 // sidecar's RPO semantics (an *unknown* age is not a breach) are what
3900 // this test is about and they are untouched.
3901 let seam = MockWal::new(4096);
3902 seam.append(1, 0, 0xAA);
3903 seam.append(2, 2, 0xBB);
3904 let mut cfg = cfg();
3905 cfg.rpo_target = Some(Duration::ZERO);
3906 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
3907 match out {
3908 StreamOutcome::Restarted { rpo, first_frame, previous_generation, .. } => {
3909 assert_eq!(first_frame, 1, "an unverifiable watermark is re-uploaded, not resumed");
3910 assert_eq!(previous_generation.salt, None);
3911 assert_eq!(rpo.watermark_age, None);
3912 assert!(!rpo.breached, "unknown age is not a breach even at zero target");
3913 }
3914 other => panic!("expected Restarted, got {other:?}"),
3915 }
3916 let upgraded = read_watermark(&target.store, &target.watermark_key())
3917 .await
3918 .unwrap()
3919 .unwrap();
3920 assert!(upgraded.written_at_nanos.is_some(), "rewrite stamps the timestamp");
3921 }
3922
3923 /// Pure rpo_status edge cases that don't need a staged store.
3924 #[test]
3925 fn rpo_status_truth_table() {
3926 let stamped = |nanos_ago: u128| PersistedWatermark {
3927 watermark: Watermark { checkpoint_seq: 0, last_frame: 1 },
3928 generation: WalGeneration::default(),
3929 written_at_nanos: Some(unix_nanos().saturating_sub(nanos_ago)),
3930 epoch: 0,
3931 pointer_generation: 0,
3932 version: None,
3933 };
3934 // No prior at all.
3935 let s = rpo_status(Some(Duration::from_secs(1)), None);
3936 assert_eq!((s.watermark_age, s.breached), (None, false));
3937 // Prior without a stamp (legacy sidecar).
3938 let legacy = PersistedWatermark {
3939 watermark: Watermark::default(),
3940 generation: WalGeneration::default(),
3941 written_at_nanos: None,
3942 epoch: 0,
3943 pointer_generation: 0,
3944 version: None,
3945 };
3946 let s = rpo_status(Some(Duration::ZERO), Some(&legacy));
3947 assert_eq!((s.watermark_age, s.breached), (None, false));
3948 // Old stamp vs tight target: breached.
3949 let s = rpo_status(Some(Duration::from_millis(1)), Some(&stamped(5_000_000_000)));
3950 assert!(s.watermark_age.unwrap() >= Duration::from_secs(4));
3951 assert!(s.breached);
3952 // Old stamp, no target: age known, never breached.
3953 let s = rpo_status(None, Some(&stamped(5_000_000_000)));
3954 assert!(s.watermark_age.is_some());
3955 assert!(!s.breached);
3956 }
3957
3958 /// R761-T1: the shipped cadence default is an arithmetic consequence of
3959 /// the 2026-08-13 `tail_sweep_harness` table, so the arithmetic is a
3960 /// test rather than a claim in a comment nobody re-checks. If a future
3961 /// harness run moves the measured points, this test is where the
3962 /// mismatch surfaces — re-derive the default, do not relax the test.
3963 #[test]
3964 fn default_tail_cadence_matches_the_measured_write_op_curve() {
3965 // Every uploading tail_frames call writes exactly two fixed objects
3966 // beyond the frames: one generation manifest, one watermark CAS.
3967 const FIXED_OBJECTS_PER_TAIL: f64 = 2.0;
3968 // Measured, flat across the whole sweep — a property of the schema
3969 // and transaction shape, not of the cadence.
3970 const MEASURED_FRAMES_PER_WRITE: f64 = 2.03;
3971 let puts_per_write =
3972 |w: f64| MEASURED_FRAMES_PER_WRITE + FIXED_OBJECTS_PER_TAIL / w;
3973
3974 // The model reproduces all four measured rows.
3975 for (writes_per_tail, measured) in [(1.0, 4.03), (5.0, 2.43), (25.0, 2.11), (100.0, 2.05)] {
3976 let modelled = puts_per_write(writes_per_tail);
3977 assert!(
3978 (modelled - measured).abs() < 0.005,
3979 "writes_per_tail={writes_per_tail}: model {modelled:.3} vs measured {measured:.3}"
3980 );
3981 }
3982
3983 // The default is tailed at half the stated bound, so one missed tick
3984 // still lands inside the promise.
3985 assert_eq!(DEFAULT_RPO_TARGET, DEFAULT_TAIL_INTERVAL * 2);
3986
3987 // At the burst rates the default was chosen against (~0.1-1 write/s
3988 // during an active session), 60s lands in the flat part of the curve:
3989 // the fixed-object term is under a fifth of the frame floor even at
3990 // the slow end, where 15s would still be paying 1.33.
3991 let slow_burst_writes_per_tail = 0.1 * DEFAULT_TAIL_INTERVAL.as_secs_f64();
3992 let fixed_term = FIXED_OBJECTS_PER_TAIL / slow_burst_writes_per_tail;
3993 assert!(
3994 fixed_term < MEASURED_FRAMES_PER_WRITE / 5.0,
3995 "fixed-object term {fixed_term:.3} at the slow-burst end is no longer small \
3996 relative to the {MEASURED_FRAMES_PER_WRITE} frame floor"
3997 );
3998 assert!(
3999 puts_per_write(slow_burst_writes_per_tail) < 2.4,
4000 "the default's worst modelled case should sit below the measured \
4001 writes_per_tail=5 point (2.43)"
4002 );
4003
4004 // R761-F2 removed the frames/write term: a tail call whose frames fit
4005 // one batch writes the batch, the manifest and the watermark, full
4006 // stop. Same curve shape, no floor.
4007 let batched_puts_per_write = |w: f64| (1.0 + FIXED_OBJECTS_PER_TAIL) / w;
4008 for (writes_per_tail, pre_batching) in [(1.0, 4.03), (5.0, 2.43), (25.0, 2.11), (100.0, 2.05)]
4009 {
4010 let cut = 1.0 - batched_puts_per_write(writes_per_tail) / pre_batching;
4011 let want = match writes_per_tail as u32 {
4012 1 => 0.256,
4013 5 => 0.753,
4014 25 => 0.943,
4015 _ => 0.985,
4016 };
4017 assert!(
4018 (cut - want).abs() < 0.005,
4019 "writes_per_tail={writes_per_tail}: batching cuts {:.1}%, expected {:.1}%",
4020 cut * 100.0,
4021 want * 100.0,
4022 );
4023 }
4024 // Which is why the cadence default did not move: what is left to win
4025 // past 60s is now a hundredth of a PUT per write at the fast-burst end
4026 // of the same band, against an RPO window that would grow 5x.
4027 assert!(
4028 batched_puts_per_write(1.0 * DEFAULT_TAIL_INTERVAL.as_secs_f64())
4029 - batched_puts_per_write(5.0 * DEFAULT_TAIL_INTERVAL.as_secs_f64())
4030 < 0.05,
4031 "batching should have flattened the cadence lever at the fast-burst end"
4032 );
4033 }
4034
4035 /// R782: `StreamOutcome::rpo()` is the accessor `tenant-streamer` pushes
4036 /// through — every variant that carries an `RpoStatus` gives it back, and
4037 /// `Fenced` (which deliberately carries none, per its own doc) gives
4038 /// `None` rather than a default/synthesized one.
4039 #[test]
4040 fn stream_outcome_rpo_accessor_covers_every_variant() {
4041 let rpo = RpoStatus { target: None, watermark_age: Some(Duration::from_secs(1)), breached: false };
4042 assert_eq!(StreamOutcome::Empty { watermark: Watermark::default(), rpo }.rpo(), Some(&rpo));
4043 assert_eq!(
4044 StreamOutcome::Streamed {
4045 generation_key: String::new(),
4046 first_frame: 1,
4047 last_frame: 1,
4048 checkpoint_seq: 0,
4049 frame_count: 1,
4050 backpressure: Default::default(),
4051 rpo,
4052 }
4053 .rpo(),
4054 Some(&rpo)
4055 );
4056 assert_eq!(
4057 StreamOutcome::Restarted {
4058 generation_key: String::new(),
4059 previous_generation: WalGeneration::default(),
4060 new_generation: WalGeneration {
4061 checkpoint_seq: 1,
4062 salt: Some(WalSalt { salt1: 1, salt2: 2 }),
4063 },
4064 first_frame: 1,
4065 last_frame: 1,
4066 frame_count: 1,
4067 backpressure: Default::default(),
4068 rpo,
4069 }
4070 .rpo(),
4071 Some(&rpo)
4072 );
4073 assert_eq!(
4074 StreamOutcome::Shed {
4075 checkpoint_seq: 0,
4076 first_frame: 1,
4077 last_frame: 1,
4078 backpressure: Default::default(),
4079 rpo,
4080 }
4081 .rpo(),
4082 Some(&rpo)
4083 );
4084 assert_eq!(
4085 StreamOutcome::Fenced {
4086 current_epoch: 2,
4087 our_epoch: 1,
4088 current_pointer_generation: 0,
4089 our_pointer_generation: 0,
4090 }
4091 .rpo(),
4092 None
4093 );
4094 }
4095
4096 /// First tail with no frames yet — Empty, no manifest, no watermark.
4097 #[tokio::test]
4098 async fn empty_when_no_frames() {
4099 let seam = MockWal::new(4096);
4100 let target = fresh_target();
4101 let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4102 match out {
4103 StreamOutcome::Empty { watermark, rpo } => {
4104 assert_eq!(watermark, Watermark::default());
4105 // No sidecar has ever been written — age is unknowable.
4106 assert_eq!(rpo, RpoStatus::default());
4107 }
4108 other => panic!("expected Empty, got {other:?}"),
4109 }
4110 // No manifest, no watermark sidecar.
4111 assert!(
4112 read_watermark(&target.store, &target.watermark_key())
4113 .await
4114 .unwrap()
4115 .is_none()
4116 );
4117 }
4118
4119 /// First tail with 3 frames — Streamed, manifest written, watermark
4120 /// recorded, every frame object retrievable at the right key.
4121 #[tokio::test]
4122 async fn streams_initial_frames_and_records_watermark() {
4123 let seam = MockWal::new(4096);
4124 seam.append(1, 0, 0xAA);
4125 seam.append(2, 0, 0xBB);
4126 seam.append(3, 3, 0xCC); // commit frame: db_size = 3 pages
4127
4128 let target = fresh_target();
4129 let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4130 let gen_key = match out {
4131 StreamOutcome::Streamed {
4132 generation_key,
4133 first_frame,
4134 last_frame,
4135 checkpoint_seq,
4136 frame_count,
4137 ..
4138 } => {
4139 assert_eq!(first_frame, 1);
4140 assert_eq!(last_frame, 3);
4141 assert_eq!(checkpoint_seq, 0);
4142 assert_eq!(frame_count, 3);
4143 assert!(generation_key.starts_with("backups/generations/gen-"));
4144 generation_key
4145 }
4146 other => panic!("expected Streamed, got {other:?}"),
4147 };
4148
4149 // R761-F2: ONE batch object holds all three frames, at the ranged key,
4150 // each frame at its own offset inside it.
4151 let frame_size = WAL_FRAME_HEADER_SIZE + 4096;
4152 let bytes = target
4153 .store
4154 .get(&target.frame_batch_key(0, 0, 1, 3))
4155 .await
4156 .unwrap()
4157 .bytes()
4158 .await
4159 .unwrap();
4160 assert_eq!(bytes.len(), 3 * frame_size);
4161 for frame_no in 1..=3u64 {
4162 let at = (frame_no as usize - 1) * frame_size;
4163 let page_no = u32::from_be_bytes(bytes[at..at + 4].try_into().unwrap());
4164 assert_eq!(page_no, frame_no as u32);
4165 }
4166 // And nothing was written per-frame.
4167 assert!(
4168 target.store.get(&target.frame_key(0, 0, 1)).await.is_err(),
4169 "the pre-R761-F2 per-frame key must not be written any more"
4170 );
4171
4172 // Manifest is parseable and round-trips.
4173 let m_bytes = target
4174 .store
4175 .get(&ObjPath::from(gen_key.clone()))
4176 .await
4177 .unwrap()
4178 .bytes()
4179 .await
4180 .unwrap();
4181 let parsed = parse_generation_manifest(&String::from_utf8_lossy(&m_bytes)).unwrap();
4182 assert_eq!(parsed.first_frame, 1);
4183 assert_eq!(parsed.last_frame, 3);
4184 assert_eq!(parsed.checkpoint_seq, 0);
4185 assert_eq!(parsed.page_size, 4096);
4186 assert_eq!(parsed.base_snapshot_key, cfg().base_snapshot_key);
4187 assert_eq!(parsed.frame_batches, vec![(1, 3)], "R761-F2: one batch, indexed");
4188
4189 // Watermark sidecar matches.
4190 let wm = read_watermark(&target.store, &target.watermark_key())
4191 .await
4192 .unwrap()
4193 .unwrap();
4194 assert_eq!(wm.watermark, Watermark { checkpoint_seq: 0, last_frame: 3 });
4195 assert!(wm.written_at_nanos.is_some(), "R574-T4: writes stamp the sidecar");
4196
4197 // The seam was told to take WAL ownership? Not directly via
4198 // tail_frames — callers do that at construction (CoreWalSeam::open).
4199 // Verify the mock invariant separately.
4200 assert!(!seam.state.borrow().auto_actions_disabled);
4201 }
4202
4203 /// Second tail with no new frames returns Empty without uploading.
4204 #[tokio::test]
4205 async fn second_tail_with_no_new_frames_is_empty() {
4206 let seam = MockWal::new(4096);
4207 seam.append(1, 1, 0xAA);
4208 let target = fresh_target();
4209 let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4210
4211 let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4212 match out {
4213 StreamOutcome::Empty { watermark, rpo } => {
4214 assert_eq!(watermark, Watermark { checkpoint_seq: 0, last_frame: 1 });
4215 // A sidecar exists from the first tail, so age is known even
4216 // with no rpo_target configured — and no target means no breach.
4217 assert!(rpo.watermark_age.is_some());
4218 assert!(!rpo.breached);
4219 }
4220 other => panic!("expected Empty, got {other:?}"),
4221 }
4222 }
4223
4224 /// Second tail with new frames uploads only the new ones and the
4225 /// generation manifest covers the incremental range. Previous frames
4226 /// remain at their prior keys (idempotent put on same key).
4227 #[tokio::test]
4228 async fn second_tail_uploads_only_new_frames() {
4229 let seam = MockWal::new(4096);
4230 seam.append(1, 0, 0x11);
4231 seam.append(2, 2, 0x22); // commit
4232 let target = fresh_target();
4233 let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4234
4235 // Append more.
4236 seam.append(3, 0, 0x33);
4237 seam.append(4, 4, 0x44); // commit
4238
4239 let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4240 match out {
4241 StreamOutcome::Streamed {
4242 first_frame,
4243 last_frame,
4244 frame_count,
4245 ..
4246 } => {
4247 assert_eq!(first_frame, 3);
4248 assert_eq!(last_frame, 4);
4249 assert_eq!(frame_count, 2);
4250 }
4251 other => panic!("expected Streamed, got {other:?}"),
4252 }
4253 assert_eq!(seam.frame_count(), 4);
4254 // Frames 1..=4 all retrievable under checkpoint_seq=0, as one batch
4255 // object per tail call — the first call's object is untouched by the
4256 // second (disjoint ranges, so no overwrite and no re-upload).
4257 for (first, last) in [(1u64, 2u64), (3, 4)] {
4258 target
4259 .store
4260 .get(&target.frame_batch_key(0, 0, first, last))
4261 .await
4262 .unwrap();
4263 }
4264 }
4265
4266 /// WAL restart bumps checkpoint_seq; the next tail uploads under the new
4267 /// sequence and reports Restarted.
4268 #[tokio::test]
4269 async fn wal_restart_emits_restarted_under_new_sequence() {
4270 let seam = MockWal::new(4096);
4271 seam.append(1, 1, 0xAA);
4272 seam.append(2, 2, 0xBB);
4273 let target = fresh_target();
4274 let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4275
4276 // Simulate the engine restarting the WAL (e.g. after a checkpoint).
4277 seam.restart();
4278 seam.append(1, 1, 0xCC);
4279
4280 let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4281 match out {
4282 StreamOutcome::Restarted {
4283 previous_generation,
4284 new_generation,
4285 first_frame,
4286 last_frame,
4287 frame_count,
4288 ..
4289 } => {
4290 assert_eq!(previous_generation.checkpoint_seq, 0);
4291 assert_eq!(new_generation.checkpoint_seq, 1);
4292 // R858-B19: an in-process restart moves the salt too, so this
4293 // regime is now caught twice over.
4294 assert_ne!(previous_generation.salt, new_generation.salt);
4295 assert_eq!(first_frame, 1);
4296 assert_eq!(last_frame, 1);
4297 assert_eq!(frame_count, 1);
4298 }
4299 other => panic!("expected Restarted, got {other:?}"),
4300 }
4301
4302 // Old seq=0 frames still in place; new seq=1 frame under its own
4303 // sequence — same key shape, different directory.
4304 target
4305 .store
4306 .get(&target.frame_batch_key(0, 0, 1, 2))
4307 .await
4308 .unwrap();
4309 target
4310 .store
4311 .get(&target.frame_batch_key(0, 1, 1, 1))
4312 .await
4313 .unwrap();
4314 }
4315
4316 /// R858-B19, the property this whole ticket turns on: **a WAL whose salt
4317 /// changed is a different WAL, even when `checkpoint_seq` did not move.**
4318 ///
4319 /// This is the writer-process-restart regime. The last connection closed,
4320 /// SQLite checkpointed and deleted the `-wal`, and the next writer built a
4321 /// fresh WAL back at checkpoint-sequence 0. Both tails therefore see
4322 /// `checkpoint_seq == 0`, which is precisely why the old
4323 /// `p.checkpoint_seq == current.checkpoint_seq` test resumed at
4324 /// `last_frame + 1` and spliced frames from two unrelated WALs into one
4325 /// generation chain — a restore that reported success and produced a stale
4326 /// or malformed image, with nothing raised anywhere.
4327 ///
4328 /// The probe (`examples/foreign_checkpoint_probe.rs -- b f`) drives the
4329 /// same fold end-to-end against the real system `sqlite3`. This test exists
4330 /// so the property is pinned here too: the probe needs an upstream sqlite3
4331 /// binary and half a minute, and a property this load-bearing should fail
4332 /// in `cargo test` when someone re-derives "the sequence is enough".
4333 #[tokio::test]
4334 async fn a_salt_change_is_a_restart_even_when_checkpoint_seq_is_unchanged() {
4335 let seam = MockWal::new(4096);
4336 seam.append(1, 1, 0xAA);
4337 seam.append(2, 2, 0xBB);
4338 seam.append(3, 3, 0xCC);
4339 let target = fresh_target();
4340 let first = tail_frames(&seam, &target, &cfg()).await.unwrap();
4341 assert!(matches!(first, StreamOutcome::Streamed { .. }), "got {first:?}");
4342
4343 // The foreign writer restarted: brand-new WAL, unrelated salt, and a
4344 // checkpoint-sequence that reads 0 both before and after.
4345 seam.recreate(WalSalt { salt1: 0x9ca8_9e29, salt2: 0x2be5_66fc });
4346 seam.append(1, 1, 0xDD);
4347 seam.append(2, 2, 0xEE);
4348 assert_eq!(
4349 seam.wal_state().unwrap().checkpoint_seq,
4350 0,
4351 "the premise: the sequence did NOT move across the recreate"
4352 );
4353
4354 let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4355 let StreamOutcome::Restarted {
4356 previous_generation,
4357 new_generation,
4358 first_frame,
4359 last_frame,
4360 ..
4361 } = &out
4362 else {
4363 panic!(
4364 "a recreated WAL must report Restarted — got {out:?}. If this is Streamed with \
4365 first_frame 4, the salt check has been removed and two WAL generations are being \
4366 spliced into one chain again (R858-B19)"
4367 );
4368 };
4369 assert_eq!(previous_generation.checkpoint_seq, new_generation.checkpoint_seq);
4370 assert_ne!(
4371 previous_generation.salt, new_generation.salt,
4372 "the salt is the only thing that moved, and it is what must be noticed"
4373 );
4374 assert_eq!((*first_frame, *last_frame), (1, 2), "re-upload from the top of the new WAL");
4375
4376 // And the chain that results is refused rather than restored: two
4377 // generations both starting at frame 1 is not a stream, and restore
4378 // must say so instead of producing a plausible-looking wrong image.
4379 let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
4380 assert_eq!(manifests.len(), 2);
4381 let err = validate_generation_chain(&manifests).unwrap_err().to_string();
4382 assert!(
4383 err.contains("WAL") || err.contains("gap in stream"),
4384 "restore must refuse the post-recreate chain loudly, got: {err}"
4385 );
4386 }
4387
4388 /// R858-B19 — the narrower door: a WAL recreate that lands **inside** one
4389 /// `tail_frames` call rather than between two. The pre-drain sample says
4390 /// generation A, the frames that got read are a mix of A and B, and no
4391 /// comparison of two watermarks can see it. The call must publish nothing.
4392 #[tokio::test]
4393 async fn a_wal_recreated_mid_drain_publishes_nothing() {
4394 let seam = MockWal::new(4096);
4395 for n in 1..=3u32 {
4396 seam.append(n, n, n as u8);
4397 }
4398 let target = fresh_target();
4399 // Read #1 is the pre-drain salt probe; reads #2..#4 are the drain. Flip
4400 // the WAL underneath us on the second frame of the drain.
4401 seam.recreate_after_reads(3);
4402
4403 let err = tail_frames(&seam, &target, &cfg()).await.unwrap_err().to_string();
4404 assert!(err.contains("recreated while this tail was uploading"), "got: {err}");
4405
4406 // The sink is untouched: no watermark to resume from, no manifest to
4407 // restore. Whatever frame objects landed are orphaned and unreferenced.
4408 assert!(read_watermark(&target.store, &target.watermark_key()).await.unwrap().is_none());
4409 assert!(list_and_parse_generation_manifests(&target).await.unwrap().is_empty());
4410 }
4411
4412 /// R858-B19 — the sidecar half of the same property. A watermark persisted
4413 /// by a pre-salt writer says nothing about which WAL it names, so the next
4414 /// tail must restart rather than resume from `last_frame + 1`. Unknown is
4415 /// not "unchanged".
4416 #[tokio::test]
4417 async fn a_saltless_legacy_watermark_forces_a_restart_rather_than_a_resume() {
4418 let target = fresh_target();
4419 // Exactly what a pre-R858-B19 writer left behind: five positional
4420 // fields, no salt pair.
4421 target
4422 .store
4423 .put(&target.watermark_key(), format!("0 2 {} 0 0\n", unix_nanos()).into_bytes().into())
4424 .await
4425 .unwrap();
4426
4427 let seam = MockWal::new(4096);
4428 seam.append(1, 1, 0xAA);
4429 seam.append(2, 2, 0xBB);
4430 seam.append(3, 3, 0xCC);
4431
4432 let out = tail_frames(&seam, &target, &cfg()).await.unwrap();
4433 let StreamOutcome::Restarted { previous_generation, first_frame, .. } = &out else {
4434 panic!(
4435 "an unverifiable watermark must restart, not resume — got {out:?}. Resuming here \
4436 trusts a position in a WAL nobody can show is the same one (R858-B19)"
4437 );
4438 };
4439 assert_eq!(previous_generation.salt, None, "the prior generation is unknown, not equal");
4440 assert_eq!(*first_frame, 1, "re-upload everything rather than trust the old position");
4441
4442 // Self-healing: the sidecar now carries a salt, so the very next tail
4443 // resumes normally. One redundant re-upload, not a permanent restart
4444 // loop.
4445 seam.append(4, 4, 0xDD);
4446 let next = tail_frames(&seam, &target, &cfg()).await.unwrap();
4447 let StreamOutcome::Streamed { first_frame, .. } = &next else {
4448 panic!("the salt is recorded now, so this must resume — got {next:?}");
4449 };
4450 assert_eq!(*first_frame, 4);
4451 }
4452
4453 /// R858-B19 — the restore-side refusal, on the shape a pre-`v4` writer
4454 /// could actually leave in a bucket: a chain that looks perfectly
4455 /// contiguous but whose manifests never recorded which WAL they came from.
4456 /// `checkpoint_seq` agrees, the frame ranges tile, and it is still
4457 /// unrestorable, because that is exactly what a splice looks like.
4458 #[test]
4459 fn validate_chain_refuses_a_multi_generation_chain_with_no_recorded_salt() {
4460 let err = validate_generation_chain(&[
4461 mk_manifest_salted("base.db", 4096, 0, None, 1, 5),
4462 mk_manifest_salted("base.db", 4096, 0, None, 6, 9),
4463 ])
4464 .unwrap_err()
4465 .to_string();
4466 assert!(err.contains("no WAL salt"), "got: {err}");
4467
4468 // One generation is not a splice — a legacy single-manifest backup
4469 // stays restorable, because there is nothing here to prove.
4470 let chain =
4471 validate_generation_chain(&[mk_manifest_salted("base.db", 4096, 0, None, 1, 5)])
4472 .unwrap();
4473 assert_eq!(chain.generation.salt, None);
4474 assert_eq!(chain.total_frames, 5);
4475 }
4476
4477 /// R858-B19 — and the same refusal when the salts are recorded and
4478 /// *disagree*: a contiguous frame range across two different WALs. This is
4479 /// the case `checkpoint_seq` can never catch, since a recreate resets it to
4480 /// the value it already had.
4481 #[test]
4482 fn validate_chain_refuses_a_chain_whose_salt_changes_mid_stream() {
4483 let other = WalSalt { salt1: 0xea10_2175, salt2: 0xb4ff_221f };
4484 let err = validate_generation_chain(&[
4485 mk_manifest_salted("base.db", 4096, 0, Some(CHAIN_SALT), 1, 30),
4486 mk_manifest_salted("base.db", 4096, 0, Some(other), 31, 57),
4487 ])
4488 .unwrap_err()
4489 .to_string();
4490 assert!(err.contains("RECREATED mid-stream"), "got: {err}");
4491 }
4492
4493 /// R858-B19 — manifest `v4` carries the salt, and the version guards move
4494 /// with it: a `v4` header without a `wal_salt` is corrupt (not legacy), and
4495 /// a pre-`v4` header *with* one is forged (not a newer writer).
4496 #[test]
4497 fn manifest_v4_round_trips_the_wal_salt_and_guards_both_directions() {
4498 let batches = [(12u64, 34u64)];
4499 let salt = WalSalt { salt1: 0xb83c_03f5, salt2: 0x0000_0001 };
4500 let text = format_generation_manifest(GenerationManifest {
4501 base_snapshot_key: "backups/snapshots/snapshot-1.db",
4502 page_size: 4096,
4503 checkpoint_seq: 7,
4504 salt: Some(salt),
4505 first_frame: 12,
4506 last_frame: 34,
4507 epoch: 9,
4508 owner: Some("node-3"),
4509 frame_batches: &batches,
4510 });
4511 assert!(text.starts_with(MANIFEST_HEADER_V4), "got: {text}");
4512 let parsed = parse_generation_manifest(&text).unwrap();
4513 assert_eq!(parsed.salt, Some(salt));
4514 assert_eq!(parsed.checkpoint_seq, 7);
4515 assert_eq!(parsed.frame_batches, batches);
4516
4517 let no_salt = text.replace(&format!("wal_salt {}-{}\n", salt.salt1, salt.salt2), "");
4518 let err = parse_generation_manifest(&no_salt).unwrap_err().to_string();
4519 assert!(err.contains("missing `wal_salt`"), "got: {err}");
4520
4521 let forged = text.replace(MANIFEST_HEADER_V4, MANIFEST_HEADER_V3);
4522 let err = parse_generation_manifest(&forged).unwrap_err().to_string();
4523 assert!(err.contains("corrupt or hand-edited"), "got: {err}");
4524 }
4525
4526 /// R858-B19 — the sidecar's salt pair is read as a pair. Half a pair is a
4527 /// torn write, and downgrading it to "unknown" would hide that.
4528 #[tokio::test]
4529 async fn watermark_sidecar_round_trips_the_salt_and_rejects_half_a_pair() {
4530 let target = fresh_target();
4531 let key = target.watermark_key();
4532 let salt = WalSalt { salt1: 0xd492_ea8a, salt2: 0x7acf_42a3 };
4533 write_watermark(
4534 &target.store,
4535 &key,
4536 Watermark { checkpoint_seq: 2, last_frame: 11 },
4537 Some(salt),
4538 6,
4539 3,
4540 None,
4541 )
4542 .await
4543 .unwrap();
4544 let read = read_watermark(&target.store, &key).await.unwrap().unwrap();
4545 assert_eq!(read.generation, WalGeneration { checkpoint_seq: 2, salt: Some(salt) });
4546
4547 target
4548 .store
4549 .put(&key, format!("2 11 {} 6 3 {}\n", unix_nanos(), salt.salt1).into_bytes().into())
4550 .await
4551 .unwrap();
4552 let err = read_watermark(&target.store, &key).await.unwrap_err().to_string();
4553 assert!(err.contains("salt1 but no salt2"), "got: {err}");
4554 }
4555
4556 /// Captured frame insert: a mock [`WalInsertSeam`] records every call so
4557 /// tests can assert ordering and content without a real turso connection.
4558 struct MockInsertSeam {
4559 log: RefCell<Vec<MockInsertEvent>>,
4560 }
4561 #[derive(Debug, Clone, PartialEq, Eq)]
4562 enum MockInsertEvent {
4563 Begin,
4564 Frame { frame_no: u64, page_no: u32, db_size: u32 },
4565 End { force_commit: bool },
4566 }
4567 impl MockInsertSeam {
4568 fn new() -> Self {
4569 Self { log: RefCell::new(Vec::new()) }
4570 }
4571 fn events(&self) -> Vec<MockInsertEvent> {
4572 self.log.borrow().clone()
4573 }
4574 }
4575 impl WalInsertSeam for MockInsertSeam {
4576 fn wal_insert_begin(&self) -> Result<()> {
4577 self.log.borrow_mut().push(MockInsertEvent::Begin);
4578 Ok(())
4579 }
4580 fn wal_insert_frame(&self, frame_no: u64, frame: &[u8]) -> Result<()> {
4581 let page_no = u32::from_be_bytes(frame[0..4].try_into().unwrap());
4582 let db_size = u32::from_be_bytes(frame[4..8].try_into().unwrap());
4583 self.log.borrow_mut().push(MockInsertEvent::Frame { frame_no, page_no, db_size });
4584 Ok(())
4585 }
4586 fn wal_insert_end(&self, force_commit: bool) -> Result<()> {
4587 self.log.borrow_mut().push(MockInsertEvent::End { force_commit });
4588 Ok(())
4589 }
4590 }
4591
4592 fn mk_manifest(
4593 base: &str,
4594 page_size: usize,
4595 checkpoint_seq: u32,
4596 first_frame: u64,
4597 last_frame: u64,
4598 ) -> OwnedGenerationManifest {
4599 mk_manifest_salted(
4600 base,
4601 page_size,
4602 checkpoint_seq,
4603 Some(CHAIN_SALT),
4604 first_frame,
4605 last_frame,
4606 )
4607 }
4608
4609 /// R858-B19: the salt every `mk_manifest` fixture shares, so a chain built
4610 /// from them is one generation unless a test deliberately says otherwise.
4611 const CHAIN_SALT: WalSalt = WalSalt { salt1: 0x1109_ca5e, salt2: 0x7acf_42a3 };
4612
4613 /// R858-B19: `mk_manifest` with the generation salt spelled out — `None`
4614 /// reproduces a manifest written by a pre-`v4` writer.
4615 fn mk_manifest_salted(
4616 base: &str,
4617 page_size: usize,
4618 checkpoint_seq: u32,
4619 salt: Option<WalSalt>,
4620 first_frame: u64,
4621 last_frame: u64,
4622 ) -> OwnedGenerationManifest {
4623 OwnedGenerationManifest {
4624 base_snapshot_key: base.to_string(),
4625 page_size,
4626 checkpoint_seq,
4627 salt,
4628 first_frame,
4629 last_frame,
4630 epoch: 0,
4631 owner: None,
4632 // Chain validation is about frame ranges, not object layout, so
4633 // these fixtures stay on the pre-R761-F2 shape — which also keeps
4634 // them exercising the legacy read path.
4635 frame_batches: Vec::new(),
4636 }
4637 }
4638
4639 /// validate_generation_chain accepts a single well-formed manifest.
4640 #[test]
4641 fn validate_chain_accepts_single_generation() {
4642 let chain = validate_generation_chain(&[mk_manifest("base.db", 4096, 0, 1, 5)]).unwrap();
4643 assert_eq!(chain.base_snapshot_key, "base.db");
4644 assert_eq!(chain.page_size, 4096);
4645 assert_eq!(chain.generation.checkpoint_seq, 0);
4646 assert_eq!(chain.total_frames, 5);
4647 }
4648
4649 /// validate_generation_chain accepts a contiguous multi-generation chain
4650 /// and sums the frame count across generations.
4651 #[test]
4652 fn validate_chain_accepts_contiguous_multi_generation() {
4653 let chain = validate_generation_chain(&[
4654 mk_manifest("base.db", 4096, 3, 1, 5),
4655 mk_manifest("base.db", 4096, 3, 6, 9),
4656 mk_manifest("base.db", 4096, 3, 10, 12),
4657 ])
4658 .unwrap();
4659 assert_eq!(chain.generation.checkpoint_seq, 3);
4660 assert_eq!(chain.total_frames, 12);
4661 }
4662
4663 /// Empty chain is the "no generations under prefix" signal.
4664 #[test]
4665 fn validate_chain_rejects_empty() {
4666 assert!(validate_generation_chain(&[]).is_err());
4667 }
4668
4669 /// A gap between manifest ranges is corruption — fail loudly.
4670 #[test]
4671 fn validate_chain_rejects_gap() {
4672 let err = validate_generation_chain(&[
4673 mk_manifest("base.db", 4096, 0, 1, 5),
4674 mk_manifest("base.db", 4096, 0, 7, 9), // missing frame 6
4675 ])
4676 .unwrap_err();
4677 assert!(format!("{err}").contains("gap in stream"), "err was {err}");
4678 }
4679
4680 /// A chain that does not start at frame 1 means the base+chain don't match
4681 /// (some early frames were never uploaded, or the chain was truncated by a
4682 /// retention sweep). Refuse.
4683 #[test]
4684 fn validate_chain_rejects_non_one_start() {
4685 let err = validate_generation_chain(&[mk_manifest("base.db", 4096, 0, 5, 9)]).unwrap_err();
4686 assert!(format!("{err}").contains("gap in stream"), "err was {err}");
4687 }
4688
4689 /// A WAL restart between generations means the source folded the WAL into
4690 /// main; the post-restart frames do not replay onto a pre-restart base.
4691 /// Refuse and direct the caller at a fresh tier-1a snapshot.
4692 #[test]
4693 fn validate_chain_rejects_restart() {
4694 let err = validate_generation_chain(&[
4695 mk_manifest("base.db", 4096, 0, 1, 5),
4696 mk_manifest("base.db", 4096, 1, 1, 3), // restart, new seq
4697 ])
4698 .unwrap_err();
4699 let s = format!("{err}");
4700 assert!(s.contains("WAL restart"), "err was {s}");
4701 assert!(s.contains("fresh tier-1a snapshot"), "err was {s}");
4702 }
4703
4704 /// Generations referencing different bases mean the chain mixed sinks /
4705 /// the base was rotated mid-stream. Refuse.
4706 #[test]
4707 fn validate_chain_rejects_base_mismatch() {
4708 let err = validate_generation_chain(&[
4709 mk_manifest("base-A.db", 4096, 0, 1, 5),
4710 mk_manifest("base-B.db", 4096, 0, 6, 9),
4711 ])
4712 .unwrap_err();
4713 assert!(format!("{err}").contains("chain spans bases"), "err was {err}");
4714 }
4715
4716 /// Page-size mismatch across manifests is a corrupt-manifest signal.
4717 #[test]
4718 fn validate_chain_rejects_page_size_mismatch() {
4719 let err = validate_generation_chain(&[
4720 mk_manifest("base.db", 4096, 0, 1, 5),
4721 mk_manifest("base.db", 8192, 0, 6, 9),
4722 ])
4723 .unwrap_err();
4724 assert!(format!("{err}").contains("page_size"), "err was {err}");
4725 }
4726
4727 /// replay_frames_into walks every manifest's range in (seq, frame_no) order
4728 /// and pushes frames into the insert seam. The mock seam records the call
4729 /// log so we can verify ordering and frame content.
4730 #[tokio::test]
4731 async fn replay_walks_manifests_in_frame_order() {
4732 // Stage two generations' worth of frames into the object store via
4733 // tail_frames, then point replay at the resulting manifests.
4734 let seam = MockWal::new(4096);
4735 seam.append(1, 0, 0xAA);
4736 seam.append(2, 2, 0xBB); // commit
4737 let target = fresh_target();
4738 let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4739 seam.append(3, 0, 0xCC);
4740 seam.append(4, 4, 0xDD); // commit
4741 let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4742
4743 let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
4744 let insert = MockInsertSeam::new();
4745 let total = replay_frames_into(&target, &insert, &manifests)
4746 .await
4747 .unwrap();
4748 assert_eq!(total, 4);
4749
4750 let events = insert.events();
4751 assert_eq!(events.len(), 4, "begin/end are the caller's responsibility");
4752 for (i, ev) in events.iter().enumerate() {
4753 let want_frame_no = (i + 1) as u64;
4754 let want_page = want_frame_no as u32;
4755 let want_db_size = if want_frame_no.is_multiple_of(2) { want_frame_no as u32 } else { 0 };
4756 assert_eq!(
4757 ev,
4758 &MockInsertEvent::Frame {
4759 frame_no: want_frame_no,
4760 page_no: want_page,
4761 db_size: want_db_size,
4762 }
4763 );
4764 }
4765 }
4766
4767 /// R761-F2 read compatibility, and the reason `frame_batches` being empty
4768 /// has to MEAN something rather than merely be absent: a sink written by a
4769 /// pre-batching writer — per-frame objects, a v2 manifest — still replays,
4770 /// and a chain that straddles the change replays as one stream. Written
4771 /// here by hand at the object level, because the writer that produced this
4772 /// shape no longer exists to produce it.
4773 #[tokio::test]
4774 async fn a_chain_straddling_the_batching_change_replays_as_one_stream() {
4775 let target = fresh_target();
4776 let frame_size = WAL_FRAME_HEADER_SIZE + 4096;
4777 let frame = |frame_no: u64, page_no: u32, db_size: u32| {
4778 let mut b = vec![0u8; frame_size];
4779 b[0..4].copy_from_slice(&page_no.to_be_bytes());
4780 b[4..8].copy_from_slice(&db_size.to_be_bytes());
4781 b[WAL_FRAME_HEADER_SIZE..].fill(frame_no as u8);
4782 b
4783 };
4784
4785 // Generation 1: the old layout — one object per frame, v2 manifest
4786 // with no batch list.
4787 for frame_no in 1..=2u64 {
4788 target
4789 .store
4790 .put(
4791 &target.frame_key(0, 0, frame_no),
4792 frame(frame_no, frame_no as u32, frame_no as u32).into(),
4793 )
4794 .await
4795 .unwrap();
4796 }
4797 let legacy = format_generation_manifest(GenerationManifest {
4798 base_snapshot_key: "b.db",
4799 page_size: 4096,
4800 checkpoint_seq: 0,
4801 salt: None,
4802 first_frame: 1,
4803 last_frame: 2,
4804 epoch: 0,
4805 owner: None,
4806 frame_batches: &[],
4807 });
4808 assert!(legacy.starts_with("TURSO-BACKUP STREAM v2\n"), "{legacy}");
4809 target
4810 .store
4811 .put(&target.generation_key(1), legacy.into_bytes().into())
4812 .await
4813 .unwrap();
4814
4815 // Generation 2: the new layout, batched, appended to the same chain.
4816 let mut batch = frame(3, 3, 0);
4817 batch.extend_from_slice(&frame(4, 4, 4));
4818 target
4819 .store
4820 .put(&target.frame_batch_key(0, 0, 3, 4), batch.into())
4821 .await
4822 .unwrap();
4823 let batched = format_generation_manifest(GenerationManifest {
4824 base_snapshot_key: "b.db",
4825 page_size: 4096,
4826 checkpoint_seq: 0,
4827 salt: None,
4828 first_frame: 3,
4829 last_frame: 4,
4830 epoch: 0,
4831 owner: None,
4832 frame_batches: &[(3, 4)],
4833 });
4834 target
4835 .store
4836 .put(&target.generation_key(2), batched.into_bytes().into())
4837 .await
4838 .unwrap();
4839
4840 let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
4841 assert!(manifests[0].frame_batches.is_empty(), "gen 1 is the legacy layout");
4842 assert_eq!(manifests[1].frame_batches, vec![(3, 4)]);
4843
4844 // R858-B19 CHANGED THIS ASSERTION, deliberately. It used to read
4845 // `validate_generation_chain(&manifests).expect("layout is not a chain
4846 // property")`. Layout still is not a chain property — that claim is
4847 // what the replay assertion below tests, and it still holds. What
4848 // changed is PROVENANCE: this fixture is a two-generation chain in
4849 // which neither manifest records which WAL its frames came from,
4850 // because both predate the `wal_salt` field. That is byte-for-byte the
4851 // shape a WAL recreate produced (probe F: contiguous ranges, one
4852 // checkpoint_seq, two different WALs), so restore can no longer tell
4853 // this chain from a spliced one and must refuse rather than hand back a
4854 // plausible wrong image. A chain this old needs a fresh tier-1a
4855 // snapshot; one generation of it would still restore.
4856 let err = validate_generation_chain(&manifests).unwrap_err().to_string();
4857 assert!(err.contains("no WAL salt"), "got: {err}");
4858
4859 let insert = MockInsertSeam::new();
4860 assert_eq!(replay_frames_into(&target, &insert, &manifests).await.unwrap(), 4);
4861 assert_eq!(
4862 insert.events(),
4863 (1..=4u64)
4864 .map(|n| MockInsertEvent::Frame {
4865 frame_no: n,
4866 page_no: n as u32,
4867 db_size: if n == 3 { 0 } else { n as u32 },
4868 })
4869 .collect::<Vec<_>>(),
4870 "both layouts deliver the same frames in the same order"
4871 );
4872 }
4873
4874 /// restore_latest_stream errors loudly when there are no manifests under
4875 /// the prefix — the caller should be using tier-1a restore instead.
4876 #[tokio::test]
4877 async fn restore_errors_when_no_generations() {
4878 let target = fresh_target();
4879 let dest = std::env::temp_dir().join(format!(
4880 "turso-backup-restore-empty-{}-{}.db",
4881 std::process::id(),
4882 unix_nanos()
4883 ));
4884 let err = restore_latest_stream(&target, dest.to_str().unwrap())
4885 .await
4886 .unwrap_err();
4887 assert!(format!("{err}").contains("no generation manifests"), "err was {err}");
4888 // The dest file is never written when validation fails up front.
4889 assert!(!dest.exists());
4890 }
4891
4892 /// A frame uploaded with a wrong byte length corrupts the chain; restore's
4893 /// replay path catches it before touching the destination's WAL.
4894 #[tokio::test]
4895 async fn replay_rejects_wrong_size_frame() {
4896 let seam = MockWal::new(4096);
4897 seam.append(1, 1, 0xAA);
4898 let target = fresh_target();
4899 let _ = tail_frames(&seam, &target, &cfg()).await.unwrap();
4900
4901 // Overwrite the one uploaded frame object with garbage of the wrong
4902 // length.
4903 target
4904 .store
4905 .put(&target.frame_batch_key(0, 0, 1, 1), b"too short".to_vec().into())
4906 .await
4907 .unwrap();
4908
4909 let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
4910 let insert = MockInsertSeam::new();
4911 let err = replay_frames_into(&target, &insert, &manifests)
4912 .await
4913 .unwrap_err();
4914 assert!(format!("{err}").contains("expected"), "err was {err}");
4915 }
4916
4917 /// Live-DB end-to-end fixture: a temp db path that cleans itself up
4918 /// (including `-wal` / `-shm` sidecars).
4919 struct TempDb(std::path::PathBuf);
4920 impl TempDb {
4921 fn new(tag: &str) -> Self {
4922 TempDb(std::env::temp_dir().join(format!(
4923 "turso-backup-stream-{tag}-{}-{}.db",
4924 std::process::id(),
4925 unix_nanos(),
4926 )))
4927 }
4928 fn path(&self) -> &str {
4929 self.0.to_str().unwrap()
4930 }
4931 }
4932 impl Drop for TempDb {
4933 fn drop(&mut self) {
4934 for sfx in ["", "-wal", "-shm"] {
4935 let _ = std::fs::remove_file(format!("{}{sfx}", self.0.display()));
4936 }
4937 }
4938 }
4939
4940 async fn seed_rows(path: &str, start: i64, count: i64) {
4941 let db = turso::Builder::new_local(path).build().await.unwrap();
4942 let conn = db.connect().unwrap();
4943 conn.execute("CREATE TABLE IF NOT EXISTS t (id INTEGER PRIMARY KEY, v TEXT)", ())
4944 .await
4945 .unwrap();
4946 conn.execute("BEGIN", ()).await.unwrap();
4947 for i in start..start + count {
4948 conn.execute("INSERT INTO t (id, v) VALUES (?, ?)", (i, format!("v{i}")))
4949 .await
4950 .unwrap();
4951 }
4952 conn.execute("COMMIT", ()).await.unwrap();
4953 }
4954
4955 async fn checkpoint_truncate(path: &str) {
4956 let db = turso::Builder::new_local(path).build().await.unwrap();
4957 let conn = db.connect().unwrap();
4958 let mut rows = conn.query("PRAGMA wal_checkpoint(TRUNCATE)", ()).await.unwrap();
4959 while rows.next().await.unwrap().is_some() {}
4960 }
4961
4962 async fn count_rows(path: &str) -> i64 {
4963 let db = turso::Builder::new_local(path).build().await.unwrap();
4964 let conn = db.connect().unwrap();
4965 let mut r = conn.query("SELECT COUNT(*) FROM t", ()).await.unwrap();
4966 let row = r.next().await.unwrap().unwrap();
4967 row.get::<i64>(0).unwrap()
4968 }
4969
4970 /// End-to-end against a real turso DB and the real `CoreWalSeam`:
4971 /// seed → snapshot → checkpoint (resets WAL, bumps checkpoint_seq) →
4972 /// append more rows (creates frames under the new seq) → tail → restore →
4973 /// row count on the dest matches the post-append source.
4974 ///
4975 /// This is the live-DB ping-pong R005-F2's handoff named: it exercises
4976 /// CoreWalSeam's `wal_state` / `wal_get_frame` AND the new
4977 /// `wal_insert_begin` / `wal_insert_frame` / `wal_insert_end` impls, plus
4978 /// the full manifest discovery + replay pipeline.
4979 #[tokio::test]
4980 async fn live_db_seed_snapshot_tail_restore_round_trips() {
4981 let src = TempDb::new("src");
4982 let dest = TempDb::new("dest");
4983
4984 // Seed 50 rows. These end up in the base snapshot via VACUUM INTO.
4985 seed_rows(src.path(), 0, 50).await;
4986
4987 let target = BackupTarget {
4988 store: Arc::new(InMemory::new()),
4989 prefix: "backups".into(),
4990 };
4991 let base_key = match crate::snapshot::snapshot_and_upload(src.path(), &target)
4992 .await
4993 .unwrap()
4994 {
4995 crate::snapshot::SnapshotOutcome::Uploaded { key, .. } => key,
4996 other => panic!("expected Uploaded base snapshot, got {other:?}"),
4997 };
4998
4999 // Checkpoint to fold prior WAL into main and reset the WAL header —
5000 // any subsequent writes land in fresh frames under a new
5001 // checkpoint_seq. This mirrors how a real orchestrator would hand off
5002 // from tier-1a (full snapshot) into tier-2 streaming.
5003 checkpoint_truncate(src.path()).await;
5004
5005 // Append 25 more rows — these become the WAL frames we stream.
5006 seed_rows(src.path(), 1000, 25).await;
5007
5008 // Tail the new frames into the sink via the real CoreWalSeam.
5009 {
5010 let seam = CoreWalSeam::open(src.path()).unwrap();
5011 let cfg = StreamConfig {
5012 base_snapshot_key: &base_key,
5013 page_size: 4096,
5014 backpressure: BackpressureConfig::default(),
5015 rpo_target: None,
5016 epoch: 0,
5017 owner: None,
5018 pointer_generation: 0,
5019 };
5020 let outcome = tail_frames(&seam, &target, &cfg).await.unwrap();
5021 match outcome {
5022 StreamOutcome::Streamed { frame_count, .. } => {
5023 assert!(frame_count > 0, "expected at least one frame uploaded");
5024 }
5025 StreamOutcome::Restarted { frame_count, .. } => {
5026 // Acceptable when this is the first tail after a checkpoint:
5027 // there's a prior implicit watermark via the WAL-restart bit.
5028 assert!(frame_count > 0);
5029 }
5030 other => panic!("expected Streamed/Restarted, got {other:?}"),
5031 }
5032 } // drop seam (and its turso_core connection) before restore opens its own.
5033
5034 // Restore into a fresh dest. Replay must produce the post-append state.
5035 let outcome = restore_latest_stream(&target, dest.path()).await.unwrap();
5036 assert_eq!(outcome.base_snapshot_key, base_key);
5037 assert!(outcome.frames_replayed > 0);
5038 assert!(outcome.generation_count >= 1);
5039
5040 // The restored DB is openable via vanilla turso and has the right rows.
5041 // (snapshot::restore_latest's tests already cover the vanilla-sqlite3
5042 // exit ramp; here the contract is "turso reopens" since wal_insert_*
5043 // is the only consumer side that knows the engine's frame layout.)
5044 let restored_rows = count_rows(dest.path()).await;
5045 assert_eq!(restored_rows, 75, "expected 50 (base) + 25 (replayed) = 75");
5046 }
5047
5048 /// Crash-consistency: the last frames captured by `tail_frames` are
5049 /// uncommitted (mid-transaction, `db_size == 0`). Restore must drop that
5050 /// suffix and end up at the previous commit's state — `wal_insert_end`
5051 /// with `force_commit = false` is the engine knob that does this.
5052 ///
5053 /// We can't easily force turso to commit-then-leak-a-partial-write in a
5054 /// unit test, so we drive the seam directly: a single `MockInsertSeam`-
5055 /// recorded restore would show the begin/end protocol, but to verify the
5056 /// engine actually truncates we need the real `CoreWalSeam`. The
5057 /// `manifest_with_uncommitted_tail_rolls_back_via_insert_end` test below
5058 /// stages frame bytes in the store by hand and aims them at a real DB.
5059 #[tokio::test]
5060 async fn manifest_with_uncommitted_tail_rolls_back_via_insert_end() {
5061 // Seed two rows via real turso so we have committed page 1 + page 2.
5062 // Then capture the WAL frames from the source. Then stage a *fake*
5063 // extra frame whose db_size = 0 (mid-transaction) in the sink, and
5064 // append it to the manifest range. Restore should drop that fake
5065 // frame and leave the dest at the 2-row state.
5066 let src = TempDb::new("ccsrc");
5067 let dest = TempDb::new("ccdest");
5068 seed_rows(src.path(), 0, 2).await;
5069 let target = BackupTarget {
5070 store: Arc::new(InMemory::new()),
5071 prefix: "backups".into(),
5072 };
5073 let base_key = match crate::snapshot::snapshot_and_upload(src.path(), &target)
5074 .await
5075 .unwrap()
5076 {
5077 crate::snapshot::SnapshotOutcome::Uploaded { key, .. } => key,
5078 other => panic!("expected Uploaded, got {other:?}"),
5079 };
5080 checkpoint_truncate(src.path()).await;
5081 seed_rows(src.path(), 100, 1).await; // one more row → at least one commit frame.
5082
5083 let (checkpoint_seq, last_committed_frame, frame_size) = {
5084 let seam = CoreWalSeam::open(src.path()).unwrap();
5085 let cfg = StreamConfig {
5086 base_snapshot_key: &base_key,
5087 page_size: 4096,
5088 backpressure: BackpressureConfig::default(),
5089 rpo_target: None,
5090 epoch: 0,
5091 owner: None,
5092 pointer_generation: 0,
5093 };
5094 let _ = tail_frames(&seam, &target, &cfg).await.unwrap();
5095 let w = seam.wal_state().unwrap();
5096 (w.checkpoint_seq, w.last_frame, WAL_FRAME_HEADER_SIZE + 4096)
5097 };
5098
5099 // Manually upload one extra "uncommitted" frame: copy the last
5100 // committed frame's bytes but zero db_size in the header. This
5101 // simulates tail_frames having captured a mid-transaction tail.
5102 // R761-F2: the frames live in one batch object, so take the last
5103 // frame's slice out of it rather than fetching a per-frame key.
5104 let batch_key = target.frame_batch_key(0, checkpoint_seq, 1, last_committed_frame);
5105 let batch = target
5106 .store
5107 .get(&batch_key)
5108 .await
5109 .unwrap()
5110 .bytes()
5111 .await
5112 .unwrap();
5113 let at = (last_committed_frame as usize - 1) * frame_size;
5114 let mut tail_bytes = batch[at..at + frame_size].to_vec();
5115 // Zero the big-endian db_size at offset 4..8 to mark this as non-commit.
5116 tail_bytes[4..8].copy_from_slice(&0u32.to_be_bytes());
5117 // Repad to ensure exact length — paranoia.
5118 assert_eq!(tail_bytes.len(), frame_size);
5119 let phantom_frame_no = last_committed_frame + 1;
5120 target
5121 .store
5122 .put(
5123 &target.frame_batch_key(0, checkpoint_seq, phantom_frame_no, phantom_frame_no),
5124 tail_bytes.into(),
5125 )
5126 .await
5127 .unwrap();
5128
5129 // Extend the latest generation manifest to claim the phantom frame.
5130 let mut keys: Vec<_> = target
5131 .store
5132 .list_with_delimiter(Some(&join_key(&target.prefix, "generations")))
5133 .await
5134 .unwrap()
5135 .objects
5136 .into_iter()
5137 .map(|o| o.location)
5138 .collect();
5139 keys.sort();
5140 let last_manifest_key = keys.last().unwrap().clone();
5141 let bytes = target
5142 .store
5143 .get(&last_manifest_key)
5144 .await
5145 .unwrap()
5146 .bytes()
5147 .await
5148 .unwrap();
5149 let mut m = parse_generation_manifest(&String::from_utf8_lossy(&bytes)).unwrap();
5150 m.last_frame = phantom_frame_no;
5151 // The phantom frame went up as its own single-frame batch object, so
5152 // the manifest's index has to name it too — the range alone is no
5153 // longer enough to find a frame (R761-F2).
5154 m.frame_batches.push((phantom_frame_no, phantom_frame_no));
5155 let new_text = format_generation_manifest(GenerationManifest {
5156 base_snapshot_key: &m.base_snapshot_key,
5157 page_size: m.page_size,
5158 checkpoint_seq: m.checkpoint_seq,
5159 salt: None,
5160 first_frame: m.first_frame,
5161 last_frame: m.last_frame,
5162 epoch: m.epoch,
5163 owner: m.owner.as_deref(),
5164 frame_batches: &m.frame_batches,
5165 });
5166 target
5167 .store
5168 .put(&last_manifest_key, new_text.into_bytes().into())
5169 .await
5170 .unwrap();
5171
5172 // Restore. The phantom uncommitted frame must be dropped by
5173 // wal_insert_end(false); the row count reflects the last *commit*.
5174 let _ = restore_latest_stream(&target, dest.path()).await.unwrap();
5175 let restored = count_rows(dest.path()).await;
5176 assert_eq!(restored, 3, "expected 2 (base) + 1 (committed) — phantom rolled back");
5177 }
5178
5179 /// `replay_wal_onto_main` with an empty WAL returns the main bytes
5180 /// untouched — the main file alone is the committed image.
5181 #[test]
5182 fn replay_empty_wal_returns_main_unchanged() {
5183 let seam = MockWal::new(4096);
5184 let mut main = vec![0u8; 4096 * 3];
5185 main[0..16].copy_from_slice(b"SQLite format 3\0");
5186 let img = replay_wal_onto_main(&seam, main.clone(), 4096).unwrap();
5187 assert_eq!(img, main);
5188 }
5189
5190 /// A single commit frame updates its page slot and the final image is
5191 /// sized to the commit's `db_size`. Unrelated pages in `main` are left
5192 /// in place.
5193 #[test]
5194 fn replay_single_commit_frame_applies_page() {
5195 let seam = MockWal::new(4096);
5196 // Frame 1: page 2, commit at db_size = 3 pages.
5197 seam.append(2, 3, 0xCC);
5198 let mut main = vec![0u8; 4096 * 3];
5199 main[0..16].copy_from_slice(b"SQLite format 3\0");
5200
5201 let img = replay_wal_onto_main(&seam, main, 4096).unwrap();
5202 assert_eq!(img.len(), 4096 * 3);
5203 // Page 1 untouched (still has the magic + zeros).
5204 assert!(img.starts_with(b"SQLite format 3\0"));
5205 // Page 2 = 0xCC fill from the WAL frame.
5206 assert!(img[4096..4096 * 2].iter().all(|&b| b == 0xCC));
5207 // Page 3 untouched.
5208 assert!(img[4096 * 2..].iter().all(|&b| b == 0));
5209 }
5210
5211 /// Uncommitted frames past the last commit are dropped and the image is
5212 /// truncated to the last commit's `db_size` — the crash-consistency story
5213 /// matching restore's `wal_insert_end(false)`.
5214 #[test]
5215 fn replay_drops_uncommitted_tail() {
5216 let seam = MockWal::new(4096);
5217 seam.append(1, 0, 0x11); // mid-txn
5218 seam.append(2, 2, 0x22); // commit at db_size = 2
5219 seam.append(3, 0, 0x33); // uncommitted (dropped)
5220 seam.append(4, 0, 0x44); // uncommitted (dropped)
5221
5222 let main = vec![0u8; 4096 * 4]; // pre-grown so a buggy replay would keep junk
5223 let img = replay_wal_onto_main(&seam, main, 4096).unwrap();
5224 assert_eq!(img.len(), 4096 * 2, "image must be truncated to db_size=2");
5225 assert!(img[0..4096].iter().all(|&b| b == 0x11));
5226 assert!(img[4096..].iter().all(|&b| b == 0x22));
5227 }
5228
5229 /// Multi-commit chain: the full commit prefix is applied and the image is
5230 /// grown to the final commit's `db_size`. Intermediate uncommitted frames
5231 /// between commits are also applied (they are part of the committed
5232 /// suffix once a later frame in the same batch becomes a commit).
5233 #[test]
5234 fn replay_applies_full_commit_prefix_and_grows_image() {
5235 let seam = MockWal::new(4096);
5236 seam.append(1, 0, 0xAA);
5237 seam.append(2, 2, 0xBB); // first commit
5238 seam.append(3, 0, 0xCC);
5239 seam.append(4, 4, 0xDD); // second commit
5240
5241 let main = vec![0u8; 4096 * 2];
5242 let img = replay_wal_onto_main(&seam, main, 4096).unwrap();
5243 assert_eq!(img.len(), 4096 * 4, "image grown to db_size=4");
5244 assert!(img[0..4096].iter().all(|&b| b == 0xAA));
5245 assert!(img[4096..4096 * 2].iter().all(|&b| b == 0xBB));
5246 assert!(img[4096 * 2..4096 * 3].iter().all(|&b| b == 0xCC));
5247 assert!(img[4096 * 3..].iter().all(|&b| b == 0xDD));
5248 }
5249
5250 /// WAL has frames but none are committed (writer mid-transaction at the
5251 /// instant we sampled). The main file alone is the image; the uncommitted
5252 /// suffix is dropped wholesale.
5253 #[test]
5254 fn replay_returns_main_when_only_uncommitted_frames_present() {
5255 let seam = MockWal::new(4096);
5256 seam.append(1, 0, 0x11);
5257 seam.append(2, 0, 0x22);
5258 let mut main = vec![0u8; 4096 * 3];
5259 main[0..16].copy_from_slice(b"SQLite format 3\0");
5260 let img = replay_wal_onto_main(&seam, main.clone(), 4096).unwrap();
5261 assert_eq!(img, main);
5262 }
5263
5264 /// End-to-end against a real turso DB and `CoreWalSeam`: seed, fold the
5265 /// first batch into main via TRUNCATE, then seed more rows so the WAL has
5266 /// uncheckpointed committed frames. `raw_consistent_copy_live` must
5267 /// reproduce the full row count without itself calling TRUNCATE — that is
5268 /// the live-writer contract a TRUNCATE-busy concurrent writer would
5269 /// otherwise block.
5270 #[tokio::test]
5271 async fn live_db_consistent_copy_without_truncate_round_trips() {
5272 let src = TempDb::new("live-cc-src");
5273 seed_rows(src.path(), 0, 25).await;
5274 checkpoint_truncate(src.path()).await; // first batch into main, WAL reset
5275 seed_rows(src.path(), 1000, 10).await; // second batch lives in WAL
5276
5277 // No TRUNCATE here — this is the live-writer path.
5278 let image = raw_consistent_copy_live(src.path(), 4096).await.unwrap();
5279 assert!(image.starts_with(b"SQLite format 3\0"));
5280
5281 // Write the image to a fresh path (no -wal sidecar — the replay
5282 // folded WAL in) and re-open to confirm all 35 rows are present.
5283 let restored = TempDb::new("live-cc-restored");
5284 std::fs::write(restored.path(), &image).unwrap();
5285 let rows = count_rows(restored.path()).await;
5286 assert_eq!(
5287 rows, 35,
5288 "expected 25 (checkpointed) + 10 (replayed from WAL)"
5289 );
5290 }
5291
5292 // ── R858-B18: reading a source nobody will lock for us ────────────────
5293 //
5294 // The property under test throughout: the copy is either a validated
5295 // point in time or a refusal, never a plausible wrong image.
5296
5297 /// The two salt readers must agree. [`read_wal_salt`] pulls it out of frame
5298 /// 1 through the [`WalSeam`] trait (the only path a `from_conn` seam has);
5299 /// [`WalFileHeader::read`] pulls it out of the 32-byte `-wal` header with no
5300 /// engine at all. They are separate code paths against separate byte
5301 /// offsets, and R858-B18 depends on them naming the same generation — if
5302 /// they could disagree, validation would compare a salt the streamer never
5303 /// recorded.
5304 #[tokio::test]
5305 async fn wal_file_header_salt_matches_the_salt_read_through_the_seam() {
5306 let src = TempDb::new("b18-salt-agree");
5307 seed_rows(src.path(), 0, 5).await;
5308
5309 let seam = CoreWalSeam::open_reader(src.path()).unwrap();
5310 let state = seam.wal_state().unwrap();
5311 assert!(state.last_frame > 0, "seed should leave frames in the WAL");
5312 let via_seam = read_wal_salt(&seam, 4096, state.last_frame).unwrap().unwrap();
5313 drop(seam);
5314
5315 let via_file = WalFileHeader::read(src.path()).unwrap().unwrap();
5316 assert_eq!(via_file.salt, via_seam, "frame 1 carries the WAL header's salt verbatim");
5317 assert_eq!(
5318 via_file.checkpoint_seq, state.checkpoint_seq,
5319 "and the on-disk sequence is the one the engine reports"
5320 );
5321 }
5322
5323 /// A `-wal` shorter than one header names no generation — an honest
5324 /// unknown, not an error (same posture as `WalGeneration { salt: None }`).
5325 #[test]
5326 fn wal_file_header_parse_needs_a_whole_header() {
5327 assert!(WalFileHeader::parse(&[0u8; WalFileHeader::SIZE - 1]).is_none());
5328 let mut hdr = [0u8; WalFileHeader::SIZE];
5329 hdr[8..12].copy_from_slice(&4096u32.to_be_bytes());
5330 hdr[12..16].copy_from_slice(&7u32.to_be_bytes());
5331 hdr[16..20].copy_from_slice(&0xb83c_03f5u32.to_be_bytes());
5332 hdr[20..24].copy_from_slice(&0x7acf_42a3u32.to_be_bytes());
5333 let parsed = WalFileHeader::parse(&hdr).unwrap();
5334 assert_eq!(parsed.page_size, 4096);
5335 assert_eq!(parsed.checkpoint_seq, 7);
5336 assert_eq!(parsed.salt, WalSalt { salt1: 0xb83c_03f5, salt2: 0x7acf_42a3 });
5337 }
5338
5339 /// The central asymmetry of [`SourceFingerprint::stable_across`]: a plain
5340 /// append is NOT movement (frames `1..=max_frame` are immutable within a
5341 /// generation, so our image is merely an earlier point in time), while a
5342 /// checkpoint IS (it rewrites the main file our copy already read).
5343 ///
5344 /// Getting this backwards is not a small error in either direction: treat
5345 /// an append as movement and every copy of a database that is actually in
5346 /// use is refused; treat a checkpoint as harmless and we hand back spliced
5347 /// state that passes `integrity_check`.
5348 #[tokio::test]
5349 async fn fingerprint_ignores_an_append_and_catches_a_checkpoint() {
5350 let src = TempDb::new("b18-fingerprint");
5351 seed_rows(src.path(), 0, 5).await;
5352
5353 let before = SourceFingerprint::read(src.path()).unwrap();
5354 seed_rows(src.path(), 100, 5).await;
5355 let appended = SourceFingerprint::read(src.path()).unwrap();
5356 assert!(
5357 appended.wal_len > before.wal_len,
5358 "the append must actually have grown the WAL, or this proves nothing \
5359 (before {}B, after {}B)",
5360 before.wal_len,
5361 appended.wal_len
5362 );
5363 assert!(
5364 before.stable_across(&appended),
5365 "an append is not movement: {} -> {}",
5366 before.describe(),
5367 appended.describe()
5368 );
5369
5370 checkpoint_truncate(src.path()).await;
5371 let folded = SourceFingerprint::read(src.path()).unwrap();
5372 assert!(
5373 !before.stable_across(&folded),
5374 "a checkpoint IS movement and must be caught: {} -> {}",
5375 before.describe(),
5376 folded.describe()
5377 );
5378 }
5379
5380 /// The protocol accepts when the source holds still, and the accepted value
5381 /// is the one the attempt produced.
5382 #[tokio::test]
5383 async fn validation_accepts_a_quiescent_source() {
5384 let src = TempDb::new("b18-quiescent");
5385 seed_rows(src.path(), 0, 3).await;
5386 let got = validated_against_source(src.path(), "test read", || async { Ok(41 + 1) })
5387 .await
5388 .unwrap();
5389 assert_eq!(got, 42);
5390 }
5391
5392 /// A source that moves under every attempt yields a REFUSAL, not a value —
5393 /// and the message names both samples so an operator can see what moved.
5394 #[tokio::test]
5395 async fn validation_refuses_when_the_source_moves_under_every_attempt() {
5396 let src = TempDb::new("b18-moving");
5397 seed_rows(src.path(), 0, 3).await;
5398 let path = src.path().to_string();
5399 let attempts = std::cell::Cell::new(0u32);
5400 let err = validated_against_source(src.path(), "test read", || {
5401 // Grow the main file inside the attempt window, which is exactly
5402 // the shape of a foreign checkpoint folding pages into it.
5403 attempts.set(attempts.get() + 1);
5404 let path = path.clone();
5405 async move {
5406 let mut f = std::fs::OpenOptions::new().append(true).open(&path)?;
5407 std::io::Write::write_all(&mut f, &[0u8; 4096])?;
5408 Ok(())
5409 }
5410 })
5411 .await
5412 .unwrap_err();
5413 assert_eq!(
5414 attempts.get(),
5415 COPY_VALIDATION_ATTEMPTS,
5416 "every attempt in the budget must be spent before refusing"
5417 );
5418 let msg = format!("{err:#}");
5419 assert!(msg.contains("refusing a test read"), "{msg}");
5420 assert!(msg.contains("before and"), "message must name both samples: {msg}");
5421 }
5422
5423 /// An attempt that fails against a source that did NOT move is a real
5424 /// error, not a race — it is returned as itself on the first attempt rather
5425 /// than retried and then reported as concurrency.
5426 #[tokio::test]
5427 async fn validation_surfaces_a_real_error_without_burning_the_budget() {
5428 let src = TempDb::new("b18-real-error");
5429 seed_rows(src.path(), 0, 3).await;
5430 let attempts = std::cell::Cell::new(0u32);
5431 let err = validated_against_source(src.path(), "test read", || {
5432 attempts.set(attempts.get() + 1);
5433 async { Err::<(), _>(anyhow::anyhow!("page 3 checksum mismatch")) }
5434 })
5435 .await
5436 .unwrap_err();
5437 assert_eq!(attempts.get(), 1, "a stable source means retrying cannot help");
5438 let msg = format!("{err:#}");
5439 assert!(msg.contains("did NOT move"), "{msg}");
5440 assert!(msg.contains("page 3 checksum mismatch"), "{msg}");
5441 }
5442
5443 /// The reader open is the one the backup path uses, and it must produce the
5444 /// same watermark the writable open does. (Cross-process non-exclusivity —
5445 /// the point of the flag — is measured in `examples/foreign_checkpoint_probe.rs`
5446 /// probes G2r/G3c, since it needs a second process.)
5447 #[tokio::test]
5448 async fn read_only_seam_reports_the_same_watermark_as_the_writable_one() {
5449 let src = TempDb::new("b18-reader-watermark");
5450 seed_rows(src.path(), 0, 4).await;
5451 let writable = CoreWalSeam::open(src.path()).unwrap().wal_state().unwrap();
5452 let reader = CoreWalSeam::open_reader(src.path()).unwrap().wal_state().unwrap();
5453 assert_eq!(writable, reader);
5454 }
5455
5456 /// A second call to `raw_consistent_copy_live` after more writes captures
5457 /// the new state — the primitive is callable repeatedly without per-call
5458 /// setup, mirroring how a snapshot loop would drive it.
5459 #[tokio::test]
5460 async fn live_db_consistent_copy_reflects_new_writes() {
5461 let src = TempDb::new("live-cc-incr-src");
5462 seed_rows(src.path(), 0, 5).await;
5463 let img1 = raw_consistent_copy_live(src.path(), 4096).await.unwrap();
5464 let r1 = TempDb::new("live-cc-incr-r1");
5465 std::fs::write(r1.path(), &img1).unwrap();
5466 assert_eq!(count_rows(r1.path()).await, 5);
5467
5468 seed_rows(src.path(), 100, 7).await;
5469 let img2 = raw_consistent_copy_live(src.path(), 4096).await.unwrap();
5470 let r2 = TempDb::new("live-cc-incr-r2");
5471 std::fs::write(r2.path(), &img2).unwrap();
5472 assert_eq!(count_rows(r2.path()).await, 12);
5473 }
5474
5475 // ── R732-F2 (W245): tenant fencing epochs on the R2 write path ────────
5476 //
5477 // The property under test throughout: a writer holding a stale fencing
5478 // token is REJECTED, not merely unlucky. Every test below names the
5479 // split-brain it rules out.
5480
5481 fn cfg_at_epoch(epoch: u64, owner: Option<&'static str>) -> StreamConfig<'static> {
5482 StreamConfig { epoch, owner, ..cfg() }
5483 }
5484
5485 /// R736-T2: like `cfg_at_epoch`, but also names the cross-cell pointer
5486 /// generation this writer believes it holds.
5487 fn cfg_at(epoch: u64, pointer_generation: u64, owner: Option<&'static str>) -> StreamConfig<'static> {
5488 StreamConfig { epoch, pointer_generation, owner, ..cfg() }
5489 }
5490
5491 /// THE canonical F2 test, and the reason the whole relay exists: two
5492 /// owners tailing the same tenant. The one at the older epoch bounces and
5493 /// writes nothing; the sink is byte-for-byte what the newer owner left.
5494 #[tokio::test]
5495 async fn a_stale_writer_is_fenced_and_writes_zero_frames() {
5496 let target = fresh_target();
5497
5498 // The real owner (epoch 2) streams three frames.
5499 let winner = MockWal::new(4096);
5500 for i in 1..=3u32 {
5501 winner.append(i, i, i as u8);
5502 }
5503 let out = tail_frames(&winner, &target, &cfg_at_epoch(2, Some("node-2")))
5504 .await
5505 .unwrap();
5506 assert!(matches!(out, StreamOutcome::Streamed { frame_count: 3, .. }));
5507
5508 let manifests_before = list_and_parse_generation_manifests(&target).await.unwrap();
5509 let watermark_before = read_watermark(&target.store, &target.watermark_key())
5510 .await
5511 .unwrap()
5512 .unwrap();
5513
5514 // The partitioned old owner (epoch 1) wakes up with its own frames and
5515 // tails the same sink, unaware it has been transferred away.
5516 let loser = MockWal::new(4096);
5517 for i in 1..=9u32 {
5518 loser.append(i, i, 0xff);
5519 }
5520 let out = tail_frames(&loser, &target, &cfg_at_epoch(1, Some("node-1")))
5521 .await
5522 .unwrap();
5523 assert_eq!(
5524 out,
5525 StreamOutcome::Fenced {
5526 current_epoch: 2,
5527 our_epoch: 1,
5528 current_pointer_generation: 0,
5529 our_pointer_generation: 0,
5530 },
5531 "the stale owner must be told it lost, with both epochs"
5532 );
5533
5534 // Zero side effects: no frames under its own epoch prefix, no extra
5535 // generation, and the watermark still names the winner.
5536 // Nothing at all under the loser's epoch prefix — asserted by listing
5537 // rather than by probing one key, so it holds whatever object layout
5538 // the writer would have used (R761-F2).
5539 let loser_prefix = "backups/frames/00000000000000000001/";
5540 assert!(
5541 !objects_under(&target)
5542 .await
5543 .iter()
5544 .any(|k| k.starts_with(loser_prefix)),
5545 "a fenced writer must not upload a single frame"
5546 );
5547 let manifests_after = list_and_parse_generation_manifests(&target).await.unwrap();
5548 assert_eq!(
5549 manifests_after, manifests_before,
5550 "a fenced writer must not write a generation manifest"
5551 );
5552 let watermark_after = read_watermark(&target.store, &target.watermark_key())
5553 .await
5554 .unwrap()
5555 .unwrap();
5556 assert_eq!(watermark_after.epoch, 2);
5557 assert_eq!(watermark_after.watermark, watermark_before.watermark);
5558 }
5559
5560 /// An UNFENCED (epoch 0) legacy streamer is fenced by a sink that has been
5561 /// claimed. This is the migration case: a node still running the old
5562 /// single-writer configuration must not be allowed to scribble over a
5563 /// tenant that yubaba has since handed to someone else.
5564 #[tokio::test]
5565 async fn an_unfenced_writer_is_fenced_by_a_claimed_sink() {
5566 let target = fresh_target();
5567 let owner = MockWal::new(4096);
5568 owner.append(1, 1, 1);
5569 tail_frames(&owner, &target, &cfg_at_epoch(1, None)).await.unwrap();
5570
5571 let legacy = MockWal::new(4096);
5572 legacy.append(1, 1, 2);
5573 assert_eq!(
5574 tail_frames(&legacy, &target, &cfg_at_epoch(0, None)).await.unwrap(),
5575 StreamOutcome::Fenced {
5576 current_epoch: 1,
5577 our_epoch: 0,
5578 current_pointer_generation: 0,
5579 our_pointer_generation: 0,
5580 }
5581 );
5582 }
5583
5584 // ── R736-T2 (W250): the second, cross-cell fence ───────────────────────
5585 //
5586 // The epoch alone is a *local* raft counter — it cannot see a tenant
5587 // that moved to a different cell's independent raft group. These tests
5588 // prove the pointer generation catches exactly the case the epoch can't:
5589 // a stale cell that is current on its own epoch.
5590
5591 /// THE canonical T2 test: a writer whose epoch is perfectly current for
5592 /// its own (now-stale) cell still bounces, because the global pointer
5593 /// says ownership moved elsewhere. Proves the epoch is necessary but not
5594 /// sufficient — this is the gap W250 exists to close.
5595 #[tokio::test]
5596 async fn a_stale_pointer_generation_fences_even_at_a_current_epoch() {
5597 let target = fresh_target();
5598
5599 // The new cell streams at generation 2, epoch 1 (its own local raft
5600 // is fresh — it just took ownership).
5601 let winner = MockWal::new(4096);
5602 for i in 1..=3u32 {
5603 winner.append(i, i, i as u8);
5604 }
5605 let out = tail_frames(&winner, &target, &cfg_at(1, 2, Some("cell-b/node-1")))
5606 .await
5607 .unwrap();
5608 assert!(matches!(out, StreamOutcome::Streamed { frame_count: 3, .. }));
5609
5610 // The old cell's writer wakes up unaware of the move. Its own local
5611 // epoch (1) is perfectly current for its own raft group — nothing
5612 // local told it to step down — but its pointer generation (1) is
5613 // behind the sink's (2).
5614 let loser = MockWal::new(4096);
5615 for i in 1..=9u32 {
5616 loser.append(i, i, 0xff);
5617 }
5618 let out = tail_frames(&loser, &target, &cfg_at(1, 1, Some("cell-a/node-1")))
5619 .await
5620 .unwrap();
5621 assert_eq!(
5622 out,
5623 StreamOutcome::Fenced {
5624 current_epoch: 1,
5625 our_epoch: 1,
5626 current_pointer_generation: 2,
5627 our_pointer_generation: 1,
5628 },
5629 "an equal, non-stale epoch must not mask a stale pointer generation"
5630 );
5631
5632 // Zero side effects, exactly like the epoch-only fence.
5633 let manifests = list_and_parse_generation_manifests(&target).await.unwrap();
5634 assert_eq!(manifests.len(), 1, "only the winner's generation was written");
5635 }
5636
5637 /// The watermark sidecar carries the pointer generation as a fifth
5638 /// positional field, and a four-field sidecar (pre-R736-T2 writer) reads
5639 /// back generation 0 — unfenced, exactly like a pre-R732 sidecar reads
5640 /// back epoch 0.
5641 #[tokio::test]
5642 async fn watermark_sidecar_round_trips_the_pointer_generation() {
5643 let target = fresh_target();
5644 let key = target.watermark_key();
5645 write_watermark(
5646 &target.store,
5647 &key,
5648 Watermark { checkpoint_seq: 2, last_frame: 11 },
5649 None,
5650 6,
5651 3,
5652 None,
5653 )
5654 .await
5655 .unwrap();
5656 let read = read_watermark(&target.store, &key).await.unwrap().unwrap();
5657 assert_eq!((read.epoch, read.pointer_generation), (6, 3));
5658
5659 // A four-field sidecar (epoch, no generation) — the R732-F2 shape.
5660 target
5661 .store
5662 .put(&key, b"2 11 12345 6\n".to_vec().into())
5663 .await
5664 .unwrap();
5665 let legacy = read_watermark(&target.store, &key).await.unwrap().unwrap();
5666 assert_eq!(legacy.epoch, 6);
5667 assert_eq!(legacy.pointer_generation, 0, "a pre-R736-T2 sidecar fences nobody on generation");
5668 }
5669
5670 /// R869: `read_fence_state` reports exactly the two comparands
5671 /// [`tail_frames`] checks — including the pre-fencing sidecar shapes, which
5672 /// must read back as unfenced rather than erroring, since a rebuild will
5673 /// meet them on any sink written before R732-F2.
5674 #[tokio::test]
5675 async fn read_fence_state_reports_what_tail_frames_would_check() {
5676 let target = fresh_target();
5677 assert_eq!(
5678 read_fence_state(&target).await.unwrap(),
5679 None,
5680 "no sidecar means no fence at all, which is not the same as a zero fence"
5681 );
5682
5683 write_watermark(
5684 &target.store,
5685 &target.watermark_key(),
5686 Watermark {
5687 checkpoint_seq: 2,
5688 last_frame: 11,
5689 },
5690 None,
5691 6,
5692 3,
5693 None,
5694 )
5695 .await
5696 .unwrap();
5697 assert_eq!(
5698 read_fence_state(&target).await.unwrap(),
5699 Some(FenceState {
5700 epoch: 6,
5701 pointer_generation: 3
5702 })
5703 );
5704
5705 // A two-field sidecar — the original pre-fencing shape. Both comparands
5706 // read back 0, so a rebuild sees "unfenced" rather than a parse error.
5707 target
5708 .store
5709 .put(&target.watermark_key(), b"2 11\n".to_vec().into())
5710 .await
5711 .unwrap();
5712 assert_eq!(
5713 read_fence_state(&target).await.unwrap(),
5714 Some(FenceState::default())
5715 );
5716 }
5717
5718 /// The property a rebuild's epoch floor rests on, stated as a test rather
5719 /// than as a comment: seeding one above `read_fence_state().epoch` is
5720 /// exactly enough to stop being fenced, and one *below* it is not.
5721 #[tokio::test]
5722 async fn a_floor_taken_from_read_fence_state_is_what_unfences_a_rebuild() {
5723 let target = fresh_target();
5724 let seam = MockWal::new(4096);
5725 for i in 1..=3u32 {
5726 seam.append(i, i, i as u8);
5727 }
5728 tail_frames(&seam, &target, &cfg_at_epoch(5, Some("dead-fleet")))
5729 .await
5730 .unwrap();
5731
5732 let fence = read_fence_state(&target).await.unwrap().unwrap();
5733 assert_eq!(fence.epoch, 5);
5734
5735 // A rebuilt cluster that restarted its epochs at 1 is refused.
5736 seam.append(4, 4, 4);
5737 let out = tail_frames(&seam, &target, &cfg_at_epoch(1, Some("rebuilt")))
5738 .await
5739 .unwrap();
5740 assert!(matches!(out, StreamOutcome::Fenced { .. }), "got {out:?}");
5741
5742 // Seeded from the fence, it is not.
5743 let out = tail_frames(
5744 &seam,
5745 &target,
5746 &cfg_at_epoch(fence.epoch + 1, Some("rebuilt")),
5747 )
5748 .await
5749 .unwrap();
5750 assert!(matches!(out, StreamOutcome::Streamed { .. }), "got {out:?}");
5751 }
5752
5753 /// The same owner resuming at the same epoch is NOT fenced — fencing is
5754 /// strictly "someone newer exists", not "someone else wrote here". An
5755 /// owner that restarts under an unchanged token must keep streaming, or
5756 /// every process restart would wedge the tenant.
5757 #[tokio::test]
5758 async fn an_equal_epoch_writer_resumes_normally() {
5759 let target = fresh_target();
5760 let seam = MockWal::new(4096);
5761 seam.append(1, 1, 1);
5762 tail_frames(&seam, &target, &cfg_at_epoch(3, None)).await.unwrap();
5763 seam.append(2, 2, 2);
5764 let out = tail_frames(&seam, &target, &cfg_at_epoch(3, None)).await.unwrap();
5765 match out {
5766 StreamOutcome::Streamed { first_frame, last_frame, .. } => {
5767 assert_eq!((first_frame, last_frame), (2, 2), "resumes after the watermark");
5768 }
5769 other => panic!("expected Streamed, got {other:?}"),
5770 }
5771 }
5772
5773 /// Frame keys are namespaced by epoch, and epoch 0 keeps the pre-fencing
5774 /// two-level layout so existing backups stay addressable.
5775 #[test]
5776 fn frame_keys_are_namespaced_by_epoch_with_zero_keeping_the_legacy_layout() {
5777 let target = fresh_target();
5778 assert_eq!(
5779 target.frame_key(0, 7, 42).to_string(),
5780 "backups/frames/0000000007/00000000000000000042",
5781 "epoch 0 must keep the original key shape"
5782 );
5783 assert_eq!(
5784 target.frame_key(5, 7, 42).to_string(),
5785 "backups/frames/00000000000000000005/0000000007/00000000000000000042"
5786 );
5787 assert_ne!(target.frame_key(5, 7, 42), target.frame_key(6, 7, 42));
5788 }
5789
5790 /// R761-F2: a batch key carries its frame range, keeps the epoch
5791 /// namespacing and the zero-padding (so lexical order is still frame
5792 /// order), and cannot be confused with a pre-batching per-frame key even
5793 /// when the batch holds exactly one frame.
5794 #[test]
5795 fn batch_keys_carry_the_range_and_never_collide_with_a_per_frame_key() {
5796 let target = fresh_target();
5797 assert_eq!(
5798 target.frame_batch_key(0, 7, 42, 99).to_string(),
5799 "backups/frames/0000000007/00000000000000000042-00000000000000000099",
5800 "epoch 0 keeps the two-level layout for batches too"
5801 );
5802 assert_eq!(
5803 target.frame_batch_key(5, 7, 42, 99).to_string(),
5804 "backups/frames/00000000000000000005/0000000007/00000000000000000042-00000000000000000099"
5805 );
5806 assert_ne!(
5807 target.frame_batch_key(0, 7, 42, 42),
5808 target.frame_key(0, 7, 42),
5809 "a one-frame batch is still a batch — the two layouts must stay distinguishable"
5810 );
5811 // Lexical order matches frame order for the batches of one stream,
5812 // which never overlap.
5813 assert!(
5814 target.frame_batch_key(0, 7, 1, 8) < target.frame_batch_key(0, 7, 9, 16),
5815 "zero-padding must keep batches lexically ordered by first frame"
5816 );
5817 }
5818
5819 /// A takeover mid-stream leaves both owners' frames intact under their own
5820 /// prefixes — the key namespacing is the backstop behind the epoch check.
5821 #[tokio::test]
5822 async fn a_takeover_writes_under_its_own_epoch_prefix_without_disturbing_the_old_one() {
5823 let target = fresh_target();
5824 let seam = MockWal::new(4096);
5825 seam.append(1, 1, 0xaa);
5826 tail_frames(&seam, &target, &cfg_at_epoch(1, None)).await.unwrap();
5827 seam.append(2, 2, 0xbb);
5828 tail_frames(&seam, &target, &cfg_at_epoch(2, None)).await.unwrap();
5829
5830 let first = target.store.get(&target.frame_batch_key(1, 0, 1, 1)).await.unwrap();
5831 assert_eq!(first.bytes().await.unwrap().len(), WAL_FRAME_HEADER_SIZE + 4096);
5832 target
5833 .store
5834 .get(&target.frame_batch_key(2, 0, 2, 2))
5835 .await
5836 .expect("the new owner's frame lives under its own epoch");
5837 assert!(
5838 target.store.get(&target.frame_batch_key(1, 0, 2, 2)).await.is_err(),
5839 "the new owner must not write into the old owner's prefix"
5840 );
5841 }
5842
5843 /// Restore refuses a chain whose epoch goes backwards: a generation
5844 /// written by an owner that had already been fenced. Replaying it would
5845 /// interleave a stale owner's frames into the live stream.
5846 #[test]
5847 fn validate_chain_refuses_an_epoch_regression() {
5848 let mut newer = mk_manifest("base.db", 4096, 0, 1, 5);
5849 newer.epoch = 4;
5850 let mut stale = mk_manifest("base.db", 4096, 0, 6, 9);
5851 stale.epoch = 3;
5852 let err = validate_generation_chain(&[newer, stale]).unwrap_err();
5853 let msg = format!("{err}");
5854 assert!(msg.contains("epoch 3"), "err was {msg}");
5855 assert!(msg.contains("fenced"), "err must name the cause: {msg}");
5856 }
5857
5858 /// …but a chain that spans an ownership TRANSFER is fine. Epochs may
5859 /// advance mid-stream; only regression is corruption.
5860 #[test]
5861 fn validate_chain_accepts_a_transfer_mid_chain() {
5862 let mut first = mk_manifest("base.db", 4096, 0, 1, 5);
5863 first.epoch = 3;
5864 let mut second = mk_manifest("base.db", 4096, 0, 6, 9);
5865 second.epoch = 4;
5866 let chain = validate_generation_chain(&[first, second]).unwrap();
5867 assert_eq!(chain.total_frames, 9);
5868 assert_eq!(chain.epoch, 4, "the chain reports the most recent owner");
5869 }
5870
5871 /// A pre-fencing chain still validates and reports epoch 0.
5872 #[test]
5873 fn validate_chain_of_pre_fencing_manifests_reports_epoch_zero() {
5874 let chain = validate_generation_chain(&[mk_manifest("base.db", 4096, 0, 1, 5)]).unwrap();
5875 assert_eq!(chain.epoch, 0);
5876 }
5877
5878 /// Manifest v2 round-trips the epoch and the owner label.
5879 #[test]
5880 fn manifest_v2_round_trips_epoch_and_owner() {
5881 let text = format_generation_manifest(GenerationManifest {
5882 base_snapshot_key: "backups/snapshots/snapshot-1.db",
5883 page_size: 4096,
5884 checkpoint_seq: 7,
5885 salt: None,
5886 first_frame: 12,
5887 last_frame: 34,
5888 epoch: 9,
5889 owner: Some("node-3"),
5890 frame_batches: &[],
5891 });
5892 assert!(text.starts_with("TURSO-BACKUP STREAM v2\n"), "{text}");
5893 let parsed = parse_generation_manifest(&text).unwrap();
5894 assert_eq!(parsed.epoch, 9);
5895 assert_eq!(parsed.owner.as_deref(), Some("node-3"));
5896
5897 // No owner label → the key is omitted entirely, not written empty.
5898 let text = format_generation_manifest(GenerationManifest {
5899 base_snapshot_key: "b.db",
5900 page_size: 4096,
5901 checkpoint_seq: 0,
5902 salt: None,
5903 first_frame: 1,
5904 last_frame: 1,
5905 epoch: 1,
5906 owner: None,
5907 frame_batches: &[],
5908 });
5909 assert!(!text.contains("owner"), "{text}");
5910 assert_eq!(parse_generation_manifest(&text).unwrap().owner, None);
5911 }
5912
5913 /// R761-F2: a batch list round-trips, and its presence is what moves the
5914 /// header to v3 — the version and the layout are one fact, so a reader can
5915 /// never see a v3 header without an index or a v2 header with one.
5916 #[test]
5917 fn manifest_v3_round_trips_the_frame_batch_list() {
5918 let batches = [(12u64, 20u64), (21, 34)];
5919 let text = format_generation_manifest(GenerationManifest {
5920 base_snapshot_key: "backups/snapshots/snapshot-1.db",
5921 page_size: 4096,
5922 checkpoint_seq: 7,
5923 salt: None,
5924 first_frame: 12,
5925 last_frame: 34,
5926 epoch: 9,
5927 owner: Some("node-3"),
5928 frame_batches: &batches,
5929 });
5930 assert!(text.starts_with("TURSO-BACKUP STREAM v3\n"), "{text}");
5931 let parsed = parse_generation_manifest(&text).unwrap();
5932 assert_eq!(parsed.frame_batches, batches.to_vec());
5933 assert_eq!(parsed.first_frame, 12);
5934 assert_eq!(parsed.last_frame, 34);
5935 assert_eq!(parsed.owner.as_deref(), Some("node-3"));
5936 }
5937
5938 /// A v3 batch list that does not exactly tile the manifest's own frame
5939 /// range is corruption, and it has to fail at parse: the list IS the frame
5940 /// index, so a hole in it becomes a 404 halfway through a replay — after
5941 /// the destination has already been overwritten with the base snapshot.
5942 #[test]
5943 fn a_v3_manifest_whose_batches_do_not_tile_its_range_is_rejected() {
5944 let head = "TURSO-BACKUP STREAM v3\nbase_snapshot b.db\npage_size 4096\ncheckpoint_seq 0\nepoch 0\n";
5945 for (batches, want, why) in [
5946 ("frame_batch 1-3\nframe_batch 5-9\n", "does not continue", "gap"),
5947 ("frame_batch 1-3\n", "but the manifest claims", "short"),
5948 ("frame_batch 2-9\n", "does not continue", "wrong start"),
5949 ("frame_batch 1-4\nframe_batch 4-9\n", "does not continue", "overlap"),
5950 ("", "no `frame_batch` lines", "missing index"),
5951 ] {
5952 let text = format!("{head}first_frame 1\nlast_frame 9\n{batches}");
5953 let err = parse_generation_manifest(&text).unwrap_err();
5954 assert!(
5955 format!("{err}").contains(want),
5956 "{why}: expected {want:?}, err was {err}"
5957 );
5958 }
5959 }
5960
5961 /// The other direction: a `frame_batch` line under a v1/v2 header is
5962 /// corrupt too. Silently honouring it would let a hand-edited manifest
5963 /// claim a layout its header says it does not have.
5964 #[test]
5965 fn a_pre_v3_manifest_carrying_a_batch_line_is_rejected() {
5966 let bad = "TURSO-BACKUP STREAM v2\nbase_snapshot b.db\npage_size 4096\ncheckpoint_seq 0\nfirst_frame 1\nlast_frame 3\nepoch 0\nframe_batch 1-3\n";
5967 let err = parse_generation_manifest(bad).unwrap_err();
5968 assert!(format!("{err}").contains("corrupt or hand-edited"), "err was {err}");
5969 }
5970
5971 /// A v1 manifest — one written before fencing existed — still parses, as
5972 /// epoch 0. Those backups have to stay restorable.
5973 #[test]
5974 fn a_v1_manifest_parses_as_epoch_zero() {
5975 let legacy = "TURSO-BACKUP STREAM v1\nbase_snapshot b.db\npage_size 4096\ncheckpoint_seq 0\nfirst_frame 1\nlast_frame 3\n";
5976 let parsed = parse_generation_manifest(legacy).unwrap();
5977 assert_eq!(parsed.epoch, 0);
5978 assert_eq!(parsed.owner, None);
5979 assert_eq!(parsed.last_frame, 3);
5980 }
5981
5982 /// A v2 manifest missing its epoch is corrupt, not legacy. Defaulting it
5983 /// to 0 would silently demote a fenced generation to unfenced — the one
5984 /// direction this mechanism must never fail in.
5985 #[test]
5986 fn a_v2_manifest_without_an_epoch_is_rejected() {
5987 let bad = "TURSO-BACKUP STREAM v2\nbase_snapshot b.db\npage_size 4096\ncheckpoint_seq 0\nfirst_frame 1\nlast_frame 3\n";
5988 let err = parse_generation_manifest(bad).unwrap_err();
5989 assert!(format!("{err}").contains("missing `epoch`"), "err was {err}");
5990 }
5991
5992 /// The watermark sidecar carries the epoch as a fourth positional field,
5993 /// and a three-field sidecar (pre-R732 writer) reads back as epoch 0.
5994 #[tokio::test]
5995 async fn watermark_sidecar_round_trips_the_epoch() {
5996 let target = fresh_target();
5997 let key = target.watermark_key();
5998 write_watermark(
5999 &target.store,
6000 &key,
6001 Watermark { checkpoint_seq: 2, last_frame: 11 },
6002 None,
6003 6,
6004 0,
6005 None,
6006 )
6007 .await
6008 .unwrap();
6009 let read = read_watermark(&target.store, &key).await.unwrap().unwrap();
6010 assert_eq!(read.epoch, 6);
6011 assert_eq!(read.watermark.last_frame, 11);
6012 assert!(read.written_at_nanos.is_some());
6013
6014 target
6015 .store
6016 .put(&key, b"2 11 12345\n".to_vec().into())
6017 .await
6018 .unwrap();
6019 let legacy = read_watermark(&target.store, &key).await.unwrap().unwrap();
6020 assert_eq!(legacy.epoch, 0, "a pre-R732 sidecar fences nobody");
6021 assert_eq!(legacy.written_at_nanos, Some(12345));
6022 }
6023
6024 /// End-to-end on a real DB: an ownership transfer happens mid-stream and
6025 /// the restore is clean — every frame from both owners replays, and the
6026 /// outcome reports the winning epoch. This is the "the N+1 writer wins;
6027 /// restore is clean" half of the ticket's ask, run against turso_core
6028 /// rather than a mock.
6029 #[tokio::test]
6030 async fn a_transfer_mid_stream_restores_cleanly_and_reports_the_new_epoch() {
6031 let src = TempDb::new("src-epoch");
6032 let dest = TempDb::new("dest-epoch");
6033 seed_rows(src.path(), 0, 50).await;
6034
6035 let target = fresh_target();
6036 let base_key = match crate::snapshot::snapshot_and_upload(src.path(), &target)
6037 .await
6038 .unwrap()
6039 {
6040 crate::snapshot::SnapshotOutcome::Uploaded { key, .. } => key,
6041 other => panic!("expected Uploaded base snapshot, got {other:?}"),
6042 };
6043 checkpoint_truncate(src.path()).await;
6044
6045 // Owner A (epoch 1) streams the first batch of writes.
6046 seed_rows(src.path(), 1000, 10).await;
6047 {
6048 let seam = CoreWalSeam::open(src.path()).unwrap();
6049 let cfg = StreamConfig {
6050 base_snapshot_key: &base_key,
6051 epoch: 1,
6052 owner: Some("node-a"),
6053 ..cfg()
6054 };
6055 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
6056 assert!(
6057 matches!(
6058 out,
6059 StreamOutcome::Streamed { .. } | StreamOutcome::Restarted { .. }
6060 ),
6061 "owner A should have streamed, got {out:?}"
6062 );
6063 }
6064
6065 // Ownership transfers. Owner B (epoch 2) picks up where A stopped.
6066 seed_rows(src.path(), 2000, 15).await;
6067 {
6068 let seam = CoreWalSeam::open(src.path()).unwrap();
6069 let cfg = StreamConfig {
6070 base_snapshot_key: &base_key,
6071 epoch: 2,
6072 owner: Some("node-b"),
6073 ..cfg()
6074 };
6075 let out = tail_frames(&seam, &target, &cfg).await.unwrap();
6076 assert!(
6077 matches!(
6078 out,
6079 StreamOutcome::Streamed { .. } | StreamOutcome::Restarted { .. }
6080 ),
6081 "owner B should have streamed, got {out:?}"
6082 );
6083 }
6084
6085 let outcome = restore_latest_stream(&target, dest.path()).await.unwrap();
6086 assert_eq!(outcome.base_snapshot_key, base_key);
6087 assert_eq!(outcome.epoch, 2, "restore reports the most recent owner");
6088 assert_eq!(
6089 count_rows(dest.path()).await,
6090 75,
6091 "50 (base) + 10 (owner A) + 15 (owner B)"
6092 );
6093 }
6094
6095 // ── R732-T3 (W245): the watermark advance is a compare-and-swap ───────
6096
6097 /// Bootstrapping a fresh sink uses `PutMode::Create`, so two writers
6098 /// racing to claim a brand-new tenant cannot both succeed. Without this
6099 /// the very first write — the one with no prior version to swap on —
6100 /// would be the one unguarded moment in the whole protocol.
6101 #[tokio::test]
6102 async fn a_second_bootstrap_of_a_fresh_sink_is_contended() {
6103 let target = fresh_target();
6104 let key = target.watermark_key();
6105 let w = Watermark { checkpoint_seq: 0, last_frame: 1 };
6106 assert_eq!(
6107 write_watermark(&target.store, &key, w, None, 1, 0, None).await.unwrap(),
6108 WatermarkCas::Advanced
6109 );
6110 assert_eq!(
6111 write_watermark(&target.store, &key, w, None, 1, 0, None).await.unwrap(),
6112 WatermarkCas::Contended,
6113 "the sink already exists — Create must not silently overwrite it"
6114 );
6115 }
6116
6117 /// An advance is conditional on the version actually read. A writer
6118 /// holding a version somebody else has already replaced loses.
6119 #[tokio::test]
6120 async fn an_advance_on_a_replaced_version_is_contended() {
6121 let target = fresh_target();
6122 let key = target.watermark_key();
6123 write_watermark(
6124 &target.store,
6125 &key,
6126 Watermark { checkpoint_seq: 0, last_frame: 1 },
6127 None,
6128 1,
6129 0,
6130 None,
6131 )
6132 .await
6133 .unwrap();
6134
6135 // Our writer reads the sidecar and holds onto that version…
6136 let stale = read_watermark(&target.store, &key).await.unwrap().unwrap();
6137 // …while somebody else advances it out from under us.
6138 write_watermark(
6139 &target.store,
6140 &key,
6141 Watermark { checkpoint_seq: 0, last_frame: 5 },
6142 None,
6143 2,
6144 0,
6145 read_watermark(&target.store, &key)
6146 .await
6147 .unwrap()
6148 .unwrap()
6149 .version
6150 .as_ref(),
6151 )
6152 .await
6153 .unwrap();
6154
6155 assert_eq!(
6156 write_watermark(
6157 &target.store,
6158 &key,
6159 Watermark { checkpoint_seq: 0, last_frame: 2 },
6160 None,
6161 1,
6162 0,
6163 stale.version.as_ref(),
6164 )
6165 .await
6166 .unwrap(),
6167 WatermarkCas::Contended,
6168 "the stale version must not be allowed to overwrite the newer one"
6169 );
6170 // And the sink still names the winner, not us.
6171 let now = read_watermark(&target.store, &key).await.unwrap().unwrap();
6172 assert_eq!((now.epoch, now.watermark.last_frame), (2, 5));
6173 }
6174
6175 /// Re-reading before advancing works: the point is the version, not the
6176 /// identity of the writer.
6177 #[tokio::test]
6178 async fn an_advance_on_the_current_version_succeeds() {
6179 let target = fresh_target();
6180 let key = target.watermark_key();
6181 write_watermark(
6182 &target.store,
6183 &key,
6184 Watermark { checkpoint_seq: 0, last_frame: 1 },
6185 None,
6186 1,
6187 0,
6188 None,
6189 )
6190 .await
6191 .unwrap();
6192 let cur = read_watermark(&target.store, &key).await.unwrap().unwrap();
6193 assert_eq!(
6194 write_watermark(
6195 &target.store,
6196 &key,
6197 Watermark { checkpoint_seq: 0, last_frame: 7 },
6198 None,
6199 1,
6200 0,
6201 cur.version.as_ref(),
6202 )
6203 .await
6204 .unwrap(),
6205 WatermarkCas::Advanced
6206 );
6207 }
6208
6209 /// A store that lets a test slip a competing writer in between our read
6210 /// of the watermark and our conditional write of it — the interleaving
6211 /// that a sequential epoch check cannot catch and the CAS must.
6212 ///
6213 /// On the first `put_opts` aimed at the watermark key it writes a rival
6214 /// sidecar (at `rival_epoch`) straight through to the inner store, then
6215 /// forwards our conditional put, which now finds a version it does not
6216 /// hold.
6217 struct RacingStore {
6218 inner: Arc<dyn ObjectStore>,
6219 watermark: ObjPath,
6220 rival_epoch: u64,
6221 /// R736-T2: the rival's pointer generation, so the same interleaving
6222 /// can be exercised for the cross-cell fence too.
6223 rival_pointer_generation: u64,
6224 fired: std::sync::atomic::AtomicBool,
6225 }
6226
6227 impl std::fmt::Display for RacingStore {
6228 fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
6229 write!(f, "RacingStore({})", self.inner)
6230 }
6231 }
6232 impl std::fmt::Debug for RacingStore {
6233 fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
6234 write!(f, "RacingStore({:?})", self.inner)
6235 }
6236 }
6237
6238 #[async_trait::async_trait]
6239 impl ObjectStore for RacingStore {
6240 async fn put_opts(
6241 &self,
6242 location: &ObjPath,
6243 payload: object_store::PutPayload,
6244 opts: PutOptions,
6245 ) -> object_store::Result<object_store::PutResult> {
6246 if location == &self.watermark
6247 && !self
6248 .fired
6249 .swap(true, std::sync::atomic::Ordering::SeqCst)
6250 {
6251 let rival =
6252 format!("0 99 1 {} {}\n", self.rival_epoch, self.rival_pointer_generation);
6253 self.inner
6254 .put(location, rival.into_bytes().into())
6255 .await?;
6256 }
6257 self.inner.put_opts(location, payload, opts).await
6258 }
6259
6260 async fn put_multipart_opts(
6261 &self,
6262 location: &ObjPath,
6263 opts: object_store::PutMultipartOptions,
6264 ) -> object_store::Result<Box<dyn object_store::MultipartUpload>> {
6265 self.inner.put_multipart_opts(location, opts).await
6266 }
6267
6268 async fn get_opts(
6269 &self,
6270 location: &ObjPath,
6271 options: object_store::GetOptions,
6272 ) -> object_store::Result<object_store::GetResult> {
6273 self.inner.get_opts(location, options).await
6274 }
6275
6276 fn delete_stream(
6277 &self,
6278 locations: futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>>,
6279 ) -> futures_util::stream::BoxStream<'static, object_store::Result<ObjPath>> {
6280 self.inner.delete_stream(locations)
6281 }
6282
6283 fn list(
6284 &self,
6285 prefix: Option<&ObjPath>,
6286 ) -> futures_util::stream::BoxStream<'static, object_store::Result<object_store::ObjectMeta>>
6287 {
6288 self.inner.list(prefix)
6289 }
6290
6291 async fn list_with_delimiter(
6292 &self,
6293 prefix: Option<&ObjPath>,
6294 ) -> object_store::Result<object_store::ListResult> {
6295 self.inner.list_with_delimiter(prefix).await
6296 }
6297
6298 async fn copy_opts(
6299 &self,
6300 from: &ObjPath,
6301 to: &ObjPath,
6302 options: object_store::CopyOptions,
6303 ) -> object_store::Result<()> {
6304 self.inner.copy_opts(from, to, options).await
6305 }
6306 }
6307
6308 fn racing_target(rival_epoch: u64) -> BackupTarget {
6309 racing_target_at(rival_epoch, 0)
6310 }
6311
6312 /// R736-T2: like `racing_target`, but also names the rival's pointer
6313 /// generation.
6314 fn racing_target_at(rival_epoch: u64, rival_pointer_generation: u64) -> BackupTarget {
6315 let inner: Arc<dyn ObjectStore> = Arc::new(InMemory::new());
6316 let prefix = "backups".to_string();
6317 let watermark = join_key(&prefix, "latest.stream-watermark");
6318 BackupTarget {
6319 store: Arc::new(RacingStore {
6320 inner,
6321 watermark,
6322 rival_epoch,
6323 rival_pointer_generation,
6324 fired: std::sync::atomic::AtomicBool::new(false),
6325 }),
6326 prefix,
6327 }
6328 }
6329
6330 /// The race the sequential check cannot see: a newer owner claims the
6331 /// sink *after* we read it and *before* we write it. The CAS catches it,
6332 /// and we report Fenced without publishing a manifest — which is why the
6333 /// manifest write had to move after the watermark advance.
6334 #[tokio::test]
6335 async fn losing_the_watermark_race_to_a_newer_owner_fences_us_before_we_publish() {
6336 let target = racing_target(9);
6337 let seam = MockWal::new(4096);
6338 for i in 1..=2u32 {
6339 seam.append(i, i, i as u8);
6340 }
6341 let out = tail_frames(&seam, &target, &cfg_at_epoch(4, Some("node-loser")))
6342 .await
6343 .unwrap();
6344 assert_eq!(
6345 out,
6346 StreamOutcome::Fenced {
6347 current_epoch: 9,
6348 our_epoch: 4,
6349 current_pointer_generation: 0,
6350 our_pointer_generation: 0,
6351 }
6352 );
6353
6354 // No manifest published — the chain stays clean, so restore is not
6355 // poisoned by a regressed generation from a writer that lost.
6356 assert!(
6357 list_and_parse_generation_manifests(&target)
6358 .await
6359 .unwrap()
6360 .is_empty(),
6361 "a writer that loses the CAS must publish nothing"
6362 );
6363 // The rival's sidecar survived untouched.
6364 let now = read_watermark(&target.store, &target.watermark_key())
6365 .await
6366 .unwrap()
6367 .unwrap();
6368 assert_eq!(now.epoch, 9);
6369 }
6370
6371 /// R736-T2: the same race, but the rival is a different cell claiming the
6372 /// tenant via a pointer CAS — our epoch is unchanged (we were never told
6373 /// to step down locally) but the generation moved under us mid-write.
6374 #[tokio::test]
6375 async fn losing_the_watermark_race_to_a_cross_cell_move_fences_us_before_we_publish() {
6376 let target = racing_target_at(4, 2);
6377 let seam = MockWal::new(4096);
6378 for i in 1..=2u32 {
6379 seam.append(i, i, i as u8);
6380 }
6381 let out = tail_frames(&seam, &target, &cfg_at(4, 1, Some("cell-a/node-loser")))
6382 .await
6383 .unwrap();
6384 assert_eq!(
6385 out,
6386 StreamOutcome::Fenced {
6387 current_epoch: 4,
6388 our_epoch: 4,
6389 current_pointer_generation: 2,
6390 our_pointer_generation: 1,
6391 },
6392 "an unchanged epoch must not mask a generation that moved under us"
6393 );
6394 assert!(
6395 list_and_parse_generation_manifests(&target)
6396 .await
6397 .unwrap()
6398 .is_empty(),
6399 "a writer that loses the CAS on generation must publish nothing"
6400 );
6401 }
6402
6403 /// Losing the race to a writer at our own epoch is NOT a fencing event —
6404 /// it means two streamers were handed the same token. That is a caller
6405 /// bug and must surface as a loud error, never be retried into success.
6406 #[tokio::test]
6407 async fn losing_the_race_to_an_equal_epoch_writer_is_a_loud_error() {
6408 let target = racing_target(4);
6409 let seam = MockWal::new(4096);
6410 seam.append(1, 1, 1);
6411 let err = tail_frames(&seam, &target, &cfg_at_epoch(4, None))
6412 .await
6413 .unwrap_err();
6414 let msg = format!("{err}");
6415 assert!(
6416 msg.contains("two streamers share one fencing token"),
6417 "err was {msg}"
6418 );
6419 }
6420
6421 /// Generation manifest round-trip and rejection of garbage.
6422 #[test]
6423 fn manifest_round_trip_and_rejection() {
6424 let text = format_generation_manifest(GenerationManifest {
6425 base_snapshot_key: "backups/snapshots/snapshot-1.db",
6426 page_size: 4096,
6427 checkpoint_seq: 7,
6428 salt: None,
6429 first_frame: 12,
6430 last_frame: 34,
6431 epoch: 0,
6432 owner: None,
6433 frame_batches: &[],
6434 });
6435 let parsed = parse_generation_manifest(&text).unwrap();
6436 assert_eq!(parsed.base_snapshot_key, "backups/snapshots/snapshot-1.db");
6437 assert_eq!(parsed.page_size, 4096);
6438 assert_eq!(parsed.checkpoint_seq, 7);
6439 assert_eq!(parsed.first_frame, 12);
6440 assert_eq!(parsed.last_frame, 34);
6441
6442 assert!(parse_generation_manifest("not a manifest").is_err());
6443 assert!(parse_generation_manifest("TURSO-BACKUP STREAM v1\nunknown 1").is_err());
6444 assert!(parse_generation_manifest("TURSO-BACKUP STREAM v1\nbase_snapshot k").is_err()); // missing fields
6445 }
6446
6447 // ---- R850-T3: tier-2 GC (gc_stream) --------------------------------
6448
6449 /// A tier-2 sink built object-by-object on an `InMemory` store.
6450 ///
6451 /// Hand-placed rather than produced by running `tail_frames`, because every
6452 /// one of these tests is about a sink in a state a *live* tailer never
6453 /// leaves behind — a superseded base, a generation's frames with no
6454 /// manifest. Driving the writer could not produce them without also
6455 /// producing the rebase that is the thing under test.
6456 struct GcSink {
6457 target: BackupTarget,
6458 }
6459
6460 impl GcSink {
6461 fn new() -> Self {
6462 GcSink {
6463 target: BackupTarget {
6464 store: Arc::new(InMemory::new()),
6465 prefix: "sink".to_string(),
6466 },
6467 }
6468 }
6469
6470 /// A base snapshot at `nanos`. Body is not a real database — nothing in
6471 /// the GC path parses it.
6472 async fn put_snapshot(&self, nanos: u128) -> String {
6473 let key = join_key(
6474 &self.target.prefix,
6475 &format!("snapshots/snapshot-{nanos:020}.db"),
6476 );
6477 self.target
6478 .store
6479 .put(&key, format!("base {nanos}").into_bytes().into())
6480 .await
6481 .unwrap();
6482 key.to_string()
6483 }
6484
6485 /// One batch object holding frames `first..=last` of generation
6486 /// `(epoch, seq)`, keyed exactly as `tail_frames` would key it.
6487 async fn put_frame_batch(&self, epoch: u64, seq: u32, first: u64, last: u64) -> String {
6488 let key = self.target.frame_batch_key(epoch, seq, first, last);
6489 self.target
6490 .store
6491 .put(&key, vec![0u8; 16].into())
6492 .await
6493 .unwrap();
6494 key.to_string()
6495 }
6496
6497 /// A generation manifest naming `base` and one batch `first..=last`.
6498 async fn put_manifest(
6499 &self,
6500 nanos: u128,
6501 base: &str,
6502 epoch: u64,
6503 seq: u32,
6504 first: u64,
6505 last: u64,
6506 ) {
6507 let batches = [(first, last)];
6508 let text = format_generation_manifest(GenerationManifest {
6509 base_snapshot_key: base,
6510 page_size: 4096,
6511 checkpoint_seq: seq,
6512 salt: Some(WalSalt {
6513 salt1: 1,
6514 salt2: 2,
6515 }),
6516 first_frame: first,
6517 last_frame: last,
6518 epoch,
6519 owner: None,
6520 frame_batches: &batches,
6521 });
6522 self.target
6523 .store
6524 .put(&self.target.generation_key(nanos), text.into_bytes().into())
6525 .await
6526 .unwrap();
6527 }
6528
6529 async fn exists(&self, key: &str) -> bool {
6530 self.target.store.get(&ObjPath::from(key)).await.is_ok()
6531 }
6532 }
6533
6534 /// Collect everything collectable: no grace, really delete.
6535 fn sweep_now() -> StreamGcConfig {
6536 StreamGcConfig {
6537 grace: Duration::ZERO,
6538 dry_run: false,
6539 }
6540 }
6541
6542 /// (a) The base a rebase superseded is reclaimed; the one a restore would
6543 /// pick is not.
6544 #[tokio::test]
6545 async fn gc_collects_a_superseded_base_snapshot_and_keeps_the_current_one() {
6546 let sink = GcSink::new();
6547 let old_base = sink.put_snapshot(1).await;
6548 let new_base = sink.put_snapshot(2).await;
6549 // The post-rebase steady state: the only chain names the new base.
6550 sink.put_manifest(10, &new_base, 0, 5, 1, 10).await;
6551 sink.put_frame_batch(0, 5, 1, 10).await;
6552
6553 let out = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6554
6555 assert_eq!(out.collected_snapshots, vec![old_base.clone()]);
6556 assert_eq!(out.retained_snapshots, 1);
6557 assert_eq!(out.live_base_snapshot_keys, vec![new_base.clone()]);
6558 assert!(!sink.exists(&old_base).await, "superseded base survived");
6559 assert!(sink.exists(&new_base).await, "live base was collected");
6560 }
6561
6562 /// The rebase-crash window: a new base is published before the old
6563 /// manifests are deleted, so the surviving chain names the OLD base. Both
6564 /// are live — the chain's pick and the newest — and the GC must refuse to
6565 /// choose between them.
6566 #[tokio::test]
6567 async fn gc_keeps_both_bases_while_a_rebase_is_half_landed() {
6568 let sink = GcSink::new();
6569 let old_base = sink.put_snapshot(1).await;
6570 let new_base = sink.put_snapshot(2).await;
6571 sink.put_manifest(10, &old_base, 0, 5, 1, 10).await;
6572
6573 let out = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6574
6575 assert!(out.collected_snapshots.is_empty());
6576 assert_eq!(out.retained_snapshots, 2);
6577 assert_eq!(out.live_base_snapshot_keys, vec![old_base, new_base]);
6578 }
6579
6580 /// (b) An orphaned generation's frames are reclaimed — under both key
6581 /// shapes — and a live generation's are not.
6582 #[tokio::test]
6583 async fn gc_collects_orphaned_frame_prefixes_and_keeps_the_live_generation() {
6584 let sink = GcSink::new();
6585 let base = sink.put_snapshot(2).await;
6586 sink.put_manifest(10, &base, 0, 5, 1, 10).await;
6587 let live = sink.put_frame_batch(0, 5, 1, 10).await;
6588 // Generation 4 was rebased away: its manifests are gone, its frames
6589 // are not. Once under the epoch-0 layout, once under the fenced one.
6590 let orphan_unfenced = sink.put_frame_batch(0, 4, 1, 6).await;
6591 let orphan_fenced = sink.put_frame_batch(7, 4, 1, 6).await;
6592
6593 let out = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6594
6595 let mut expected = vec![orphan_unfenced.clone(), orphan_fenced.clone()];
6596 expected.sort();
6597 assert_eq!(out.collected_frame_objects, expected);
6598 assert_eq!(out.retained_frame_objects, 1);
6599 assert_eq!(out.collected_bytes, 32, "two 16-byte batch objects");
6600 assert!(sink.exists(&live).await, "live generation's frames collected");
6601 assert!(!sink.exists(&orphan_unfenced).await);
6602 assert!(!sink.exists(&orphan_fenced).await);
6603 }
6604
6605 /// (c) Nothing inside the grace window is touched, however unreachable it
6606 /// looks. Everything an `InMemory` store holds was written moments ago, so
6607 /// a non-zero grace must spare the whole sweep.
6608 #[tokio::test]
6609 async fn gc_spares_everything_inside_the_grace_window() {
6610 let sink = GcSink::new();
6611 let stale_base = sink.put_snapshot(1).await;
6612 let base = sink.put_snapshot(2).await;
6613 sink.put_manifest(10, &base, 0, 5, 1, 10).await;
6614 sink.put_frame_batch(0, 5, 1, 10).await;
6615 let orphan = sink.put_frame_batch(0, 4, 1, 6).await;
6616
6617 let cfg = StreamGcConfig {
6618 grace: Duration::from_secs(3600),
6619 dry_run: false,
6620 };
6621 let out = gc_stream(&sink.target, &cfg).await.unwrap();
6622
6623 assert!(out.collected_frame_objects.is_empty());
6624 assert!(out.collected_snapshots.is_empty());
6625 assert_eq!(out.collected_bytes, 0);
6626 assert_eq!(out.spared_by_grace, 2, "the orphan batch and the stale base");
6627 assert!(sink.exists(&orphan).await);
6628 assert!(sink.exists(&stale_base).await);
6629 }
6630
6631 /// (d) A dry run reports exactly what the real sweep would collect, and
6632 /// deletes none of it.
6633 #[tokio::test]
6634 async fn gc_dry_run_reports_without_deleting() {
6635 let sink = GcSink::new();
6636 let stale_base = sink.put_snapshot(1).await;
6637 let base = sink.put_snapshot(2).await;
6638 sink.put_manifest(10, &base, 0, 5, 1, 10).await;
6639 sink.put_frame_batch(0, 5, 1, 10).await;
6640 let orphan = sink.put_frame_batch(0, 4, 1, 6).await;
6641
6642 let dry = gc_stream(
6643 &sink.target,
6644 &StreamGcConfig {
6645 grace: Duration::ZERO,
6646 dry_run: true,
6647 },
6648 )
6649 .await
6650 .unwrap();
6651
6652 assert!(dry.dry_run);
6653 assert_eq!(dry.collected_snapshots, vec![stale_base.clone()]);
6654 assert_eq!(dry.collected_frame_objects, vec![orphan.clone()]);
6655 assert!(dry.collected_bytes > 0);
6656 assert!(sink.exists(&stale_base).await, "dry run deleted a snapshot");
6657 assert!(sink.exists(&orphan).await, "dry run deleted a frame object");
6658
6659 // The proposal is what the real sweep then executes, key for key.
6660 let wet = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6661 assert!(!wet.dry_run);
6662 assert_eq!(wet.collected_snapshots, dry.collected_snapshots);
6663 assert_eq!(wet.collected_frame_objects, dry.collected_frame_objects);
6664 assert_eq!(wet.collected_bytes, dry.collected_bytes);
6665 assert!(!sink.exists(&stale_base).await);
6666 assert!(!sink.exists(&orphan).await);
6667 }
6668
6669 /// The pin `gc_stream`'s docs promise: its live set is the exact complement
6670 /// of what a restore selects, on BOTH restore branches. If either selection
6671 /// rule is ever changed without changing the other, this fails.
6672 #[tokio::test]
6673 async fn gc_liveness_is_the_complement_of_restore_selection() {
6674 // Branch 1 — no generations. `hydrate::restore_subject` falls back to
6675 // `snapshot::restore_latest`, which takes the lexically-greatest key.
6676 // That key is live; the older ones are not collected either, because a
6677 // chainless prefix is indistinguishable from a tier-1a sink whose older
6678 // snapshots are history (see `base_snapshots_skipped`).
6679 let sink = GcSink::new();
6680 let _older = sink.put_snapshot(1).await;
6681 let newest = sink.put_snapshot(2).await;
6682
6683 let restore_would_pick = crate::snapshot::latest_snapshot_key(&sink.target)
6684 .await
6685 .unwrap()
6686 .unwrap();
6687 assert_eq!(restore_would_pick, newest);
6688
6689 let out = gc_stream(
6690 &sink.target,
6691 &StreamGcConfig {
6692 grace: Duration::ZERO,
6693 dry_run: true,
6694 },
6695 )
6696 .await
6697 .unwrap();
6698 assert_eq!(out.live_base_snapshot_keys, vec![restore_would_pick]);
6699 assert!(out.base_snapshots_skipped);
6700 assert!(out.collected_snapshots.is_empty());
6701 assert_eq!(out.retained_snapshots, 2);
6702
6703 // Branch 2 — a validating chain. `restore_stream_from_manifests` takes
6704 // the chain's base, which here is the OLDER snapshot. The GC must keep
6705 // it, and keep the newest as well, since the fallback branch is still
6706 // reachable for any reader that finds no manifests.
6707 let chained = GcSink::new();
6708 let chain_base = chained.put_snapshot(1).await;
6709 let unreferenced_newer = chained.put_snapshot(2).await;
6710 chained.put_manifest(10, &chain_base, 0, 5, 1, 10).await;
6711 chained.put_frame_batch(0, 5, 1, 10).await;
6712
6713 let manifests = list_and_parse_generation_manifests(&chained.target)
6714 .await
6715 .unwrap();
6716 let chain = validate_generation_chain(&manifests).unwrap();
6717 assert_eq!(chain.base_snapshot_key, chain_base);
6718
6719 let out = gc_stream(
6720 &chained.target,
6721 &StreamGcConfig {
6722 grace: Duration::ZERO,
6723 dry_run: true,
6724 },
6725 )
6726 .await
6727 .unwrap();
6728 assert!(!out.base_snapshots_skipped, "a chain makes this a tier-2 sink");
6729 assert!(out.live_base_snapshot_keys.contains(&chain.base_snapshot_key));
6730 assert!(out.live_base_snapshot_keys.contains(&unreferenced_newer));
6731 assert!(out.collected_snapshots.is_empty());
6732
6733 // And the retained frame set is exactly what a replay would fetch.
6734 let replayed: Vec<String> = manifests
6735 .iter()
6736 .flat_map(|m| chained.target.frame_objects_of(m))
6737 .map(|(key, _, _)| key.to_string())
6738 .collect();
6739 assert_eq!(out.retained_frame_objects, replayed.len());
6740 assert!(out.collected_frame_objects.is_empty());
6741 }
6742
6743 /// A prefix with no chain keeps every snapshot — its older ones could be a
6744 /// tier-1a sink's history — but its `frames/` are unreachable by
6745 /// construction and still go.
6746 #[tokio::test]
6747 async fn gc_without_a_chain_keeps_snapshots_but_still_collects_frames() {
6748 let sink = GcSink::new();
6749 let older = sink.put_snapshot(1).await;
6750 let newest = sink.put_snapshot(2).await;
6751 let orphan = sink.put_frame_batch(0, 4, 1, 6).await;
6752
6753 let out = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6754
6755 assert!(out.base_snapshots_skipped);
6756 assert!(out.collected_snapshots.is_empty());
6757 assert_eq!(out.collected_frame_objects, vec![orphan.clone()]);
6758 assert!(sink.exists(&older).await);
6759 assert!(sink.exists(&newest).await);
6760 assert!(!sink.exists(&orphan).await);
6761 }
6762
6763 /// An empty prefix is a no-op, not an error — a GC pointed at a sink that
6764 /// has never been written to must not fail.
6765 #[tokio::test]
6766 async fn gc_on_an_empty_prefix_collects_nothing() {
6767 let sink = GcSink::new();
6768 let out = gc_stream(&sink.target, &sweep_now()).await.unwrap();
6769 assert_eq!(
6770 out,
6771 StreamGcOutcome {
6772 base_snapshots_skipped: true,
6773 ..Default::default()
6774 }
6775 );
6776 }
6777}