1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
//! Concurrent index build coordination.
//!
//! Concurrent `DEFINE INDEX` can run asynchronously while user writes continue.
//! In a multi-node deployment every node must make the same decisions about
//! which builder owns the work, whether writers should queue mutations, and
//! when queries may use the index. This module keeps those decisions in durable
//! table-scoped keys instead of process-local memory.
//!
//! The durable protocol uses four key families:
//!
//! - `!bs`: one build-state record per index, including phase, owner, generation, report counters,
//! the initial-scan continuation checkpoint, and error reason.
//! - `!bt`: the generation-scoped writer-admission ticket counter, kept off `!bs` so an admitted
//! write never invalidates the builder's in-flight batch.
//! - `!bg`: generation-scoped queued mutations that the builder replays.
//! - `!bp`: per-record pointers to the first queued mutation seen during the initial scan, so the
//! scan indexes the writer-observed old state.
//! - `!br`: writer reservations that keep `Closing` from publishing `Online` until every admitted
//! writer has either committed its `!bg` entry, released its ticket after transaction close, or
//! died.
//!
//! Generation numbers fence stale queued work. Builder owner heartbeats fence
//! stale builders. Query planning only sees durable-`Online` indexes, while
//! document writes still see building indexes so they can enqueue mutations.
//! A build in durable `Error` keeps admitting writes the same way, so a failed
//! background build never blocks user writes: the errored generation's queue
//! is never replayed — `REBUILD INDEX` wipes it and rescans the table.
//! Legacy `!ig`/`!ip` appendings are still drained for committed work from older
//! code paths, but new writes use the durable queue.
use Duration;
pub use ;
pub use ;
// Only the frozen-fixture corpus names this directly; the builder reaches a
// primary appending through its ticket.
pub use PrimaryAppending;
// What a build persists is part of the keyspace, so it is declared below this
// layer; the coordination protocol above reads it from there.
pub use ;
use Instant;
use crateIndexBuildReservationRelease;
/// How long a writer admission reservation is considered owned by the writer.
const BUILD_RESERVATION_TTL_SECS: i64 = 30;
/// How long a builder may go without heartbeating its durable state before
/// another builder may take ownership of the same generation. This assumes
/// bounded clock skew between nodes; ownership transitions are still fenced by
/// CAS on `(generation, owner)`, so a stale owner cannot publish progress after
/// takeover.
const BUILD_OWNER_LEASE_SECS: i64 = 60;
/// Poll cadence while writer admission waits for `Closing` to become `Online`
/// or `Error`. The caller's context deadline is the only timeout budget.
const BUILD_CLOSING_SLEEP: Duration = from_millis;
/// Total time one closing transaction may spend waiting for the local builders
/// it aborted to stop, before deleting their durable state anyway.
///
/// A builder polls its abort flag per record during the initial scan and once
/// per iteration in every replay, drain and retry loop, so the wait normally
/// ends in microseconds. The budget only binds where a single uninterruptible
/// stretch runs long, and the longest of those is one HNSW/DiskANN compaction
/// plan, whose apply and commit carry no internal checkpoint. Overrunning it
/// falls back to the compare-and-swap on `!bs` as the only ordering against the
/// builder, which is exactly what the wait exists to replace, so the budget is
/// deliberately generous: a cancelled statement stalls for at most this long,
/// whereas falling short strands durable state that nothing collects.
const BUILD_ABORT_STOP_BUDGET: Duration = from_secs;
type IndexBuilding = Arc;
/// Deadline for an abort-wait in a close drain that began at `drain_started_at`.
///
/// The budget is per drain, not per build: a schema transaction that defined
/// several indexes and then rolled back queues one cleanup per index, they run
/// one at a time, and their waits must not multiply into a close the client
/// reads as a hang. Measuring every one of them from the same origin caps the
/// whole drain at a single budget.
pub
pub
const LEGACY_BATCH_ID: BatchId = 0;