1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
//! Cross-process exclusive advisory lock around a whole-file critical section.
//!
//! Why: [`crate::json_rmw`] already owned this lock, but only for JSON
//! documents it also serialises and publishes itself. `trusty-search`'s
//! `indexes.toml` is a TOML registry with its own loader, its own
//! `skip_serializing_if` shape, and its own fail-closed parse contract, so it
//! cannot route through `json_rmw::update` — yet it has the identical failure
//! mode: several independent PROCESSES (the daemon, `trusty-search prune`,
//! `trusty-search prune-orphans`) each run load → mutate → save-the-whole-file,
//! and a write landing between another writer's load and its save is discarded
//! with both callers reporting success (#5344). Extracting the lock here means
//! there is still exactly ONE implementation of the critical section; `json_rmw`
//! now calls it rather than owning it.
//! What: [`with_exclusive_lock`] runs a closure while holding an exclusive
//! `flock(2)`-style advisory lock on a `<path>.lock` sidecar, releasing it by
//! RAII on every exit path including a panic. [`with_exclusive_lock_timeout`]
//! is the same entry point with the wait bound chosen by the caller, and
//! [`lock_path`] names that sidecar.
//! Test: `cargo test -p trusty-common --features unconditional-only --
//! file_lock::tests`.
//!
//! # Contract
//!
//! - **Serialisation.** The lock is held by the open file description, so it
//! serialises separate PROCESSES and separate threads that each call
//! [`with_exclusive_lock`], on Unix and Windows alike.
//! - **Advisory, not mandatory.** A process that writes the guarded file
//! without going through this entry point is not blocked. Every writer of a
//! given file must use it.
//! - **Never fail open.** A lock that cannot be created or acquired — including
//! one whose acquisition times out — is an `Err`; the closure never runs.
//! Proceeding unlocked is the lost-update bug this module exists to remove.
//! - **Bounded, never indefinite — for local contention.** Acquisition retries
//! until [`DEFAULT_LOCK_TIMEOUT`] (or the caller's own bound) and then fails
//! with an [`std::io::ErrorKind::TimedOut`] error wrapping a [`LockTimeout`].
//! A holder that is wedged rather than dead — SIGSTOP'd, stopped in a
//! debugger — therefore costs the waiter a bounded delay and a diagnosable
//! error instead of a silent hang (#7762). The bound is enforced BETWEEN
//! `try_write` attempts, not inside one: a `$HOME` wedged on a stalled
//! network filesystem can leave a single `flock(2)` call blocked past the
//! timeout, and the deadline is never reached because the loop never gets
//! control back.
//! - **Diagnostic pid in the sidecar.** The holder writes its pid into the
//! sidecar after acquiring, which is the sidecar's only content and the only
//! thing a waiter reads out of it. It is best-effort: it goes stale on
//! release, and an empty or unparsable sidecar simply yields "holder pid
//! unknown". It never gates acquisition and never turns a timeout into a
//! success.
//! - **Not reentrant.** Nesting two [`with_exclusive_lock`] calls on the same
//! path cannot succeed — the second acquisition uses a different descriptor,
//! so it waits out its whole timeout and then errors.
//! - **Blocking.** Acquisition blocks the calling thread for up to the timeout.
//! Async callers must run it on a blocking-safe thread (e.g.
//! `tokio::task::spawn_blocking`).
//!
//! [`with_exclusive_lock`]: crate::file_lock::with_exclusive_lock
//! [`with_exclusive_lock_timeout`]: crate::file_lock::with_exclusive_lock_timeout
//! [`lock_path`]: crate::file_lock::lock_path
//! [`DEFAULT_LOCK_TIMEOUT`]: crate::file_lock::DEFAULT_LOCK_TIMEOUT
//! [`LockTimeout`]: crate::file_lock::LockTimeout
use ;
use ;
use ;
use ;
/// How long [`with_exclusive_lock`] waits before giving up.
///
/// Why: the default has to suit the interactive path — every
/// `.claude/settings.json` writer behind `tm launch` goes through this lock
/// (#7762) — where a user watching a silent terminal is the failure. Ten
/// seconds is far longer than any real critical section here (a few hundred
/// kilobytes of load → mutate → save) and short enough to read as "something is
/// wrong" rather than as a hang.
/// What: the bound [`with_exclusive_lock`] passes to
/// [`with_exclusive_lock_timeout`]. A batch or daemon path that genuinely
/// wants to wait longer calls the latter directly.
/// Test: `with_exclusive_lock_default_path_is_unchanged_when_uncontended`.
pub const DEFAULT_LOCK_TIMEOUT: Duration = from_secs;
/// Gap between acquisition attempts.
///
/// Short enough that an uncontended-after-a-moment lock is taken promptly,
/// long enough that a long wait is not a spin.
const ACQUIRE_POLL_INTERVAL: Duration = from_millis;
/// The bounded wait expired without the lock.
///
/// Why: "could not acquire" must be actionable without a debugger. The two
/// facts a waiter can act on are WHICH lock it waited for and WHO held it, so
/// both travel in the error rather than in a log line the caller may not emit.
/// What: carried as the inner error of an [`std::io::Error`] with kind
/// [`std::io::ErrorKind::TimedOut`], so callers that already propagate
/// `io::Error` need no change and callers that want the fields recover them
/// with `err.get_ref().and_then(|e| e.downcast_ref::<LockTimeout>())`.
/// `holder_pid` is `None` when the sidecar is empty or unparsable — it is a
/// diagnostic, never a precondition.
/// Test: `with_exclusive_lock_timeout_errors_while_another_descriptor_holds_it`,
/// `with_exclusive_lock_timeout_reports_unknown_pid_for_a_garbage_sidecar`,
/// `with_exclusive_lock_timeout_names_the_pid_of_a_holding_process`.
/// Sidecar lock-file path for `path`.
///
/// Why: locking the guarded file itself would mean opening it for write before
/// we know whether the update will succeed, and the lock would be lost across
/// the `rename` that publishes a new version (the renamed-over inode, and any
/// lock on it, is discarded). A stable sidecar survives every publish.
/// What: appends `.lock` to the file name, keeping it in the same directory.
/// Test: `lock_path_is_a_sidecar`.
/// Run `f` while holding the exclusive cross-process lock guarding `path`.
///
/// Why: see the module docs — this is the one place the load → mutate → save
/// critical section is made safe against writers in other processes.
/// What: [`with_exclusive_lock_timeout`] at [`DEFAULT_LOCK_TIMEOUT`]. `f`'s own
/// return value — commonly a `Result` — passes through untouched; the `Err`
/// returned here is only ever a lock-acquisition failure, so a caller can never
/// confuse "could not lock" with "the work failed".
/// Test: `with_exclusive_lock_serialises_separate_descriptors`,
/// `with_exclusive_lock_releases_on_panic`, `with_exclusive_lock_unopenable_errors`,
/// `with_exclusive_lock_default_path_is_unchanged_when_uncontended`.
/// [`with_exclusive_lock`] with the caller's own wait bound.
///
/// Why: #7762 — the blocking acquisition this replaced hung `tm launch` with no
/// output whenever a holder was wedged rather than dead. A non-interactive
/// writer may legitimately want to wait longer than the interactive default, so
/// the bound is a parameter rather than a constant.
/// What: creates (if needed) and opens the [`lock_path`] sidecar, retries the
/// non-blocking exclusive acquisition every [`ACQUIRE_POLL_INTERVAL`] until
/// `timeout` elapses, records the acquiring pid in the sidecar, runs `f`, and
/// releases the lock by RAII. One attempt always happens, so a zero `timeout`
/// is a single try. On expiry the return is an [`std::io::ErrorKind::TimedOut`]
/// error wrapping [`LockTimeout`] and `f` is never run. The bound covers the
/// gap BETWEEN `try_write` calls, not the inside of one: on a wedged network
/// filesystem, a single `try_write` can block in `flock(2)` past `timeout`,
/// and this function does not observe the deadline until that call returns.
/// Test: `with_exclusive_lock_timeout_errors_while_another_descriptor_holds_it`,
/// `with_exclusive_lock_timeout_reports_unknown_pid_for_a_garbage_sidecar`,
/// `with_exclusive_lock_timeout_names_the_pid_of_a_holding_process`,
/// `with_exclusive_lock_records_the_acquiring_pid`.
/// Stamp this process's pid into the held sidecar.
///
/// Why: a later waiter can only name the holder if the holder left its name;
/// `flock(2)` itself exposes no owner, and `F_GETLK` describes `fcntl` record
/// locks, not these.
/// What: truncate-and-rewrite under the lock we already hold, so the pid is the
/// sidecar's whole content. Failures are swallowed: the pid is diagnostic, and
/// an unwritable sidecar must not undo an acquisition that succeeded. A
/// `set_len(0)` that succeeds followed by a rewind or write that fails leaves
/// the sidecar empty rather than restoring its prior content; a later
/// [`read_holder_pid`] then reports "holder pid unknown" exactly as it would
/// for a sidecar that was never written.
/// Test: `with_exclusive_lock_records_the_acquiring_pid`,
/// `with_exclusive_lock_acquires_over_a_garbage_sidecar`.
/// The pid the sidecar names, if it names one.
///
/// Why: read on the failure path only, where any answer is better than none and
/// no answer must still be an error.
/// What: parses the sidecar's whole trimmed content as a pid. Missing, empty,
/// non-UTF-8 and unparsable content all read as `None`. This read is not
/// synchronised with [`record_holder_pid`]'s truncate-then-write: a waiter can
/// land in the gap between them, so `None` here also covers a live holder
/// caught mid-write, not only an empty or garbage sidecar.
/// Test: `with_exclusive_lock_timeout_reports_unknown_pid_for_a_garbage_sidecar`.