1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
//! Single-instance guard for the trusty-memory daemon.
//!
//! Why: macOS launchd `KeepAlive { SuccessfulExit: false }` respawns the daemon
//! whenever it exits with a non-zero code. A second instance that fails to bind
//! exits non-zero, launchd reads that as a crash and spawns another copy, and
//! the resulting zombie herd (69 observed in the wild) exhausts file
//! descriptors on top of the existing fd-limit bug.
//!
//! The fix: before binding, probe the socket. If something is already serving
//! it, exit **0**. Launchd treats exit-0 as a clean shutdown and does not
//! respawn, which collapses the herd on the next invocation without touching
//! the launchd config.
//!
//! #6286 changed what is probed, not the decision. It used to read the
//! `http_addr` discovery file and GET `/health` at whatever address that named
//! — a file that goes stale after a SIGKILL, which is why the probe had to
//! tolerate one pointing at a dead port. The socket path is derived rather than
//! published, so there is nothing to be stale and the probe is a bare connect
//! through `trusty_common::uds::socket_is_serving`.
//!
//! What: [`single_instance_check`] (async, for real daemon startups) and
//! [`StartupAction`] (a pure enum, so the decision is unit-testable without
//! I/O).
//!
//! Test: `startup_action_*` for the decision, `single_instance_check_*` for the
//! probe.
use Path;
use Duration;
/// How long a liveness connect may take before the path is called dead.
///
/// A local socket accepts or refuses in microseconds; this is headroom for a
/// loaded machine, not a latency budget. It matches trusty-analyze's
/// `daemon_guard::PROBE_TIMEOUT` so the two daemons wait the same.
const PROBE_TIMEOUT: Duration = from_millis;
/// What the daemon startup should do after the single-instance check.
///
/// Why: separating the decision from the I/O lets us unit-test the logic
/// with injected probe results rather than spinning up real TCP listeners.
/// What: three variants covering the full decision tree.
/// Test: `startup_action_from_probe_result_*` tests in this module.
/// Decide what to do based on the result of a liveness probe.
///
/// Why: the single-instance check reduces to "did the health probe succeed?".
/// Encoding the decision as a pure function (rather than embedding it in the
/// async probe body) makes the logic unit-testable without actual network I/O.
/// What: `probe_ok = true` → [`StartupAction::ExitAlreadyRunning`];
/// `probe_ok = false` → [`StartupAction::Proceed`].
/// Test: `startup_action_from_probe_result_when_alive`,
/// `startup_action_from_probe_result_when_dead`.
/// Perform the single-instance check at daemon startup.
///
/// Why: launchd respawns any non-zero exit, so a second instance that fails to
/// bind causes an endless respawn storm. Exiting 0 when another healthy
/// instance is detected short-circuits it.
///
/// What: a bare connect to `socket`. It deliberately does NOT call
/// `memory.health`: the question is whether the endpoint is live, and a daemon
/// that is up but degraded must not be reported absent and spawned on top of
/// itself. An absent or dead socket returns [`StartupAction::Proceed`], so a
/// cold start is never blocked.
///
/// Test: `single_instance_check_proceeds_when_nothing_is_serving`,
/// `single_instance_check_exits_when_something_is_serving`.
pub async
/// Single-instance check with up to `max_retries` additional probes.
///
/// Why (issue #1152, Tier 3): a single probe can miss a daemon that is
/// mid-boot — it has not bound the socket yet. Retrying with a short sleep lets
/// a slow-boot daemon be detected and this caller exit 0, rather than
/// proceeding to open redb and triggering `DatabaseAlreadyOpen`.
/// What: calls `single_instance_check` repeatedly up to `1 + max_retries`
/// times, sleeping `delay_ms` between each call, stopping on the first
/// non-`Proceed` result. Returns the final `StartupAction`.
/// Test: covered by the unit tests for `startup_action_from_probe_result`;
/// the retry path is exercised by the integration guard in `main.rs`.
pub async