1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
//! Starting the `trusty-search` daemon the whole audit stack stands on (#5670).
//!
//! Why: the audit's prerequisite chain is trusty-search → per-repository index →
//! trusty-analyze, and [`super::analyze`] closed only the last link. Starting
//! `trusty-analyze` is not enough on its own, in two different ways.
//!
//! On a cold machine `trusty-analyze serve` exits at its own trusty-search check
//! before it ever binds a port, so the analyze preflight spawns a process that is
//! already gone by the first poll and refuses the run. The operator's remedy was
//! to run `trusty-search start` by hand, which DOC-67 §2 does not allow: the
//! sweep gets one non-interactive shot and owes no manual prerequisite.
//!
//! The second case is the one a fresh-spawn fix misses. `trusty-analyze` answers
//! its own `/health` with `503 degraded` whenever trusty-search is unreachable,
//! and `probe_once` counts only a 2xx — so the analyze preflight re-reads
//! trusty-search's LIVE status on every run. An analyze daemon that has been up
//! for days on top of a trusty-search that died an hour ago fails the probe, the
//! spawned replacement exits at its own search check, the original keeps
//! answering 503, and the readiness poll refuses. Nothing about that run is a
//! cold start.
//!
//! What: [`ensure_search_daemon`], the address and binary rules it applies, and
//! [`SearchDaemonUnavailable`]. It runs BEFORE the analyze preflight in
//! `crate::commands::audit::run`, which is the whole point — analyze cannot boot
//! without it. The probe/spawn/poll loop is `trusty_common::daemon_guard`, the
//! workspace's one daemon-lifecycle entry point, and the address comes from
//! [`DaemonAddrLayout::TRUSTY_SEARCH`] rather than a second copy of
//! trusty-search's discovery rules.
//!
//! Nothing here is fail-open. Both failure arms — the binary would not spawn, and
//! the daemon never answered — return `Err`, for the same reason the analyze
//! preflight refuses: a report with its findings, complexity and health sections
//! empty reads as a clean bill of health rather than as an outage.
//! Test: `super::tests` against stubs; `super::real_binary_tests` (`#[ignore]`d)
//! against the real `trusty-search` binary.
//!
//! # Spec References
//! - [`SPEC-TGAUDIT-06~draft`](../../../../docs/specs/DOC-67-tga-audit-mode.md#SPEC-TGAUDIT-06~draft)
//! - [`SPEC-TGAUDIT-09~draft`](../../../../docs/specs/DOC-67-tga-audit-mode.md#SPEC-TGAUDIT-09~draft)
use Duration;
use ;
use ;
/// Wall-clock budget for a freshly-spawned `trusty-search` to answer `/health`.
///
/// Why: 60s, matching `trusty-search`'s own guard
/// (`crates/trusty-search/src/commands/daemon_guard.rs`'s `READY_TIMEOUT`) rather
/// than `daemon_guard::DEFAULT_STARTUP_TIMEOUT`'s 30s. The HTTP port binds in
/// about a second, but a first run on a machine with no model cache spends 15–30s
/// in ONNX load before it answers, and refusing an audit at 30s would turn a slow
/// cold start into a failed engagement.
pub const SEARCH_STARTUP_TIMEOUT: Duration = from_secs;
/// The audit cannot proceed without the trusty-search daemon.
///
/// Why: this is refused before the sweep rather than reported after it, so the
/// message is the operator's whole remedy. It names trusty-search as the FIRST
/// link rather than describing the analyze symptom, because an operator reading
/// "trusty-analyze is degraded" reaches for the wrong daemon.
/// What: the address probed, the binary tried, and the underlying cause.
/// Test: `super::tests::an_unspawnable_search_binary_refuses_the_audit`.
/// Where and how [`ensure_search_daemon`] looks for the daemon.
///
/// Why: the address and binary come from the environment and the two budgets are
/// fixed, which makes the whole guard untestable if it reads them itself. Taking
/// them as a value is what lets a test drive the spawn-and-poll path against a
/// stub executable on an ephemeral port — the same split [`super::AnalyzeGuard`]
/// uses.
/// What: the daemon address, the binary to spawn, and the readiness budget.
/// Test: `super::tests::a_reachable_search_daemon_is_not_restarted`.
/// The health endpoint for a base address, tolerating a trailing slash.
/// The exact argument vector the guard hands `trusty-search`.
///
/// Why: this list IS the tga→trusty-search contract, and `--foreground` is
/// load-bearing rather than decorative — a bare `trusty-search start` re-spawns
/// itself as a background daemon and the parent exits, which is why
/// trusty-search's own guard passes the flag
/// (`crates/trusty-search/src/commands/daemon_guard.rs`'s
/// `spawn_daemon_with_device`). We already detach the child ourselves, so the
/// second fork would only cost the guard its view of the process it started.
/// Building the vector in a pure function is what lets a test assert its contents
/// without spawning anything.
/// What: `start --foreground`.
/// Test: `super::tests::the_search_spawn_arguments_are_start_in_the_foreground`.
pub
/// Ensure the trusty-search daemon is up before anything else in the audit.
///
/// Why/What: see the module docs. Resolves the guard from the environment and
/// delegates to [`ensure_search_daemon_with`], so the environment and the
/// discovery files are read exactly once, at the public entry point.
///
/// # Errors
///
/// [`SearchDaemonUnavailable`] when the daemon is absent and cannot be started.
///
/// Test: `super::tests::a_reachable_search_daemon_is_not_restarted`.
pub async
/// [`ensure_search_daemon`] with the address, binary and budgets already fixed.
///
/// Why: taking them as a value is what lets a test drive the whole spawn-and-poll
/// path against a stub executable on an ephemeral port, without touching the
/// process environment or leaving a daemon behind.
/// What: probes `<url>/health`; on a hit, returns without spawning anything. On a
/// miss, spawns `<binary> start --foreground` detached and polls until it answers.
/// Neither failure is downgraded — a spawn that fails and a daemon that never
/// answers both return `Err`.
///
/// The address is resolved once, before the spawn, and the poll reuses it. That
/// is what `trusty-search`'s own CLI does at every call site
/// (`commands::add`'s `ensure_daemon_running_or_exit(&daemon_base_url())`), so a
/// daemon this guard starts is found the same way the daemon's own client would
/// find it.
///
/// # Errors
///
/// [`SearchDaemonUnavailable`] carrying the spawn error, or the readiness timeout.
///
/// Test: `super::tests::{a_reachable_search_daemon_is_not_restarted,
/// an_unspawnable_search_binary_refuses_the_audit,
/// a_search_daemon_that_never_comes_up_refuses_the_audit,
/// the_search_guard_recovers_a_stale_degraded_analyze_daemon}`.
pub async