trusty-common 0.46.5

Shared utilities and provider-agnostic streaming chat (ChatProvider, OllamaProvider, OpenRouter, tool-use) for trusty-* projects
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
//! The serving half of [`crate::uds::rpc`]: a generic framed JSON-RPC server
//! over a hardened Unix socket (#6277, ADR-0032).
//!
//! Why: `uds::rpc` is the client — it dials, writes one frame, reads one back —
//! and nothing in this workspace owned the other end generically. The only
//! server-side accept loop was `webhook_relay::serve`, which is a single-purpose
//! durable-delivery contract: one method, an ack conditioned on an fsync, an
//! inbox. ADR-0032 puts every trusty-* service's own transport on UDS, so the
//! next service to migrate needed either a fourth hand-rolled accept loop or
//! this. `webhook_relay` is deliberately left untouched; its ordering rule is
//! the reason it exists and is not a thing to generalise.
//!
//! What, in the order a daemon uses them.
//!   - [`RpcRouter`] is the caller's half: method names mapped to handlers over
//!     the caller's own request and response types (see [`RpcRouter::typed`]).
//!     A service that already has a generic `(method, params)` dispatcher
//!     mounts it whole through [`RpcRouter::fallback`] instead (#6286). A method
//!     that answers in many frames rather than one is registered with
//!     [`RpcRouter::typed_stream`] — see the wire contract below.
//!   - [`RpcServer::run`] is the whole body of a daemon — bind, serve, unlink.
//!   - [`serve_until`] and [`handle_connection`] are that body's two halves,
//!     public so a caller with its own bind (say
//!     [`crate::uds::bind_singleton_hardened`], which takes over a stale socket
//!     file where [`crate::uds::bind_hardened`] refuses) drives the loop itself.
//!
//! **The trust boundary is the socket, not the payload.** [`bind_hardened`]
//! puts the socket at `0600` inside a `0700` directory, and every accepted
//! connection runs [`ensure_peer_is_self`] before a single byte is read. No
//! CSRF, origin, or token machinery is ported from the HTTP shape those checks
//! replace — on a UDS socket it would guard nothing (#6277 design review).
//!
//! ## The wire contract, including streams (#6286)
//!
//! A request is one newline-terminated JSON-RPC frame, unchanged:
//!
//! ```text
//! {"jsonrpc":"2.0","id":7,"method":"chat","params":{…}}
//! ```
//!
//! plus ONE optional field, `"stream": true`. Absent means false, and false is
//! exactly the protocol as it stood — one request frame, one response frame,
//! connection closed. An old client and a new server therefore behave as they
//! always did, byte for byte, and so does a new client calling a method that
//! does not stream.
//!
//! A request that asks for a stream, against a method registered with
//! [`RpcRouter::typed_stream`], is answered with a SEQUENCE of frames on the
//! same connection, each newline-terminated and each carrying a `"stream"`
//! discriminant a plain response never has:
//!
//! ```text
//! {"jsonrpc":"2.0","id":7,"stream":"item","result":"Hel"}
//! {"jsonrpc":"2.0","id":7,"stream":"item","result":"lo"}
//! {"jsonrpc":"2.0","id":7,"stream":"end"}
//! ```
//!
//! **A stream terminates on a frame, never on EOF.** Exactly one terminal frame
//! is written on every path — `"stream":"end"` when the producer finishes, and
//! `"stream":"error"` carrying an [`RpcError`] when the handler fails mid-stream,
//! when it fails to open at all, or when an item does not fit the frame budget.
//! A client that reaches EOF without one reports it rather than returning what
//! it happened to receive: a truncated token stream read as a complete answer is
//! the Fail-Open branch this contract exists to close, and
//! `stream_reports_a_truncated_stream_rather_than_an_empty_success` is its
//! regression test.
//!
//! **The two mismatches both fail immediately, in the shape the caller reads.**
//! A request WITHOUT the flag against a streaming method gets one ordinary
//! response frame carrying [`CODE_STREAM_REQUIRED`]. A request WITH the flag
//! against anything that does not stream — a unary method, a fallback-served
//! name, an unknown name — gets one terminal `"stream":"error"` frame carrying
//! [`CODE_STREAM_UNSUPPORTED`], naming the methods this listener does stream.
//! Neither hangs, and neither leaves the caller decoding a frame shape it does
//! not expect.
//!
//! **The socket file is unlinked explicitly.** [`bind_hardened`] binds and
//! chmods; neither it nor `tokio::net::UnixListener`'s `Drop` removes the path,
//! so a server that just returned would leave a file the next start fails to
//! bind. [`RpcServer::run`] removes it — before dropping the listener, for the
//! reason `webhook_relay::listener` records: with the order reversed there is a
//! window where nothing answers the path but the file is still there, and a
//! successor that rebinds in that window has its fresh socket deleted by this
//! process's `remove_file`.
//!
//! Test: `tests.rs` — `dispatch_*` for the decision, `serve_*` for the socket.

mod idle;
mod router;
mod stream;
mod wire;

#[cfg(test)]
#[path = "tests.rs"]
mod tests;

use std::path::{Path, PathBuf};
use std::sync::Arc;
use std::time::Duration;

use tokio::io::{AsyncBufReadExt, AsyncReadExt as _, AsyncWriteExt, BufReader};
use tokio::net::{UnixListener, UnixStream};

use crate::uds::{MAX_FRAME_BYTES, UdsSecurityError, bind_hardened, ensure_peer_is_self};

pub use idle::{IdleGuard, IdleTracker};
pub use router::{RpcFallback, RpcMethod, RpcRouter, typed_method};
pub use stream::{RpcOutcome, RpcStreamItems, RpcStreamMethod, typed_stream_method, write_stream};
pub use wire::{
    CODE_INTERNAL_ERROR, CODE_INVALID_PARAMS, CODE_INVALID_REQUEST, CODE_METHOD_NOT_FOUND,
    CODE_PARSE_ERROR, CODE_STREAM_REQUIRED, CODE_STREAM_UNSUPPORTED, JSONRPC_VERSION, RpcError,
    RpcRequest, RpcResponse, RpcStreamFrame, StreamPhase,
};

/// Everything that can stop this server, or stop one of its connections.
///
/// `#[non_exhaustive]` for the same reason [`UdsSecurityError`] carries it: the
/// list grows as the transport tightens, and no consumer matches it
/// exhaustively — they log it or convert it.
#[derive(Debug, thiserror::Error)]
#[non_exhaustive]
pub enum RpcServerError {
    /// The socket could not be bound or hardened.
    #[error("bind the rpc socket at {path}: {source}")]
    Bind {
        /// Socket that could not be bound.
        path: PathBuf,
        /// Why the bind failed.
        #[source]
        source: UdsSecurityError,
    },

    /// The connected peer failed the uid check, or its credentials could not be
    /// read. Fatal to the connection by design — see the module docs.
    #[error("refuse an accepted connection: {source}")]
    Peer {
        /// Which check failed.
        #[source]
        source: UdsSecurityError,
    },

    /// Reading the request frame failed.
    #[error("read request frame: {source}")]
    Read {
        /// Underlying OS error.
        #[source]
        source: std::io::Error,
    },

    /// The peer sent `limit` bytes without a newline.
    #[error("request frame exceeded {limit} bytes without a newline")]
    FrameTooLarge {
        /// The budget, in bytes.
        limit: u64,
    },

    /// The response frame could not be serialised.
    #[error("serialize response frame: {source}")]
    Encode {
        /// Underlying serde error.
        #[source]
        source: serde_json::Error,
    },

    /// Writing the response frame failed.
    #[error("write response frame: {source}")]
    Write {
        /// Underlying OS error.
        #[source]
        source: std::io::Error,
    },
}

/// Per-connection budgets for [`serve_until`].
///
/// Test: `serve_rejects_an_oversized_frame`.
#[derive(Debug, Clone, Copy)]
pub struct RpcServeOptions {
    /// Longest one connection may take to deliver a complete request frame.
    ///
    /// Bounds the READ only. A handler is not covered, deliberately: a service
    /// whose method runs for minutes (`trusty-review`'s is the case #6277 is
    /// migrating) would otherwise need this raised to a figure that also lets a
    /// peer which connects and never writes hold a task for that long.
    ///
    /// **It does not apply between streamed response frames** (#6286). The gap
    /// between two frames of a stream is the handler's own production latency —
    /// an LLM deciding on the next token — and a budget there would kill a
    /// legitimately slow stream. It bounds one thing: how long a peer may take
    /// to deliver its REQUEST frame. A service that wants an idle-stream budget
    /// enforces it in its producer, where the difference between "thinking" and
    /// "stuck" is knowable; the client applies its own per-frame budget in
    /// [`crate::uds::send_framed_stream_request`].
    pub read_timeout: Duration,

    /// Largest request frame accepted, counting its terminating newline.
    ///
    /// Precisely: a frame is accepted when its JSON body plus the `\n` is at
    /// most `max_frame_bytes`, so the usable payload is `max_frame_bytes - 1`.
    /// The terminator counting against the budget is not an accident to be
    /// tidied away — [`crate::uds::rpc`]'s `read_one_frame` reads with the same
    /// `take(max_frame_bytes)` and the same comparison, and the two ends have to
    /// draw the line at the same byte. Moving it here alone would produce a
    /// frame one end sends happily and the other refuses.
    ///
    /// Defaults to [`MAX_FRAME_BYTES`], the same control-plane budget
    /// [`crate::uds::send_framed_request`] applies. A service exchanging bulk
    /// payloads raises it here, mirroring
    /// [`crate::uds::send_framed_request_capped`] — and must raise the client's
    /// to match, or it has only moved which side refuses.
    ///
    /// **Per frame, in both directions** (#6286). On a streaming response it
    /// governs each item frame separately, not the stream's total: an item that
    /// would not fit is refused with a terminal error frame rather than
    /// truncated, because the client applies the same budget to each frame it
    /// reads and a frame it cannot buffer would desynchronise the rest of the
    /// connection.
    ///
    /// Not settable per method: the method name lives inside the frame this
    /// budget governs the reading of, so it is not known until after the budget
    /// has already been enforced.
    ///
    /// Test: `frame_of_exactly_the_budget_including_its_newline_is_accepted`,
    /// `serve_rejects_an_oversized_frame`,
    /// `stream_refuses_an_item_larger_than_the_frame_budget`.
    pub max_frame_bytes: u64,
}

impl Default for RpcServeOptions {
    /// Thirty seconds, and [`MAX_FRAME_BYTES`].
    ///
    /// The read covers a local socket writing one already-serialised frame, so
    /// thirty seconds is headroom for a stalled writer rather than a latency
    /// budget.
    fn default() -> Self {
        Self {
            read_timeout: Duration::from_secs(30),
            max_frame_bytes: MAX_FRAME_BYTES,
        }
    }
}

/// What one accepted connection turned out to be.
///
/// Why: a peer that connects and closes without writing is a liveness probe —
/// [`crate::uds::probe::socket_is_serving`] and `UdsServiceSupervisor` both do
/// exactly that. Collapsing it into the failure arm makes the one warning an
/// operator greps for fire on every successful health check.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Served {
    /// A frame arrived and a response frame was written.
    Answered {
        /// Whether that response carried a JSON-RPC error.
        errored: bool,
    },
    /// The peer connected and closed without sending a byte.
    LivenessProbe,
}

/// Serve one accepted connection: verify the peer, read one frame, answer one.
///
/// Why: split out so a test can drive the wire behaviour against a plain
/// `UnixStream` pair without an accept loop.
/// What: refuses a peer whose uid is not our own, reads bytes up to the first
/// newline or EOF under [`RpcServeOptions::max_frame_bytes`], dispatches through
/// `router`, and writes the answer — one response frame for a unary call, or the
/// frame sequence the module docs' wire contract describes for a streaming one
/// (#6286).
///
/// # Errors
///
/// Any [`RpcServerError`] variant except `Bind`. An error here means no complete
/// answer was written and the client sees a transport failure — which is why
/// every failure the router can reason about is a frame instead. On a stream,
/// [`RpcServerError::Write`] mid-sequence is the client having gone away: the
/// producer stops when its receiver drops here, and the accept loop is
/// unaffected.
///
/// Test: `serve_round_trips_a_request_over_a_real_socket`,
/// `serve_rejects_an_oversized_frame`,
/// `handle_connection_reports_a_liveness_probe_rather_than_a_failure`,
/// `stream_round_trips_many_frames_over_a_real_socket`,
/// `stream_survives_a_client_that_disconnects_mid_stream`.
pub async fn handle_connection(
    mut stream: UnixStream,
    router: Arc<RpcRouter>,
    options: RpcServeOptions,
) -> Result<Served, RpcServerError> {
    ensure_peer_is_self(&stream).map_err(|source| RpcServerError::Peer { source })?;

    let mut frame: Vec<u8> = Vec::new();
    {
        let mut reader = BufReader::new((&mut stream).take(options.max_frame_bytes));
        let read = tokio::time::timeout(options.read_timeout, reader.read_until(b'\n', &mut frame))
            .await
            .map_err(|_| RpcServerError::Read {
                source: std::io::Error::new(
                    std::io::ErrorKind::TimedOut,
                    format!("no complete frame within {:?}", options.read_timeout),
                ),
            })?
            .map_err(|source| RpcServerError::Read { source })?;
        if read == 0 && frame.is_empty() {
            return Ok(Served::LivenessProbe);
        }
        if !frame.ends_with(b"\n") && frame.len() as u64 >= options.max_frame_bytes {
            return Err(RpcServerError::FrameTooLarge {
                limit: options.max_frame_bytes,
            });
        }
    }

    let errored = match router.dispatch_streaming(&frame).await {
        RpcOutcome::Single(response) => {
            let errored = response.is_error();
            // One owned, already-newline-terminated buffer, so the response
            // leaves in a single write rather than a serialise-then-append pair.
            let bytes = crate::uds::encode_frame(&response)
                .map_err(|source| RpcServerError::Encode { source })?;
            stream
                .write_all(&bytes)
                .await
                .map_err(|source| RpcServerError::Write { source })?;
            stream
                .flush()
                .await
                .map_err(|source| RpcServerError::Write { source })?;
            errored
        }
        // #6286: many frames, ending in exactly one terminal frame. A write
        // failure part-way through is a dead client, not a server fault — it
        // surfaces as `Write`, the connection is dropped, and the accept loop
        // keeps serving. The handler's producer stops on its own when the
        // receiver `write_stream` owns is dropped here.
        RpcOutcome::Stream { id, items } => {
            write_stream(&mut stream, id, items, options.max_frame_bytes)
                .await
                .map_err(|source| RpcServerError::Write { source })?
        }
    };
    Ok(Served::Answered { errored })
}

/// Accept and serve connections until `shutdown` resolves.
///
/// Why: the whole loop of a UDS daemon, so a service supplies a [`RpcRouter`]
/// and nothing else.
///
/// What: each connection is handed to `tokio::spawn` rather than served inline.
/// That is a REQUIREMENT, not a throughput preference: `uds::probe`'s
/// `SocketVerdict` docs record that on macOS a bound listener with a saturated
/// accept queue answers ECONNREFUSED, which a prober classifies as `NotServing`.
/// A server that dispatched inline would be read as dead under exactly the load
/// it was handling.
///
/// The listener is borrowed, not consumed, so a caller can unlink the socket
/// while it is still bound — see the module docs for why that order matters.
///
/// A connection that errors — or whose handler panics — is logged and dropped
/// without answering, and the loop keeps accepting.
///
/// Test: `serve_round_trips_a_request_over_a_real_socket`,
/// `serve_stops_on_shutdown`,
/// `serve_handles_concurrent_connections_without_serialising`,
/// `serve_survives_a_panicking_handler_and_answers_the_next_connection`.
pub async fn serve_until(
    listener: &UnixListener,
    router: Arc<RpcRouter>,
    options: RpcServeOptions,
    shutdown: impl std::future::Future<Output = ()> + Send,
) {
    let _ = serve_until_idle(listener, router, options, shutdown, None).await;
}

/// Why a serve loop ended.
///
/// Why this is returned rather than logged: an on-demand service's caller
/// distinguishes the two — a shutdown is an operator or supervisor stopping the
/// process, an idle exit is the process reclaiming itself and is the normal end
/// of a successful lifetime. `trusty-analyze` prints a different line for each.
///
/// Test: `serve_until_idle_exits_when_the_window_elapses`,
/// `serve_stops_on_shutdown`.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum ServeExit {
    /// The caller's shutdown future resolved (SIGTERM/SIGINT, or a test).
    Shutdown,
    /// No connection answered anything for the whole idle window.
    Idle,
}

/// How long an idle exit waits for a connection already in the kernel backlog.
///
/// Why (#6350): `connect(2)` against a listening socket succeeds the moment the
/// kernel queues it — the server need not have called `accept` yet. So a client
/// that dialled microseconds before the idle window elapsed is sitting in the
/// backlog, and a loop that returned straight from the idle arm would drop the
/// listener and unlink the socket with that client's connection still queued.
/// The client sees a reset, not a refusal, and only `trusty-review`'s adapter
/// retries one; `trusty-analyze deep` and `tctl`'s probe report it as a failure
/// the operator cannot act on.
///
/// What: 50ms is chosen against the cost of being wrong in each direction. A
/// queued connection is already in the backlog, so it is accepted on the first
/// poll and the window is never actually spent; the only case that spends it is
/// a genuinely idle service, which pays 50ms once in its whole lifetime.
const IDLE_EXIT_DRAIN: Duration = Duration::from_millis(50);

/// The future [`serve_until_idle`] races against `accept`.
///
/// Why a named function rather than an inline `async` block: two `async` blocks
/// have two different anonymous types, and the loop re-arms this one in place
/// after a drain. `Pin::set` needs the replacement to be the SAME type, which
/// only one `impl Future` origin gives.
///
/// What: [`IdleTracker::expired`] when a policy is configured. With none it
/// never resolves, so the arm is inert and the loop behaves as [`serve_until`]
/// always has.
async fn idle_expiry(idle: Option<Arc<IdleTracker>>) {
    match idle {
        Some(tracker) => tracker.expired().await,
        None => std::future::pending::<()>().await,
    }
}

/// What [`drain_backlog`] found in the backlog.
enum Drained {
    /// A client was queued and its request has been answered. The exit is
    /// cancelled and the idle window restarts.
    Answered,
    /// Nothing was queued, or what was queued asked nothing. The exit stands.
    Nothing,
}

/// Which arm of [`serve_until_idle`]'s `select!` won.
///
/// Why an enum rather than the drain running inside the arm: the drain re-arms
/// the pinned idle future afterwards, and `Pin::set` cannot run while the
/// `select!` still holds the `&mut` borrow of it. Naming the winner ends that
/// borrow at the `select!`'s closing brace.
enum Step {
    /// `accept` won; this is its result, error included.
    Accepted(std::io::Result<(UnixStream, tokio::net::unix::SocketAddr)>),
    /// The idle window elapsed.
    IdleElapsed,
}

/// Serve one connection queued before the idle window elapsed, if any.
///
/// Why: see [`IDLE_EXIT_DRAIN`]. Resolving the race in the CLIENT's favour is
/// the only safe direction — serving one extra request costs a round trip,
/// while resetting a connection the client believes it made costs that client a
/// failure it has no way to distinguish from a broken service.
///
/// 🔴 Why only an ANSWERED connection cancels the exit: a bare connect-and-close
/// is what [`crate::uds::socket_is_serving`] does on a poll loop, and a drain
/// that treated one as activity would hand the poller exactly the power
/// [`IdleGuard`] exists to deny it — an observed livelock, not a hypothetical.
/// A probe caught in the drain is served and the exit proceeds.
///
/// Why the connection is served INLINE rather than spawned: this is the last
/// thing that happens before the caller unlinks the socket and drops the
/// listener, and a spawned task would race the process exit. Awaiting it is what
/// guarantees the client has its answer first. One connection, once per
/// lifetime.
///
/// Test: `a_client_queued_when_the_idle_window_elapses_is_served_not_reset`,
/// `serve_until_idle_ignores_liveness_probes`.
async fn drain_backlog(
    listener: &UnixListener,
    router: &Arc<RpcRouter>,
    options: RpcServeOptions,
    idle: Option<&Arc<IdleTracker>>,
) -> Drained {
    let Ok(accepted) = tokio::time::timeout(IDLE_EXIT_DRAIN, listener.accept()).await else {
        return Drained::Nothing;
    };
    let stream = match accepted {
        Ok((stream, _)) => stream,
        Err(e) => {
            tracing::warn!(error = %e, "uds rpc listener accept failed during idle drain");
            return Drained::Nothing;
        }
    };
    let mut guard = idle.map(IdleTracker::connection_opened);
    let served = handle_connection(stream, Arc::clone(router), options).await;
    let answered = matches!(served, Ok(Served::Answered { .. }));
    if let Err(e) = &served {
        tracing::warn!(error = %e, "uds rpc connection failed during idle drain");
    }
    if answered && let Some(g) = guard.as_mut() {
        g.answered();
    }
    if let Some(g) = guard {
        g.release().await;
    }
    if answered {
        Drained::Answered
    } else {
        Drained::Nothing
    }
}

/// [`serve_until`], plus an optional idle-exit policy (#6350).
///
/// Why: ADR-0032 makes `trusty-analyze` an on-demand service — clients spawn it
/// and nothing supervises it — so the accept loop is the only place that knows
/// enough to end the process. It has to be here rather than in a timer the
/// service arms itself, because "idle" means no connection is OPEN and none has
/// answered recently, and only this loop observes both.
///
/// What: identical to [`serve_until`] with `idle` as `None`. With a tracker, the
/// loop additionally races [`IdleTracker::expired`] against `accept` and returns
/// [`ServeExit::Idle`] when it wins. Each accepted connection holds an
/// [`IdleGuard`] for its lifetime, so the window can never elapse under an open
/// connection; the guard is marked answered only for a connection that produced
/// a response, which is what keeps a liveness-probe poll loop from pinning the
/// process alive.
///
/// #6350: an expired window does not exit immediately. [`drain_backlog`] first
/// gives the kernel backlog [`IDLE_EXIT_DRAIN`] to yield a connection that was
/// queued before the window elapsed; one that appears is served like any other
/// and the loop continues, so the exit only stands when nobody was waiting.
///
/// Test: `serve_until_idle_exits_when_the_window_elapses`,
/// `serve_until_idle_is_reset_by_an_answered_request`,
/// `serve_until_idle_ignores_liveness_probes`,
/// `a_client_queued_when_the_idle_window_elapses_is_served_not_reset`.
pub async fn serve_until_idle(
    listener: &UnixListener,
    router: Arc<RpcRouter>,
    options: RpcServeOptions,
    shutdown: impl std::future::Future<Output = ()> + Send,
    idle: Option<Arc<IdleTracker>>,
) -> ServeExit {
    tokio::pin!(shutdown);
    let idle_expired = idle_expiry(idle.clone());
    tokio::pin!(idle_expired);

    loop {
        let step = tokio::select! {
            biased;
            () = &mut shutdown => return ServeExit::Shutdown,
            () = &mut idle_expired => Step::IdleElapsed,
            accepted = listener.accept() => Step::Accepted(accepted),
        };
        // #6350: the drain runs HERE, not inside the arm — the `select!`'s
        // borrow of `idle_expired` has ended, so the re-arm below can run. The
        // future is re-armed ONLY on this path: an `async` block panics if
        // polled after it completes, and rebuilding it on every iteration
        // instead would restart `expired`'s open-connection branch, which sleeps
        // a whole window whenever it observes a connection in flight — a client
        // polling faster than the window would then hold the process open
        // forever.
        let accepted = match step {
            Step::Accepted(accepted) => accepted,
            Step::IdleElapsed => {
                match drain_backlog(listener, &router, options, idle.as_ref()).await {
                    Drained::Answered => {
                        idle_expired.set(idle_expiry(idle.clone()));
                        continue;
                    }
                    Drained::Nothing => return ServeExit::Idle,
                }
            }
        };
        let stream = match accepted {
            Ok((stream, _)) => stream,
            Err(e) => {
                tracing::warn!(error = %e, "uds rpc listener accept failed");
                continue;
            }
        };
        // Taken BEFORE the connection task is spawned: taking it inside would
        // leave a window in which the count is zero while a connection is
        // already accepted, and the idle future could win the race in it.
        let guard = idle.as_ref().map(IdleTracker::connection_opened);
        let router = Arc::clone(&router);
        tokio::spawn(async move {
            // #6277 review: the connection runs in an INNER task whose
            // `JoinHandle` is awaited, because a panicking handler is otherwise
            // invisible. Tokio stores the panic payload in the handle and drops
            // it with the handle, so a dropped handle turns a caller-visible
            // hang-up into a server that logged nothing at all. Awaiting it
            // costs one task per connection and buys the one log line that says
            // which method took the process down a path nobody expected.
            let inner = tokio::spawn(handle_connection(stream, router, options));
            let mut guard = guard;
            match inner.await {
                Ok(Ok(Served::Answered { .. })) => {
                    if let Some(g) = guard.as_mut() {
                        g.answered();
                    }
                }
                Ok(Ok(Served::LivenessProbe)) => {
                    tracing::debug!("liveness probe connected and closed without a frame");
                }
                Ok(Err(e)) => {
                    tracing::warn!(error = %e, "uds rpc connection failed; nothing was answered");
                }
                // A panic is the handler's bug, not the client's, so it is an
                // ERROR rather than the WARN a refused or broken connection
                // gets. `JoinError`'s `Display` carries the panic message.
                Err(join) if join.is_panic() => {
                    tracing::error!(
                        error = %join,
                        "uds rpc handler panicked; the connection was dropped unanswered"
                    );
                }
                Err(join) => {
                    tracing::warn!(error = %join, "uds rpc connection task was cancelled");
                }
            }
            if let Some(g) = guard {
                g.release().await;
            }
        });
    }
}

/// A framed JSON-RPC server bound to one hardened socket.
///
/// Why: the shape a daemon's `serve` command wants — one value carrying the
/// path, the methods and the budgets, with [`run`] as its whole body.
/// What: [`RpcServer::run`] binds through [`bind_hardened`], serves until the
/// caller's shutdown future resolves, then unlinks the socket file.
/// Test: `server_round_trips_and_removes_its_socket_on_shutdown`.
///
/// [`run`]: RpcServer::run
#[derive(Debug)]
pub struct RpcServer {
    socket: PathBuf,
    router: Arc<RpcRouter>,
    options: RpcServeOptions,
}

impl RpcServer {
    /// A server that will bind `socket` and serve `router`.
    pub fn new(socket: impl Into<PathBuf>, router: RpcRouter) -> Self {
        Self {
            socket: socket.into(),
            router: Arc::new(router),
            options: RpcServeOptions::default(),
        }
    }

    /// Override the per-connection budgets.
    pub fn with_options(mut self, options: RpcServeOptions) -> Self {
        self.options = options;
        self
    }

    /// Socket this server binds.
    pub fn socket(&self) -> &Path {
        &self.socket
    }

    /// Bind and serve until `shutdown` resolves, then unlink the socket.
    ///
    /// [`bind_hardened`] refuses a path a socket file already occupies rather
    /// than clobbering what might be a live owner. A service that is respawned
    /// on demand, and so has to take over the corpse a killed predecessor left,
    /// binds with [`crate::uds::bind_singleton_hardened`] itself and drives
    /// [`serve_until`] directly — that decision is the caller's, which is the
    /// same split `bind_hardened`'s own docs make.
    ///
    /// # Errors
    ///
    /// [`RpcServerError::Bind`] when the path cannot be bound or hardened.
    /// Per-connection failures never reach here; they are logged and the loop
    /// continues.
    ///
    /// Test: `server_round_trips_and_removes_its_socket_on_shutdown`.
    pub async fn run(
        self,
        shutdown: impl std::future::Future<Output = ()> + Send,
    ) -> Result<(), RpcServerError> {
        let listener = bind_hardened(&self.socket).map_err(|source| RpcServerError::Bind {
            path: self.socket.clone(),
            source,
        })?;

        tracing::info!(
            socket = %self.socket.display(),
            methods = ?self.router.method_names().collect::<Vec<_>>(),
            streams = ?self.router.stream_names().collect::<Vec<_>>(),
            "uds rpc server bound"
        );

        serve_until(&listener, Arc::clone(&self.router), self.options, shutdown).await;

        // Unlink BEFORE dropping the listener — see the module docs. A failure
        // is not worth an exit code: the file is either already gone or belongs
        // to whoever rebound the path.
        if let Err(e) = std::fs::remove_file(&self.socket) {
            tracing::debug!(socket = %self.socket.display(), error = %e, "socket already gone");
        }
        drop(listener);
        Ok(())
    }
}