1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
//! Sizing `SO_SNDBUF` / `SO_RCVBUF` on every Unix socket this module owns
//! (#6896).
//!
//! Why: macOS ships `net.local.stream.sendspace` and `.recvspace` at 8192
//! bytes, and nothing in this workspace ever raised them. Every frame a daemon
//! exchanges is multi-KiB at least and multi-MiB at the budget, so it moves
//! through an 8 KiB pipe in hundreds or thousands of write-then-drain round
//! trips, each one a task park and unpark. Measured off the #6876 fix on this
//! host, the two frame-budget exchanges cost ~235 ms and ~204 ms unloaded and
//! inflated about 45x, to ~11 s and ~9 s, under 480 concurrent processes, while
//! connection setup stayed at ~1 ms. The cost is the round trips, not the bytes.
//!
//! What: two entry points, differing only in what they do when the socket's
//! peer has already hung up. Both set both buffers to [`SOCKET_BUFFER_BYTES`]
//! and read back what the kernel granted, because the kernel clamps rather than
//! refuses.
//! - [`tune_listener_buffers`] is for a socket that has no peer and never
//! will: any failure is a failure. [`crate::uds::bind_hardened`] calls it.
//! - [`tune_connected_buffers`] is for a socket that has a peer, which may
//! have gone away. [`crate::uds::connect_hardened`] calls it for the
//! dialling end and [`crate::uds::accept_sized`] for the accepted end.
//!
//! 🔴 **Why the split is not a stylistic one (#6896 review).** The benign
//! classification rests on `getpeername` failing, and `getpeername` ALWAYS
//! fails on a listening socket — it has no peer to name. A single forgiving
//! entry point therefore degrades, on the listener, to trusting the errno
//! alone: any `EINVAL` from `setsockopt` would be recorded as a hung-up peer,
//! and `bind_hardened` would return a listener silently left at the platform
//! default. The strict form has no benign outcome to reach, in the type as well
//! as in the code.
//!
//! 🔴 **Both the listener AND each accepted socket are sized, because only
//! macOS inherits (#6896 follow-up).** macOS builds the server-side socket in
//! `sonewconn`, which copies the listener's `sb_hiwat` onto it, so sizing the
//! listener there covers everything `accept` returns. Linux does not: AF_UNIX
//! builds the server-side socket from scratch in `unix_stream_connect`, so it
//! comes back at `net.core.wmem_default` / `rmem_default` however the listener
//! was sized. Sizing only the listener therefore left every server-side socket
//! on Linux at the platform default. [`crate::uds::accept_sized`] closes that,
//! through the forgiving entry point: an accepted socket can have lost its peer
//! before the server touches it — a liveness probe connects and closes, which
//! is exactly what [`crate::uds::probe::socket_is_serving`] does — and that
//! peer-hangup is what [`SocketBufferOutcome::PeerHungUp`] absorbs rather than
//! turning every probe into a connection failure.
//!
//! **A failure is propagated, never defaulted — with one named exception.** A
//! socket left at the platform default still works, at the cost this module
//! exists to remove and silently, so a real `setsockopt` failure is an error.
//! The exception is a CONNECTED socket whose peer has already hung up: it is
//! reported as [`SocketBufferOutcome::PeerHungUp`] rather than sized, because
//! it will carry no frame at all and the caller's next read or write is what
//! tells it so.
//!
//! **Linux is unaffected or better.** `net.core.wmem_default` is 212992 there,
//! 26x the macOS figure, a request above `wmem_max` is clamped silently rather
//! than refused, and `setsockopt` does not consult the connection state — so
//! this raises the buffer where the host allows it and changes nothing where it
//! does not.
//!
//! 🟡 **Nothing here compares a read-back against what was requested.** Linux
//! stores twice the accepted value and `getsockopt` returns that doubled
//! figure, so a request of 1 MiB reads back as 2 MiB — or as 425984 on a host
//! whose `wmem_max` is 212992, the clamp applied first and the doubling after.
//! The read-back is logged and returned for observation, never asserted
//! against [`SOCKET_BUFFER_BYTES`]. A caller comparing two sockets on the same
//! host is comparing like with like; a caller comparing either against the
//! request is not.
//!
//! Test: `tune_connected_buffers_raises_both_buffers_on_a_socketpair`,
//! `tune_connected_buffers_reports_a_peer_that_already_hung_up`,
//! `a_listener_can_never_classify_a_failure_as_a_hung_up_peer`,
//! `tune_listener_buffers_refuses_where_the_connected_form_tolerates`,
//! `accept_sized_raises_the_accepted_socket_to_the_listeners_sizing`,
//! `hardened_sockets_hold_far_more_in_flight_than_the_platform_default`.
use io;
use AsRawFd;
use UdsSecurityError;
use MAX_FRAME_BYTES;
/// Bytes requested for `SO_SNDBUF` and `SO_RCVBUF` on every socket this module
/// creates or accepts.
///
/// Why an eighth of [`MAX_FRAME_BYTES`] rather than the whole frame: the buffer
/// is charged to the kernel per socket, and the load this issue was measured
/// under is 480 concurrent processes. A full 8 MiB frame budget on both ends of
/// each of those is memory the machine does not have to spare, for a saving
/// that has already been taken — 1 MiB is 128x the macOS default and drops an
/// 8 MiB frame from 1024 round trips to 8. Neither platform refuses a larger
/// request (macOS clamps to `kern.ipc.maxsockbuf`, Linux to
/// `net.core.wmem_max`), so this is a memory decision, not a compatibility one.
///
/// 🔴 **One size for every consumer, and it is not proportional to any
/// consumer's own frame budget (#6896 review).** A service that raises its
/// budget above [`MAX_FRAME_BYTES`] gets this same 1 MiB: `trusty-memory`
/// serves 32 MiB frames, and a 32 MiB frame here still costs about 32 round
/// trips rather than the 4 a proportional buffer would give. That is the
/// deliberate trade — 32 round trips against the ~4,100 the 8 KiB default
/// charged — not an oversight, and not something a per-consumer parameter
/// should be added to address.
///
/// Test: `socket_buffer_request_stays_within_the_frame_budget`.
pub const SOCKET_BUFFER_BYTES: usize = as usize;
/// What the kernel actually granted, after its own clamping.
///
/// Why read back at all: `setsockopt` succeeds while granting less than was
/// asked for — Linux clamps to `net.core.wmem_max`, macOS to
/// `kern.ipc.maxsockbuf`, and both round. A caller that assumed the request was
/// honoured would report a buffer size the socket does not have.
///
/// The two figures are not comparable across platforms: Linux doubles the value
/// it reports to account for its own bookkeeping, macOS does not.
///
/// Test: `tune_connected_buffers_raises_both_buffers_on_a_socketpair`.
/// Whether [`tune_connected_buffers`] had a live socket to size.
///
/// Why this is not folded into an error: a socket whose peer has hung up is not
/// a failure to report — the caller's next read or write reports it, in the
/// shape the caller already handles (`UdsRpcError::NoResponse`,
/// `Served::LivenessProbe`). Naming the case keeps it distinguishable from a
/// socket that was genuinely sized, without turning a routine liveness probe
/// into a connection error.
///
/// [`tune_listener_buffers`] does not return this type at all — see the module
/// docs for why a listener must not be able to reach this outcome.
///
/// Test: `tune_connected_buffers_reports_a_peer_that_already_hung_up`.
/// Whether `fd` still has a peer.
///
/// 🔴 Always false for a LISTENING socket, which has no peer to name — which is
/// why [`hangup_is_benign`] takes `connected` rather than relying on this alone.
/// Whether a `setsockopt` failure on `fd` may be recorded as a hung-up peer
/// rather than reported.
///
/// Why a named function: this is the whole fail-closed decision, and it is the
/// one the #6896 review found could not be made from the errno and
/// `getpeername` alone. `connected` is the caller's statement about what kind
/// of socket it holds; a listener passes `false` and can never reach the benign
/// arm, whatever the errno and whatever `getpeername` says.
///
/// Test: `a_listener_can_never_classify_a_failure_as_a_hung_up_peer`.
pub
/// Ask for `bytes` on one buffer option.
/// Read one buffer option back.
/// Read a socket's effective buffer sizes without changing them.
///
/// Why: the only way to observe what a socket carries. What the kernel granted
/// is not what was asked for — Linux doubles it and clamps it first — so a
/// caller checking that a socket really is sized has to read it back rather
/// than assume (#6896).
///
/// # Errors
///
/// [`UdsSecurityError::SocketBufferRead`] when either option cannot be read.
///
/// Test: `accept_sized_raises_the_accepted_socket_to_the_listeners_sizing`.
Sized>
/// Set both buffers, tolerating a hung-up peer only when `connected` says the
/// socket has one.
/// Size the buffers of a socket that has no peer, failing on any error.
///
/// Why: a listener is the one socket for which the hung-up-peer classification
/// cannot be evaluated — see the module docs. There is no benign outcome in the
/// return type, so a caller cannot proceed on an unsized listener without
/// seeing an error (#6896 review).
///
/// What: `setsockopt` for both options, then `getsockopt` for both, at debug
/// level. macOS copies what this sets onto every socket `accept` returns; Linux
/// does not, which is why [`crate::uds::accept_sized`] sizes the accepted end
/// as well.
///
/// # Errors
///
/// [`UdsSecurityError::SocketBuffer`] when either option cannot be set,
/// [`UdsSecurityError::SocketBufferRead`] when either cannot be read back.
///
/// Test: `a_listener_can_never_classify_a_failure_as_a_hung_up_peer`,
/// `tune_listener_buffers_refuses_where_the_connected_form_tolerates`,
/// `accept_sized_raises_the_accepted_socket_to_the_listeners_sizing`.
Sized>
/// Size the buffers of a connected socket, reporting a peer that has already
/// hung up rather than failing on it.
///
/// Why: the dialling end of a pair the server may have dropped between
/// `connect` returning and this call — the #6601 shutdown drain does exactly
/// that. Such a socket will carry no frame, so there is nothing to size and
/// nothing to report as broken.
///
/// What: as [`tune_listener_buffers`], except that a `setsockopt` failure whose
/// errno says the socket buffer is torn down AND whose `getpeername` confirms
/// there is no peer returns [`SocketBufferOutcome::PeerHungUp`]. Every other
/// failure is an error.
///
/// # Errors
///
/// [`UdsSecurityError::SocketBuffer`] when either option cannot be set on a
/// still-connected socket, [`UdsSecurityError::SocketBufferRead`] when either
/// cannot be read back.
///
/// Test: `tune_connected_buffers_raises_both_buffers_on_a_socketpair`,
/// `tune_connected_buffers_reports_a_peer_that_already_hung_up`,
/// `tune_listener_buffers_refuses_where_the_connected_form_tolerates`.
Sized>