1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
//! LRU-bounded cache of per-palace `ChatSessionStore` handles (issue #4639).
//!
//! Why: `AppState::session_stores` was a plain `DashMap<String,
//! Arc<ChatSessionStore>>` with no `remove`, no TTL, and no cap — every palace
//! the daemon ever touched leaked one `chat_sessions.redb` file descriptor for
//! the process lifetime. A live daemon was measured holding 844 such handles
//! (all 844 pointing at files already unlinked from disk) against an 8 192 fd
//! ceiling, growing ~250-300/day. `PalaceRegistry` already solved exactly this
//! failure class for kg/usearch/recall via an LRU (issue #463); this module
//! applies the same shape to the one file type that registry never tracked.
//!
//! What: [`SessionStoreCache`] — a `parking_lot::Mutex<LruCache<..>>` (opened
//! unbounded, trimmed manually to [`SessionStoreCache::capacity`]) that opens a
//! store on miss, promotes on hit, and evicts from the cold end once resident
//! entries exceed the cap. Dropping the last `Arc` closes the redb `Database`
//! and releases its fd; the next request reopens from disk transparently.
//!
//! Two invariants make eviction safe, both enforced under the single cache
//! mutex:
//! 1. **Never evict a store a caller still holds.** redb takes an exclusive
//! `flock`, and `ChatSessionStore::open` has no snapshot fallback, so a
//! second open of a file that is still open *in this same process* fails
//! hard with `Database already open. Cannot acquire lock.` (reproduced
//! directly while diagnosing #4639). The chat streaming handler holds an
//! `Arc` across a whole SSE response (`chat::handler`), so unconditional
//! LRU eviction would turn a leak into an outage. Eviction therefore skips
//! any entry whose `Arc::strong_count() > 1` and overshoots the cap rather
//! than closing a store out from under its user.
//! 2. **Exactly one store per palace.** The open runs while the cache mutex
//! is held, so two concurrent callers for the same id cannot both open the
//! file (the `DatabaseAlreadyOpen` path `PalaceRegistry` guards with
//! per-id mutexes). The critical section is pure blocking I/O with no
//! `.await`, so `parking_lot::Mutex` is safe here.
//!
//! Test: `tests::open_handles_are_bounded_by_cap`,
//! `session_store_fd_count_is_bounded_by_cap` (in
//! `tests/session_store_fd_bound.rs` — needs its own process to measure real
//! fds; see that file's header),
//! `tests::evicted_store_reopens_with_data_intact`,
//! `tests::in_use_store_is_never_evicted`,
//! `tests::concurrent_callers_share_one_store`,
//! `tests::remove_drops_cached_handle`.
use Result;
use LruCache;
use Mutex;
use Path;
use Arc;
use ChatSessionStore;
/// Environment variable overriding the resident chat-session-store cap.
///
/// Why: mirrors `TRUSTY_MEMORY_MAX_OPEN_PALACES` so an operator close to the fd
/// ceiling can shrink the chat cache — or a host with a high limit can raise
/// it — without a rebuild.
/// What: parsed by [`max_open_session_stores_from_env`].
/// Test: `tests::env_override_is_honoured`.
pub const MAX_OPEN_SESSION_STORES_ENV: &str = "TRUSTY_MEMORY_MAX_OPEN_SESSION_STORES";
/// Default maximum number of `chat_sessions.redb` handles held open at once.
///
/// Why: deliberately half of `DEFAULT_MAX_OPEN_PALACES` (64), for two reasons.
/// (a) Access pattern: chat is a far colder path than KG recall — the measured
/// production workload opens a palace's session store once, appends a short
/// burst of turns, and never returns to it (~250-300 *distinct* palaces/day),
/// so cache hit rate past a small working set is ~0 and a larger cap buys
/// nothing but resident fds. (b) fd budget: 64 palaces × 3 registry files + 32
/// chat files = 224, which still fits inside the 256-fd macOS soft limit that
/// motivated issues #462/#463 — the fix holds even when the daemon runs outside
/// launchd's 8 192 ceiling. 32 also leaves ample headroom over the only
/// entries eviction cannot reclaim: concurrently-streaming chat sessions.
/// What: a compile-time constant, overridable per-instance via
/// [`SessionStoreCache::with_max_open`] or by [`MAX_OPEN_SESSION_STORES_ENV`].
/// Test: `tests::open_handles_are_bounded_by_cap`.
pub const DEFAULT_MAX_OPEN_SESSION_STORES: usize = 32;
/// Resolve the effective cap from the environment.
///
/// Why: centralises the parse so construction and diagnostics agree on both the
/// value and the fallback.
/// What: reads [`MAX_OPEN_SESSION_STORES_ENV`]; returns its parsed `usize` when
/// set to a value `>= 1`, else [`DEFAULT_MAX_OPEN_SESSION_STORES`].
/// Test: `tests::env_override_is_honoured`.
/// LRU-bounded, thread-safe cache of open per-palace chat-session stores.
///
/// Why/What: see the module docs. Cloning is cheap — callers share one instance
/// behind an `Arc` on `AppState`.
/// Test: see the module-level test list.