1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
//! Bounded, streaming palace fan-out for the recall-all family (issue #7125).
//!
//! Why: the three recall-all entry points opened every palace on disk into one
//! `Vec<Arc<PalaceHandle>>` and held all of them for the whole query. On a
//! 94-palace estate that hydrates the drawer table, HNSW graph, KG adjacency
//! and redb page cache of all 94 at once — roughly 13x the on-disk bytes — and
//! when the local `Vec` drops the LRU is left holding its 64 most recent
//! victims, a multi-GB floor that persists until the idle sweep runs. The
//! 64-slot cap bounded the cache, never the peak, and never the residue. Peak
//! residency must track the batch size, not the palace count, and the daemon's
//! steady state after a recall-all must match its steady state before one.
//! What: [`recall_streamed`] walks the palace list in [`RECALL_PALACE_BATCH`]
//! batches. Each batch is opened on the blocking pool, handed to the caller's
//! search, then released — so at most one batch is resident at a time. Palaces
//! the registry already held when the call started are never released, so an
//! operator's working set survives a recall-all untouched. Per-batch results
//! are merged with the same dedup-by-drawer-id/sort/truncate rule
//! `recall_across_palaces` applies within a batch, so the global top-k is
//! unchanged: it is always a subset of the union of the per-batch top-ks.
//! Test: `recall_all_returns_open_palaces_to_baseline` and
//! `recall_all_keeps_palaces_that_were_already_open` in
//! `tests/recall_all_residency.rs`;
//! `recall_all_releases_every_batch_when_the_search_fails` and
//! `recall_streamed_visits_every_palace_in_bounded_batches` in
//! `recall_stream_tests`.
use crateAppState;
use Result;
use ;
use Future;
use Arc;
use ;
use CrossPalaceResult;
use PalaceHandle;
use Uuid;
use open_palaces_blocking;
/// How many palaces a recall-all holds open at once (issue #7125).
///
/// Why: this is the peak-residency knob. Each open palace costs its drawer
/// table, HNSW graph, KG adjacency and redb page cache — on the estate that
/// prompted #7125, ~13x its on-disk bytes across 94 palaces. Eight keeps the
/// fan-out's own concurrency (`recall_across_palaces` joins the batch) worth
/// having while bounding peak residency at roughly an eighth of the old
/// whole-estate open. Raising it trades RAM for wall-clock; lowering it does
/// the reverse.
/// What: the chunk width [`recall_streamed`] slices the palace list into.
/// Test: `recall_all_returns_open_palaces_to_baseline` exercises an estate
/// wider than one batch, so the batching loop actually iterates.
pub const RECALL_PALACE_BATCH: usize = 8;
/// Run `search` over every palace in `palaces`, one bounded batch at a time.
///
/// Why (#7125): see the module doc — the old shape's peak and its residue both
/// scaled with the palace count. This is the shared replacement, so all three
/// recall-all entry points (`MemoryService::recall_all`,
/// `handle_memory_recall_all`, `chat::tools::execute_recall_all`) inherit the
/// bound rather than each growing their own.
/// What: snapshots which palaces the registry already holds, then for each
/// batch opens the handles (on the blocking pool, via `open_palaces_blocking`),
/// awaits `search`, and releases every handle the batch itself brought in.
/// Release runs BEFORE the caller's error is propagated, so a failing search
/// cannot leak a batch. `search` takes the handles by value and its future is
/// dropped before the release, so the registry sees `strong_count == 1` and
/// `release_if_unreferenced` can actually pop; a palace another task is holding
/// stays put by that same guard. Results are merged across batches by drawer
/// id (highest score wins), sorted by score descending, and truncated to
/// `top_k`.
/// Test: `recall_all_returns_open_palaces_to_baseline`,
/// `recall_all_keeps_palaces_that_were_already_open`,
/// `recall_all_releases_every_batch_when_the_search_fails`,
/// `recall_streamed_visits_every_palace_in_bounded_batches`.
pub async
/// Ids the registry holds open right now.
///
/// Why (#7125): the baseline a recall-all must return to. Taken once, before
/// the first batch opens anything, so it names the caller's working set rather
/// than the query's own footprint.
/// What: snapshots `PalaceRegistry::list` into a set for O(1) membership.
/// Test: `recall_all_keeps_palaces_that_were_already_open`.
/// Hand back the handles this batch opened, keeping the pre-call working set.
///
/// Why (#7125): without this the LRU keeps whatever the fan-out last touched,
/// so the daemon's floor after a recall-all is `min(palace_count, max_open)`
/// fully hydrated palaces regardless of what the operator was actually using.
/// What: for each id the batch opened that was NOT already resident, calls
/// `PalaceRegistry::release_if_unreferenced` (#7106), which pops only when the
/// cache holds the sole reference — so a concurrent recall, remember, or dream
/// cycle holding the same `Arc` is never disturbed. redb stays the source of
/// truth; the next access transparently reopens.
/// Test: `recall_all_returns_open_palaces_to_baseline`,
/// `recall_all_keeps_palaces_that_were_already_open`.
/// Fold one batch's hits into the running cross-palace result set.
///
/// Why (#7125): batching splits a fan-out that used to merge once, so the
/// dedup-by-drawer-id rule `recall_across_palaces` applies within a batch has
/// to be reapplied across batches or the same drawer could appear twice.
/// What: keeps the highest-scoring occurrence of each drawer id, indexing into
/// `merged` through `by_drawer` so a later higher score replaces in place.
/// Test: `recall_all_returns_open_palaces_to_baseline` drives the merge across
/// more than one batch.