1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
//! Relevance floor for recall results (#5037).
//!
//! Why: every retrieval path in [`super::layers`] ends in `truncate(top_k)` and
//! nothing else — `top_k` is the only length control, so a query with no good
//! answer still returns a full `top_k` of whatever ranked highest. This module
//! is the missing half: a minimum score a candidate must clear to be shown at
//! all, plus a count of what was dropped so the caller can say so out loud.
//! What: [`DEFAULT_RELEVANCE_FLOOR`], [`FloorOutcome`], and
//! [`apply_relevance_floor`]. Pure — no I/O, no allocation beyond the returned
//! vector.
//! Test: `relevance_floor_drops_l1_penalty_noise`,
//! `relevance_floor_keeps_genuine_hit`, `relevance_floor_keeps_unscored_items`,
//! `relevance_floor_counts_withheld`.
//!
//! Lives beside `layers.rs` rather than inside it because `layers.rs` sits at
//! 464 of its 500 SLOC budget; the repo's cap forces the split
//! (`docs/reference/sloc-cap.md`).
/// Minimum score a recall candidate must reach to be worth showing.
///
/// Why: `rescore_l1_by_similarity` assigns an L1 drawer the HNSW search never
/// returned a score of `importance * L1_NO_SIMILARITY_PENALTY` — at most
/// `0.15`. That is a nonzero score, so it competes for a `top_k` slot and
/// renders identically to a real hit. Measured against the live `trusty-tools`
/// palace (1,332 drawers) over four query populations:
///
/// | population | n | min | median | max |
/// |---|---|---|---|---|
/// | 15 off-topic prompts × top-10 ("capital of France", "go", "ok", …) | 150 | 0.1500 | 0.1510 | **0.3439** |
/// | 60 self-retrieval queries, score of the *correct* drawer | 57 | **0.4844** | 0.6950 | 0.9743 |
/// | 120 real logged hook prompts × top-10 | 1200 | 0.4042 | 0.5422 | 0.7527 |
/// | 12 paraphrased on-topic questions, best hit per query | 12 | 0.3293 | 0.4317 | 0.5359 |
///
/// 75 of the 150 off-topic candidates scored *exactly* 0.1500 — the L1 penalty
/// itself. The noise and signal distributions are nearly disjoint, and the gap
/// between them (0.3439 … 0.4042) is where this constant sits.
/// What: `0.35` — the smallest swept value at which zero off-topic candidates
/// survive, rounded up from the highest observed noise score (0.3439). At this
/// floor the real logged corpus loses 0 of 1200 candidates and 0 of 120
/// queries; self-retrieval keeps 57 of 57 correct drawers; 10 of 12 paraphrased
/// questions still return something. The 2 that stop returning are short
/// keyword-bearing queries ("what is the MSRV for this workspace", 0.3293),
/// which is dense retrieval's known weak spot and BM25's strength — #5036 is
/// the named fix for that class. Their loss is reported, never silent: see
/// [`FloorOutcome::withheld`].
/// Test: `relevance_floor_drops_l1_penalty_noise` pins the 0.15 case;
/// `relevance_floor_keeps_genuine_hit` pins the 0.56 case.
pub const DEFAULT_RELEVANCE_FLOOR: f32 = 0.35;
/// What cleared a relevance floor, and how much did not.
///
/// Why: dropping weak candidates silently turns "we found nothing good" into
/// something indistinguishable from "this palace is empty". Carrying the count
/// alongside the survivors is what lets a caller render the difference.
/// What: `kept` in the input's order; `withheld` is how many were dropped.
/// Test: `relevance_floor_counts_withheld`.
/// Drop candidates scoring below `floor`, counting what went.
///
/// Why: the one implementation of "below the floor, it is not shown" — the
/// prompt-context injection and any future recall consumer must agree on the
/// comparison and on how an unscored candidate is treated, or the two drift.
/// What: keeps every item whose `score_of` is `>= floor`. `score_of` returning
/// `None` means the score is unknown, and an unknown-score item is **kept** and
/// not counted as withheld — dropping what cannot be judged would let a wire
/// format that stops emitting `score` silently empty the result set. A
/// non-positive or NaN `floor` disables the gate entirely.
/// Test: `relevance_floor_drops_l1_penalty_noise`,
/// `relevance_floor_keeps_genuine_hit`, `relevance_floor_keeps_unscored_items`,
/// `relevance_floor_counts_withheld`.