Expand description
Relevance-first lexical machinery: subtoken tokenizer, BM25, query shape.
This module is a lens, not the production ranker. PageRank answers structural importance; BM25 answers “does this text match the query”. Those scores stay separately inspectable. Do not fuse them here.
Structs§
- Bm25
Corpus Stats - Persisted corpus-wide BM25 statistics (Ripwire warm lexindex). IDF and
average length come from the full indexed corpus, not the candidate slice.
Stored in store meta (
bm25_corpus), never on FILE entity attributes. - LexDoc
- A document scored by the BM25 lens (symbol, component, route, …).
- LexField
- One weighted text field of a BM25 document.
- Query
Locus - FILE:LINE (and optional symbol) extracted from a stack/error query.
- Query
Mention - One explicit mention extracted from query/task text.
- Relevance
Hit - A relevance hit. BM25 and exact-anchor are separate numbers/flags.
- Retrieval
Plan - How the retrieval lens should spend its budget. Flags, not a fused score.
Enums§
- Query
Shape - Query shape from text alone. Conservative: under-fire rather than demote documents for a prose question that merely mentions a path.
- Ranking
Arm - Experimental ranking arms. Production default is blended importance (today’s surface). Other arms exist so evals can compare them.
Constants§
- BM25_B
- BM25 length normalization.
- BM25_K1
- Robertson/Sparck Jones BM25 term saturation.
- MIN_
SUBTOKEN_ LEN - Subtokens shorter than this are dropped (Ripwire ≥2-byte rule).
- QUERY_
MENTION_ MAX_ RAW - Cap on extracted mention tokens (Ripwire
kMentionMaxRawTokens). - WEIGHT_
BODY - Body / signature / callee vocabulary weight (Ripwire body=1).
- WEIGHT_
DOC - Doc-comment field weight (Ripwire doc=2).
- WEIGHT_
NAME - Identifier / symbol-name field weight (Ripwire name=3).
- WEIGHT_
PATH - File/path field weight.
Functions§
- bm25_
rank - Rank documents by BM25 descending, id ascending. Deterministic.
- bm25_
scores - BM25 scores, one per document, in input order. Empty query → all zeros.
- bm25_
scores_ with_ stats - BM25 using persisted corpus IDF/
avgdl. Term frequencies still come fromdocs. Whendocsis the same corpusstatswas built from, scores matchbm25_scores. - classify_
query - Classify query text. Pure; two runs agree by construction.
- extract_
loci - Extract
path:lineloci. A lonehost:portis ignored (no file extension). - extract_
query_ mentions - Extract path / dotted / backtick mentions from task text. Plain prose words never qualify.
- is_
exact_ anchor - True when
nameis an exact anchor forqueryafter identifier normalization (case-insensitive, subtoken-set equal, or last path segment). - mention_
matches_ doc - True when a mention names this document’s entity name or path.
- path_
matches_ locus - Indexed path matches a stack-frame path: equal, or a
/-anchored suffix.extra.pydoes not match locusa.py. - ranking_
arm_ ids - All ranking-arm ids. Tests fail if an arm is added to the enum and omitted here — same SoT pattern as language/predicate registries.
- relevance_
hits - Score documents with BM25 and mark exact anchors. Does not blend PPR.
Uses candidate-set IDF (cold). Prefer
relevance_hits_with_statswhen corpus-wide stats are persisted. - relevance_
hits_ with_ stats - Same as
relevance_hitsbut BM25 uses persisted corpus IDF/avgdl. Experimental lens only — productionbuild_surfacemust not call this. - route_
query - Route a query to a retrieval plan. Does not score the corpus.
- subtokens
- One tokenizer for query and documents. Port of Ripwire
forEachLexSubtoken: alphanumeric runs, split on camelCase and ACRONYMWord (HTTPServer→HTTP+Server), keep all-caps acronyms (MCP→mcp).