Skip to main content

Module lex

Module lex 

Source
Expand description

Relevance-first lexical machinery: subtoken tokenizer, BM25, query shape.

This module is a lens, not the production ranker. PageRank answers structural importance; BM25 answers “does this text match the query”. Those scores stay separately inspectable. Do not fuse them here.

Structs§

Bm25CorpusStats
Persisted corpus-wide BM25 statistics (Ripwire warm lexindex). IDF and average length come from the full indexed corpus, not the candidate slice. Stored in store meta (bm25_corpus), never on FILE entity attributes.
LexDoc
A document scored by the BM25 lens (symbol, component, route, …).
LexField
One weighted text field of a BM25 document.
QueryLocus
FILE:LINE (and optional symbol) extracted from a stack/error query.
QueryMention
One explicit mention extracted from query/task text.
RelevanceHit
A relevance hit. BM25 and exact-anchor are separate numbers/flags.
RetrievalPlan
How the retrieval lens should spend its budget. Flags, not a fused score.

Enums§

QueryShape
Query shape from text alone. Conservative: under-fire rather than demote documents for a prose question that merely mentions a path.
RankingArm
Experimental ranking arms. Production default is blended importance (today’s surface). Other arms exist so evals can compare them.

Constants§

BM25_B
BM25 length normalization.
BM25_K1
Robertson/Sparck Jones BM25 term saturation.
MIN_SUBTOKEN_LEN
Subtokens shorter than this are dropped (Ripwire ≥2-byte rule).
QUERY_MENTION_MAX_RAW
Cap on extracted mention tokens (Ripwire kMentionMaxRawTokens).
WEIGHT_BODY
Body / signature / callee vocabulary weight (Ripwire body=1).
WEIGHT_DOC
Doc-comment field weight (Ripwire doc=2).
WEIGHT_NAME
Identifier / symbol-name field weight (Ripwire name=3).
WEIGHT_PATH
File/path field weight.

Functions§

bm25_rank
Rank documents by BM25 descending, id ascending. Deterministic.
bm25_scores
BM25 scores, one per document, in input order. Empty query → all zeros.
bm25_scores_with_stats
BM25 using persisted corpus IDF/avgdl. Term frequencies still come from docs. When docs is the same corpus stats was built from, scores match bm25_scores.
classify_query
Classify query text. Pure; two runs agree by construction.
extract_loci
Extract path:line loci. A lone host:port is ignored (no file extension).
extract_query_mentions
Extract path / dotted / backtick mentions from task text. Plain prose words never qualify.
is_exact_anchor
True when name is an exact anchor for query after identifier normalization (case-insensitive, subtoken-set equal, or last path segment).
mention_matches_doc
True when a mention names this document’s entity name or path.
path_matches_locus
Indexed path matches a stack-frame path: equal, or a /-anchored suffix. extra.py does not match locus a.py.
ranking_arm_ids
All ranking-arm ids. Tests fail if an arm is added to the enum and omitted here — same SoT pattern as language/predicate registries.
relevance_hits
Score documents with BM25 and mark exact anchors. Does not blend PPR. Uses candidate-set IDF (cold). Prefer relevance_hits_with_stats when corpus-wide stats are persisted.
relevance_hits_with_stats
Same as relevance_hits but BM25 uses persisted corpus IDF/avgdl. Experimental lens only — production build_surface must not call this.
route_query
Route a query to a retrieval plan. Does not score the corpus.
subtokens
One tokenizer for query and documents. Port of Ripwire forEachLexSubtoken: alphanumeric runs, split on camelCase and ACRONYMWord (HTTPServer → HTTP+Server), keep all-caps acronyms (MCP → mcp).