# rq architecture (target model)
This is the **design we are building toward**. No code exists yet; this
document is the contract the implementation should satisfy.
[ROADMAP.md](ROADMAP.md) tracks what ships in which phase.
## Core principle
`rq` is a **navigation engine**. It optimizes for reaching the one result a
developer most likely wants, fast — not for enumerating every match. Three
ranked priorities resolve every design tension:
1. relevance over completeness
2. navigation over discovery
3. speed over exhaustiveness
The latency target is **< 50 ms perceived** for index-backed results, then
*progressive improvement* — slower layers stream in behind the fast first
answer. This forces one early commitment: **results are a stream, not a
synchronous list.** Everything below assumes that.
## Implementation language
Rust. The latency target effectively requires a compiled language with
near-zero startup cost; Rust also has first-class Tree-sitter bindings and
ships as a single static binary (the `rg`/`fd`/`fzf` feel we are matching). A
scripting-language runtime's startup alone would consume the whole 50 ms
budget.
## The common symbol model
Every language plugin emits the same shape. The core never sees a
language-specific concept.
```text
Symbol {
repository # which repo it belongs to
language # ruby, go, ts, ...
name # RefundProcessor, perform, User
line # 1-based
parent # enclosing symbol (cheap nesting, NOT a call graph)
}
```
`parent` records lexical nesting only (`Foo::Bar#baz`). It is **not** reference
tracking or inheritance — those are explicit non-goals for the MVP.
## Repository identity — two levels
Identity answers two different questions, so it is modeled at two levels:
- **Logical project** — `github.com/org/repo` (from the upstream remote) or
`local:/abs/path` fallback. Used to dedupe symbols across checkouts. Robust
to forks/clones being the "same" project.
- **Local checkout** — a root path plus current branch. Used for indexing
coverage state and git-aware ranking. One project may have several checkouts
(multiple clones, all valid). A checkout whose path no longer exists is pruned
when the repo is next indexed/warmed (not on every search — stale rows are
cheap, since reads route around them), so a moved repo self-heals; symbols
are keyed by identity, so pruning a checkout only forgets a *location*.
The system is designed for **many** repositories and millions of symbols from
day one. It never assumes a single repository.
## Module layout
Language-agnostic core; language specifics quarantined under `lang/`.
```text
src/
cli/ # `rq <query>` default command, arg parsing, output
core/ # symbol model, repository identity, scoring — NO language specifics
store/ # SQLite schema, migrations, queries (WAL mode)
index/ # walker, incremental indexer, coverage tracking
search/ # staged pipeline, scorer, --explain
lang/ # Tree-sitter plugins: ruby, rust, go, python, typescript
ruby/ # the first plugin
rust/ # what rq dogfoods on its own source
```
A `LanguagePlugin` trait is the only seam languages plug into:
```rust
trait LanguagePlugin {
fn extensions(&self) -> &[&str];
fn extract(&self, source: &str) -> Vec<Symbol>;
}
```
A registry maps file extension → plugin. Adding Java/C# is a new
plugin. The one shared thing a language may extend is the `core::Kind`
vocabulary — Rust added `struct`/`enum`/`trait`, then `type`/`macro`/`variant` — which generalizes the model
rather than leaking a language into `index`/`search`/scoring.
## SQLite schema
WAL mode is mandatory — the background indexer writes while searches read.
```sql
PRAGMA journal_mode = WAL;
-- a logical project
repositories (
id INTEGER PRIMARY KEY,
created_at INTEGER, updated_at INTEGER
);
-- a local clone of a repository
checkouts (
id INTEGER PRIMARY KEY,
repository_id INTEGER NOT NULL REFERENCES repositories(id),
root_path TEXT NOT NULL UNIQUE,
current_branch TEXT
);
files (
id INTEGER PRIMARY KEY,
repository_id INTEGER NOT NULL REFERENCES repositories(id),
path TEXT NOT NULL, -- repo-relative
language TEXT,
mtime INTEGER, -- unix *nanoseconds* (racy-edit protection)
content_hash TEXT, -- staleness detection
indexed_at INTEGER,
generated INTEGER NOT NULL DEFAULT 0, -- header declares it generated (v19)
UNIQUE(repository_id, path)
);
symbols (
id INTEGER PRIMARY KEY,
repository_id INTEGER NOT NULL REFERENCES repositories(id),
file_id INTEGER NOT NULL REFERENCES files(id),
name TEXT NOT NULL,
name_lower TEXT NOT NULL, -- prefix / ranking
line INTEGER NOT NULL,
end_line INTEGER, -- 1-based last line of the definition body
-- (NULL for rows indexed before v4)
parent TEXT, -- enclosing symbol's qualified NAME
-- (lexical nesting only), e.g. Foo::Bar
-- backfill lazily)
stub INTEGER NOT NULL DEFAULT 0 -- declares what is defined elsewhere
-- (a .d.ts entry, an overload signature)
);
-- exact and prefix recall: every name query is scoped by repository, even
-- unscoped (`-a`) ones, which seek it once per repo through `repositories`
-- (also serves the per-repo counts the old repository_id index did). The
-- trigram FTS table and the name_lower-only index went in v18 (D26).
CREATE INDEX idx_symbols_repo_name ON symbols(repository_id, name_lower);
-- the name index (NAME_INDEX.md, D23): each repo's distinct symbol names
-- (kind 0) and file paths (kind 1, signed by their stem) as 40-byte
-- signatures, in append-order chunks of up to 512
name_sigs (
repository_id INTEGER NOT NULL,
kind INTEGER NOT NULL,
chunk INTEGER NOT NULL,
n INTEGER NOT NULL, -- keys in the chunk
sigs BLOB NOT NULL, -- n signatures, back to back
keys BLOB NOT NULL, -- n end offsets (u32), then the keys' bytes
PRIMARY KEY (repository_id, kind, chunk)
);
-- a repo's name index is read only while current: built under this format
-- (score::NAME_INDEX_FORMAT) and maintained since. Missing or another format
-- is rebuilt before recall reads it; -1 while a cold pass suspends it (built
-- then holds the pass's pid: recall rebuilds one whose pass has died).
name_index (
repository_id INTEGER PRIMARY KEY,
format INTEGER NOT NULL,
built INTEGER NOT NULL -- keys the last rebuild wrote
);
-- partial-indexing state, per repo (or directory scope)
coverage (
id INTEGER PRIMARY KEY,
repository_id INTEGER NOT NULL REFERENCES repositories(id),
scope TEXT NOT NULL DEFAULT 'full', -- 'full' or a directory prefix
files_seen INTEGER, files_indexed INTEGER,
UNIQUE(repository_id, scope)
);
-- cumulative usage counters, read by `--usage`, never by ranking.
usage_daily (
day TEXT NOT NULL, -- local date, YYYY-MM-DD
source TEXT NOT NULL,
flags TEXT NOT NULL,
searches INTEGER NOT NULL,
misses INTEGER NOT NULL, -- answered nothing, against a ready index
warming INTEGER NOT NULL, -- answered nothing because it wasn't ready
on_complete INTEGER NOT NULL, -- ran against a fully indexed repo
live INTEGER NOT NULL, -- answered from a live scan, not the index
PRIMARY KEY (day, source, flags)
);
-- small key/value store (indexed HEAD, warm lock, warm verdict, branch-file
-- cache, and the files the index holds as uncommitted edits)
meta ( key TEXT PRIMARY KEY, value TEXT NOT NULL );
```
Decisions worth calling out:
- **The name index** holds, per repo, a signature for every distinct symbol
name and file stem: which characters it has and which pairs of them a query
could step across under `align`'s rules. Fuzzy recall screens every
signature, verifies the survivors with the scorer's own match chain, and
fetches rows only for the names it accepts, so its candidates are exactly
what `score` would take from any row (NAME_INDEX.md, D23). The default
since D24 fixed the ranking weaknesses complete recall exposed, and the only
fuzzy recall since D26 removed the trigram FTS nets it replaced. A repo whose
index is missing or from another format is rebuilt before recall reads it,
and one suspended by a live cold pass is verified from its rows (D25), as is one
whose rebuild finds another writer holding the lock past the busy timeout
(D26).
- **`content_hash`** detects staleness so partial/old indexes don't silently
point at moved lines.
- **An extraction change re-extracts by migration.** When a plugin starts
emitting something new, a schema step clears the language's `mtime` and
`content_hash` so neither skip keeps the old rows (the hash to `''`, not
NULL, which the write path can't read), and demotes its repos' coverage to
`warming` so the next search sweeps them. v14 did this for the Go, Python and
TS/JS constants, v19's re-read of every file (for the `generated` flag)
also picked up Rust's variants, aliases, macros and macro-body items, and v20
re-reads Go, Python and TS/JS for their types, variants and nested defs, and v21
re-reads TS/JS for ambient declarations and overload signatures and Python
for its local classes: users
upgrade and the symbols appear, with no `--drop`. Old
symbols stay readable until each file is rewritten.
- **`coverage`** lets search know its own confidence and decide whether to
append a live-scan tail.
- **A miss and a not-yet are counted apart.** rq already separates them in its
exit codes (1 = absent, 2 = index still warming); netting them into one
number would overstate how often it truly finds nothing, and the two call for
opposite responses — index more, versus the symbol isn't there.
- **`usage_daily` is observability, not ranking input.** Nothing reads it
back into scoring, so counting a search can never move a result. Counters
are incremented on write rather than kept as a raw log: the question "how
much is rq used, and by whom" needs a total, and a bounded log can only give
a ceiling.
- **Counted after the answer.** A search's usage write runs once its results
are printed, so a slow write never delays them (DECISIONS D13). `--show`,
`--open` and `--web` count before they fork, since `--open` `exec`s.
## Indexing model
Indexing is **decoupled** from search — a background worker parses and writes;
search only reads.
- **One core, two entry points** — explicit (`index_under`, unbounded) and
opportunistic (`index_budgeted`, time-bounded) both call `run_index`, which
differs only by parameters (active files, subtrees, deadline): collect
candidates serially → parse the changed/new ones → write a batch.
- **Incremental** — a cheap `mtime` match short-circuits before any read; the
content `hash` then guards the write. The walker respects `.gitignore`, and
no pass indexes a hidden path (a `.`-prefixed file or directory): a warm
enumerates with `git ls-files`, which lists tracked ones, so one filter
(`is_source`) keeps the indexed set the same whichever pass finishes.
- **Parallel parse, batched write** — parsing (the expensive Tree-sitter step)
fans out across CPUs; the parsed files are written in **one** transaction (one
`fsync` per batch, not per file). Writes stay serialized; parsing doesn't.
A pass over a cold repo (explicit or a first search's warm) suspends the
name index and rebuilds it at the end, before coverage is recorded, and
meanwhile fuzzy recall verifies the repo's committed names directly. Every other write appends the names and files new to the repo in
the transaction that writes them, and every pass ends by rebuilding an index
that is missing, from another format, or holding a quarter more keys than its
last rebuild wrote.
- **Opportunistic + time-bounded** (`index_budgeted`) — the first query warms the
index without blocking on a full walk: a small inline budget indexes the active
(branch) files first and answers, then the deferred pass warms more per query
until a full sweep marks coverage `complete` (reconciling deletions + capturing
commit times). Explicit `rq --index` is the same path, unbounded.
- **Block-until-answered (cold start)** — the time-boxed warm exists so a query
never hangs, but on a *huge, cold* repo it can expire before the symbol is
indexed, turning a real hit into a false "no matches". Correctness beats the
first query's latency (and once warm the repo answers fast), so a query against
a genuinely warming repo keeps indexing until the answer appears or the sweep
completes — for humans **and** programs alike. Small/medium repos finish inside
the normal budget and are unaffected; only a large cold repo waits, once.
- **Humans** (a TTY, plain text) also get a one-line "indexing…" progress
heads-up on stderr after ~500 ms and a graceful **Ctrl-C** (a `SIGINT` handler
over `libc`, installed only on this path) that aborts and prints the best
partial results. Interactive waits are unbounded — Ctrl-C is the escape.
- **Programs** (`--json`/`--ndjson` or any pipe) block silently, bounded by a
wait budget (`RQ_WAIT_BUDGET_MS`, default 1 min; `0` = non-blocking) since
there's no one to interrupt. **`--wait <dur>`** (`50ms`/`2s`/`1m`/bare ms)
overrides that budget per-call. A caller that prefers *fail-fast over
block-until-answered* passes **`--no-wait`** (shorthand for `--wait 0`): it
answers from the committed index immediately — never blocking, and skipping
the in-process warm so a query issued mid-rebuild neither waits on nor contends
with the writer — while leftover warming still detaches to a background child.
A `--no-wait` miss on an incomplete index still reports `warming` (exit 2).
- The poll that watches the warming index re-queries every `POLL_INTERVAL`
(100 ms) — coarse enough that these read transactions don't steal CPU or
read-lock churn from the active writer, fine enough that an early answer or a
completed sweep surfaces within a frame.
- A miss distinguishes **definitive** (index `complete` → exit 1) from
**indeterminate** (still `warming`, e.g. the wait budget was hit on a huge
repo → exit 2 + a one-line stderr note), so a caller isn't misled into
treating "not yet" as "absent". Both are non-zero, so `rq … && …` is
unchanged. Committed batches persist, so a re-run resumes.
- `index_budgeted_cancellable` carries the abort flag (Ctrl-C, a wait timeout,
or an early answer) down into the walk so the pass stops promptly without
losing committed work.
- **Discovery vs tracking** — a *git work tree* is auto-discovered (a stray query
may warm it); a *non-git* dir is only indexed when asked (`rq --index`), after
which it's **tracked** (has coverage) and treated like any repo. Git-ness gates
auto-discovery and branch-awareness; tracking gates the current-repo boost and
self-healing warm.
- **Prioritized** — active (branch) files first, so the working set is indexed
and kept fresh ahead of the rest of the repo. A search's warm then parses the
files that contain the query's leaf name (a read-and-substring pass,
uncapped, several times cheaper than parsing), because an exact or prefix
match — the only answer a warming search accepts — must live in one. Then
everything else in walk order, files whose name resembles the query first.
Parsed files commit at least every 50 ms, since a warming search only sees
committed rows. See D11 for why this is two tiers rather than a priority heap.
- **Coverage-aware** — every walk updates `coverage` (`warming` until a full
sweep completes, then `complete`). A subtree index (`--index --path`) is a
*seed*, not a fence: it gets the named files in first and leaves coverage
`warming`, so normal warming continues over the rest of the repo through use.
- **Git off the hot path** — `is_git_repo` is native (walk up for `.git`),
identity is cached by checkout root, and the `git log` for commit-time recency
runs only when a sweep actually (re)indexed something. The one remaining
per-search question — has the worktree moved since it was indexed? — forks
`git status`, which grows with the worktree; a hit hands it to the detached
warm child rather than wait on it, so a hit on an indexed repo forks no `git`.
A miss still asks inline, since its exit code (absent vs. still warming)
depends on the answer. "Moved" means a new HEAD or a dirty source file whose
mtime differs from the indexed one — dirty-but-indexed is unchanged. The
index also remembers which files it took in as edits (the dirty set at each
check and at the end of a sweep, plus any file revalidated singly), and checks
those too: a discarded edit (`git checkout -- f`) is clean, so status no
longer names it, yet the index still holds the edit until it's reindexed. A child
that finds nothing moved records the verdict with a git-state stamp (HEAD
commit + `.git/index` mtime), and for 10 s (`RQ_WARM_RECHECK_MS`) a hit whose
stamp still matches skips the spawn too. Staging, commits, checkouts and pulls
change the stamp; an unstaged edit doesn't, so it waits out the window (D16).
- **Language-isolated** — the indexer is blind to language; plugins emit the
common symbol model.
Tree-sitter parsing is the expensive step and is kept **off the search critical
path**: the inline warm is time-boxed, and the bulk of extraction persists for
the *next* query rather than blocking the current one.
## Search / ranking pipeline
Staged, streaming, early-exit on confidence:
| 0 | parse query | case, separators, looks-like-a-path? |
| 1 | exact / prefix symbol | indexed `name_lower`; fastest, highest confidence |
| 2 | fuzzy symbol | the name index's exact candidate set (D23) → abbreviation-aware scorer; an fst over names was slower (D21), and the trigram FTS nets it replaced are gone (D26) |
| 3 | path / filename | |
| 4 | live scan | async, streamed when coverage is low |
| 5 | opportunistic extraction | parse newly-seen files, persist for next time |
**Confidence gate:** a strong exact match in the current repo returns
immediately and stops the pipeline. Otherwise return the top-N from layers 1–3
now and stream refinements from 4–5.
### Scoring — simple, additive, explainable
Ranking is an additive sum of named features so `--explain` can print exactly
why a result ranked where it did:
- **match quality** — exact > prefix > camel-hump abbreviation > subsequence
A query may leave out the word joiners `_`, `-` and `.` and still match
exactly, 50 behind the spelled-out name (`separators`). Any other character is
part of the name, so `save` is a prefix of `save!`, not an exact match (D19).
A fuzzy query that begins with a sigil (`_dshrz`) favours the names that
begin with it: it scores, but never decides what matches (D24)
- **case** — a query carrying any uppercase rewards the candidate spelled the
same way, so `Symbol` finds the type rather than a `symbol` method that
matches case-insensitively. An all-lowercase query is how people type
casually, so it stays case-agnostic and neither spelling is favoured. Large
enough to outweigh `recency`, or which of two same-named symbols won would
come down to file mtimes
- **kind weight** — tunable (e.g. class/module slightly above method)
- **definition shape** — `kind`, `extent` (log of body lines), `path`, `depth`
(nesting past two levels) and `private` pick among names that answer the
query about equally well. Each is scaled by the match quality, so they keep
full weight between two exact matches and shrink to about a third on a fuzzy
one, where they used to outweigh the name itself (D24)
- **test path** — a definition under `test/`, `spec/` and the like, or in a
`_test`/`_spec` file, takes −400 on a literal match, enough to cross from exact
to prefix. A fuzzy or typo match gives up 0.4 × its name evidence instead, so a
test definition that reads as the query clearly better still outranks a weak
match elsewhere (D24)
- **generated** — a file whose header declares it generated (`Code generated …
DO NOT EDIT`, `@generated`, read at index time into `files.generated`) takes
the test penalty, under its own name: secondary code the same way (D28)
- **example path** — so does a definition under an example, demo or docs app
(`examples/`, `example/`, `_examples/`, `demo/`, `demos/`, `docs/`,
`dev-docs/`; not `doc/`, often a library's own package) — `example_path` (D29).
None of the three applies in the `--anchor`'s own file: asked from inside a
test, that file's definitions are the context (D31)
- **visibility** — a definition its language marks private/protected takes a
small penalty (public API over internal helpers; a tiebreaker, never a
filter — and unknown visibility carries no signal). Sourced per language:
Rust `pub`, Ruby access sections, Python underscore convention, Go
capitalization, TypeScript member modifiers and ESM `export`. A `local`
definition (Python's nested `def`) takes a larger one, `local`, which ranks
it below every same-named definition outside a function body (D36)
- **stub** — a declaration whose body is elsewhere (TypeScript's ambient
`declare` and `.d.ts` values, overload signatures; `symbols.stub`; a
declared interface or type is the definition, not a stub) takes the
same size as `local`: the implementation ranks first when it's indexed, and
the declaration is the answer when it isn't (D38)
- **qualifier** — a scoped query (`Foo::Bar`, `Foo::Bar#baz`, `Foo.baz`; `::`,
`#` and `.` are all scope separators) matches its leaf against the name and
requires a `parent` ending with the named scope chain (`Bar` inside `Foo`) —
a candidate outside it drops out. Scopes no parent records (a Go package, a
Python module, a Rust `mod` file) are read off the file's path instead: the
segments the parent doesn't hold must appear in order among the repo's name,
the file's directories and its stem (`path_scope`, halved per directory
between the scope and the file). Only the best-scoped results stay: a parent
over a path, a scope's own directory over its subdirectories (D27). Godoc's
`(*T).M` reads as `T.M`. `Foo.new` also matches the constructor a
plugin names via `LanguagePlugin::constructor` (Ruby `initialize`, Python
`__init__`, JS/TS `constructor`); when `Foo` declares none (inherited or
implicit — rq doesn't track inheritance), the class itself answers, flagged
`constructor_owner` at 0.75 confidence. The typo retry forgives up to two edits in
the scope too (a `scope_typo` feature, typo-level confidence). A `.` query
that no scope answers tries `.` as a one-char wildcard before any typo retry
- **path** — query also matches the file's name (Layer 3)
- **current-repo scope + boost** — results are restricted to the repo you're in
by default (a search there answers about *that* repo, never leaking another
indexed one; `--all-repos` opts into cross-repo), and within it the current
repo's rows still carry the boost
- **recency** — symbols in recently-active files (~14-day half-life), sourced
from the more recent of file mtime and last git commit time (captured once per
index, not on the search path)
- **branch** — on a feature branch, symbols in files that differ from the trunk
(committed since divergence + uncommitted) get a strong boost; symbols in
those files' directories a smaller one. This is the one git signal computed *at
search time* (a few `git diff --name-only` calls) because it tracks live
working state; it's gated to feature branches, so the trunk pays nothing.
The active-file set also drives proactive pre-indexing — `index_budgeted`
warms those files first.
- **anchor** — `--anchor FILE:LINE[:COL]` names where the query is asked from.
Two features, both boosts, never filters. `enclosing`: the candidate's
`parent` is a leading run of the scope chain of the innermost definition
whose `line..end_line` span holds the anchor line, 60 per shared level, capped
at 180. `proximity`: 90 in the anchor's own file, else 60 in its directory,
halving per directory step and dropped below 5; anchor's repo only. Built
only from stored spans and parents (or a live parse of the anchor file when
the index doesn't hold its current version), so it is language-blind. No
inheritance, so an inherited method earns no `enclosing` (D18).
Match quality and the static features live in the pure `score()` function. The
dynamic, context-dependent signals (`recency`, `branch`, `enclosing`,
`proximity`) are computed by the search layer — which owns the clock, the
branch state and the anchor — and passed in via a `Boosts` struct, so a new git signal (recent commit, branch, ownership) is a new
field, not a new parameter. Prefer understandable scoring over sophisticated
algorithms; tuning a weight must never require re-indexing.
### Abbreviation matching
`refundproc → RefundProcessor`, `usr → User`, `perf → perform`:
1. Tokenize the candidate on camel-case / underscore boundaries
(`RefundProcessor` and `refund_processor` both → `[refund, processor]`).
2. Greedily match the query against token prefixes and initials.
3. Score by contiguity and token-boundary alignment. Letters matched before
the alignment reaches its first word start earn no credit (D15).
Intra-token fuzz (`paymnt → Payments`) falls back to subsequence matching with
a penalty. A near miss (`sleect → Select`, up to two edits) competes with those
fuzzy matches whenever nothing matched literally. It is scored by the letters it
keeps (D14) and ranks beside them on its score (D24). Quality of ranking matters more than the cleverness of the
algorithm.
## Partial indexing
The index is **never assumed complete**.
- `coverage.status` tells search its own confidence (`warming | complete`,
or no row until a pass finishes). `warming` is indexing in progress — whether
opportunistic or seeded by a subtree `--index --path`. `--status` reports a
repo with no row as `warming` too: only a pass registers a repo, so it's one
whose first pass is running (or was cut short).
- A `warming` repo **blocks until answered** (see the indexing model), so
incomplete coverage yields a delayed-but-correct answer rather than a
confident-looking wrong one. A dir with no finished pass that this query
isn't warming — untracked and non-git, or a git repo asked with `--no-wait` —
gets a bounded in-memory live scan, merged with whatever the index offered. Each
result carries its `source` (`index` or `live`), so a blended answer says
which parts were never persisted.
- **Opportunistic extraction** grows coverage through normal use.
- **Staleness:** a `content_hash` mismatch marks a file's symbols stale; search
lazily validates only the **top-N** results (stat, re-parse if changed) before
presenting — cheap because it touches a handful of files, not the index.
Degradation ladder:
```text
zero index → pure live scan (works, slower)
warming index → index results, blocking until the answer is trustworthy
complete + fresh → index only, sub-50 ms
```
The user never needs to know which layer a result came from.
## Behavioral learning — removed
Ranking once learned from which definition got used: `--open`, `--show`, and a
`--record` hook logged picks, a rollup aggregated them into `selection_stats`,
and a decaying `learned` boost fed the scorer. It was deleted after six weeks
of real use left the table empty — see [DECISIONS](DECISIONS.md) D10 for the
numbers, and why `--show` could only ever have confirmed static ranking. Search
is now a function of the index, recency, and the branch.
### No daemon — detached post-interaction work
Proactive work like warming the index is **not** a resident daemon. Each `rq`
invocation prints results first; leftover index warming is handed to a
**detached child**: after results print, the search re-execs `rq --warm <root>`
with null stdio in its own process group and exits — the shell only ever waits
on the answer. The
child runs niced (and with throttled disk I/O on macOS) on a seconds-scale
budget (`RQ_WARM_BUDGET_MS`), sweeping until coverage completes, and is
single-flighted per repo via a pid-stamped lock in `meta`, so a burst of
queries runs at most one warmer. On a complete repo the child is spawned after
a hit and first asks whether anything moved; usually nothing has, and it
exits after one `git status`, recording that verdict so hits over the next few
seconds don't spawn at all (D16). Still no daemon: the child does one job and
exits. `RQ_WARM_DETACH=0` reverts to finishing the (small) warm in-process —
the hermetic mode tests and debugging use.
Git-awareness (current branch, recent commits, ownership, recently-modified
areas) enters later as additional **ranking hints — never hard filters**.
## Editor integration
No editor-specific coupling in the core. Result locations are `path:line`, so
any editor can jump to them, and `rq -o/--open` hands the best match to a
launcher. VS Code, Neovim, and JetBrains are all just result openers — see
[EDITORS](EDITORS.md).
## Open risks (tracked, not yet resolved)
1. **Fuzzy-over-millions latency** — mitigated by the name index's signature
screen (D23); needs measurement against the 50 ms budget at scale.
2. **Cross-repo ranking** — resolved for the common case by scoping to the
current repo by default (`--all-repos` opts out); cross-repo ranking priors
(recency) still matter under `--all-repos`.
3. **Ranking explainability** — `--explain` from day one is the mitigation.
4. **Scope creep** — Layers 4–5 are a streamed tail, not a second search engine;
keep them lean for the MVP.