scryer-engine 0.3.0

Tree-sitter AST indexing and reference resolution engine for Scryer code intelligence
docs.rs failed to build scryer-engine-0.3.0
Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.

scryer-engine

Core AST parsing, syntax analysis, and symbol resolution engine for Scryer.

Responsibilities

  • Workspace Traversal & Gitignore Filtering (WorkspaceScanner): High-throughput file tree walking via ignore::WalkBuilder respecting .gitignore, .ignore, global git filters, and hidden files. Source files over MAX_SOURCE_FILE_BYTES (2 MiB; generated and minified bundles) are skipped, and a watched file that grows past it is dropped from the index.
  • Incremental Change Detection (ChangeDetector): 256-bit BLAKE3 content hashing compared against cached hashes in Turso (SourceFile.content_hash) to avoid redundant parsing.
  • Multi-Language Tree-sitter AST Extraction:
    • Rust (RustAstParser): Extracts declarations (functions, structs, enums, unions, traits, impls, modules, type aliases, constants, statics), condensed signatures, visibility modifiers (public, crate, private), doc comments, and nested lexical scopes.
    • Python (PythonAstParser): Extracts functions, async functions, classes, methods, top-level constants/variables, signatures with type annotations, docstrings (triple/single-quoted), visibility by convention (public, private), and indentation-based scopes.
    • TypeScript / TSX (TypeScriptAstParser): Extracts interfaces, type aliases, classes, enums, functions, arrow function assignments, methods, JSDoc comments, access modifiers (public, private, protected), and lexical scopes.
  • Decoupled Batch Ingestion (BatchIngestionActor): Streams parsed entities from Rayon CPU workers into Tokio writer tasks over bounded MPSC channels, executing transactional Turso writes with batch grouping and status-aware record cleanup to prevent dangling references or orphaned symbols. Linking is independent of batching: see Deterministic linking.
  • Debounced Workspace Watcher (WorkspaceWatcher): One notify-debouncer-full watcher (200ms debounce, one inotify instance) shared by any number of projects (add_project / remove_project). Only non-ignored directories are watched, each non-recursively (WorkspaceScanner::scan_dirs), and watches are reference-counted across nested projects. A file event re-indexes or removes that file in every project containing it. A directory create, delete or rename, or an ignore-file change, re-walks the project's directories and runs an incremental index_project. Events pass through IgnoreFilter, which applies the scanner's rules (hidden files, nested .ignore/.gitignore, .git/info/exclude, global excludes) per path and reloads a directory's rules when its ignore file changes.
  • Dependency Discovery (DependencyProvider, CargoDependencyProvider): Pluggable ecosystem abstraction for package discovery. CargoDependencyProvider executes cargo_metadata to discover workspace direct and transitive dependencies, classifies sources (crates.io, git, path), extracts package checksums from Cargo.lock, maps active features, and locates unpacked crate root directories in $CARGO_HOME/registry/src/ or $CARGO_HOME/git/checkouts/.
  • Index Health (IndexStats, index_health): Each indexing run records what it could not index (files that failed to parse, files with syntax errors, files over the size cap or unreadable, unresolved references, and the Cargo dependency state), and index_health combines that with COUNT(*) queries into one report shared by get_indexing_status and scryer doctor.
  • High-Level Coordination (EngineService): Unified facade coordinating scanning, hashing, parallel parsing, and database updates.
  • Lexical Search (search): Deterministic BM25F symbol search behind EngineService::search_symbols (see below). The Bm25Index and tokenizer are generic and reused by scryer-mcp to rank ADRs.

Lexical Search

ADR 0010 chose in-memory lexical ranking over neural embeddings: turso_core 0.7.2 has no FTS or vector index, and embeddings would need a model or network API.

  • Tokenizer (search::tokenize): splits on non-alphanumerics, camelCase / acronym (HTTPServerError → http server error) and letter↔digit boundaries; lowercases; strips a trailing s from tokens longer than 3 chars; drops language keywords and common English stopwords. Identifiers also emit their whole lowercased form, never stopword-filtered, so exact names rank highest. A query made only of stopwords ("impl", "to the") falls back to its raw lowercased split.
  • Index (search::bm25): Bm25Index<D> with weighted fields, BM25F scoring (k1 = 1.2, b = 0.75), deterministic ties (insertion order), and matched_terms per hit.
  • Symbol documents: name ×3, module path ×2, signature ×1, docstring ×1, file path ×1. reexport symbols are skipped.
  • Cache (search::cache::SearchIndexCache): LRU of 16 indexes keyed by (project_id, dependency_package_id). Every index write bumps the project's generation (touch_project, purge_project, and dependency indexing for project 0 and for new project_dependency links); a stale index is rebuilt wholesale on its next query, since symbol IDs change on every re-ingest. Rows load under the DB lock; tokenization runs on the Rayon pool under spawn_blocking; a per-key lock prevents duplicate builds.
  • Scopes (search::SearchScope): Project, or Dependencies { crate_name }, which is limited to the project's linked packages (a named crate falls back to any cached version). One package uses its own index; several share the project-0 index with a package filter.
use scryer_engine::search::{SearchQuery, SearchScope};

let results = engine
    .search_symbols(project_id, SearchQuery {
        text: Some("retry backoff".into()),
        kinds: vec!["fn".into()],
        scope: SearchScope::Project,
        limit: 20,
        ..SearchQuery::default()
    })
    .await?;
for hit in results.hits {
    println!("{:.2} {} {}:{}", hit.score, hit.qualified_name, hit.file_path, hit.start_line);
}

Architecture

Filesystem (Scanner / Watcher)
      │
      ▼
BLAKE3 ChangeDetector (vs Turso SourceFile.content_hash)
      │ (Added / Modified)
      ▼
Rayon Worker Pool ──► RustAstParser (Tree-sitter)
      │
      ▼ (ParsedFilePayload)
Bounded MPSC Channel (128 capacity)
      │
      ▼
Tokio Ingestion Actor ──► Turso Database Transaction
                           - Purge the file's own rows by file id
                           - Insert SourceFile and Scope; keep the ids of
                             symbols the file still defines, insert the rest
                           - Link same-file references; defer the rest
                             (and links into symbols the file dropped)
      │ (all files and dependencies stored)
      ▼
Final link pass ──► resolve deferred links against the full symbol table
                    - link what earlier runs left unresolved to new symbols
                    - keep what is still unresolved (unresolved_reference)
                    - record unresolved count and index health (index_stats)

Deterministic linking

Indexing the same tree always stores the same references and edges, however the writer batches its work and whatever order files are processed in (tests/determinism_test.rs; IndexOptions { writer_batch_cap, file_order_seed, dependencies } varies the scheduling for tests and benchmarks).

  1. First pass (per writer batch): symbols, scopes and files are stored; only references whose target is in the same file are linked, because the batch cannot see files stored later. Everything else is kept as a deferred link instead of being dropped.
  2. Final pass (BatchIngestionActor::link_deferred): once every file and the Cargo dependencies are stored, the deferred links are resolved against the complete symbol table in one transaction. What still matches nothing is kept in the unresolved_reference table and counted (IndexReport::unresolved_references is the project's whole count, stored in index_stats). The same pass first takes the persisted rows whose bare name matches a symbol this run inserted out of the table and links them, so a function defined by a later edit, watcher event or run links to a caller that is not touched again. index_file, refresh_files and remove_file finish with the same pass.
  3. Tie-break among same-named symbols (ADR 0013): exact qualified name, else a path suffix, else the ordered-segments match for a path through a re-export (never exact; three or more segments, or two with the definition's own crate root); then a symbol in the referencing file, then the same directory, then the first by (file path, byte offset), never by database id. A reference never resolves into a file of another language.
  4. Everything is processed in path order: symbols are tabled and files resolved in sorted order, so a reference with several candidate definitions resolves the same way every run.

What a reference records

  • role: call, type_annotation, import, reexport, and value (a free function named without being called: let f = target;, register(target); skipped for locals, struct-pattern bindings, methods and text the parser could not parse).
  • via (schema version 3): exact (a normalized path, an import, self.m()/this.m()/Self::m(), a trait-qualified call, the project's one definition of the name, or the file's own symbols, to the only matching symbol), name (by name alone: any method call on a receiver of unknown type, or a name several symbols share) or macro (found inside a macro invocation by scanning its tokens).
  • Beyond plain calls: calls inside macro invocations (println!, assert_eq!, vec!), Trait::m(x), <T as Trait>::m(x) and T::m(x) (through T's bounds) resolving to the trait's method (required trait methods are symbols), parameter types, and default-argument expressions. Calls in const/var/class initialisers get a call edge from that symbol; a reference held by no symbol reports its module (module_name_for).

Benchmark and recall tests

  • cargo run --release -p scryer-engine --example index_bench -- <root>... --runs 3 --batch-cap 1 --batch-cap 500 --probe NAME --recall crates/scryer-engine/tests/recall --baseline prev.json --json out.json indexes each root into fresh databases and prints time, counts, references by via, whether every run stored the same references and edges (exit status 1 if not), and recall on the annotated fixtures. Method, results and findings per phase: docs/learnings/index-benchmark.md.
  • tests/recall/{rust,python,typescript} are fixtures whose lines carry @ref target role=… phase=N markers (the references an index must contain); tests/recall_test.rs fails on a missing one.
  • python3 testing-harness/harness.py verify cargo --recall compares Scryer's reference files with grep hits in code for real repositories (see the harness README).

Usage Example

use std::path::Path;
use scryer_db::ScryerDb;
use scryer_engine::EngineService;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let db = ScryerDb::connect("turso::memory:").await?;
    let engine = EngineService::new(db);

    let project_id = 1;
    let root = Path::new("/path/to/project");

    // Full or incremental indexing pass
    let report = engine.index_project(project_id, root).await?;
    println!(
        "Indexed {} files (added: {}, modified: {}, unchanged: {}, references: {}, edges: {})",
        report.scanned_files,
        report.added_files,
        report.modified_files,
        report.unchanged_files,
        report.total_references,
        report.total_edges,
    );

    // Re-index one file (what the watcher does): its references are resolved and linked
    engine.index_file(project_id, root, Path::new("src/main.rs")).await?;

    // Start background file watcher (200ms debounce); add more with watcher.add_project
    let watcher_handle = engine.watch_project(project_id, root)?;

    // ... application runs ...

    watcher_handle.stop().await;
    Ok(())
}

Reference Resolution

Every reference candidate (a call, a type, a name used as a value, an import) is resolved against the symbols of the files being indexed. There is no stack-graph engine: the rules it was loaded with only ever bound names globally, and its crate is archived and pinned the tree-sitter core to 0.24 (ADR 0017).

  1. Path, receiver and bounds first. Path-qualified Rust references (crate::a::f, other_crate::m::f, Type::new) are rewritten by normalize_rust_path (crate/self/super) and ingest matches the path against qualified names (find_symbol_for_target). self.m()/this.m()/Self::m() resolve to the surrounding type's method, T::m() through T's trait bounds.
  2. The one definition of a bare name. Methods and use stubs are not definitions a bare name can mean; a project with a single function, type, module or constant of that name in the same language gives an exact link. A method call (x.m()) is never matched this way.
  3. A symbol of the referencing file (ScmFallbackResolver): its own definition, or the use that imports it. With several definitions of the name, the import is the precise link; resolve_definition follows it to the definition.
  4. By name (via = name): the ADR 0013 pick (same file, same directory, first by path), or, with no candidate in this pass, the bare name, linked against the persisted index at ingest time so references into files outside the pass still link.
  5. A call through a local binding (a parameter, closure or let, count_tokens(&json)) links only to a function declared inside the function that holds the call; otherwise it is dropped. A symbol is a declaration (ADR 0018): module- and class-level bindings, and functions and classes at any depth; what a function body binds is not indexed.

resolve_definition reads the exact reference stored at the cursor and falls back to the name lookup.

Parsing

Grammars: tree-sitter 0.27, tree-sitter-rust 0.24, tree-sitter-python 0.25 (ABI 15) and tree-sitter-typescript 0.23 (ABI 14). Rust source goes through parsers::rust::prepare_source first: the front matter of a cargo -Zscript file is blanked and snapbox's str![..] is parsed as Str![..], both without changing a byte offset or a line, because the grammar rejects them and a file with a parse error loses code around it. get_type_contract reads Rust fields, variants and trait impls, and TypeScript fields, interface members and enum variants, from the source at query time (extract_type_members); Python and TypeScript methods and Python class attributes are symbols of the class. search_symbols scores a hit in test code (is_test_symbol: a tests or test directory, test_*.py, *.test.ts, an inline mod tests, ...) at half.

Rust qualified names follow the module path (crates/my-pkg/src/a.rs → my_pkg::a; a root-level src/ uses crate). ingest_batch inserts every file's symbols before linking references. Re-indexing a file keeps the ids of the symbols it still defines (matched by qualified name, kind and position among duplicates), so references and edges from other files into them stay valid and cost nothing; links into a symbol the file no longer defines go back to the deferred links and re-link to a symbol of the same qualified path, or wait as unresolved rows until one exists (ADR 0016). A reference that is already linked never moves to a newly added same-named symbol.