Weavatrix Scan
weavatrix-scan is a deterministic, read-only repository scanner for static
analysis, code intelligence, indexing, and AI tooling.
It does more than walk a directory. A scan produces a stable manifest with
normalized paths, file sizes, optional content hashes, an aggregate revision,
and explicit evidence explaining why files were skipped. Linux and macOS builds
have zero mandatory runtime dependencies; Windows uses only winapi-util for
native volume and file identities.
Why another repository walker?
walkdir and
jwalk are
excellent traversal libraries.
ignore adds
mature Git-style filtering. Weavatrix Scan exposes deliberately separate
layers:
Walker: iterative, streaming, lossless low-level traversal;WalkBuilder: multi-root traversal, native sorting, directory filters, and contents-first ordering;Scanner: ignore-aware deterministic manifest, hashes, revision, typed evidence, and bounded one-pass selected-content callbacks;CompactScanReport: the same deterministic selection and revision with one retained root path and optional boxed rich evidence for million-file manifests;MultiScanner: ordered concurrent manifests and a globally bounded multi-root content pipeline;ParallelMultiWalker: collected or direct-streaming traversal across concurrent raw roots with ordered per-root reports;RepositoryMatcher: cached path selection for incremental consumers;SelectionMatcher: the complete scanner selection policy as a reusable, typed standalone matcher;ParallelWalker: bounded adaptive traversal for broad or skewed trees.ParallelRuntime: process-global, dedicated, or application-owned execution shared by walker and scanner APIs.
Competitive position
Against the versions tested in this repository (ignore 0.4.31, walkdir
2.5.0, and jwalk 0.8.1), Weavatrix Scan is the strongest overall fit when the
output must be a deterministic, explainable code-scanner manifest rather than
only a stream of directory entries. It is not the universal winner for every
walker workload:
| Workload | Strongest fit in this comparison | Why |
|---|---|---|
| Deterministic code-scanner manifest | Weavatrix Scan | Only entry with normalized paths, hashes, aggregate revision, typed skips, portable evidence, incremental cache, and changed-path updates |
| Raw parallel streaming | Weavatrix ParallelWalker |
264.7 ms and 7.6 MiB peak on the measured 1,000,000-file Windows fixture |
| Minimal serial traversal | walkdir / Weavatrix Walker |
walkdir remains the small established primitive; Walker measured 546.3 ms versus 584.5 ms with comparable 5 MiB-class memory |
| Memory-efficient deterministic manifest | Weavatrix CompactScanReport |
Exact output parity at 1,019.4 ms / 63.2 MiB; ignore used 56.4 MiB but took 2,106.7 ms and does not produce revision/evidence |
| Reusable selection matcher | Weavatrix / ignore | Weavatrix combines ignore, overrides, types, depth, size, symlink and filesystem policy with typed outcomes; ignore exposes modular matcher builders |
| Host-owned parallel scheduling | Weavatrix / jwalk | Weavatrix accepts any fallible executor, optionally wraps existing/new Rayon pools directly, and preserves busy-timeout policy |
| Capability | weavatrix-scan | ignore | walkdir | jwalk |
|---|---|---|---|---|
| Iterative traversal | Yes | Yes | Yes | Yes |
| Single-file root | Yes | Yes | Yes | Yes |
| Lossless native paths | Yes | Yes | Yes | Yes |
| Continue after local errors | Configurable | Yes | Yes | Yes |
| Depth / open-handle control | Depth + configurable max_open |
Depth + internally bounded | Depth + configurable max_open |
Depth + directory scheduler |
| Same-filesystem boundary | Yes | Yes | Yes | No |
.gitignore hierarchy |
Yes | Yes | No | No |
| Custom ignore files | Yes | Yes | No | No |
| Repository / Git-compatible ignore modes | Yes | Yes | No | No |
| Override globs / source switches | Yes | Yes | No | No |
| Reusable cached full selection matcher | Yes | Yes | No | No |
| Multi-root / custom sort | Serial + parallel / full DirEntry |
Yes / name or path, serial only | No / full DirEntry |
No / mutable directory batch |
| Directory callback / contents-first | Parallel typed batch / Yes | Filter only / No | Filter / Yes | Parallel typed batch / No |
| Built-in types / composition / negation | 265 / Yes / Yes | 224 / Yes / Yes | No | No |
| Stable normalized paths | Yes | No | No | No |
| Path-safe portable report | Yes | No | No | No |
| Snapshot-verified content provider | Yes | No | No | No |
| File sizes and SHA-256 hashes | Yes | No | No | No |
| Versioned compact incremental cache | Yes | No | No | No |
| Watcher events to changed-path manifest update | Yes | No | No | No |
Optional direct notify adapter |
Yes | No | No | No |
| Concurrent-mutation evidence | Yes | No | No | No |
| Aggregate deterministic revision | Yes | No | No | No |
| Typed manifest delta / rename evidence | Yes | No | No | No |
| Binary and oversized-file policy | Yes | No | No | No |
| Typed skip reasons and warnings | Yes | No | No | No |
| Symlinks skipped by default / loop detection | Yes | Yes | Yes | Configurable |
| Parallel collected / streaming traversal | Yes / visitor + bounded pull | No / callback | No / serial iterator | No / ordered iterator |
| Deterministic backpressured scan sink | Yes | No | No | No |
| Parallel one-pass verified content callback | Yes | No | No | No |
| Multi-root verified content callback | Yes | No | No | No |
| Changed-path-only content callback | Yes | No | No | No |
| Streaming content without retained manifest | Yes | No | No | No |
| Parallel pull iterator | Bounded unordered / ordered DFS | No (callback API) | No | Ordered DFS |
| Parallel multi-root raw traversal | Collected + streaming | Streaming callback | No | No |
| Parallel multi-root manifest scanner | Yes | No | No | No |
| Stateful per-directory batch | Parallel ordered, typed | No | No | Parallel ordered, typed |
| Redirected-stdout protection | Yes | Yes | No | No |
| Separate root-symlink policy | Yes | No | Yes | No |
| Cancellation and whole-scan budgets | Yes | Quit only | No | No |
| Minimum depth / hidden policy | Yes / Yes | Yes / Yes | Yes / No | Yes / Yes |
| Existing/dedicated worker pool | Generic external / owned | Internal threads | Not applicable | Rayon existing / new |
| Busy timeout / fallible submission | External contract / Yes | No / No | Not applicable | Yes / Yes |
| Measured 1M raw time / peak | 264.7 ms / 7.6 MiB | Not raw-equivalent | 584.5 ms / 4.6 MiB | 313.1 ms / 159.7 MiB |
| Measured 833k manifest time / peak | Compact: 1,019.4 ms / 63.2 MiB | 2,106.7 ms / 56.4 MiB | Not a scanner | Not a scanner |
| Default runtime dependencies | 0 Unix / 1 Windows | Multiple | 2 platform helpers | Rayon stack |
Use Walker when you only need paths. Use Scanner when downstream results
must be reproducible and explainable.
Remaining competitive boundaries
The functional gaps in retained-manifest memory, embeddable scheduling, parallel stateful batches, multi-root streaming, and traversal-free changed content are now closed. The remaining differences are evidence and ecosystem boundaries:
- Matcher production history.
ignoreremains the established Git-ignore implementation. Weavatrix checks representative, randomized, arbitrary-byte, and million-file exact-manifest parity, but does not claim equal ecosystem age. - Cross-platform million-file evidence. CI and normal benchmarks cover Linux, Windows, and macOS, while the opt-in million-file RSS result below has so far been measured only on Windows.
The million-file result establishes top-tier performance on the measured Windows fixture, not a universal cross-platform ranking. The regular benchmark workflow still validates smaller output-equivalent corpora on Linux, Windows, and macOS; the million-file profile remains opt-in.
Install
[]
= "0.4"
Enable serialization only when needed:
[]
= { = "0.4", = ["serde"] }
Enable direct conversion from notify::Event without making a watcher runtime
mandatory for other users:
= { = "0.4", = ["notify"] }
Enable direct existing/new Rayon pool integration without changing the default scheduler:
= { = "0.4", = ["rayon"] }
The default build has no third-party runtime dependency on Unix. Windows uses
the small winapi-util safe wrapper for file, volume, and redirected-stdout
identity because the equivalent std by-handle identity APIs remain unstable
on the Rust 1.88 MSRV. Reimplementing that layer locally would require unsafe
WinAPI FFI and still depend on Windows bindings; this crate keeps
unsafe_code = "forbid".
Quick start
use ;
let options = default
.with_extensions
.with_parallelism;
let report = new.options.scan?;
println!;
for file in &report.files
for skipped in &report.skipped
# Ok::
For the fastest path-only discovery, disable content reads:
use ;
let report = new
.options
.scan_compact?;
assert!;
# Ok::
scan_compact discovers directly into root-shared records; it does not first
build and convert a full report. Use report.absolute_path(file) only for the
entries that need an owned absolute path. Rich compact scans retain hashes in
optional boxed content evidence, accessible with file.content_hash().
Low-level walkers
Walker is a streaming iterator. It keeps paths as native PathBuf/OsStr
values, uses iterative DFS, bounds open directory handles, and yields local
errors according to policy:
use ;
let options = default
.with_max_depth
.with_max_open
.with_same_file_system
.with_error_policy;
let mut walker = with_options?;
while let Some = walker.next
# Ok::
The root is depth zero. After receiving a directory, callers can invoke
skip_current_dir() before requesting the next item. Symbolic links are not
followed by default; enabling .with_follow_links(true) keeps traversal inside
the root and reports loops as typed skip reasons. The root itself follows a
separate RootSymlinkPolicy: Follow is the compatibility default and
Reject prevents an explicitly supplied symlink root from being opened.
WalkBuilder adds flexible low-level policies without changing the minimal
streaming Walker:
use WalkBuilder;
let entries = new
.add_root
.sort_by_file_name
.filter_directories
.contents_first
.build
.?;
# Ok::
sort_by receives complete std::fs::DirEntry values, including path,
file type, and metadata access. sort_by_name remains the allocation-free
native OsStr comparator, so sorting never requires lossy UTF-8 conversion.
Low-level walkers accept either a directory or one file as the root.
skip_stdout(true) prevents a redirected output file inside the tree from
feeding back into a command that is scanning it. ParallelWalker applies the
same option to collected, visitor, unordered-pull, and ordered-pull traversal;
ParallelMultiWalker applies it to every root. Directory filters run before
descent.
filter_directories_stateful accepts FnMut, serializes callback access, and
keeps one captured state across every root in the builder. For batch mutation
and typed state propagation, StatefulWalkBuilder<R, E>::process_read_dir
receives all immediate children before they are yielded. It can reorder or
retain the batch, mutate R inherited by child directories, attach E to
entries, and disable descent per entry.
use StatefulWalkBuilder;
let entries = new
.with_parallelism
.process_read_dir
.build_parallel_ordered?
.?;
# Ok::
The parallel form runs each complete directory batch on the configured
runtime, propagates callback-mutated state to child tasks, and yields strict
DFS order under bounded backpressure. build() remains the zero-coordinator
serial iterator.
ParallelWalker adapts between low-overhead frontier lanes and dynamic
scheduling below narrow top-level trees:
use ParallelWalker;
let report = new
.with_parallelism
.walk?;
println!;
# Ok::
For pipelines that should parse entries immediately instead of collecting
them, visit invokes a thread-safe callback directly on traversal workers:
use ;
let summary = new.visit?;
# Ok::
Applications can isolate scans or reuse their own scheduler:
use ;
let runtime = dedicated?;
let raw = new
.runtime
.walk?;
let manifest = new
.runtime
.scan_compact?;
assert!;
println!;
# Ok::
ParallelRuntime::external accepts an Arc<dyn ParallelExecutor>. The
executor receives each boxed job and the optional busy timeout, and can reject
submission with io::Error; traversal reports that as
WalkOperation::ScheduleWorker without waiting for an unsubmitted worker.
With the optional rayon feature,
ParallelRuntime::{rayon_existing, rayon_new} provides the same contract
without requiring an application wrapper. Busy timeout cancels a queued job if
the Rayon pool does not start it before the deadline.
Consumers that prefer pull semantics can use a bounded iterator. A full buffer applies backpressure to traversal workers, and dropping the iterator cancels and joins its coordinator:
use ParallelWalker;
for entry in new.into_iter_bounded
# Ok::
Use into_iter_ordered_bounded when consumers require strict deterministic DFS
ordering. It prefetches directory reads in parallel while preserving the
configured output capacity and max_open bound. Both pull modes cancel and
join their coordinator when dropped. try_into_iter_bounded and
try_into_iter_ordered_bounded report coordinator thread creation failures
instead of panicking before traversal starts.
Larger bounded buffers improve throughput without changing the memory bound. Very small capacities are useful when minimum buffered state matters more than raw traversal speed.
ParallelMultiWalker::visit applies the same direct callback contract across
multiple roots. Callback order is intentionally concurrent, every event is
tagged with its root insertion index, and the returned reports stay in root
insertion order:
use ;
let summary = new
.add_root
.with_root_parallelism
.visit?;
println!;
# Ok::
WalkControl::Quit cooperatively cancels every active root. The cancellable
form shares one CancellationToken; skip_stdout, traversal limits, error
policy, and the selected ParallelRuntime apply to every root.
Parallel callbacks may start another walk using the same runtime. Such reentrant walks fall back to the iterative serial engine instead of waiting on workers that they already occupy. A callback panic stops and wakes the dynamic scheduler, is resumed on the caller, and leaves global, dedicated, or external execution reusable.
Scan modes
The same scanner supports three useful cost levels:
| Mode | Configuration | Reads content | Skip evidence | Hashes content |
|---|---|---|---|---|
| Rich manifest | ScanOptions::default() |
Yes | Complete | Yes |
| Safe discovery | hash_file_contents = false |
First 8 KiB | Complete | No |
| Metadata only | .metadata_only() |
No | Complete | No |
| Selected manifest | .metadata_only().selected_files_only() |
No | Omitted | No |
Traversal and content inspection use bounded available parallelism by default.
Set
.with_parallelism(1) for a serial run or pass a fixed worker count when a
host application owns the wider scheduling policy. Use
.with_traversal_parallelism(...) and .with_content_parallelism(...) when
directory latency and content hashing need separate budgets.
For many independent roots, MultiScanner scans roots concurrently while
returning reports in insertion order:
use ;
let reports = new
.add_root
.options
.with_root_parallelism
.scan?;
assert_eq!;
# Ok::
Scanner::scan_into keeps deterministic discovery metadata, then inspects and
hands off one selected file at a time. The synchronous sink provides
backpressure without an unbounded channel, and selected records are not retained
by ScanStreamReport:
use ;
let summary = new.scan_into?;
assert_eq!;
# Ok::
For Search, Clone, and language adapters that need bytes, visit_content
connects ignore-aware traversal to bounded parallel content workers. A
worker-local callback receives borrowed chunks from the same read used for
binary detection and optional SHA-256 evidence:
use ;
let options = default
.with_extensions
.with_content_discovery
.with_content_validation;
let summary = new
.options
.visit_content?;
println!;
# Ok::
ContentVisitControl::SkipFile suppresses remaining chunks for one file while
allowing required hash/binary evidence to finish. Quit cancels every worker.
Events include root_index, root path, normalized relative path, and a
monotonic sequence. Sort durable results by (root_index, relative).
ContentValidationPolicy::Strict verifies native file evidence before and
after the read; Fast keeps the safe opened-handle check but omits the
post-read check for latency-sensitive local search. A deterministic total-byte
budget automatically uses the compact two-phase path so budget selection
remains path-order stable.
ContentDiscoveryMode::Streaming is the constant-memory default: one serial
producer overlaps discovery with bounded content readers.
BufferedParallel uses the parallel ignore-aware walker, retains only compact
candidate evidence, then dispatches the same verified readers. Choose it for
minimum latency on wide or warm repositories; Search and index builders can
sort their durable results after the callback. Both modes use the same
selection, validation, binary, error, and cancellation contracts.
visit_content_streaming keeps the same byte, validation, cancellation, hash,
and binary contracts but does not retain compact selected-file evidence or
compute a revision. With selected_files_only() it has memory bounded by the
queue, worker state, and one 64 KiB buffer per worker rather than selected-file
count. ContentVisitReport::mode makes the empty streaming revision explicit.
A deterministic max_total_bytes budget still requires the two-phase
selection path.
MultiScanner::visit_content and visit_content_streaming use the selected
runtime across all roots; the factory receives (root_index, worker_index) and
reports remain in root insertion order. Scanner::visit_changed_content and
visit_changed_content_streaming accept a safe file-only WatchPlan, read
only existing changed paths, return removed paths separately, and yield
FullRescanRequired before callbacks for structural or selection-changing
plans.
Output contract
ScanReport contains:
root: canonical absolute repository root;files: stable, lexicographically sortedScannedFilevalues;skipped: stable, sorted evidence for excluded entries;warnings: non-fatal ignore-file and local I/O diagnostics;ignore_sources: typed location and hash of every loaded selection input;revision: SHA-256 digest over ignore inputs, selected paths, optional content hashes, portability, and partial-termination state;complete: false when local errors made the evidence partial.termination: typed reason for a bounded or cancelled partial scan;portable: false when host-level Git configuration affected selection.cache: content reads, strict-validation fingerprint reads, and strong hashes reused by an incremental scan.
Each ScannedFile contains an absolute path, slash-normalized repository path,
byte size, optional sha256: content hash, whole-content cache fingerprint, and
file-version evidence used to validate persistent cache reuse. The scanner
compares size, timestamps, native file identity where available, and metadata
before/after content reads. Native
paths remain lossless in the walker and absolute PathBuf; invalid Unicode
units in normalized manifest names are escaped (%XX on Unix, %uXXXX on
Windows) instead of being replaced with the lossy Unicode replacement marker.
With the serde feature, invalid native path units use a tagged byte/wide-unit
representation and round-trip without loss; ordinary Unicode paths remain
plain JSON strings.
Portable evidence and verified content
Use ScanReport::to_portable before sending scan evidence to another process,
writing public logs, or attaching it to an AI request. PortableScanReport
omits the absolute root, absolute file paths, file identities, timestamps,
cache statistics, and free-form diagnostic text. Repository-relative paths and
typed skip kinds remain; diagnostic details are represented only by stable
SHA-256 values. External ignore-source locations are removed. Its
selection_portable field separately records whether host-level Git
configuration affected file selection.
Consumers reopening an existing snapshot can bind bytes back to the full local report:
use ;
let report = new.scan?;
let portable = report.to_portable;
let content = report
.content_provider?
.read_bounded?;
assert!;
assert_eq!;
# Ok::
SnapshotContentProvider accepts only sorted entries belonging to its report,
rejects path escapes and symlinks, enforces an optional byte limit, and compares
size plus native file-version evidence before and after reading. When a content
hash exists, returned bytes must also match the recorded SHA-256. A missing or
changed file produces typed SnapshotReadError::Stale evidence.
Weavatrix package boundary
Scan owns repository selection, safe bounded content delivery, hashes, revision, cancellation, and incremental deltas. Search owns literal/regex matching, line context, encodings, compressed inputs, indexes, and result formatting. Clone owns token normalization, Moss/winnowing, MinHash/LSH, Aho-Corasick, and clone grouping. Graph consumes normalized facts and keeps graph algorithms; content-search and clone algorithms do not belong there.
Incremental consumers
Two completed reports produce a stable changed-file set without filesystem access:
use ;
let previous = new.scan?;
// Apply repository changes, then scan again.
let current = new.scan_incremental?;
let delta = current.delta_from;
assert!;
println!;
# Ok::
Unchanged files reuse prior SHA-256 values without reopening their content.
Reports from another root, legacy reports without version evidence, and files
whose size/version changed are read normally. Rename evidence is emitted only
when the same content hash is unique in both
manifests; duplicate-content moves remain explicit add/remove pairs instead of
being guessed. Metadata-only deltas compare size plus available file-version
evidence; callers that need content certainty should keep hashes enabled.
Partial scans always produce DeltaQuality::Partial.
Persistent consumers should store ScanReport::to_cache() instead of the full
report and pass it to Scanner::scan_cached. ScanCache has an explicit format
version and contains only the canonical root plus relative path, size, version,
hash, whole-content fingerprint, and binary-check evidence for reusable files.
CacheValidationPolicy::Fast trusts matching file-version evidence.
CacheValidationPolicy::Strict additionally reads a compact whole-content
fingerprint before reusing the cached SHA-256, protecting coarse-timestamp and
network filesystems from same-size, same-timestamp changes.
use ;
let options = default
.with_cache_validation;
let first = new.options.scan?;
let second = new.options.scan_cached?;
assert_eq!;
# Ok::
Long-lived file watchers that only need ignore decisions can keep a
RepositoryMatcher. Consumers that need the exact Scanner selection contract
can use SelectionMatcher; it additionally applies depth, symlink, standard
directory, named type, extension, maximum-size, and filesystem-boundary policy:
use ;
let options = default.with_extensions;
let mut matcher = with_options?;
let decision = matcher.matched?;
assert!;
# Ok::
SelectionMatcher::matched_entry reuses metadata already present in a
Weavatrix WalkEntry, while matched safely classifies an isolated existing
path and its ancestors. Matchers are cloneable for worker-local incremental
queries. Both matcher types expose refresh(). Refresh builds a replacement
ignore matcher first, so a failure leaves the existing matcher usable.
WatcherEventAdapter converts create/modify/remove/rename notifications from
any watcher library into sorted relative WatchPlan invalidations. Events
outside the root are rejected, while directory, ignore-source, and explicit
rescan events request a full scan. ScanCache::apply_watch_plan removes only
affected entries or clears the cache when selection may have changed.
Scanner::scan_watch_plan goes further: for a safe file-only plan it re-matches
and inspects only changed paths, removes deleted paths from the previous
manifest, keeps unchanged evidence, and recomputes the deterministic revision
without traversing the tree. Structural, unsafe, partial, or selection-changing
plans automatically use a complete scan.
For indexes that consume bytes directly, visit_changed_content performs the
same safe path matching and one-pass verified content delivery without walking
unchanged directories. Its revision covers only the changed subset;
visit_changed_content_streaming omits that subset manifest and revision.
With the optional notify feature, plan_notify maps notify::Event batches
directly. Access-only events are ignored; imprecise, rescan, and possibly
structural events conservatively request a complete scan.
SkipKind distinguishes:
BinaryConcurrentModificationExtensionFileSystemBoundaryHiddenIgnoredIoErrorMaxDepthOverrideOversizedPathEscapeScanLimitStandardDirectorySymlinkSymlinkLoop
This distinction matters to analyzers: "not selected by policy" is different from "unreadable" or "outside the repository."
Configuration
ScanOptions exposes:
| Option | Default | Purpose |
|---|---|---|
max_file_bytes |
1,500,000 | Reject oversized source candidates |
extensions |
Empty | Empty accepts every extension |
file_types |
Empty | Named file-name/repository-relative glob groups |
ignore_files |
.gitignore, .ignore, .weavatrixignore |
Hierarchical local ignore files |
ignore_policy |
Repository-only | Optional parents, .git/info/exclude, global Git and explicit files |
override_rules |
Empty | Request-level include/exclude globs above ignore sources |
ignore_case_insensitive |
false |
Optional ASCII case-insensitive ignore matching |
skip_hidden |
false |
Skip dot-prefixed and Windows-hidden entries unless included |
standard_skips |
Enabled | Skip generated/vendor directories |
hash_file_contents |
true |
Attach per-file hashes and content-sensitive revision |
cache_validation |
Fast |
Trust file-version evidence, or verify a whole-content fingerprint in Strict mode |
content_validation |
Strict |
Verify newly opened content before and after reading, or omit the post-read check in Fast mode |
content_discovery |
Streaming |
Constant-memory overlapped discovery, or compact BufferedParallel discovery for minimum latency |
detect_binary_files |
true |
Reject files containing a NUL byte |
evidence |
Complete |
Keep all typed exclusions, or only selected files |
parallelism |
0 |
Traversal/content workers; zero uses bounded available parallelism |
traversal_parallelism |
None | Optional traversal-only worker override |
content_parallelism |
None | Optional content-inspection worker override |
limits.max_entries |
None | Bound examined filesystem entries |
limits.max_total_bytes |
None | Deterministically bound selected content bytes |
limits.timeout |
None | Stop traversal/content inspection after a duration |
cancellation |
None | Cooperative cross-thread cancellation token |
walk.max_depth |
None | Limit entry depth; root is zero |
walk.min_depth |
0 |
Suppress shallower results while still traversing them |
walk.max_open |
64 |
Bound live directory handles/workers |
walk.same_file_system |
false |
Stop at filesystem boundaries when enabled |
walk.follow_links |
false |
Follow only in-root links and detect cycles |
walk.root_symlink_policy |
Follow |
Follow or reject the explicitly supplied root symlink |
walk.error_policy |
Continue |
Continue with partial typed evidence or abort |
walk.collect_metadata |
true in ScanOptions |
Reuse directory-entry metadata without reopening selected paths |
The standard directory policy skips:
.git .hg .svn .venv __pycache__ build coverage dist
node_modules target vendor
Disable it when another layer owns generated-directory policy:
use ;
let mut options = default;
options.standard_skips = Disabled;
NamedFileTypes::defaults() provides 265 deterministic language, markup, data,
build, configuration, and infrastructure definitions backed by 678 patterns.
That is a strict name-and-pattern superset of the 224 definitions and 594
patterns in ignore 0.4.31. len() and names() expose the catalog without
activating it. Types can be composed, selected, and negated; later matching
selections win:
use NamedFileTypes;
let types = defaults
.with_composed_type
.select
.negate;
Ignore semantics
Ignore files are loaded hierarchically with source precedence
.weavatrixignore/custom > .ignore > .gitignore >
.git/info/exclude > global Git. Deeper files win within the same source
class. Supported Git-style constructs include:
- comments and escaped leading
#/!; - negation with
!; - root-anchored patterns;
- directory-only patterns;
*,**, and?;- character classes, negated classes, and ranges;
- brace alternatives such as
{foo,bar}; - escaped literals and escaped trailing spaces.
The default scanner intentionally does not read global Git configuration,
parent rules outside the scan root, or .git/info/exclude; repository-local
selection therefore stays portable. IgnorePolicy::git_compatible() enables
all three explicitly inside Git repositories, records their content hashes,
honors unconditional and matching includeIf Git config includes, and marks
host-dependent reports non-portable. Local .gitignore, .ignore,
and custom sources can be toggled independently. Request-level override globs
use ignore::Override semantics: ordinary patterns include and leading !
patterns exclude. Explicit includes can opt paths back into standard-directory
and extension filtering, but never bypass size or binary safety checks.
RepositoryMatcher::matched exposes the winning typed
decision without requiring a full walk. Differential tests compare
exact selected path sets against the ignore crate for anchored, nested,
negated, wildcard, and character-class fixtures plus 96-seed deterministic
randomized rule sets and direct comparison with git check-ignore. A scheduled
workflow also runs 100,000 arbitrary-byte grammar cases and deterministic
directory-read fault injection. Stress cases cover
deep trees, permission errors, raw non-UTF8 ignore rules/names, percent escapes,
and followed symlink loops. The
differential suite and competitor crates are dev-only.
Safety model
- never executes repository code;
- never starts subprocesses or accesses the network;
- canonicalizes and validates the root before traversal;
- does not follow symlink entries by default;
- rejects followed links outside the canonical root and detects cycles;
- follows in-root links in parallel using per-task ancestry, without a serial mode switch;
- can enforce a same-filesystem boundary;
- continues after independent local errors by default and marks the report partial;
- caps selected file size before content reads;
- rejects repository-local ignore-file symlinks and path traversal;
- supports entry, total-byte, timeout, and cooperative cancellation bounds;
- exports path-safe portable evidence without host paths or diagnostic text;
- revalidates snapshot content before and after bounded consumer reads;
- forbids unsafe Rust.
The scanner is read-only. Concurrent filesystem changes between discovery and
the final metadata check are surfaced as ConcurrentModification warnings and
skips under Continue, or as the first error under Abort.
Benchmarks
Run all included benchmarks:
Run the competitor comparison:
The Competitor benchmarks workflow runs the same output-equivalent comparison
on Ubuntu, Windows, and macOS for scanner or benchmark changes.
Run exact selected-path parity on a real repository:
$env:WEAVATRIX_BENCH_ROOT = "C:\path\to\repository"
cargo bench --locked --bench real_repository
Run skewed, deep, first-touch, bounded-handle, large-content, and incremental profiles:
Run root-policy, stateful-callback, bounded-pull, and watcher-adapter profiles:
Create, verify, and measure an opt-in synthetic scale fixture outside the repository:
The command refuses to populate an existing unmarked directory. The fixture
uses 500 empty .rs files per directory; one sixth of its directories are
excluded by a root .ignore. verify asserts the exact sorted path/size
manifest from both full and compact scanners against ignore, not only the
selected count. The profile is
intentionally opt-in because creating and removing hundreds of thousands or
millions of filesystem entries is itself expensive.
The synthetic comparison uses 6,000 source files across Rust, Go, and TypeScript in 80 sibling directories. It runs two warmups and 11 interleaved measured samples, then reports the median. Raw walkers must produce the same fully sorted native relative-path set; the ignore-aware comparison additionally checks the same normalized path-and-size manifest.
Sample result on Windows 11, Rust 1.97.1, warm filesystem cache, measured
2026-07-26 against ignore 0.4.31, walkdir 2.5.0, and jwalk 0.8.1:
| Mode | Library | Files | Median |
|---|---|---|---|
| Raw paths | weavatrix Walker |
6,004 | 16.5 ms |
| Raw paths | weavatrix ParallelWalker |
6,004 | 10.2 ms |
| Raw paths | ignore | 6,004 | 17.3 ms |
| Raw paths | walkdir | 6,004 | 15.4 ms |
| Raw paths | jwalk | 6,004 | 10.8 ms |
| Ignore-aware manifest | weavatrix Scanner serial |
6,001 | 35.2 ms |
| Ignore-aware manifest | weavatrix Scanner parallel |
6,001 | 20.7 ms |
| Ignore-aware manifest | ignore | 6,001 | 37.4 ms |
| Rich SHA-256 manifest | weavatrix Scanner |
6,000 | 146.2 ms |
On Windows, 0.4.1 reuses the file metadata already collected by the walker when applying the hidden-attribute policy. This removes a redundant metadata query per selected entry; hidden, ignore, override, and manifest results are unchanged.
Each row is the median of five independent process medians. Every process runs
11 interleaved output-equivalent samples after two warmups. On this measurement
ParallelWalker was 5.9% faster than jwalk; the parallel
selected-manifest Scanner was 44.7% faster than ignore. The rich row
additionally reads content, detects binaries, computes SHA-256 hashes, captures
snapshot evidence, and records typed exclusions. Absolute timings vary by
filesystem, cache, antivirus, CPU, and operating system; the benchmark workflow
reruns the same checks on Ubuntu, Windows, and macOS.
The separate scale profile was measured on the same Windows host with a warm filesystem cache. Raw streaming rows are the median of seven independent process medians with seven measured runs after warmup. Metadata rows use five process medians with five runs, and the rich SHA-256 row uses three process medians with three runs. Peak working set is sampled in a fresh process containing one warmup and one measured run.
| Work | Implementation | Files | Median | Peak |
|---|---|---|---|---|
| Raw streaming count | ParallelWalker::visit |
300,000 | 90.7 ms | 6.9 MiB |
| Raw streaming count | jwalk | 300,000 | 96.6 ms | 56.5 MiB |
| Raw collected paths | ParallelWalker::walk |
300,000 | 208.8 ms | 162.5 MiB |
| Ignore-aware path/size manifest | Scanner metadata-only |
250,000 | 346.9 ms | 97.3 MiB |
| Ignore-aware path/size manifest | ignore | 250,000 | 432.4 ms | 20.5 MiB |
| No-ignore metadata manifest | Scanner metadata-only |
300,000 | 368.9 ms | 157.5 MiB |
| Ignore-aware content manifest | Scanner SHA-256 |
250,000 | 5,693.0 ms | 235.9 MiB |
| Ignore-aware path emission | rg --files to null |
250,000 | 1,117.7 ms | 35.3 MiB |
The scanner row is a stronger contract than the comparison manifest: it also
captures native version evidence, hashes ignore inputs, normalizes paths,
sorts deterministically, and computes a revision. The ripgrep row is a
whole-command throughput guardrail, not a library-equivalent benchmark:
ripgrep formats and writes every path, while the library rows count or retain
typed entries in-process. Content hashing is necessarily compared separately
because rg --files does not open and hash every selected file. On this sample
the bounded streaming walker stayed below jwalk in both median and peak memory,
and metadata scanning stayed below both the output-equivalent ignore manifest
and the ripgrep guardrail.
The same profile was then expanded to 1,000,000 files in 2,000 directories;
833,000 files passed .ignore. Raw rows below are the median of five
independent process medians with five measured runs after warmup. The new
compact/full/ignore manifest rows are the median of three independent process
medians, each with one warmup and five measured runs; peak is the median of the
three sampled process peaks. Serial, collected, no-ignore, rich, and ripgrep
rows retain their earlier methodology described in the preceding revision of
this benchmark.
| Work | Implementation | Files | Median | Peak |
|---|---|---|---|---|
| Raw serial count | Walker |
1,000,000 | 546.3 ms | 5.0 MiB |
| Raw streaming count | ParallelWalker::visit |
1,000,000 | 264.7 ms | 7.6 MiB |
| Raw streaming count | jwalk | 1,000,000 | 313.1 ms | 159.7 MiB |
| Raw serial count | walkdir | 1,000,000 | 584.5 ms | 4.6 MiB |
| Raw collected paths | ParallelWalker::walk |
1,000,000 | 650.2 ms | 529.0 MiB |
| Ignore-aware compact path/size manifest | Scanner::scan_compact metadata-only |
833,000 | 1,019.4 ms | 63.2 MiB |
| Ignore-aware full path/size manifest | Scanner::scan metadata-only |
833,000 | 1,665.9 ms | 309.6 MiB |
| Ignore-aware path/size manifest | ignore | 833,000 | 2,106.7 ms | 56.4 MiB |
| No-ignore metadata manifest | Scanner metadata-only |
1,000,000 | 1,125.2 ms | 369.6 MiB |
| Ignore-aware content manifest | Scanner SHA-256 |
833,000 | 24,716.9 ms | 776.2 MiB |
| Ignore-aware path emission | rg --files to null |
833,000 | 2,156.3 ms | 42.3 MiB |
| No-ignore path emission | rg --no-ignore --files to null |
1,000,000 | 2,535.6 ms | 95.1 MiB |
On this million-file sample, streaming traversal was 15.5% faster than jwalk
and used 95.2% less peak working set. The compact metadata Scanner was 51.6%
faster than the output-equivalent ignore manifest while using 6.8 MiB more
peak working set. Compared with the compatibility-oriented full report, it
reduced peak memory by 79.6% and median time by 38.8%. Use the full report when
every entry needs an owned absolute path and file-version evidence; use the
compact report for large retained manifests, streaming traversal for raw
consumers, and rich hashing only when content evidence is required.
The API benchmark uses the same corpus and methodology. Five process medians
measured the ordered bounded DFS iterator at 6.4 ms versus 8.2 ms for jwalk.
The new parallel typed stateful batch iterator measured 3.4 ms versus 4.3 ms
for jwalk::process_read_dir, 22.2% faster on this sample while preserving the
same batch mutation, child-state propagation, pruning, and ordered output
contract. Output-equivalent streaming over two roots and 12,008 files measured
3.375 ms for Weavatrix versus 7.408 ms for ignore::build_parallel, 54.4%
faster on this sample. Each process result is itself the median of 11
interleaved measured runs after two warmups, not a single best run. Watcher
planning and changed-path scan profiles remain in the same reproducible
benchmark.
The one-pass content profile selects and reads the same 6,001 tiny files in
every case. Fast and the corresponding ignore baseline both validate the
opened handle once; Strict and its baseline validate before and after the
read. The raw ignore row intentionally omits snapshot validation and is the
lower-contract throughput floor:
| Content pipeline | Retention / validation | Median |
|---|---|---|
Weavatrix visit_content_streaming |
None / opened handle | 62.793 ms |
Weavatrix visit_content_streaming |
None / before and after | 73.872 ms |
Weavatrix visit_content |
Revision / opened handle | 83.307 ms |
Weavatrix visit_content |
Revision / before and after | 78.475 ms |
ignore + File::read |
None / opened handle | 76.139 ms |
ignore + File::read |
None / before and after | 80.529 ms |
ignore + File::read |
None / unchecked | 70.748 ms |
These are five independent process medians from 11 interleaved samples after
two warmups, measured on the Windows host above. Streaming Fast was 11.2%
faster than unchecked ignore despite validating the opened handle; Streaming
Strict was 8.3% faster than the equivalent before/after baseline. The
two-root streaming profile processed 12,002 files in 152.013 ms versus
163.588 ms for verified ignore, 7.1% faster. A safe 1,024-file changed plan
completed in 47.879 ms versus 109.694 ms for a full 6,000-file scan.
The scale memory check used a temporary 100,000-file fixture with 83,000 selected empty files. A fresh-process sample measured 7.3 MiB peak for streaming versus 39.4 MiB for revision retention, an 81.5% reduction; three-run medians were 2,109.956 and 2,182.466 ms respectively. This validates the scanner-to-consumer handoff, not future literal or regex matching: final comparison with ripgrep belongs to the Search package and must include matched output, line handling, and encoding policy.
Source review explains the remaining differences:
walkdirstreams unsorted directory entries and bounds open descriptors;jwalkschedulesread_dirwork through Rayon and restores ordered output;ignorecompiles patterns intoGlobSetmatchers and shares inherited matchers;- Weavatrix
Walkerstreams iterative DFS, bounds live handles and buffers the oldest remaining frame only whenmax_openis reached; its plain-entry fast path and consumingWalkEntry::into_pathavoid universal-policy checks and long-path clones in raw traversal; - Weavatrix
ParallelWalkerexpands a small shallow frontier for narrow roots, then uses up to 16 Windows or 8 Unix workers without serially over-expanding small trees; bounded lanes keep report order independent of worker completion; - Weavatrix
Scannerreuses inherited rules, indexes exact literals, specializes prefix/suffix globs, prefilters complex patterns, and sorts only the final report.
The optional real-repository benchmark keeps repository identities, paths, and
individual measurements local. Published documentation contains only synthetic
and hosted-runner corpus results. Every local comparison first asserts the exact
same sorted (normalized path, bytes) manifest.
The synthetic stress profile measured a skewed raw tree at 6.7 ms
(ParallelWalker), 8.2 ms (jwalk), and 14.1 ms (walkdir). The expanded
synthetic deep-tree profile contains 60 levels and 7,680 files; five independent
process medians measured 18.9 ms for Walker and 20.0 ms for walkdir, making
Walker 5.6% faster on that sample. An unchanged synthetic 12 MiB SHA-256
manifest fell from 164.4 ms full scan to 1.9 ms with incremental hash reuse.
Treat these as reproducible samples, not universal constants.
Correctness checks
The test suite covers:
- deterministic results and revisions;
- ignore-rule precedence and nested ignore files;
- repository-only, Git-exclude, parent, explicit and reusable-matcher policies;
- representative and randomized parity with
ignore; - raw entry parity with
walkdirandjwalk; - opt-in exact path/size manifest parity with
ignoreat arbitrary scale, including the measured 1,000,000-file fixture; - iterative deep trees, bounded handles, local error continuation, non-UTF8 paths, and symlink loops;
- concurrent mutation detection and same-size incremental changes;
- strict cache validation under simulated size/timestamp collisions;
- multi-root walking, named file types, custom native sorting, directory filtering, and contents-first ordering;
- single-file roots, full-entry sorting, built-in type composition/negation,
strict catalog superset parity with
ignore, redirected-stdout protection, parallel raw roots, andnotifyconversion; - binary, oversized, extension, generated-directory, and symlink policies;
- serial/parallel content-inspection equivalence;
- streaming parallel pruning and cancellation;
- ordered bounded parallel DFS and parallel followed-link cycle handling;
- global, dedicated, and rejecting external runtimes, typed submission failures, reentrant callbacks, panic propagation, and pool reuse;
- parallel multi-root callback tagging, subtree pruning, global quit, cancellation, and deterministic per-root reports;
- serial/parallel stateful batch order, pruning, entry state, inherited child state, and multi-worker execution;
- full/compact exact manifest and revision equivalence;
- changed-path watcher manifests, arbitrary-byte ignore grammar, and injected directory-read failures;
- manifest delta evidence and live matcher refresh;
- optional Serde support.
The real-repository benchmark compares the complete normalized selected-path
set against ignore. Its comparison policy disables Weavatrix's file-size cap
so an oversized file cannot masquerade as an ignore-rule mismatch.
Development
The MSRV is Rust 1.88. CI checks Rust 1.88 on Linux, Windows, and macOS, with stable test coverage on all three platforms.
Relationship to Weavatrix
weavatrix-scan owns repository discovery. It does not parse languages or
build graphs. weavatrix-graph
owns typed graph primitives. Higher-level Weavatrix crates can compose both
without coupling either library to MCP, a CLI, or language-specific parsers.
License
MIT © 2026 Sergii Ziborov.