dry4rust
Created by Matjaz Domen Pecan as cargo-dupes,
forked and extended by Umberto Gotti under the MIT licence.
A cargo subcommand for detecting duplicated code patterns in Rust -- DRY (Don't Repeat
Yourself) analysis.
Provenance
dry4rust is a fork of cargo-dupes by
Matjaz Domen Pecan, used under the MIT licence. The duplicate-detection
engine documented below — AST normalisation, fingerprint hashing and Dice-coefficient
similarity — is his work, and both copyright notices are carried in LICENSE and in every
source file.
The fork keeps the Rust analyser and drops the tree-sitter and Python backends, so this is a
single-language tool by intent rather than a general one. What is added beyond that point is
noted in CHANGELOG.md.
Family
Part of the same family as grip (testability),
braintax (cognitive load), and
crap4rust (change-risk complexity × coverage).
Where those three measure how safe, understandable, and risky a codebase is, dry4rust
measures how much of it is needlessly repeated.
Status
The engine is upstream's, with four corrections it does not have — an exact size bound, aligned sequence children, a fingerprint format that survives a toolchain upgrade, and near-duplicate detection that no longer hides what exact detection reports. Added on top: baseline mode, and configuration that cannot hold an impossible threshold.
The surrounding repository is under this family's house rules and gates: cargo stern4rust
reports no offences with all twenty-one rules applied and no baseline of its own,
cargo crap4rust finds no function at or above 15, every source file has a mirrored test
file, and no file reaches 10 on cargo iceberg4rust.
Commands below are cargo dry4rust; the upstream cargo dupes spelling is gone.
Install
cargo install cargo-dry4rust
License
MIT — see LICENSE, which carries both copyright notices.
How It Works
dry4rust parses Rust source files into ASTs using syn, then normalizes each function, method, and closure into a canonical form where:
- Identifiers are replaced with positional placeholders (so
foo(x)andbar(y)are identical) - Literal values are erased but types preserved (
42and99are both "integer literal") - Control flow structure is preserved exactly
- Macro invocations become opaque nodes
This normalized AST is hashed into a fingerprint for exact duplicate detection, and compared tree-by-tree using the Dice coefficient for near-duplicate detection.
Usage
cargo dry4rust [OPTIONS] [COMMAND]
Commands:
stats Show duplication statistics only
report Show full duplication report (default)
check Check for duplicates and exit with non-zero if thresholds exceeded
ignore Add a fingerprint to the ignore list
ignored List all ignored fingerprints
cleanup Remove ignore entries that no longer match anything
baseline Record the duplication that is already there
Options:
-p, --path <PATH> Path to analyze (defaults to current directory)
--min-nodes <MIN_NODES> Minimum AST node count for analysis
--min-lines <MIN_LINES> Minimum source line count for analysis
--threshold <THRESHOLD> Similarity threshold (0.0-1.0)
--format <FORMAT> Output format [default: text] [possible values: text, json]
--exclude <EXCLUDE> Exclude patterns (can be repeated)
--exclude-tests Exclude test code (#[test] functions and #[cfg(test)] modules)
-s, --sub-function Also analyse if-branches, match arms, loop bodies and closure bodies
--min-sub-nodes <N> Minimum AST node count for a sub-function unit [default: 5]
--baseline <PATH> Judge the run against a recorded baseline of inherited duplication
-h, --help Print help
-V, --version Print version
Every example below is real output from fixture/exact_dupes, which ships in the
repository -- cargo run -p cargo-dry4rust -- --path fixture/exact_dupes report reproduces
it.
--threshold and the two --max-*-percent ceilings are checked against their ranges. A
threshold outside 0.0..=1.0, or a percentage outside 0.0..=100.0, fails the run with a
message naming the field rather than being accepted and quietly finding nothing.
Examples
Full report:
=====================
)
)
)
)
)
================
)
)
)
)
Statistics only:
=====================
)
)
)
)
)
JSON output:
{
}
Every --format json run is a single JSON document, so jq and any ordinary parser
read it in one call. report and check name their sections:
|
|
[
| command | keys |
|---|---|
stats |
the summary fields, at the top level |
report |
stats, exact, near, and sub_exact/sub_near when sub-function analysis found any |
check |
stats, passed, breaches, and exact/near when a ceiling was breached |
exact and near are always present on report because those are always analysed. The
sub-function sections are absent rather than empty when the analysis was not asked for, so
[] never stands in for "not looked at". On check, the groups behind a breach are listed
once even when two ceilings on the same set are breached together.
CI check (fail if any exact duplicates exist):
# Exits with code 1 if exact duplicate groups > 0
# Exits with code 0 if within thresholds
CI check with percentage thresholds (fail if >5% of lines are exact duplicates):
# Exits with code 1 if exact duplicate lines exceed 5% of total lines
Exclude test code (inline #[cfg(test)] modules and #[test] functions):
Exclude test directories by path:
Only report duplicates that are at least 10 lines long:
Lower the similarity threshold:
Sub-function Analysis
By default a code unit is a whole function, method or closure. Two functions that share a
copy-pasted match arm but differ elsewhere are not duplicates of each other, and nothing
is reported.
--sub-function (or -s) also treats each if-branch, match arm, loop body and closure
body as a unit in its own right:
)
)
=============================
)
)
)
)
)
)
Each member names the function it came from, and the line range shown is that parent
function's — not the branch's. Sub-function units are grouped separately from top-level
ones and counted under their own headings, so a function never shares a group with its own
branch. --min-sub-nodes (default 5) is the floor a branch must reach to be considered
at all; raise it when small branches produce noise.
Two limits are worth knowing before trusting the output, both in docs/OPEN_POINTS.md: sub-function findings restate a function-level one when two functions are already duplicates of each other, and the parent line range means a report cannot be read straight to the branch.
Configuration
Configuration can be provided in three ways (in order of precedence):
- CLI flags (highest priority)
dry4rust.tomlin the project rootCargo.tomlunder[package.metadata.dry4rust]
dry4rust.toml
= 15
= 5
= 0.85
= ["tests", "benches"]
= true
= 0
= 10
= 5.0
= 10.0
Cargo.toml
[]
= 15
= 0.85
= ["tests"]
Configuration Options
| Option | Default | Description |
|---|---|---|
min_nodes |
10 |
Minimum AST node count for a code unit to be analyzed. Increase to skip trivial functions. |
min_lines |
0 |
Minimum source line count for a code unit to be analyzed. 0 means disabled. |
similarity_threshold |
0.9 |
Minimum similarity score for near-duplicate detection. Must be within 0.0..=1.0. |
exclude |
[] |
Path patterns to exclude from scanning (substring match). |
exclude_tests |
false |
Exclude #[test] functions and #[cfg(test)] modules from analysis. |
sub_function |
false |
Also analyse if-branches, match arms, loop bodies and closure bodies. |
min_sub_nodes |
5 |
Minimum AST node count for a sub-function unit to be analyzed. |
baseline |
None |
Path to a recorded baseline of inherited duplication, relative to the analysed root. |
max_exact_duplicates |
None |
For check subcommand: maximum allowed exact duplicate groups. |
max_near_duplicates |
None |
For check subcommand: maximum allowed near-duplicate groups. |
max_exact_percent |
None |
For check subcommand: maximum allowed exact duplicate line percentage. Must be within 0.0..=100.0. |
max_near_percent |
None |
For check subcommand: maximum allowed near-duplicate line percentage. Must be within 0.0..=100.0. |
A value outside the range its field allows fails the run with a message naming the field, and exits 2. A config file that cannot be read or parsed is passed over, because a project with no configuration looks the same.
Ignoring Duplicates
Some duplicates are intentional (e.g., test helpers, trait implementations). You can ignore them by fingerprint:
# Add a fingerprint to the ignore list
# List ignored fingerprints
)
# Ignored groups are automatically filtered from reports and checks
# The ignored group will not appear
The ignore list is stored in .dry4rust-ignore.toml in the project root. When cleanup
prunes the last entry it removes the file rather than leaving an empty one behind: an empty
suppression list says exactly what no file says.
Entries whose fingerprint no longer matches anything go stale — after a refactor, or after
an upgrade that changed the fingerprint format. cleanup prunes them:
Adopting on a Codebase That Already Has Duplication
ignore is for duplication that is meant to be there. Duplication nobody has got to yet
is a different thing, and recording it as intentional — one fingerprint at a time, with a
reason invented to fill the field — writes a lie into a file that outlives whoever wrote
it.
Record it as a baseline instead. check then fails on what is added, not on what was
inherited:
# Record what is already there
)
# From now on, a zero ceiling is a gate rather than a wall
Commit dry4rust-baseline.json, or put baseline = "dry4rust-baseline.json" in
dry4rust.toml so every run picks it up without the flag. baseline --dry-run lists what
would be recorded without writing anything.
A baseline records a group's fingerprint and its member count. A third copy of an already-recorded duplicate makes the group larger than what was recorded, so it is reported — an exact group keeps its fingerprint when a copy joins it, and a baseline keyed on the fingerprint alone would inherit every future copy. Deleting a copy is admitted: progress is not something to fail on.
A baseline that cannot be read — missing, malformed, or written by a version with a different format — fails the run and says so, rather than being treated as empty. Every summary a baseline touched states how many groups it suppressed, so a stale baseline is visible rather than silent.
The reasoning is in ADR-BaselineIsInheritedNotForgiven.
CI Integration
Use the check subcommand in CI pipelines:
# GitHub Actions example
- name: Check for code duplication
run: cargo dry4rust check --max-exact 0 --max-exact-percent 5.0
On a codebase with existing duplication, add a baseline so the gate measures what the change introduced:
- name: Check for new code duplication
run: cargo dry4rust --baseline dry4rust-baseline.json check --max-exact 0
Exit codes:
- 0 — Check passed (within thresholds)
- 1 — Check failed (thresholds exceeded)
- 2 — Error (no source files, invalid path, etc.)
What Gets Analyzed
| Code Unit | Description |
|---|---|
| Functions | Top-level fn items |
| Methods | fn items inside impl blocks |
| Trait impls | fn items inside impl Trait for Type blocks |
| Closures | Closure expressions (above the min node threshold) |
The scanner automatically:
- Skips
target/directories - Skips hidden directories (starting with
.) - Respects exclude patterns
- Handles parse errors gracefully (skips unparseable files with a warning)
Documentation
| Where | What |
|---|---|
| docs/ARCHITECTURE.md | The pipeline, the components, and the data model |
| docs/FORMULA.md | Fingerprinting and the similarity score, precisely |
| docs/ADRs/ | The load-bearing decisions and why they were forced |
| docs/OPEN_POINTS.md | Where the model is known to be thin, including two silent false-negative sources |
| docs/ROADMAP.md | Direction, current baseline, and what is left |
| docs/IMPLEMENTED-FEATURES.md | What each version added, upstream's included |
| docs/dry4rust-dossier.md | Research on where this tool is going, and what already holds the ground |
Development
Requirements: Rust 1.91+ (edition 2024), matching the rust-version every
workspace crate declares.
The repository is a workspace: core/ is the published crate, validation/ holds the
end-to-end tests, xtask/ runs the stage 2 gates, and fixture/ is the corpus both
are pointed at.
Both gates are mandatory after any change under src/ or tests/, and both run
the same way on Windows, Linux and macOS:
Stage 1 is cargo built-ins only, so it works on a fresh checkout with none of the
tools below installed. Note that it lints test targets too
(--all-targets) — combined with the pedantic and nursery groups that core
and xtask enable, that holds the test files to the same bar as the sources.
Stage 2 is cargo xtask stage2, a real crate rather than a script, so the gate
argument lists and failure messages are covered by its own integration tests.
xtask is itself a workspace member and is gated like everything else — the
crate that runs the gates is not exempt from them.
Everything the two stages need, none of which ships with cargo:
| Tool | Install | Needed by |
|---|---|---|
just |
cargo install just |
both stages |
cargo-llvm-cov |
cargo install cargo-llvm-cov |
stage 2 |
llvm-tools rustup component |
rustup component add llvm-tools |
stage 2 |
cargo-stern4rust |
cargo install cargo-stern4rust |
stage 2 |
cargo-crap4rust |
cargo install cargo-crap4rust |
stage 2 |
cargo-twin4rust |
cargo install cargo-twin4rust |
stage 2 |
cargo-iceberg4rust |
cargo install cargo-iceberg4rust |
stage 2 |
cargo-llvm-cov and llvm-tools are what the CRAP gate needs; without them it
fails with a bare exit code that says nothing about a missing install.
CI (.github/workflows/ci.yml) runs both stages on Ubuntu, Windows and macOS
for every pull request and every push to main.