dry4rust
Created by Matjaz Domen Pecan as cargo-dupes,
forked and extended by Umberto Gotti under the MIT licence.
A cargo subcommand for detecting duplicated code patterns in Rust -- DRY (Don't Repeat
Yourself) analysis.
Provenance
dry4rust is a fork of cargo-dupes by
Matjaz Domen Pecan, used under the MIT licence. The duplicate-detection
engine documented below — AST normalisation, fingerprint hashing and Dice-coefficient
similarity — is his work, and both copyright notices are carried in LICENSE and in every
source file.
The fork keeps the Rust analyser and drops the tree-sitter and Python backends, so this is a
single-language tool by intent rather than a general one. What is added beyond that point is
noted in CHANGELOG.md.
Family
Part of the same family as grip (testability),
braintax (cognitive load), and
crap4rust (change-risk complexity × coverage).
Where those three measure how safe, understandable, and risky a codebase is, dry4rust
measures how much of it is needlessly repeated.
Status
The engine is upstream's, with four corrections it does not have — an exact size bound, aligned sequence children, a fingerprint format that survives a toolchain upgrade, and near-duplicate detection that no longer hides what exact detection reports. Added on top: baseline mode, and configuration that cannot hold an impossible threshold.
The surrounding repository is under this family's house rules and gates: cargo stern4rust
reports no offences with all twenty-two rules applied to every workspace member and none
skipped -- its one baseline holds the ten end-to-end scenario files that have no source
file to be named after -- cargo crap4rust finds no function at or above 15, every source
file has a mirrored test file, no file reaches 10 on cargo iceberg4rust, and dry4rust
run on its own source finds nothing beyond the two near-duplicate groups its baseline records.
Commands below are cargo dry4rust; the upstream cargo dupes spelling is gone.
Install
cargo install cargo-dry4rust
License
MIT — see LICENSE, which carries both copyright notices.
How It Works
dry4rust parses Rust source files into ASTs using syn, then normalizes each function, method, and closure into a canonical form where:
- Identifiers are replaced with positional placeholders (so
foo(x)andbar(y)are identical) - Literal values are erased but types preserved (
42and99are both "integer literal") - Control flow structure is preserved exactly
- Macro invocations become opaque nodes
This normalized AST is hashed into a fingerprint for exact duplicate detection, and compared tree-by-tree using the Dice coefficient for near-duplicate detection.
Usage
cargo dry4rust [OPTIONS] [COMMAND]
Commands:
stats Show duplication statistics only
report Show full duplication report (default)
check Check for duplicates and exit with non-zero if thresholds exceeded
ignore Add a fingerprint to the ignore list
ignored List all ignored fingerprints
cleanup Remove ignore entries that no longer match anything
baseline Record the duplication that is already there
Options:
-p, --path <PATH> Path to analyze (defaults to current directory)
--min-nodes <MIN_NODES> Minimum AST node count for analysis
--min-lines <MIN_LINES> Minimum source line count for analysis
--threshold <THRESHOLD> Similarity threshold (0.0-1.0)
--format <FORMAT> Output format [default: text] [possible values: text, json]
--exclude <EXCLUDE> Exclude patterns (can be repeated)
--exclude-tests Exclude test code (#[test] functions and #[cfg(test)] modules)
-s, --sub-function Also analyse if-branches, match arms, loop bodies and closure bodies
--min-sub-nodes <N> Minimum AST node count for a sub-function unit [default: 5]
--baseline <PATH> Judge the run against a recorded baseline of inherited duplication
-h, --help Print help
-V, --version Print version
Every example below is real output from fixture/exact_dupes, which ships in the
repository -- cargo run -p cargo-dry4rust -- --path fixture/exact_dupes report reproduces
it.
--threshold and the two --max-*-percent ceilings are checked against their ranges. A
threshold outside 0.0..=1.0, or a percentage outside 0.0..=100.0, fails the run with a
message naming the field rather than being accepted and quietly finding nothing.
Examples
Full report:
=====================
)
)
)
)
)
================
)
)
)
)
Statistics only:
=====================
)
)
)
)
)
JSON output:
{
}
Every --format json run is a single JSON document, so jq and any ordinary parser
read it in one call. report and check name their sections:
|
|
[
| command | keys |
|---|---|
stats |
the summary fields, at the top level |
report |
stats, exact, near, and sub_exact/sub_near when sub-function analysis found any |
check |
stats, passed, breaches, and exact/near when a ceiling was breached |
exact and near are always present on report because those are always analysed. The
sub-function sections are absent rather than empty when the analysis was not asked for, so
[] never stands in for "not looked at". On check, the groups behind a breach are listed
once even when two ceilings on the same set are breached together.
CI check (fail if any exact duplicates exist):
# Exits with code 1 if exact duplicate groups > 0
# Exits with code 0 if within thresholds
CI check with percentage thresholds (fail if >5% of lines are exact duplicates):
# Exits with code 1 if exact duplicate lines exceed 5% of total lines
Exclude test code (inline #[cfg(test)] modules and #[test] functions):
Exclude test directories by path:
Only report duplicates that are at least 10 lines long:
Lower the similarity threshold:
Sub-function Analysis
By default a code unit is a whole function, method or closure. Two functions that share a
copy-pasted match arm but differ elsewhere are not duplicates of each other, and nothing
is reported.
--sub-function (or -s) also treats each if-branch, match arm, loop body and closure
body as a unit in its own right:
)
)
=============================
)
)
)
)
)
)
Each member names the function it came from, and the line range shown is that parent
function's — not the branch's. Sub-function units are grouped separately from top-level
ones and counted under their own headings, so a function never shares a group with its own
branch. --min-sub-nodes (default 5) is the floor a branch must reach to be considered
at all; raise it when small branches produce noise.
Two limits are worth knowing before trusting the output, both in docs/OPEN_POINTS.md: sub-function findings restate a function-level one when two functions are already duplicates of each other, and the parent line range means a report cannot be read straight to the branch.
Configuration
Configuration can be provided in three ways (in order of precedence):
- CLI flags (highest priority)
dry4rust.tomlin the project rootCargo.tomlunder[package.metadata.dry4rust]
dry4rust.toml
= 15
= 5
= 0.85
= ["tests", "benches"]
= true
= 0
= 10
= 5.0
= 10.0
Cargo.toml
[]
= 15
= 0.85
= ["tests"]
Configuration Options
| Option | Default | Description |
|---|---|---|
min_nodes |
10 |
Minimum AST node count for a code unit to be analyzed. Increase to skip trivial functions. |
min_lines |
0 |
Minimum source line count for a code unit to be analyzed. 0 means disabled. |
similarity_threshold |
0.9 |
Minimum similarity score for near-duplicate detection. Must be within 0.0..=1.0. |
exclude |
[] |
Path patterns to exclude from scanning (substring match). |
exclude_tests |
false |
Exclude #[test] functions and #[cfg(test)] modules from analysis. |
sub_function |
false |
Also analyse if-branches, match arms, loop bodies and closure bodies. |
min_sub_nodes |
5 |
Minimum AST node count for a sub-function unit to be analyzed. |
baseline |
None |
Path to a recorded baseline of inherited duplication, relative to the analysed root. |
max_exact_duplicates |
None |
For check subcommand: maximum allowed exact duplicate groups. |
max_near_duplicates |
None |
For check subcommand: maximum allowed near-duplicate groups. |
max_exact_percent |
None |
For check subcommand: maximum allowed exact duplicate line percentage. Must be within 0.0..=100.0. |
max_near_percent |
None |
For check subcommand: maximum allowed near-duplicate line percentage. Must be within 0.0..=100.0. |
A value outside the range its field allows fails the run with a message naming the field, and exits 2. A config file that cannot be read or parsed is passed over, because a project with no configuration looks the same.
Ignoring Duplicates
Some duplicates are intentional (e.g., test helpers, trait implementations). You can ignore them by fingerprint:
# Add a fingerprint to the ignore list
# List ignored fingerprints
)
# Ignored groups are automatically filtered from reports and checks
# The ignored group will not appear
The ignore list is stored in .dry4rust-ignore.toml in the project root. When cleanup
prunes the last entry it removes the file rather than leaving an empty one behind: an empty
suppression list says exactly what no file says.
Entries whose fingerprint no longer matches anything go stale — after a refactor, or after
an upgrade that changed the fingerprint format. cleanup prunes them:
Adopting on a Codebase That Already Has Duplication
ignore is for duplication that is meant to be there. Duplication nobody has got to yet
is a different thing, and recording it as intentional — one fingerprint at a time, with a
reason invented to fill the field — writes a lie into a file that outlives whoever wrote
it.
Record it as a baseline instead. check then fails on what is added, not on what was
inherited:
# Record what is already there
)
# From now on, a zero ceiling is a gate rather than a wall
Commit dry4rust-baseline.json, or put baseline = "dry4rust-baseline.json" in
dry4rust.toml so every run picks it up without the flag. baseline --dry-run lists what
would be recorded without writing anything.
A baseline records a group's fingerprint and its member count. A third copy of an already-recorded duplicate makes the group larger than what was recorded, so it is reported — an exact group keeps its fingerprint when a copy joins it, and a baseline keyed on the fingerprint alone would inherit every future copy. Deleting a copy is admitted: progress is not something to fail on.
A baseline that cannot be read — missing, malformed, or written by a version with a different format — fails the run and says so, rather than being treated as empty. Every summary a baseline touched states how many groups it suppressed, so a stale baseline is visible rather than silent.
The reasoning is in ADR-BaselineIsInheritedNotForgiven.
CI Integration
Use the check subcommand in CI pipelines:
# GitHub Actions example
- name: Check for code duplication
run: cargo dry4rust check --max-exact 0 --max-exact-percent 5.0
On a codebase with existing duplication, add a baseline so the gate measures what the change introduced:
- name: Check for new code duplication
run: cargo dry4rust --baseline dry4rust-baseline.json check --max-exact 0
Exit codes:
- 0 — Check passed (within thresholds)
- 1 — Check failed (thresholds exceeded)
- 2 — Error (no source files, invalid path, etc.)
What Gets Analyzed
| Code Unit | Description |
|---|---|
| Functions | Top-level fn items |
| Methods | fn items inside impl blocks |
| Trait impls | fn items inside impl Trait for Type blocks |
| Closures | Closure expressions (above the min node threshold) |
The scanner automatically:
- Skips
target/directories - Skips hidden directories (starting with
.) - Respects exclude patterns
- Handles parse errors gracefully (skips unparseable files with a warning)
Documentation
| Where | What |
|---|---|
| docs/ARCHITECTURE.md | The pipeline, the components, and the data model |
| docs/FORMULA.md | Fingerprinting and the similarity score, precisely |
| docs/ADRs/ | The load-bearing decisions and why they were forced |
| docs/OPEN_POINTS.md | Where the model is known to be thin, including two silent false-negative sources |
| docs/ROADMAP.md | Direction, current baseline, and what is left |
| docs/IMPLEMENTED-FEATURES.md | What each version added, upstream's included |
| docs/dry4rust-dossier.md | Research on where this tool is going, and what already holds the ground |
Development
Requirements: Rust 1.91+ (edition 2024), matching the rust-version every
workspace crate declares.
The repository is a workspace: core/ is the published crate, validation/ holds the
end-to-end tests, xtask/ runs the stage 2 gates, and fixture/ is the corpus both
are pointed at.
Both gates are mandatory after any change under src/ or tests/, and both run
the same way on Windows, Linux and macOS:
Stage 1 is cargo built-ins only, so it works on a fresh checkout with none of the
tools below installed. Note that it lints test targets too
(--all-targets) — combined with the pedantic and nursery groups that core
and xtask enable, that holds the test files to the same bar as the sources.
Stage 2 is cargo xtask stage2, a real crate rather than a script, so the gate
argument lists and failure messages are covered by its own integration tests.
xtask is itself a workspace member and is gated like everything else — the
crate that runs the gates is not exempt from them.
Its second gate is dry4rust run on itself: this checkout's binary, built from
source rather than installed, over core/src with zero ceilings against
dry4rust-baseline.json, at a floor of 25 AST nodes. A tool that enforces a
check it does not pass is not worth installing. It needs nothing beyond cargo,
which is why it is not in the table below.
Everything the two stages need, none of which ships with cargo:
| Tool | Install | Needed by |
|---|---|---|
just |
cargo install just |
both stages |
cargo-llvm-cov |
cargo install cargo-llvm-cov |
stage 2 |
llvm-tools rustup component |
rustup component add llvm-tools |
stage 2 |
cargo-stern4rust |
cargo install cargo-stern4rust |
stage 2 |
cargo-crap4rust |
cargo install cargo-crap4rust |
stage 2 |
cargo-twin4rust |
cargo install cargo-twin4rust |
stage 2 |
cargo-iceberg4rust |
cargo install cargo-iceberg4rust |
stage 2 |
cargo-llvm-cov and llvm-tools are what the CRAP gate needs; without them it
fails with a bare exit code that says nothing about a missing install.
CI (.github/workflows/ci.yml) runs both stages on Ubuntu, Windows and macOS
for every pull request and every push to main.