cargo-crap 0.5.0

Change Risk Anti-Patterns (CRAP) metric for Rust projects
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
# cargo-crap

[![v0.5.0](https://img.shields.io/badge/v0.5.0-2563eb?style=for-the-badge)](https://github.com/minikin/cargo-crap/releases/tag/v0.5.0)
[![crates.io](https://img.shields.io/badge/crates.io-E57300?style=for-the-badge&logo=rust&logoColor=white)](https://crates.io/crates/cargo-crap)
[![docs.rs](https://img.shields.io/badge/docs.rs-000000?style=for-the-badge&logo=docsdotrs&logoColor=white)](https://docs.rs/cargo-crap/0.5.0/cargo_crap/)
[![CRAP](https://img.shields.io/endpoint?url=https%3A%2F%2Fraw.githubusercontent.com%2Fminikin%2Fcargo-crap%2Fbadges%2Fcrap-badge.json&style=for-the-badge)](https://github.com/minikin/cargo-crap/actions/workflows/ci.yml)

> [!TIP]
> For more context on the motivation behind this crate, read:
> [cargo-crap: Finding Untested Complexity in AI-Generated Rust Code](https://minikin.me/blog/cargo-crap) or watch [
Your AI Code Might Be CRAP! (Here's How To Fix It)](https://www.youtube.com/watch?v=XuMR1pgc6pc).

Compute the **CRAP** (Change Risk Anti-Patterns) metric for Rust projects.

CRAP combines cyclomatic complexity and test coverage into a single number
that is high when code is both hard to understand and poorly tested — i.e.
where bugs love to hide. The metric was introduced by Savoia & Evans in
2007 and was originally implemented for Java (Crap4j) and .NET (NDepend).
`cargo-crap` brings it to the Rust ecosystem.

```text
CRAP(m) = comp(m)² × (1 − cov(m)/100)³ + comp(m)
```

A few properties worth internalizing before you use the output:

- A trivial function (CC=1, 100% covered) scores exactly 1.0. That's the
  lower bound.
- At 100% coverage the quadratic term collapses and **CRAP equals CC**.
  When you see matching values in those two columns, that function is
  fully covered — tests are capping the damage, but the complexity itself
  remains. It's a good sign, not a bug.
- Above CC ≈ 30 no amount of coverage keeps you under the default
  threshold of 30. That's not a bug in the formula — it's the formula
  saying "this function is too big to certify as clean, regardless of
  tests."

## Install

**Via `cargo binstall`** (downloads the right pre-built binary automatically):

```bash
cargo binstall cargo-crap
```

**From source** (requires Rust stable ≥ 1.88):

```bash
cargo install cargo-crap
```

**From the AUR**:

```bash
paru -S cargo-crap
```

**Pre-built binary** (manual download):

```bash
# macOS (Apple Silicon)
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/minikin/cargo-crap/releases/latest/download/cargo-crap-aarch64-apple-darwin.tar.gz | tar xz -C ~/.cargo/bin

# macOS (Intel)
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/minikin/cargo-crap/releases/latest/download/cargo-crap-x86_64-apple-darwin.tar.gz | tar xz -C ~/.cargo/bin

# Linux (x86_64)
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/minikin/cargo-crap/releases/latest/download/cargo-crap-x86_64-unknown-linux-gnu.tar.gz | tar xz -C ~/.cargo/bin

# Linux (aarch64)
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/minikin/cargo-crap/releases/latest/download/cargo-crap-aarch64-unknown-linux-gnu.tar.gz | tar xz -C ~/.cargo/bin
```

Windows: download `cargo-crap-x86_64-pc-windows-msvc.zip` from the [latest release](https://github.com/minikin/cargo-crap/releases/latest) and extract `cargo-crap.exe` into a directory on your `PATH`.

## Quick start

```bash
# 0. cargo-llvm-cov is a separate tool; install it once.
cargo install cargo-llvm-cov

# 1. Generate an LCOV coverage report.
cargo llvm-cov --lcov --output-path lcov.info

# 2. Score every function.
cargo crap --lcov lcov.info

# 3. Gate CI on the threshold.
cargo crap --lcov lcov.info --fail-above

# 4. Whole-workspace analysis (monorepos).
cargo llvm-cov --workspace --lcov --output-path lcov.info
cargo crap --workspace --lcov lcov.info

# 5. Quick aggregate summary (no table).
cargo crap --workspace --lcov lcov.info --summary

# 6. Only selected workspace members (changed-file CI).
cargo crap -p backend_core -p backend_identity --lcov lcov.info
```

Example output:

```
┌───┬───────┬────┬───────────────────┬──────────┬───────────────┐
│   ┆  CRAP ┆ CC ┆ Coverage          ┆ Function ┆ Location      │
╞═══╪═══════╪════╪═══════════════════╪══════════╪═══════════════╡
│ ✗ ┆ 156.0 ┆ 12 ┆ ░░░░░░░░░░   0.0% ┆ crappy   ┆ src/lib.rs:24 │
├╌╌╌┼╌╌╌╌╌╌╌┼╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤
│ ▲ ┆   6.7 ┆  4 ┆ ████░░░░░░  44.4% ┆ moderate ┆ src/lib.rs:12 │
├╌╌╌┼╌╌╌╌╌╌╌┼╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤
│ ✓ ┆   1.0 ┆  1 ┆ ██████████ 100.0% ┆ trivial  ┆ src/lib.rs:8  │
└───┴───────┴────┴───────────────────┴──────────┴───────────────┘
✗ 1/3 function(s) exceed CRAP threshold 30.
```

## Flags

| Flag                                                             | Default       | Purpose                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| ---------------------------------------------------------------- | ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--lcov <FILE>`                                                  | —             | LCOV file from `cargo llvm-cov` or `cargo tarpaulin`. **Omitting it is not a coverage-free run**: every function is scored as if it had 0% coverage, so CRAP collapses to `CC² + CC` and the whole table reads red. That is useful for a first look at the complexity distribution — it is not a CRAP run.                                                                                                                                                                            |
| `--path <DIR>`                                                   | `.`           | Root to walk for `.rs` files (respects `.gitignore`).                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| `--threshold <N>`                                                | `30`          | Score above which a function is flagged.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| `--min <SCORE>`                                                  | —             | Hide entries below this score.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                        |
| `--top <N>`                                                      | —             | Show only the N worst offenders.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `--sort {crap,file}`                                             | `crap`        | Final ordering of entries. `crap` sorts by score descending (best for reading top-down). `file` sorts by `(file, function, line)` ascending — stable across score changes, so a committed JSON baseline produces minimal diffs. `--top` always selects the N highest-CRAP functions first, then `--sort` reorders them. Applies to every format.                                                                                                                                                                                                                                                                      |
| `--missing {pessimistic,optimistic,skip}`                        | `pessimistic` | How to score a function with no coverage data.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                        |
| `--exclude <GLOB>`                                               | —             | Skip files matching this pattern (repeatable). `**` crosses directories. Appends to the default exclusions.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| `--no-default-excludes`                                          | off           | Disable the built-in default exclusions (`tests/**`, `benches/**`, `examples/**`, matched relative to each analyzed root). By default these standard Cargo target directories are skipped — integration tests exist to cover production code, and benches/examples are not executed during a coverage run, so they only add 0%-coverage noise.                                                                                                                                                                                                                                                                        |
| `--allow <GLOB>`                                                 | —             | Suppress matching functions (repeatable). An entry containing `/` or `**` is a path glob and matches the file the function is in (e.g. `src/generated/**`); otherwise it matches the function name and `*` crosses `::` (e.g. `Foo::*`). Path globs analyze the file but hide its functions — distinct from `--exclude`, which skips files at walk time.                                                                                                                                                                                                                                                              |
| `--duplicates`                                                    | off           | Also report candidate duplicate functions — functions whose normalized structure is close enough to be worth review. A second analysis over the same AST; off by default because it costs time on a large tree. Needs no `--lcov`.                                                                                                                                                                                                          |
| `--dup-threshold <SCORE>`                                        | `0.82`        | Similarity at or above which a pair is reported, in `0.0..=1.0`. Named apart from `--threshold`, which is the CRAP score threshold.                                                                                                                                                                                                                                                                                                          |
| `--format {human,json,github,markdown,pr-comment,sarif,shields}` | `human`       | Output format. `json` emits a versioned envelope (see [JSON output schema](#json-output-schema) below). `github` emits `::warning` annotations. `markdown` emits a GFM table (exhaustive). `pr-comment` is the opinionated PR-bot variant: hides unchanged rows, caps each section, collapses non-critical info into `<details>` blocks. `sarif` emits SARIF 2.1.0 JSON for upload to GitHub Code Scanning, VS Code, and other static-analysis tooling (see [SARIF output](#sarif-output) below). `shields` emits Shields.io endpoint-badge JSON for a README badge (see [Shields.io badge](#shieldsio-badge) below). |
| `--summary`                                                      | off           | Print only aggregate stats (total, crappy count, worst offender) — no per-function table. In `--workspace` mode this becomes the per-crate summary plus the aggregate line. `json` and `github` remain machine-readable and are unaffected.                                                                                                                                                                                                                                                                                                                                                                             |
| `--workspace`                                                    | off           | Analyze all Cargo workspace members (discovered via `cargo metadata`). Ignores `--path`. Adds a *Per-crate summary* table to human and markdown output, and a `crate` field to JSON entries.                                                                                                                                                                                                                                                                                                                                                                                                                          |
| `-p, --package <NAME>`                                           | —             | Analyze only the named workspace member(s); repeatable, cargo-style (`-p core -p api`). One invocation, one LCOV parse, one report, one gate decision over exactly the selected members — ideal for changed-file CI that already knows which packages a PR touches. Unknown names fail before analysis (exit 2). Ignores `--path`; conflicts with `--workspace`. A selected member's walk never descends into another member's nested root.                                                                                                                                                                            |
| `--fail-above`                                                   | off           | Exit 1 if any function exceeds `--threshold`.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| `--baseline <FILE>`                                              | —             | JSON from a previous `--format json` run. Enables delta mode (shows Δ column). Functions that moved between files (same name, body unchanged) are detected and reported as `Moved` rather than as separate New + Removed entries; renderers show `← <previous_file>` next to the new location. Baseline entries that the current run's `--exclude`/`--allow`/default exclusions would drop are filtered out before comparison, so changing the exclusion set does not flood the report with phantom `removed` entries.                                                                                                |
| `--fail-regression`                                              | off           | Exit 1 if any function's score increased since `--baseline`. `Moved` (pure relocation, no score change) is not a regression.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                          |
| `--show-unchanged`                                               | off           | In `--baseline` mode, also list `Unchanged` rows in the human and markdown tables. By default only changed functions (`Regressed`/`Improved`/`New`/`Moved`) are shown; when everything is unchanged the table is replaced with `No changes since baseline.`. The summary line always counts every entry. Requires `--baseline`. Does not affect `json` (always exhaustive) or `pr-comment` (keeps its own row policy).                                                                                                                                                                                                |
| `--epsilon <VALUE>`                                              | `0.01`        | Tolerance for the regression detector. Score deltas with absolute value at or below this count as `Unchanged`. Set to `0.0` to flag every increase, or higher to tolerate noisy coverage. Must be non-negative.                                                                                                                                                                                                                                                                                                                                                                                                       |
| `--jobs <N>`                                                     | host CPUs     | Cap parallel source-file analysis at `N` threads. Useful in memory-constrained CI/Docker environments. Must be a positive integer.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `--output <FILE>`                                                | —             | Write output to FILE instead of stdout (useful for saving JSON baselines).                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `--repo-url <URL>`                                               | —             | Base URL of the source-hosting repo (e.g. `https://github.com/owner/repo`). With `--commit-ref`, makes Function and Location cells in `markdown` / `pr-comment` output clickable links to the source. Defaults from `GITHUB_SERVER_URL` + `GITHUB_REPOSITORY` inside GitHub Actions.                                                                                                                                                                                                                                                                                                              |
| `--commit-ref <REF>`                                             | —             | Commit SHA or branch to deep-link into. Defaults from `GITHUB_SHA` when set. No effect unless `--repo-url` is also set.                                                                                                                                                                                                                                                                                                                                                                                                                                                                          |

Colour in the `human` and `--summary` formats is automatic: it is enabled only when writing to a terminal, and never into an `--output` file or a pipe. Set `NO_COLOR=1` to disable colour unconditionally, or `FORCE_COLOR=1` to force it on (e.g. for `| less -R`); `NO_COLOR` wins when both are set.

### JSON output schema

`--format json` produces a versioned envelope with a `$schema` URL pointing
at the published JSON Schema. Consumers can validate output offline or
generate types directly from the schema.

| Variant                    | Schema                                                                                                       |
| -------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Absolute (no `--baseline`) | [`schemas/report-v1.json`](https://raw.githubusercontent.com/minikin/cargo-crap/main/schemas/report-v1.json) |
| Delta (with `--baseline`)  | [`schemas/delta-v2.json`](https://raw.githubusercontent.com/minikin/cargo-crap/main/schemas/delta-v2.json)   |

```jsonc
// cargo crap --format json
{
  "$schema": "https://raw.githubusercontent.com/minikin/cargo-crap/main/schemas/report-v1.json",
  "version": "0.5.0",     // the cargo-crap version that produced the report
  "entries": [
    {
      "file": "src/lib.rs",
      "function": "do_thing",
      "line": 12,
      "cyclomatic": 4.0,
      "coverage": 75.0,        // null when no coverage data was found
      "crap": 5.5625,
      "crate": "my-crate",     // present only with --workspace
      "uncovered": [           // uncovered line ranges; omitted when empty
        { "start": 15, "end": 16 }
      ]
    }
  ]
}

// cargo crap --format json --baseline baseline.json
{
  "$schema": "https://raw.githubusercontent.com/minikin/cargo-crap/main/schemas/delta-v2.json",
  "version": "0.5.0",     // the cargo-crap version that produced the report
  "entries": [ /* DeltaEntry — current + baseline_crap + delta + status (+ optional previous_file when moved) */ ],
  "removed": [ /* RemovedEntry — function, file, baseline_crap */ ]
}
```

`--baseline` only reads files in this envelope shape; bare-array baselines
from older runs must be regenerated.

Both envelopes also carry an optional `diagnostics` object whenever `--lcov`
was supplied: analyzed/LCOV/matched file counts plus bounded example lists
of files present only on one side. When the analyzed tree and the coverage
run describe different scopes (the classic cause of a delta full of
unrelated 0%-coverage entries), a warning with the same numbers is printed
to stderr before the report, and CI wrappers can gate on the JSON counts.

On `--duplicates` runs both envelopes grow a `duplicates` array — one object
per candidate pair, in the same order as the human section:

```jsonc
"duplicates": [
  {
    "first_file": "src/report/markdown.rs",
    "first_function": "write_markdown_absolute_heading",
    "first_start_line": 18,
    "first_end_line": 33,
    "second_file": "src/report/pr_comment.rs",
    "second_function": "write_pr_comment_delta_headline",
    "second_start_line": 275,
    "second_end_line": 286,
    "score": 0.92          // Jaccard similarity, in [0, 1]
  }
]
```

The key is absent — not empty — when detection was not requested, so "not
asked" stays distinguishable from "asked, found nothing".

### SARIF output

`--format sarif` emits a [SARIF 2.1.0](https://docs.oasis-open.org/sarif/sarif/v2.1.0/sarif-v2.1.0.html)
JSON document — the format consumed by GitHub Code Scanning, VS Code,
rust-analyzer, and most static-analysis tooling.

- Each crappy function (entry above `--threshold`) becomes one
  `result` with `level: "warning"` and a physical location pointing at
  the function's start line.
- Functions below the threshold are not included.
- An empty result set still produces a valid SARIF document with the
  full `runs[0].tool.driver` envelope.
- `--baseline` is rejected with `--format sarif`; SARIF describes
  findings, not deltas. Use `--format json` for delta output.

### Shields.io badge

`--format shields` emits a single JSON object following the
[Shields.io endpoint schema](https://shields.io/badges/endpoint-badge).
Serve the file at a stable URL (GitHub Pages, raw blob) and embed it as a
normal badge image:

```markdown
![CRAP](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/owner/repo/main/crap-badge.json)
```

The label embeds the *effective* threshold — `CRAP > 30` by default, or
whatever `--threshold` was given (`CRAP > 15` in this repo's own run) — so the badge reads
as a complete statement. The message is `passing` (brightgreen) when no
function exceeds `--threshold`, `N crappy` in yellow for 1–5 offenders,
and red for 6 or more. `--baseline` is silently ignored — the badge
always reflects absolute current scores. See [Badge generation](#badge-generation) for a
CI recipe.

## Finding duplicates

`--duplicates` answers a different question from the CRAP table: *is this
already implemented somewhere else in this codebase?* Copy-pasted logic is
invisible to CRAP — two identical 40-line functions each score exactly what
one of them would.

```bash
cargo crap --path src --duplicates
```

```
2 duplicate candidates:

DUPLICATE score=1.00
  src/report/types.rs:167-178  write_abs_gfm_header
  src/report/types.rs:182-193  write_delta_gfm_header
DUPLICATE score=0.92
  src/report/markdown.rs:18-33  write_markdown_absolute_heading
  src/report/pr_comment.rs:275-286  write_pr_comment_delta_headline
```

How it works: every function is parsed with `syn`, normalized into a
structural tree, and fingerprinted — one fingerprint per subtree. Two
functions are compared by Jaccard similarity over their fingerprint sets
(`|A ∩ B| / |A ∪ B|`), and pairs scoring at or above the threshold are
reported.

Normalization **drops** what a function is called and what values it
mentions — function and method names, parameter and local names, field and
path names, literal values. It **keeps** structure: control flow, statement
order, receiver shape, type structure, and operators. `x + y` and `x * y`
are different shapes; `xs`/`items` and `1`/`0` are not:

```rust
fn alpha(xs: &[i32]) -> Vec<i32> {          fn beta(items: &[i32]) -> Vec<i32> {
    let mut ys = Vec::new();                    let mut kept = Vec::new();
    for x in xs {                               for item in items {
        if x % 2 == 1 {                             if item % 2 == 0 {
            ys.push(x + 1);                             kept.push(item + 1);
        }                                           }
    }                                           }
    ys                                          kept
}                                           }
```

These score **1.00**.

The tool does not decide whether duplication should be removed — two
functions scoring 1.00 may be a bug waiting to happen or two unrelated trait
impls that share a shape. It names the candidate; you judge it.

Three things worth knowing:

- **Test code is not compared.** `#[test]` functions and `#[cfg(test)]`
  modules are skipped, the same way the complexity pass skips them. Test
  bodies are repetitive by construction and would bury everything else.
- **Trivial functions are not compared.** Functions below
  `duplicates.min-nodes` (default 20 normalized nodes) are skipped, because
  every pair of one-line accessors is a genuine structural match and reports
  as one. Set it to `0` to compare everything.
- **Macro bodies are opaque.** `syn` does not parse a macro's token stream,
  so every `println!`/`vec!` is one node. Two different macro invocations
  look identical to this analysis.

`--format json` carries the pairs in a `duplicates` array inside the usual
envelope; the key is absent entirely when detection was not requested, so
"not asked" is distinguishable from "asked, found nothing".

## Configuration file

Most flags can be set persistently in `.cargo-crap.toml` at the project root
(or any parent directory — the tool walks up until it finds one). CLI flags
always take precedence. The per-run selectors have no config key —
`--path`, `--format`, `--output`, `--summary`, `--workspace`, `-p`/`--package`,
`--baseline`, `--no-default-excludes`, `--repo-url` and `--commit-ref` are
flags only. `uncovered-hints` is the mirror case: config only, no flag.

```toml
# .cargo-crap.toml
threshold = 30.0
fail-above = true
missing = "pessimistic"   # pessimistic | optimistic | skip
# `exclude` appends to the default exclusions.
exclude = ["src/generated/**"]
# `default-excludes` replaces the built-in default list
# (tests/**, benches/**, examples/**). Set to [] to disable it;
# list a subset to re-include some directories; extend it freely.
default-excludes = ["benches/**", "examples/**", "fuzz/**"]
# `allow` accepts both function-name globs and path globs (any entry
# containing `/` or `**` is a path glob).
allow   = ["generated::*", "src/generated/**"]
epsilon = 0.01            # regression-detector tolerance
jobs    = 4               # cap parallel analysis at 4 threads
sort    = "file"          # entry ordering: crap (default) | file
min     = 5.0             # hide entries scoring below this
top     = 20              # keep only the N worst offenders
fail-regression = true    # exit 1 when a score rises against --baseline
show_unchanged = false    # list Unchanged rows in --baseline mode
# Append an Uncovered column (per-function uncovered line ranges, e.g.
# `142–158, 171, 180–184 +2 more`) to the human, markdown, and pr-comment
# outputs.
# Config-only — there is deliberately no CLI flag. JSON always carries
# the full ranges regardless of this key.
uncovered-hints = false
[duplicates]
enabled   = false   # same as passing --duplicates
threshold = 0.82    # similarity at or above which a pair is reported
min-nodes = 20      # skip functions smaller than this; 0 compares everything
```

All keys are optional. Unknown keys are rejected to catch typos.

Every key above accepts both the kebab-case house spelling and its
snake_case alias (`show-unchanged` / `show_unchanged`). Where a key and
its flag are not spelled the same, the mapping is:

| Flag                  | Config key             |
| --------------------- | ---------------------- |
| `--threshold`         | `threshold`            |
| `--fail-above`        | `fail-above`           |
| `--missing`           | `missing`              |
| `--exclude`           | `exclude` (appends)    |
| `--allow`             | `allow`                |
| `--min`               | `min`                  |
| `--top`               | `top`                  |
| `--sort`              | `sort`                 |
| `--epsilon`           | `epsilon`              |
| `--jobs`              | `jobs`                 |
| `--fail-regression`   | `fail-regression`      |
| `--show-unchanged`    | `show-unchanged`       |
| `--duplicates`        | `duplicates.enabled`   |
| `--dup-threshold`     | `duplicates.threshold` |
| *(no flag)*           | `duplicates.min-nodes` |
| *(no flag)*           | `uncovered-hints`      |
| *(no key)*            | `--path`, `--format`, `--output`, `--summary`, `--workspace`, `-p`/`--package`, `--baseline`, `--no-default-excludes`, `--repo-url`, `--commit-ref` |

`default-excludes` is the one key with no direct flag equivalent: it
*replaces* the built-in default list, where `--no-default-excludes`
empties it.

## Design

The tool has seven orthogonal modules. Each is testable in isolation; the
join between them has its own integration test.

```
  cargo llvm-cov                  syn
  (LCOV file)                 (Rust AST)
        │                         │
        ▼                         ▼
  ┌───────────┐            ┌────────────┐
  │ coverage  │            │ complexity │
  │  module   │            │   module   │
  └─────┬─────┘            └──────┬─────┘
        │                         │
        └──────────┬──────────────┘
                   ▼
             ┌──────────┐
             │  merge   │  ← path normalization lives here
             └─────┬────┘
                   ▼
             ┌──────────┐     ┌───────┐
             │  score   │ ──▶ │ delta │  ← baseline comparison (optional)
             └─────┬────┘     └───────┘
                   ▼
             ┌──────────┐
             │  report  │  ← human / JSON / GitHub / Markdown /
             └──────────┘     pr-comment / SARIF / Shields

                syn
            (Rust AST)
                 │
                 ▼
          ┌────────────┐
          │ duplicates │  ← second pass, only on --duplicates
          └────────────┘
```

### The path-matching problem

This is where silent failures happen. Complexity analysis produces
absolute paths (whatever was passed to the walker). LCOV files contain
whatever the coverage tool decided to write:

1. Absolute paths — `/home/alice/project/src/foo.rs`
2. Workspace-relative paths — `src/foo.rs`
3. Crate-relative paths in a workspace — `crates/core/src/foo.rs`
4. Paths with `./` or `../` components

A naïve `HashMap<PathBuf, _>` lookup silently returns `None` for 100% of
files when the two don't agree, and every function reports as 0% covered.
`cargo-crap` handles this with a two-level index:

- Absolute coverage paths → direct canonical-path hash lookup.
- Relative coverage paths → suffix match on path components (not bytes —
  `/foo/bar.rs` must not match `oofoo/bar.rs`).

Ambiguous inputs resolve deterministically (spec 26): when several
relative keys suffix-match one file (`src/lib.rs` vs
`vendor/dep/src/lib.rs`), the longest — most specific — suffix wins, and
different spellings of the same file (symlinked roots, `lcov -a`-merged
legs, `./`-prefixed variants) merge their line data instead of racing on
map order.

Relative paths are **never** canonicalized against the process's CWD, which
would otherwise silently bind them to whatever file happened to exist
under the tool's working directory. The regression test
`relative_coverage_paths_are_not_resolved_against_cwd` in `src/merge.rs`
pins this.

### What gets a score, and what counts as complexity

Two things surprise people the first time their numbers do not match
another tool.

**Test code is never scored.** A function carrying `#[test]` and every
item inside a `#[cfg(test)] mod` is skipped by the complexity pass, so it
never reaches the table. (Only that exact spelling — `#[cfg(not(test))]`
and `#[cfg(any(test, …))]` are left alone.) On top of that, `tests/**`,
`benches/**` and `examples/**` are excluded at walk time; pass
`--no-default-excludes` to get them back. So "my test helpers are
missing" is the tool working as intended, not a path-matching failure.

**Cyclomatic complexity starts at 1 and adds one for each of:** `if`
(including every `else if`), `for`, `while`, `loop`, **every** `match`
arm, each `&&` and `||`, and each `?`. Two consequences worth knowing:

- A three-arm `match` adds 3, giving CC 4 — textbook McCabe counts N−1
  branch points and would say 3. Scores are internally consistent and
  comparable across runs of this tool; they are not comparable to
  another tool's numbers.
- `?` counts as a decision point, so an idiomatic `Result` chain scores
  higher than its branching suggests.

Closures and items nested inside a function body (a local `fn`, `impl`,
`mod`, …) are *not* folded into the enclosing function — each is its own
scope, and a closure's branches belong to the closure.

### The `--missing` policy

Some functions have complexity data but no coverage data — the coverage
tool didn't instrument them, or they were excluded via `#[cfg(test)]`, or
the coverage run was scoped to a subset of the workspace. Three policies:

- **pessimistic** (default): treat as 0% covered. Surfaces unmapped code as
  a red flag. Correct for CI gates.
- **optimistic**: treat as 100% covered. Useful during local development
  when you're iterating on a specific module.
- **skip**: drop the row entirely.

## Integrating with CI

### Exit codes

The exit code distinguishes a finished CRAP verdict from a broken run,
so wrappers never need file-size or log-parsing heuristics:

| Code | Meaning                                                                                        |
|------|------------------------------------------------------------------------------------------------|
| 0    | Analysis completed; no requested gate tripped.                                                 |
| 1    | Analysis completed and the report was fully written; `--fail-above` / `--fail-regression` tripped. |
| 2    | The run did not complete: usage, input, analysis, or output error.                             |

### Absolute threshold gate

```yaml
- run: cargo llvm-cov --lcov --output-path lcov.info
- run: cargo crap --lcov lcov.info --fail-above --threshold 30
```

### Regression gate (recommended for teams)

Save a baseline on `main`, then fail on any PR that makes a score go up.
This works regardless of the absolute threshold and catches regressions as
they are introduced, not weeks later.

```yaml
# On main branch — upload baseline as a CI artifact
- run: cargo llvm-cov --lcov --output-path lcov.info
- run: cargo crap --lcov lcov.info --format json --output baseline.json
- uses: actions/upload-artifact@v4
  with:
    name: crap-baseline
    path: baseline.json

# On pull requests — download baseline and compare
# NOTE: actions/download-artifact@v4 extracts to a subfolder named after the
# artifact by default — pin `path:` so the file lands somewhere predictable.
- uses: actions/download-artifact@v4
  with:
    name: crap-baseline
    path: baseline
- run: cargo llvm-cov --lcov --output-path lcov.info
- run: cargo crap --lcov lcov.info --baseline baseline/baseline.json --fail-regression
```

Two flags make this workflow nicer:

- **Commit the baseline to git instead of uploading it as an artifact.**
  Add `--sort file` when generating it so entries are ordered by
  `(file, function, line)` rather than by score. The order then stays put
  across runs, so a code change touches only the affected entry's fields —
  minimal, reviewable diffs:

  ```bash
  cargo crap --lcov lcov.info --format json --sort file --output crap_baseline.json
  ```

- **The comparison output is changed-only by default.** In `--baseline`
  mode the human and markdown tables list just the functions that
  `Regressed` / `Improved` / are `New` / `Moved`; when nothing changed you
  get `No changes since baseline.`. The summary line still counts every
  function. Pass `--show-unchanged` for the full exhaustive table. (JSON
  stays exhaustive either way, so machine consumers are unaffected.)

### GitHub Code Scanning (SARIF)

Upload `--format sarif` output to surface crappy functions in the
repository's **Security → Code scanning** tab. The job needs
`security-events: write`.

```yaml
self_score:
  permissions:
    security-events: write
  steps:
    - run: cargo llvm-cov --lcov --output-path lcov.info
    - run: cargo crap --lcov lcov.info --format sarif --output crap.sarif
    - uses: github/codeql-action/upload-sarif@v3
      with:
        sarif_file: crap.sarif
        category: cargo-crap
```

### Badge generation

Regenerate the badge JSON on every push to the default branch and commit
it back so the README embed stays current:

```yaml
- name: Generate CRAP badge
  run: |
    cargo crap \
      --lcov lcov.info \
      --workspace \
      --threshold 30 \
      --format shields \
      --output crap-badge.json

- name: Commit badge
  run: |
    git config user.name "github-actions[bot]"
    git config user.email "github-actions[bot]@users.noreply.github.com"
    git add crap-badge.json
    git diff --cached --quiet || git commit -m "chore: update CRAP badge"
    git push
```

This repository does it differently, and the badge at the top of this file
points at that: `ci.yml` uploads `crap-badge.json` as an artifact and a
separate `badge` job pushes it to a dedicated `badges` branch, so the
default branch never carries a generated file. See
[`.github/workflows/ci.yml`](.github/workflows/ci.yml) if you want that
shape instead.

### PR comment bot

`--format pr-comment` produces a sticky comment that surfaces regressions
and new functions in the primary table and tucks improvements / removed
functions / above-threshold hot-spots into collapsed `<details>` blocks.
A hidden marker (`<!-- cargo-crap-report -->`) lets the script update an
existing comment instead of posting duplicates. The job needs
`pull-requests: write`.

```yaml
self_score:
  permissions:
    pull-requests: write
  steps:
    # ...generate lcov.info and download the baseline as above...

    - name: Generate PR comment
      if: github.event_name == 'pull_request'
      run: |
        cargo crap \
          --lcov lcov.info \
          --baseline baseline.json \
          --format pr-comment \
          --output crap-comment.md

    - name: Post or update PR comment
      if: github.event_name == 'pull_request'
      uses: actions/github-script@v7
      with:
        script: |
          const fs = require('fs');
          const body = fs.readFileSync('crap-comment.md', 'utf8');
          const marker = '<!-- cargo-crap-report -->';
          const { data: comments } = await github.rest.issues.listComments({
            owner: context.repo.owner,
            repo: context.repo.repo,
            issue_number: context.issue.number,
          });
          const existing = comments.find(c => c.body.startsWith(marker));
          const args = {
            owner: context.repo.owner,
            repo: context.repo.repo,
            body,
          };
          if (existing) {
            await github.rest.issues.updateComment({ ...args, comment_id: existing.id });
          } else {
            await github.rest.issues.createComment({ ...args, issue_number: context.issue.number });
          }
```

## Troubleshooting

**Every function shows `—` or 0% coverage.** Either no `--lcov` was
passed (see the flag table — that run scores everything as uncovered by
design), or the LCOV file and the analyzed tree describe different
scopes. `cargo-crap` detects the second case and prints analyzed / LCOV /
matched file counts to stderr before the report, with examples of files
present on only one side; the same numbers are in the `diagnostics`
object of both JSON envelopes, so CI can gate on them. The usual cause is
a coverage run scoped to one crate and an analysis scoped to the
workspace, or vice versa.

**My test helpers are not listed.** They are skipped on purpose — see
[What gets a score](#what-gets-a-score-and-what-counts-as-complexity).

**CC is higher than another tool reports.** Every `match` arm and every
`?` counts; the same section explains why.

**`--baseline` reports functions as `removed` that clearly still exist.**
Baseline entries are filtered through the current run's exclusions before
comparison, so this is usually a scope change instead: a `-p` subset run
against a whole-workspace baseline, or an `--exclude` added since. Match
the scope, or regenerate the baseline once.

## Prior art and references

- [Savoia, A. & Evans, B. (2007). *The CRAP Metric.*](https://www.artima.com/weblogs/viewpost.jsp?thread=210575)
- [Crap4j](http://www.crap4j.org/) — the original Java implementation.
- [dry4go](https://github.com/unclebob/dry4go) — the Go duplicate detector whose
  normalize/fingerprint/Jaccard approach `--duplicates` follows.

## License

This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.