fdu 0.4.0

Fastest du replacement, with .gitignore-aware sizes and code and document counts, for the command line, Python, and Rust
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
---
name: fdu
description: >-
  Inspect disk usage and codebases: directory sizes, file counts, recency, file types,
  lines of code by language, and words by document type. Use to find large or stale
  build directories, size up the source and prose in a codebase, or collect structured
  filesystem roll-ups for scripts and agents.
---
<!-- generated by fdu; re-run fdu --install-skill to update -->

# fdu Directory Roll-Ups

This is the complete fdu usage contract; it needs no setup chat or prior session
context.

Use `fdu` to summarize a directory tree without modifying files in that tree.
It reports size, file count, recency, and file kinds for every directory at once, and on
request lines of code by language and words by document type.
It walks the tree on several threads through each platform’s native directory interface,
and on a generated million-entry tree (875,000 files) it finished ahead of `du` and the
seven other disk-usage tools measured on Linux and macOS. Every report requires an
explicit `PATH`, `.` for the current directory; bare `fdu` prints help instead of
scanning. `fdu --docs` prints common commands, cache behavior, and the full usage
contract without a PATH and without scanning.

## Run fdu

Start with the report that answers the question:

```bash
fdu .                          # directory-size tree; metadata only
fdu . --view=code,documents    # code lines by language, words by document type
fdu . --format=json --view=tree --depth=2 --limit=20
```

The tree reads only metadata.
`--view=code,documents` also reads file contents, both analyses in one scan, and caches
the results, so a repeated run reads only changed files.
The third is a shallower tree for a script: `--format=json` (or `jsonl` or `yaml`) gives
exact values under a versioned `schema`, so use it whenever the output is parsed.

Prefer an `fdu` already on `PATH`. When there is none, run the latest release through
`uvx`, which needs no install step:

```bash
if command -v fdu >/dev/null 2>&1; then
  fdu . --format=json --view=tree --depth=2 --limit=20
else
  uvx --no-build fdu@latest . --format=json --view=tree --depth=2 --limit=20
fi
```

`--no-build` makes the fallback use a published wheel instead of compiling Rust.
To keep fdu installed, run `uv tool install --no-build fdu` (later
`uv tool upgrade --no-build fdu`) or `cargo install --locked fdu`.

Generated by fdu `__FDU_VERSION__`; if `fdu --version` differs, re-run
`fdu --install-skill` so this file describes the installed command.

More reports, each from one scan:

```bash
fdu . --view=code                          # code overview with population and coverage
fdu . --view=documents                     # words and pages by document format
fdu . --view=code --ignored=exclude        # source overview without ignored content
fdu . --ignored=exclude                    # skip ignored trees and their contents
fdu . --view=summary                       # one total with its ignored share
fdu . --view=languages                     # detected language sizes; metadata only
fdu . --view=families,types,extensions     # three file-kind breakdowns
fdu . --view=largest --limit=10            # ten largest files
fdu . --view=recent --limit=10             # ten most recently modified files
```

`--view` chooses what is printed, and `--analyze` adds what may be read beyond it.
`code` and `documents` are the views that read file contents: they show nothing without
analysis, so naming one runs its analyzer.
Every other view reads only metadata unless `--analyze` names an analyzer, as in these
two controls:

```bash
fdu . --analyze=code --view=languages      # code lines in the language rows
fdu . --analyze=lines --view=languages     # physical lines and raw words by language
```

Without code analysis, language percentages are byte shares; with it, the rows add code,
comment, and blank-line metrics and use code-line shares.
Use `--size apparent` when logical file lengths are wanted instead of allocated bytes.

## Read the Result and Its Notes

Stdout holds only the result: rows, column headings, and, when several views are shown,
an all-caps header per view, such as `LANGUAGES (3 of 15)` when a row limit hid some.
A grouped row (`families`, `types`, `languages`, `documents`) ends its first line at its
file count; its other measures, such as lines, words, and files not analyzed, are
indented lines under that count, not rows of their own.
Every explanation follows on stderr, one prefixed line each, in this order: `note:` for
what the result covers and how to read it, `warn:` for an operation that failed or an
answer nothing verified, `tip:` for a runnable change, and, after text reports, `perf:`
for the work done and the elapsed time.
Notes are consolidated: one says what totals include, one lists every display limit that
hid rows, and one tip lifts them all:

```text
note: totals include gitignored sizes and descendants
note: display limits: below 1% of root, depth 5
tip: show more: --min-share=0% --depth=all
```

Add the tip’s flags to the same command to see what was hidden, or use `--full` to lift
every display bound.
Display limits never change totals.
Notes also name a percentage denominator other than bytes, with the section it applies
to when several views are shown
(`note: percentages are shares of code lines (CODE), document words (DOCUMENTS)`), and
the code view’s coverage (`note: 15 languages analyzed`,
`note: not analyzed: 2 unsupported`). Machine formats send the same lines to stderr, so
stdout stays parseable; check their structured fields, not the notes, before trusting a
result. Use `--quiet` (`-q`) to hide notes, tips, performance lines, and progress.
Result stdout, warnings, errors, and exit status are unchanged; structured facts are
retained.

The default is `list` in `tree` format, in allocated bytes, largest first, to depth 5.
It shows directory subtrees and file leaves contributing at least 1% of the selected
root. Breadth and total rows are unbounded unless requested; `--depth`, `--min-share`,
`--breadth`, and `--limit` compose independently.
Hidden and ignored entries are included.
`.gitignore` is read to label ignored shares, not to exclude matching entries.
A parenthetical amount such as `(73 MiB gitignored)` is included in the row total.
Directories get a `/` suffix, except `.` and `..`; file names and structured paths do
not change.

Color is on only when the output is a terminal or `--color=always` is given, so piped
tree bars are plain: `█` for usage and `░` for unused width.
Colored bars use green `█` for non-gitignored usage, green `▓` for gitignored usage, and
green `▒` when classification is unknown; dim green `░` fills unused width.
Zero sizes and percentages below 1% are gray; sizes at least 1 GiB are bold.
Cyan names are bright and bold.
Gitignored directories use regular, nonbold cyan; containing ignored files is not
enough. The directory `/` suffix is gray.
`--bar-size=20` widens tree bars; zero or negative values hide the bar column.
The default is 10 characters; the maximum is 4,096. This does not change structured
output or measurements.

## Complete Inventories and File Search

```bash
fdu . --view tree --full --format json             # full recursive hierarchy
fdu . --kind dir --full --sort name --format json  # directory rows with recursive usage
fdu . --view files --kind file --full --format json # individual regular files
fdu . --kind file --include '*.rs' --full --format paths
fdu . --kind dir --include node_modules --full --long
```

`--full` means `--depth=all --breadth=all --limit=all --min-share=0%`. Explicit bounds
override it, regardless of order.
`--view full` chooses multiple views; `--full` expands whichever views were chosen.
Neither changes scan scope or analysis.

Use these commands for find/fd-style filename searches and complete usage inventories.
fdu includes hidden and ignored content; fd needs `--unrestricted` for that population.
fdu selection is glob-based.
It does not implement find expressions or `-exec`, and its ignore sources differ from
fd’s. Paths output escapes control characters; use structured output for arbitrary
native filenames. Directory rows include descendants and overlap; add
`--view list,summary` for a path-union total rather than summing rows.

A tree’s `remainder` contains recursive `files`, `bytes`, `allocated`, and applicable
`reasons` outside its displayed root-level rows; `null` means nothing is hidden there.
A displayed directory already represents its whole subtree, including descendants whose
rows were bounded away.
Per-boundary `entries` counts hidden roots, while `files` counts regular files
recursively. Unknown amounts are null.
The text equivalent is one root-level line with the same bar, percentage, and size
columns as the tree rows, followed by `… and N more files`. These columns represent the
remaining share of the selected root.
Unknown size or count stays unknown; an unknown hidden size has no numeric share.
Fully expanded, fully observed trees have no remainder line and no display-limit notes.
Check completeness separately: unreadable directories and scan-depth restrictions still
apply.

## Compose the Request From Six Axes

The six axes separate discovery, measurement, selection, and presentation.
Unsupported combinations fail before scanning.
There are no subcommands: the grammar is always “report on a path”.

| Axis | Question | Options |
| --- | --- | --- |
| Scope | What is scanned and cached? | `PATH`, `--scan-depth N`, `--one-filesystem`, `--gitignore-budget SIZE\|all`, `--gitignore-line-limit SIZE\|all`, `--no-gitignore`, `--ignored=include\|exclude\|only` |
| Content | Which file bodies are read beyond what the views imply? | `--analyze none\|lines\|code\|words\|all` |
| Selection | Which entries does this query consider? | `--include`, `--exclude`, `--min-size`, `--modified-since`, `--modified-before`, `--kind`, `--depth`, `--min-share`, `--breadth`, `-n/--limit`, `--full`, `--sort`, `--reverse`, `--size` |
| View | Which roll-up is reported? | `--view list,summary,tree,families,types,extensions,languages,code,documents,largest,recent,files`, or `--view full` |
| Format | How is it serialized? | `--format text\|tree\|paths\|long\|json\|jsonl\|yaml`, `--color`, `--progress` |
| Mode | How is work performed? | `--cache auto\|on\|off`, `--stale-ok`, `--cache-dir DIR`, `--watch`, `--workers N` |

Scope determines what is scanned and cached.
In a one-shot report, the ignored population also determines which subtrees and file
bodies may be skipped.
A retained index can answer narrower queries when it holds the required facts.

Work has three layers.
A single unfiltered `--view summary PATH` is the one exact composition that retains only
aggregate tallies and no index, except under `--cache on` and `--stale-ok`, whose
contracts are about the snapshot itself.
Otherwise a snapshot cannot save the walk that request is already doing, so it neither
reads nor writes one.
Reading `.gitignore`, the summary keeps the rules and classifies each entry as it counts
it, so its ignored share needs no index either.
An unfiltered tree with a share floor, such as the default `fdu PATH`, keeps every
directory but only the files large enough to show as rows; other metadata requests
retain the reusable index.
None of them reads regular-file contents.
One-shot metadata reports under `auto` neither load a snapshot, which cannot avoid the
current metadata walk, nor write one; `--cache on` writes one.
Any analyzer, named by `--analyze` or implied by `--view code` or `--view documents`,
opts into a separate content sidecar.
fdu reads eligible files whose requested result is absent or stale; a compatible
repeated run reuses unchanged records.
Coverage is scoped to the analyzers too: an unsupported deeper analyzer leaves byte
metadata visible but does not retain a separate lower-level metric record for that file.

## Pick the View, Then Shape It

- `--view list` (default), with `--format tree` for directory roll-ups, `--format paths`
  for complete flat matching paths, or `--long` for size, age, and path.
- `--view tree` for the directory hierarchy, in machine output too.
- `--view extensions` for the raw-extension breakdown.
  Rows partition the tree and so sum to its total; a derived extension always carries a
  leading dot, and names having none are tallied under the literal `(none)`.
- `--view types` for stable detected file types and exact byte shares.
- `--view families` for code, prose, markup, data, binary, and unknown roll-ups.
- `--view languages` for code-family rows and byte shares from path-only detection.
- `--view code` for source-line totals, coverage, and a language breakdown whose TOTAL
  row includes languages a display bound hid.
  It runs code analysis itself and is the default view for that analyzer.
- `--view documents` for words and pages by document format; it runs words analysis
  itself and is the default view for that analyzer.
- `--view largest` for the 20 largest regular files, and `--view recent` for the 20 most
  recently modified. Both are presets over `files`, not separate machinery: `largest` is
  `files --sort size --limit 20` and `recent` is `files --sort mtime --limit 20`, each
  restricted to regular files, and `--sort` and `--limit` still override them.
- `--view files` for a complete flat listing: every matching entry, in name order.
  One-shot text adds the performance footer described below; use a machine format when
  output is consumed programmatically.
- `--view summary` for one aggregate row.
- `--view full` for the bounded digest of every view but List and Files.
- Several views in one run share one scan: `--view summary,types,families`. Text then
  labels each block with an all-caps header naming its view; a single-view text report
  has no header. Machine formats tag every report with `view` either way.

`--analyze` names a set of analyzers, comma-separated, from `lines`, `code`, and
`words`; `none` and `all` are totals and cannot be combined with anything else.
It runs in union with what the views imply, so `--analyze none --view code` still runs
code analysis. Any analyzer may open eligible files and adds content work beyond the
metadata walk. Compatible cached results prevent unchanged bodies from being reread.

Add `--analyze lines` to stream physical, blank, and nonblank lines and raw word counts.
`--view code`, or `--analyze code` with another view, gives standard LOC, comment, and
code-blank partitions across supported common languages.
The Code table aligns code, comment, and blank lines with analyzed-file coverage for
each language and a bold TOTAL row.
TOTAL includes languages hidden by display bounds, and a note says so.
Each row’s gray parenthetical, such as `(0 gitignored)`, is its gitignored contribution,
already included in the row; coverage follows the table as notes.
Languages retains byte sizes and uses code-line shares when code analysis is enabled.
`--view documents`, or `--analyze words` with another view, gives normalized word
volume, paragraphs, aggregate-derived pages, and reader-visible Markdown that excludes
destinations and code.
The `documents` percentage column is document-word share, and a note says so in text.
`--view code,documents` — like `--analyze code,words` or `all` — computes both in one
streaming pass.
Both include `lines`, so adding it explicitly changes neither the metrics
nor the work. `lines` alone measures physical text volume without code counting or word
normalization.
Unsupported code languages still have physical-line metrics; their SLOC is
unavailable.

Requesting analysis without naming a view selects one that displays it: `code` selects
`code`, `words` selects `documents`, and `lines` selects `families`. `code,words` or
`all` selects both `code` and `documents`. Naming `--view` overrides that.
The converse holds only for the two views with no metadata meaning: `--view code`
implies `code` and `--view documents` implies `words`. Every other view, `full`
included, never enables an analyzer: a display choice with a cheap meaning must not
quietly authorize reading every file.
Headers name views; columns name metrics.
`words` is an analyzer, while `documents` is the prose/markup population.
Use `--analyze=words --view=types` to include word metrics for other text types.
Use `--analyze=lines --view=languages` for physical line counts across code languages.

When the selected views show none or only part of the requested analysis, a note names
what is not shown and a tip gives the views that show it:
`--analyze=code --view=summary` ends with `note: code analysis not shown by summary` and
`tip: show it: --view code`, and `--analyze=code --view=documents` with
`note: code analysis not shown by documents` and `tip: show it: --view documents,code`.
`--view full` includes Code only with code analysis and Documents only with words
analysis, naming the views it skipped and `--analyze all` to include them.
Use `--workers` to bound concurrent reads and `--words-per-page` to control page
derivation. Analysis never truncates a file or excludes it because of size.
One fixed bound changes a method: `words` counts a Markdown file over 64 MiB as plain
text, read whole but not rendered, and says so (`counted as text`, `text_only`, a note).
Invalid UTF-8, binary data, and unsupported SLOC languages remain visible as normal
coverage outcomes. Only I/O failures, files changed during a read, or stale commits make
analysis operationally partial.
Content analysis is one-shot and cannot be combined with `--watch`.

One-shot text reports end with a `perf:` line on stderr.
It reports regular files walked and their represented bytes, ignore files and accepted
rules, actual content bytes read, fresh and cached analysis, the cache tier, and elapsed
time. Total files/s and binary GiB/s use that elapsed time; represented GiB/s is not
content-read bandwidth.
Content-read throughput uses the analysis duration.
Known binary files can contribute walked bytes but zero read bytes.
`--stale-ok` runs report zero walked files because they never consult the tree.
The line is gray only when color is active and has no ANSI escapes otherwise.
Paths, Long, JSON, JSONL, YAML, skill output, lifecycle output, and watch streams omit
it.

A progress line can appear on stderr for a person at a terminal.
It is never drawn when stderr is not a terminal, `TERM` is `dumb`, or `CI` is set, so
agents need no flag; `fdu --docs` states the full rule.

Common shapes are compositions rather than dedicated flags:

```bash
fdu --view largest -n 100 PATH                        # the 100 largest files
fdu --kind file --modified-since 2h PATH              # files changed in the last two hours
fdu --view files --include '*.{rs,toml}' PATH         # by pattern
fdu --view tree --sort mtime PATH                     # an activity map
```

Tree output defaults to depth 5 and a minimum share of 1% of the selected root size.
Significant files appear alongside directories.
`--depth`, `--min-share`, `--breadth`, and `--limit` compose: depth bounds levels, share
hides smaller branches, breadth caps children per directory, and limit caps data rows
per section. Breadth and limit default to `all`; largest and recent retain their 20-row
presets. A `note: display limits:` line names each bound that hid rows, and one
`tip: show more:` names the flags that lift them; neither changes totals.

Use `--full`, or `--min-share=0% --depth=all --breadth=all --limit=all`, to remove every
display bound. Numeric content sorts such as `--sort=code_lines` require their analyzer.
Extensions supports metadata sorts only; use a metric-capable view for content ranking.
`--scan-depth` changes what is scanned and retained; display bounds only shorten output.

## Find Environments and Build Outputs

List all matching directories, then tally the covered paths from the same snapshot:

```bash
fdu PATH --kind dir --include .venv --include node_modules --include target --long --cache on
fdu PATH --kind dir --include .venv --include node_modules --include target --view summary --stale-ok
fdu PATH --kind dir --include .venv --include node_modules --include target --view files,summary --sort size --format json
```

The second command reads the snapshot the first one left with `--cache on`, without
revalidating it. Keep the root, cache destination, and scan population the same.
The third command combines flat rows and Summary in one machine report.
Ignored directories are included by default.

To select old environments by modification activity:

```bash
fdu PATH --kind dir --include .venv --modified-before 7d --long
fdu PATH --kind dir --include node_modules --modified-before 30d --long
fdu PATH --kind dir --include target --modified-before 30d --format paths
fdu PATH --kind dir --include .venv --include venv --include node_modules --include target --modified-before 30d --long --sort mtime --reverse
fdu PATH --kind dir --include .venv --modified-before 30d --format json
```

Kind, basename/relative-path patterns, size, and modification age are filters.
Repeated includes form a union.
Directory bytes sum eligible regular-file contents; recency is the newest root or
eligible descendant mtime, including directories and symlinks.
Empty directories use their own mtime.
This is modification activity, not access or last use.
Directory names such as Cargo’s `target` are conventions, not proof of ownership.
Allocated bytes are counted per path; separate hard links are not deduplicated and
clones can share physical blocks.
Windows currently reports apparent size as allocated.
These measurements do not estimate space freed by deletion.
Symlinks are not followed.

Exclusions win throughout selected subtrees before size/age bounds; ignored-only queries
traverse structural ancestors.
Flat output lists matching entries, including nested roots whose sizes overlap.
Aggregate views count the covered union once.
The default is the directory tree; `--tree` makes its format explicit.
List’s flat formats are complete and size-ranked by default; Files retains name order
unless `--sort` changes it.
A row limit caps each section, including a tree; breadth caps each directory and depth
bounds displayed levels.
None changes measured totals.
Paths escapes control characters only and keeps stdout to paths; bound and rule notices
go to stderr. Long adds size and signed age.
Machine rows retain exact `mtime_ns`, `age_ns`, directory `files`/`dirs`, and the
report’s `age_reference_ns`; unknown ages are null.

Tree/Paths/Long require one compatible list view.
Use automatic Text or machine output for grouped/mixed views and Full.
Largest/recent retain regular-file ranks with Paths or Long.
Explicit Paths/Long overrides the Tree view’s presentation.
Format flags conflict.
Rust/Python callers select format on the query before reading; a detached Report cannot
turn a folded tree into a complete flat inventory.
Request another report from the retained index for that change, without scanning again.

## Read What `.gitignore` Covers

Fresh scans read applicable per-directory `.gitignore` files by default.
`--stale-ok` reports use retained rule state, and `--no-gitignore` disables the rules.
Summary, tree, and extension rows show ignored size as a gray parenthetical such as
`(128 B gitignored)` when color is enabled.
The performance line counts ignore files and accepted rules.

```bash
fdu PATH --ignored=exclude                              # folders by what the rules leave
fdu PATH --view=files --ignored=only --format=jsonl     # every entry the rules cover
fdu PATH --no-gitignore                                 # read no rules, show no share
```

Selecting a side changes sizes, ordering, and `--min-size` together, because they follow
the entries shown. `--ignored=exclude` avoids enumerating safely ignored subtrees and
reading ignored bodies.
`--ignored=only` traverses the directories needed to discover ignored entries and
analyzes only ignored bodies.
The default `include` measures both populations and reports their contributions
separately. Unknown classifications cannot justify pruning.
`--no-gitignore` with either selection is a usage error.
Only per-directory `.gitignore` files apply, not `core.excludesFile`,
`.git/info/exclude`, or a global ignore file.
Each is found as git opens it, so a `.GITIGNORE` counts on a case-insensitive volume;
matching itself is case-sensitive.
Unignored does not mean tracked: `.git` is unignored unless a rule names it.
For recent working files, add both `--ignored=exclude` and `--exclude='.git/**'`. An
unreadable `.gitignore` makes the result partial (exit 2), while one past
`--gitignore-budget` or `--gitignore-line-limit` is refused whole and named in a note:
sizes stay exact, the ignored shares under that directory do not.

Under `--watch`, a rule edit that moves an entry into either selection streams the
upsert that draws it and one that moves it out streams the removal, so the stream holds
the entry set the flag names.
An upsert carries `ignored`, and so does a removal a rule edit caused; an ordinary
removal, an invalidation, and every record of a run that read no rules omit it.
Without either flag the stream maintains membership rather than each row’s bit, so
re-read a listing after a rule edit if the bit matters.

## Value Grammars

- Sizes: `512`, `10k`, `10M`, `1.5GiB`. Decimal and binary units, case-insensitive.
- Times: `now`, a compound age (`200ms`, `45s`, `2h`, `1h30m`), an RFC 3339 timestamp
  with an offset (`2026-08-10T18:22:31Z`), or `@` epoch seconds.
  Calendar units and fractional ages are rejected with the spelling to use instead; a
  bare local date-time is rejected because resolving it needs a time-zone database.
- `--modified-since` is inclusive and `--modified-before` is exclusive.

## Use Timestamps as a Sync Watermark

Every report carries `provenance.scan_started_at`. Feeding it back selects exactly what
changed after that scan began, which is what makes incremental follow-up sound:

```bash
fdu --view summary --format json PATH              # record provenance.scan_started_at
fdu --view files --kind file --format jsonl --modified-since <that> PATH
```

Use the scan’s *start*, not its end: a file modified mid-scan may have been observed
before the modification, so only the start bound is conservative.

## Validate Every Automated Result

Check the process exit status and these fields:

- `schema` before parsing anything else: a report carries `fdu.report/10`, a `--watch`
  stream carries `fdu.stream/2`, and `--cache-status` carries `fdu.cache/3`. Treat an
  unrecognized value as a version you cannot parse rather than guessing at the fields.
- Integer fields that exceed 2^53 (fingerprints, option hashes, nanosecond timestamps)
  lose precision in IEEE 754 binary64 parsers such as JavaScript `JSON.parse`
- `status.complete`, `status.errors`, and `status.errors_omitted` before trusting totals
- `provenance.freshness` and `provenance.source` before presenting data as current
- `truncated` on a tree node before treating it as exhaustive
- Each requested unit in a metric row’s `coverage` before presenting its metrics as
  complete
- `ignored` on a row before calling anything ignored or not: an object, or `true` and
  `false` on a file row, where rules were read, and `null` where none were, which never
  means nothing is ignored
- `ignore_rules.refused` before trusting an ignored share: below a refused `.gitignore`
  the split is not exact, though sizes are
- `detection.sources`, `detection.confidence`, and `detection.flags` before treating a
  deep-detected type or origin label as exact

`provenance.source` is `cold_scan`, `warm_revalidate`, or `cache_only`. Only
`--stale-ok` can return `provenance.freshness: stale`, and it says so rather than
implying currency: every format also prints a `warn: stale answer` line on stderr, which
`--quiet` keeps. It fails outright when no usable snapshot exists rather than silently
scanning.

Exit 0 is accepted success, exit 1 is a fatal failure, and exit 2 is incomplete data or
invalid usage. Do not discard useful stdout from exit 2; inspect the completeness fields
and use `--allow-partial` only when incomplete totals are acceptable.

## Cache Behavior

No ordinary view requires a preexisting cache.
`--cache auto`, the default, uses the cache only where the kind of run gains from it.
Metadata-only one-shot reports include current sizes or timestamps, so they must inspect
every entry; under `auto` they neither load a snapshot, which cannot make that
verification cheaper, nor write one, which no later report reads.
Content analysis, `--watch`, and a retained index from the Rust or Python `open` read,
revalidate, and write it.
`--cache on` also writes after a one-shot report, which is how to leave a snapshot for a
later `--stale-ok` answer.

Content analysis is where ordinary repeated runs benefit most.
The first compatible run reads eligible bodies; a later run restores unchanged records
from the content sidecar and reads only changed or newly eligible files.
A stored analyzer set answers only the same set: a different one, wider or narrower,
reads the files again and replaces it.
The performance footer reports fresh and cached analysis separately.

`--stale-ok` is a distinct contract: it never verifies the source tree, labels the
answer stale, and fails unless compatible metadata and any requested content analysis
already exist. `--cache=off` neither reads nor writes fdu cache data.

The cache defaults to `~/.cache/fdu` on macOS and Linux, and `%LOCALAPPDATA%/fdu` on
Windows. Set an exact destination with `--cache-dir DIR` or `FDU_CACHE_DIR`; the flag
wins. Otherwise `XDG_CACHE_HOME/fdu` overrides the platform default.
Status and clear use the same destination.

Each root has a `<16-hex-key>.metadata.bin` filesystem snapshot and, when analyzed, a
matching `<16-hex-key>.analysis.bin` file of derived metrics.
The key identifies the canonical root.
Analysis files store counts and classifications, not copies of source bodies.
These binary files are disposable; use cache status to inspect them.
`--cache-status` maps a hash-named file back to the tree it describes, and
`--cache-clear` removes it; both run without scanning.
Cache status is its own document, carrying the `fdu.cache/3` schema in every machine
format rather than a report schema.
A current snapshot’s row carries the `identity` of the entry and `.gitignore` tiers it
holds, and every row a `content` object for the sidecar beside it, with its own `state`
and a current sidecar’s `identity`, or `null` when there is none.
Each status row carries a `state`: `current`, `stale` for a snapshot another fdu version
wrote or one this build cannot read, `leftover` for a file fdu left behind, with a
`leftover_kind`, `unrecognized` for a file that is not fdu’s, or `absent`. Clearing
removes current and stale snapshots, so `--cache-clear=all` reclaims what an upgrade
leaves behind; it also reclaims leftovers, and it never removes an unrecognized file.

All shipped metadata views need current per-file attributes because they report sizes or
timestamps. An in-place edit changes no directory timestamp, so a cached directory
fingerprint cannot prove those answers current.
Content reuse still pays because an unchanged file fingerprint avoids the much more
expensive body read and analysis.

Exact names and ordinary extensions remain path-only classifications.
When analysis is enabled, unresolved files and ambiguous `.h` headers may use bounded
shebang, modeline, literal, or signature probes.
Do not collapse their provenance into an unqualified language claim; retain the report’s
source and confidence fields when summarizing or transforming machine output.

Run `fdu --help` for the complete flag, cache, color, scope, and exit contract.

<!-- This document follows common-doc-guidelines.md.
See github.com/jlevy/practical-prose and review guidelines before editing.
-->