---
name: fdu
description: >-
Inspect disk usage and codebases: directory sizes, file counts, recency, file types,
lines of code by language, and words by document type. Use to find large or stale
build directories, size up the source and prose in a codebase, or collect structured
filesystem roll-ups for scripts and agents.
---
# fdu Directory Roll-Ups
This is the complete fdu usage contract; it needs no setup chat or prior session
context.
Use `fdu` to summarize a directory tree without modifying files in that tree.
It reports size, file count, recency, and file kinds for every directory at once, and on
request lines of code by language and words by document type.
It walks the tree on several threads through each platform’s native directory interface,
and on a generated million-entry tree (875,000 files) it finished ahead of `du` and the
seven other disk-usage tools measured on Linux and macOS. Every report requires an
explicit `PATH`, `.` for the current directory; bare `fdu` prints help instead of
scanning. `fdu --docs` prints common commands, cache behavior, and the full usage
contract without a PATH and without scanning.
## Run fdu
Start with the report that answers the question:
```bash
fdu . # directory-size tree; metadata only
fdu . --view=code,documents # code lines by language, words by document type
fdu . --format=json --view=tree --depth=2 --limit=20
```
The tree reads only metadata.
`--view=code,documents` also reads file contents, both analyses in one scan, and caches
the results, so a repeated run reads only changed files.
The third is a shallower tree for a script: `--format=json` (or `jsonl` or `yaml`) gives
exact values under a versioned `schema`, so use it whenever the output is parsed.
Prefer an `fdu` already on `PATH`. When there is none, run the latest release through
`uvx`, which needs no install step:
```bash
if command -v fdu >/dev/null 2>&1; then
fdu . --format=json --view=tree --depth=2 --limit=20
else
uvx --no-build fdu@latest . --format=json --view=tree --depth=2 --limit=20
fi
```
`--no-build` makes the fallback use a published wheel instead of compiling Rust.
To keep fdu installed, run `uv tool install --no-build fdu` (later
`uv tool upgrade --no-build fdu`) or `cargo install --locked fdu`.
Generated by fdu `__FDU_VERSION__`; if `fdu --version` differs, re-run
`fdu --install-skill` so this file describes the installed command.
More reports, each from one scan:
```bash
fdu . --view=code # code overview with population and coverage
fdu . --view=documents # words and pages by document format
fdu . --view=code --ignored=exclude # source overview without ignored content
fdu . --ignored=exclude # skip ignored trees and their contents
fdu . --view=summary # one total with its ignored share
fdu . --view=languages # detected language sizes; metadata only
fdu . --view=families,types,extensions # three file-kind breakdowns
fdu . --view=largest --limit=10 # ten largest files
fdu . --view=recent --limit=10 # ten most recently modified files
```
`--view` chooses what is printed, and `--analyze` adds what may be read beyond it.
`code` and `documents` are the views that read file contents: they show nothing without
analysis, so naming one runs its analyzer.
Every other view reads only metadata unless `--analyze` names an analyzer, as in these
two controls:
```bash
fdu . --analyze=code --view=languages # code lines in the language rows
fdu . --analyze=lines --view=languages # physical lines and raw words by language
```
Without code analysis, language percentages are byte shares; with it, the rows add code,
comment, and blank-line metrics and use code-line shares.
Use `--size apparent` when logical file lengths are wanted instead of allocated bytes.
## Read the Result and Its Notes
Stdout holds only the result: rows, column headings, and, when several views are shown,
an all-caps header per view, such as `LANGUAGES (3 of 15)` when a row limit hid some.
A grouped row (`families`, `types`, `languages`, `documents`) ends its first line at its
file count; its other measures, such as lines, words, and files not analyzed, are
indented lines under that count, not rows of their own.
Every explanation follows on stderr, one prefixed line each, in this order: `note:` for
what the result covers and how to read it, `warn:` for an operation that failed or an
answer nothing verified, `tip:` for a runnable change, and, after text reports, `perf:`
for the work done and the elapsed time.
Notes are consolidated: one says what totals include, one lists every display limit that
hid rows, and one tip lifts them all:
```text
note: totals include gitignored sizes and descendants
note: display limits: below 1% of root, depth 5
tip: show more: --min-share=0% --depth=all
```
Add the tip’s flags to the same command to see what was hidden, or use `--full` to lift
every display bound.
Display limits never change totals.
Notes also name a percentage denominator other than bytes, with the section it applies
to when several views are shown
(`note: percentages are shares of code lines (CODE), document words (DOCUMENTS)`), and
the code view’s coverage (`note: 15 languages analyzed`,
`note: not analyzed: 2 unsupported`). Machine formats send the same lines to stderr, so
stdout stays parseable; check their structured fields, not the notes, before trusting a
result. Use `--quiet` (`-q`) to hide notes, tips, performance lines, and progress.
Result stdout, warnings, errors, and exit status are unchanged; structured facts are
retained.
The default is `list` in `tree` format, in allocated bytes, largest first, to depth 5.
It shows directory subtrees and file leaves contributing at least 1% of the selected
root. Breadth and total rows are unbounded unless requested; `--depth`, `--min-share`,
`--breadth`, and `--limit` compose independently.
Hidden and ignored entries are included.
`.gitignore` is read to label ignored shares, not to exclude matching entries.
A parenthetical amount such as `(73 MiB gitignored)` is included in the row total.
Directories get a `/` suffix, except `.` and `..`; file names and structured paths do
not change.
Color is on only when the output is a terminal or `--color=always` is given, so piped
tree bars are plain: `█` for usage and `░` for unused width.
Colored bars use green `█` for non-gitignored usage, green `▓` for gitignored usage, and
green `▒` when classification is unknown; dim green `░` fills unused width.
Zero sizes and percentages below 1% are gray; sizes at least 1 GiB are bold.
Cyan names are bright and bold.
Gitignored directories use regular, nonbold cyan; containing ignored files is not
enough. The directory `/` suffix is gray.
`--bar-size=20` widens tree bars; zero or negative values hide the bar column.
The default is 10 characters; the maximum is 4,096. This does not change structured
output or measurements.
## Complete Inventories and File Search
```bash
fdu . --view tree --full --format json # full recursive hierarchy
fdu . --kind dir --full --sort name --format json # directory rows with recursive usage
fdu . --view files --kind file --full --format json # individual regular files
fdu . --kind file --include '*.rs' --full --format paths
fdu . --kind dir --include node_modules --full --long
```
`--full` means `--depth=all --breadth=all --limit=all --min-share=0%`. Explicit bounds
override it, regardless of order.
`--view full` chooses multiple views; `--full` expands whichever views were chosen.
Neither changes scan scope or analysis.
Use these commands for find/fd-style filename searches and complete usage inventories.
fdu includes hidden and ignored content; fd needs `--unrestricted` for that population.
fdu selection is glob-based.
It does not implement find expressions or `-exec`, and its ignore sources differ from
fd’s. Paths output escapes control characters; use structured output for arbitrary
native filenames. Directory rows include descendants and overlap; add
`--view list,summary` for a path-union total rather than summing rows.
A tree’s `remainder` contains recursive `files`, `bytes`, `allocated`, and applicable
`reasons` outside its displayed root-level rows; `null` means nothing is hidden there.
A displayed directory already represents its whole subtree, including descendants whose
rows were bounded away.
Per-boundary `entries` counts hidden roots, while `files` counts regular files
recursively. Unknown amounts are null.
The text equivalent is one root-level line with the same bar, percentage, and size
columns as the tree rows, followed by `… and N more files`. These columns represent the
remaining share of the selected root.
Unknown size or count stays unknown; an unknown hidden size has no numeric share.
Fully expanded, fully observed trees have no remainder line and no display-limit notes.
Check completeness separately: unreadable directories and scan-depth restrictions still
apply.
## Compose the Request From Six Axes
The six axes separate discovery, measurement, selection, and presentation.
Unsupported combinations fail before scanning.
There are no subcommands: the grammar is always “report on a path”.
| Scope | What is scanned and cached? | `PATH`, `--scan-depth N`, `--one-filesystem`, `--gitignore-budget SIZE\|all`, `--gitignore-line-limit SIZE\|all`, `--no-gitignore`, `--ignored=include\|exclude\|only` |
| Content | Which file bodies are read beyond what the views imply? | `--analyze none\|lines\|code\|words\|all` |
| Selection | Which entries does this query consider? | `--include`, `--exclude`, `--min-size`, `--modified-since`, `--modified-before`, `--kind`, `--depth`, `--min-share`, `--breadth`, `-n/--limit`, `--full`, `--sort`, `--reverse`, `--size` |
| View | Which roll-up is reported? | `--view list,summary,tree,families,types,extensions,languages,code,documents,largest,recent,files`, or `--view full` |
| Format | How is it serialized? | `--format text\|tree\|paths\|long\|json\|jsonl\|yaml`, `--color`, `--progress` |
| Mode | How is work performed? | `--cache auto\|on\|off`, `--stale-ok`, `--cache-dir DIR`, `--watch`, `--workers N` |
Scope determines what is scanned and cached.
In a one-shot report, the ignored population also determines which subtrees and file
bodies may be skipped.
A retained index can answer narrower queries when it holds the required facts.
Work has three layers.
A single unfiltered `--view summary PATH` is the one exact composition that retains only
aggregate tallies and no index, except under `--cache on` and `--stale-ok`, whose
contracts are about the snapshot itself.
Otherwise a snapshot cannot save the walk that request is already doing, so it neither
reads nor writes one.
Reading `.gitignore`, the summary keeps the rules and classifies each entry as it counts
it, so its ignored share needs no index either.
An unfiltered tree with a share floor, such as the default `fdu PATH`, keeps every
directory but only the files large enough to show as rows; other metadata requests
retain the reusable index.
None of them reads regular-file contents.
One-shot metadata reports under `auto` neither load a snapshot, which cannot avoid the
current metadata walk, nor write one; `--cache on` writes one.
Any analyzer, named by `--analyze` or implied by `--view code` or `--view documents`,
opts into a separate content sidecar.
fdu reads eligible files whose requested result is absent or stale; a compatible
repeated run reuses unchanged records.
Coverage is scoped to the analyzers too: an unsupported deeper analyzer leaves byte
metadata visible but does not retain a separate lower-level metric record for that file.
## Pick the View, Then Shape It
- `--view list` (default), with `--format tree` for directory roll-ups, `--format paths`
for complete flat matching paths, or `--long` for size, age, and path.
- `--view tree` for the directory hierarchy, in machine output too.
- `--view extensions` for the raw-extension breakdown.
Rows partition the tree and so sum to its total; a derived extension always carries a
leading dot, and names having none are tallied under the literal `(none)`.
- `--view types` for stable detected file types and exact byte shares.
- `--view families` for code, prose, markup, data, binary, and unknown roll-ups.
- `--view languages` for code-family rows and byte shares from path-only detection.
- `--view code` for source-line totals, coverage, and a language breakdown whose TOTAL
row includes languages a display bound hid.
It runs code analysis itself and is the default view for that analyzer.
- `--view documents` for words and pages by document format; it runs words analysis
itself and is the default view for that analyzer.
- `--view largest` for the 20 largest regular files, and `--view recent` for the 20 most
recently modified. Both are presets over `files`, not separate machinery: `largest` is
`files --sort size --limit 20` and `recent` is `files --sort mtime --limit 20`, each
restricted to regular files, and `--sort` and `--limit` still override them.
- `--view files` for a complete flat listing: every matching entry, in name order.
One-shot text adds the performance footer described below; use a machine format when
output is consumed programmatically.
- `--view summary` for one aggregate row.
- `--view full` for the bounded digest of every view but List and Files.
- Several views in one run share one scan: `--view summary,types,families`. Text then
labels each block with an all-caps header naming its view; a single-view text report
has no header. Machine formats tag every report with `view` either way.
`--analyze` names a set of analyzers, comma-separated, from `lines`, `code`, and
`words`; `none` and `all` are totals and cannot be combined with anything else.
It runs in union with what the views imply, so `--analyze none --view code` still runs
code analysis. Any analyzer may open eligible files and adds content work beyond the
metadata walk. Compatible cached results prevent unchanged bodies from being reread.
Add `--analyze lines` to stream physical, blank, and nonblank lines and raw word counts.
`--view code`, or `--analyze code` with another view, gives standard LOC, comment, and
code-blank partitions across supported common languages.
The Code table aligns code, comment, and blank lines with analyzed-file coverage for
each language and a bold TOTAL row.
TOTAL includes languages hidden by display bounds, and a note says so.
Each row’s gray parenthetical, such as `(0 gitignored)`, is its gitignored contribution,
already included in the row; coverage follows the table as notes.
Languages retains byte sizes and uses code-line shares when code analysis is enabled.
`--view documents`, or `--analyze words` with another view, gives normalized word
volume, paragraphs, aggregate-derived pages, and reader-visible Markdown that excludes
destinations and code.
The `documents` percentage column is document-word share, and a note says so in text.
`--view code,documents` — like `--analyze code,words` or `all` — computes both in one
streaming pass.
Both include `lines`, so adding it explicitly changes neither the metrics
nor the work. `lines` alone measures physical text volume without code counting or word
normalization.
Unsupported code languages still have physical-line metrics; their SLOC is
unavailable.
Requesting analysis without naming a view selects one that displays it: `code` selects
`code`, `words` selects `documents`, and `lines` selects `families`. `code,words` or
`all` selects both `code` and `documents`. Naming `--view` overrides that.
The converse holds only for the two views with no metadata meaning: `--view code`
implies `code` and `--view documents` implies `words`. Every other view, `full`
included, never enables an analyzer: a display choice with a cheap meaning must not
quietly authorize reading every file.
Headers name views; columns name metrics.
`words` is an analyzer, while `documents` is the prose/markup population.
Use `--analyze=words --view=types` to include word metrics for other text types.
Use `--analyze=lines --view=languages` for physical line counts across code languages.
When the selected views show none or only part of the requested analysis, a note names
what is not shown and a tip gives the views that show it:
`--analyze=code --view=summary` ends with `note: code analysis not shown by summary` and
`tip: show it: --view code`, and `--analyze=code --view=documents` with
`note: code analysis not shown by documents` and `tip: show it: --view documents,code`.
`--view full` includes Code only with code analysis and Documents only with words
analysis, naming the views it skipped and `--analyze all` to include them.
Use `--workers` to bound concurrent reads and `--words-per-page` to control page
derivation. Analysis never truncates a file or excludes it because of size.
One fixed bound changes a method: `words` counts a Markdown file over 64 MiB as plain
text, read whole but not rendered, and says so (`counted as text`, `text_only`, a note).
Invalid UTF-8, binary data, and unsupported SLOC languages remain visible as normal
coverage outcomes. Only I/O failures, files changed during a read, or stale commits make
analysis operationally partial.
Content analysis is one-shot and cannot be combined with `--watch`.
One-shot text reports end with a `perf:` line on stderr.
It reports regular files walked and their represented bytes, ignore files and accepted
rules, actual content bytes read, fresh and cached analysis, the cache tier, and elapsed
time. Total files/s and binary GiB/s use that elapsed time; represented GiB/s is not
content-read bandwidth.
Content-read throughput uses the analysis duration.
Known binary files can contribute walked bytes but zero read bytes.
`--stale-ok` runs report zero walked files because they never consult the tree.
The line is gray only when color is active and has no ANSI escapes otherwise.
Paths, Long, JSON, JSONL, YAML, skill output, lifecycle output, and watch streams omit
it.
A progress line can appear on stderr for a person at a terminal.
It is never drawn when stderr is not a terminal, `TERM` is `dumb`, or `CI` is set, so
agents need no flag; `fdu --docs` states the full rule.
Common shapes are compositions rather than dedicated flags:
```bash
fdu --view largest -n 100 PATH # the 100 largest files
fdu --kind file --modified-since 2h PATH # files changed in the last two hours
fdu --view files --include '*.{rs,toml}' PATH # by pattern
fdu --view tree --sort mtime PATH # an activity map
```
Tree output defaults to depth 5 and a minimum share of 1% of the selected root size.
Significant files appear alongside directories.
`--depth`, `--min-share`, `--breadth`, and `--limit` compose: depth bounds levels, share
hides smaller branches, breadth caps children per directory, and limit caps data rows
per section. Breadth and limit default to `all`; largest and recent retain their 20-row
presets. A `note: display limits:` line names each bound that hid rows, and one
`tip: show more:` names the flags that lift them; neither changes totals.
Use `--full`, or `--min-share=0% --depth=all --breadth=all --limit=all`, to remove every
display bound. Numeric content sorts such as `--sort=code_lines` require their analyzer.
Extensions supports metadata sorts only; use a metric-capable view for content ranking.
`--scan-depth` changes what is scanned and retained; display bounds only shorten output.
## Find Environments and Build Outputs
List all matching directories, then tally the covered paths from the same snapshot:
```bash
fdu PATH --kind dir --include .venv --include node_modules --include target --long --cache on
fdu PATH --kind dir --include .venv --include node_modules --include target --view summary --stale-ok
fdu PATH --kind dir --include .venv --include node_modules --include target --view files,summary --sort size --format json
```
The second command reads the snapshot the first one left with `--cache on`, without
revalidating it. Keep the root, cache destination, and scan population the same.
The third command combines flat rows and Summary in one machine report.
Ignored directories are included by default.
To select old environments by modification activity:
```bash
fdu PATH --kind dir --include .venv --modified-before 7d --long
fdu PATH --kind dir --include node_modules --modified-before 30d --long
fdu PATH --kind dir --include target --modified-before 30d --format paths
fdu PATH --kind dir --include .venv --include venv --include node_modules --include target --modified-before 30d --long --sort mtime --reverse
fdu PATH --kind dir --include .venv --modified-before 30d --format json
```
Kind, basename/relative-path patterns, size, and modification age are filters.
Repeated includes form a union.
Directory bytes sum eligible regular-file contents; recency is the newest root or
eligible descendant mtime, including directories and symlinks.
Empty directories use their own mtime.
This is modification activity, not access or last use.
Directory names such as Cargo’s `target` are conventions, not proof of ownership.
Allocated bytes are counted per path; separate hard links are not deduplicated and
clones can share physical blocks.
Windows currently reports apparent size as allocated.
These measurements do not estimate space freed by deletion.
Symlinks are not followed.
Exclusions win throughout selected subtrees before size/age bounds; ignored-only queries
traverse structural ancestors.
Flat output lists matching entries, including nested roots whose sizes overlap.
Aggregate views count the covered union once.
The default is the directory tree; `--tree` makes its format explicit.
List’s flat formats are complete and size-ranked by default; Files retains name order
unless `--sort` changes it.
A row limit caps each section, including a tree; breadth caps each directory and depth
bounds displayed levels.
None changes measured totals.
Paths escapes control characters only and keeps stdout to paths; bound and rule notices
go to stderr. Long adds size and signed age.
Machine rows retain exact `mtime_ns`, `age_ns`, directory `files`/`dirs`, and the
report’s `age_reference_ns`; unknown ages are null.
Tree/Paths/Long require one compatible list view.
Use automatic Text or machine output for grouped/mixed views and Full.
Largest/recent retain regular-file ranks with Paths or Long.
Explicit Paths/Long overrides the Tree view’s presentation.
Format flags conflict.
Rust/Python callers select format on the query before reading; a detached Report cannot
turn a folded tree into a complete flat inventory.
Request another report from the retained index for that change, without scanning again.
## Read What `.gitignore` Covers
Fresh scans read applicable per-directory `.gitignore` files by default.
`--stale-ok` reports use retained rule state, and `--no-gitignore` disables the rules.
Summary, tree, and extension rows show ignored size as a gray parenthetical such as
`(128 B gitignored)` when color is enabled.
The performance line counts ignore files and accepted rules.
```bash
fdu PATH --ignored=exclude # folders by what the rules leave
fdu PATH --view=files --ignored=only --format=jsonl # every entry the rules cover
fdu PATH --no-gitignore # read no rules, show no share
```
Selecting a side changes sizes, ordering, and `--min-size` together, because they follow
the entries shown. `--ignored=exclude` avoids enumerating safely ignored subtrees and
reading ignored bodies.
`--ignored=only` traverses the directories needed to discover ignored entries and
analyzes only ignored bodies.
The default `include` measures both populations and reports their contributions
separately. Unknown classifications cannot justify pruning.
`--no-gitignore` with either selection is a usage error.
Only per-directory `.gitignore` files apply, not `core.excludesFile`,
`.git/info/exclude`, or a global ignore file.
Each is found as git opens it, so a `.GITIGNORE` counts on a case-insensitive volume;
matching itself is case-sensitive.
Unignored does not mean tracked: `.git` is unignored unless a rule names it.
For recent working files, add both `--ignored=exclude` and `--exclude='.git/**'`. An
unreadable `.gitignore` makes the result partial (exit 2), while one past
`--gitignore-budget` or `--gitignore-line-limit` is refused whole and named in a note:
sizes stay exact, the ignored shares under that directory do not.
Under `--watch`, a rule edit that moves an entry into either selection streams the
upsert that draws it and one that moves it out streams the removal, so the stream holds
the entry set the flag names.
An upsert carries `ignored`, and so does a removal a rule edit caused; an ordinary
removal, an invalidation, and every record of a run that read no rules omit it.
Without either flag the stream maintains membership rather than each row’s bit, so
re-read a listing after a rule edit if the bit matters.
## Value Grammars
- Sizes: `512`, `10k`, `10M`, `1.5GiB`. Decimal and binary units, case-insensitive.
- Times: `now`, a compound age (`200ms`, `45s`, `2h`, `1h30m`), an RFC 3339 timestamp
with an offset (`2026-08-10T18:22:31Z`), or `@` epoch seconds.
Calendar units and fractional ages are rejected with the spelling to use instead; a
bare local date-time is rejected because resolving it needs a time-zone database.
- `--modified-since` is inclusive and `--modified-before` is exclusive.
## Use Timestamps as a Sync Watermark
Every report carries `provenance.scan_started_at`. Feeding it back selects exactly what
changed after that scan began, which is what makes incremental follow-up sound:
```bash
fdu --view summary --format json PATH # record provenance.scan_started_at
fdu --view files --kind file --format jsonl --modified-since <that> PATH
```
Use the scan’s *start*, not its end: a file modified mid-scan may have been observed
before the modification, so only the start bound is conservative.
## Validate Every Automated Result
Check the process exit status and these fields:
- `schema` before parsing anything else: a report carries `fdu.report/10`, a `--watch`
stream carries `fdu.stream/2`, and `--cache-status` carries `fdu.cache/3`. Treat an
unrecognized value as a version you cannot parse rather than guessing at the fields.
- Integer fields that exceed 2^53 (fingerprints, option hashes, nanosecond timestamps)
lose precision in IEEE 754 binary64 parsers such as JavaScript `JSON.parse`
- `status.complete`, `status.errors`, and `status.errors_omitted` before trusting totals
- `provenance.freshness` and `provenance.source` before presenting data as current
- `truncated` on a tree node before treating it as exhaustive
- Each requested unit in a metric row’s `coverage` before presenting its metrics as
complete
- `ignored` on a row before calling anything ignored or not: an object, or `true` and
`false` on a file row, where rules were read, and `null` where none were, which never
means nothing is ignored
- `ignore_rules.refused` before trusting an ignored share: below a refused `.gitignore`
the split is not exact, though sizes are
- `detection.sources`, `detection.confidence`, and `detection.flags` before treating a
deep-detected type or origin label as exact
`provenance.source` is `cold_scan`, `warm_revalidate`, or `cache_only`. Only
`--stale-ok` can return `provenance.freshness: stale`, and it says so rather than
implying currency: every format also prints a `warn: stale answer` line on stderr, which
`--quiet` keeps. It fails outright when no usable snapshot exists rather than silently
scanning.
Exit 0 is accepted success, exit 1 is a fatal failure, and exit 2 is incomplete data or
invalid usage. Do not discard useful stdout from exit 2; inspect the completeness fields
and use `--allow-partial` only when incomplete totals are acceptable.
## Cache Behavior
No ordinary view requires a preexisting cache.
`--cache auto`, the default, uses the cache only where the kind of run gains from it.
Metadata-only one-shot reports include current sizes or timestamps, so they must inspect
every entry; under `auto` they neither load a snapshot, which cannot make that
verification cheaper, nor write one, which no later report reads.
Content analysis, `--watch`, and a retained index from the Rust or Python `open` read,
revalidate, and write it.
`--cache on` also writes after a one-shot report, which is how to leave a snapshot for a
later `--stale-ok` answer.
Content analysis is where ordinary repeated runs benefit most.
The first compatible run reads eligible bodies; a later run restores unchanged records
from the content sidecar and reads only changed or newly eligible files.
A stored analyzer set answers only the same set: a different one, wider or narrower,
reads the files again and replaces it.
The performance footer reports fresh and cached analysis separately.
`--stale-ok` is a distinct contract: it never verifies the source tree, labels the
answer stale, and fails unless compatible metadata and any requested content analysis
already exist. `--cache=off` neither reads nor writes fdu cache data.
The cache defaults to `~/.cache/fdu` on macOS and Linux, and `%LOCALAPPDATA%/fdu` on
Windows. Set an exact destination with `--cache-dir DIR` or `FDU_CACHE_DIR`; the flag
wins. Otherwise `XDG_CACHE_HOME/fdu` overrides the platform default.
Status and clear use the same destination.
Each root has a `<16-hex-key>.metadata.bin` filesystem snapshot and, when analyzed, a
matching `<16-hex-key>.analysis.bin` file of derived metrics.
The key identifies the canonical root.
Analysis files store counts and classifications, not copies of source bodies.
These binary files are disposable; use cache status to inspect them.
`--cache-status` maps a hash-named file back to the tree it describes, and
`--cache-clear` removes it; both run without scanning.
Cache status is its own document, carrying the `fdu.cache/3` schema in every machine
format rather than a report schema.
A current snapshot’s row carries the `identity` of the entry and `.gitignore` tiers it
holds, and every row a `content` object for the sidecar beside it, with its own `state`
and a current sidecar’s `identity`, or `null` when there is none.
Each status row carries a `state`: `current`, `stale` for a snapshot another fdu version
wrote or one this build cannot read, `leftover` for a file fdu left behind, with a
`leftover_kind`, `unrecognized` for a file that is not fdu’s, or `absent`. Clearing
removes current and stale snapshots, so `--cache-clear=all` reclaims what an upgrade
leaves behind; it also reclaims leftovers, and it never removes an unrecognized file.
All shipped metadata views need current per-file attributes because they report sizes or
timestamps. An in-place edit changes no directory timestamp, so a cached directory
fingerprint cannot prove those answers current.
Content reuse still pays because an unchanged file fingerprint avoids the much more
expensive body read and analysis.
Exact names and ordinary extensions remain path-only classifications.
When analysis is enabled, unresolved files and ambiguous `.h` headers may use bounded
shebang, modeline, literal, or signature probes.
Do not collapse their provenance into an unqualified language claim; retain the report’s
source and confidence fields when summarizing or transforming machine output.
Run `fdu --help` for the complete flag, cache, color, scope, and exit contract.