---
name: fdu
description: >-
Inspect directory trees with hierarchical file counts, apparent and allocated sizes,
recency, and extension tallies. Use when investigating disk usage, finding large
directories, summarizing file types, listing files by size or age, or collecting stable
JSON filesystem roll-ups for scripts and coding agents.
---
# fdu Directory Roll-Ups
Use `fdu` to summarize a directory tree without modifying files in that tree.
`fdu --docs` prints common commands, cache behavior, and the full usage contract without
a PATH and without scanning.
Every report requires an explicit `PATH`; bare `fdu` prints help instead of scanning the
current directory.
## Run fdu
Use `.` for the current directory.
Start with the report that answers the question:
```bash
fdu . # directory-size tree; metadata only
fdu . --exclude-ignored # select what .gitignore rules leave
fdu . --view=summary # one total with its ignored share
fdu . --view=languages # detected language sizes; metadata only
fdu . --view=families,types,extensions # three file-kind breakdowns
fdu . --view=recent --limit=10 # ten most recently modified files
fdu . --analyze=lines # physical lines and raw words
fdu . --analyze=lines --view=languages # those metrics by language
fdu . --analyze=code # standard LOC by language
fdu . --analyze=words # prose volume by document type
```
The default is `list` in `tree` format, in allocated bytes, largest first, to depth 2,
with at most ten children per directory.
Hidden and ignored entries are included.
`.gitignore` is read to label ignored shares, not to exclude matching entries.
`--analyze` chooses what may be read and `--view` chooses what is printed.
The language commands differ only on the analysis axis: without code analysis
percentages are byte shares; with it, the report adds code, comment, and blank-line
metrics and uses code-line shares.
Text labels a non-byte percentage denominator explicitly.
Use `--size apparent` when logical file lengths are wanted instead of allocated bytes.
For a bounded machine-readable tree, use:
```bash
fdu --format json --view tree --depth 2 --limit 20 PATH
```
Prefer an `fdu` already on `PATH`. When there is none, run the latest release through
`uvx`, which needs no install step:
```bash
if command -v fdu >/dev/null 2>&1; then
fdu --format json --view tree --depth 2 --limit 20 PATH
else
uvx fdu@latest --format json --view tree --depth 2 --limit 20 PATH
fi
```
To keep it installed, run `uv tool install fdu` (later `uv tool upgrade fdu`) or
`cargo install --locked fdu`.
Generated by fdu `__FDU_VERSION__`; if `fdu --version` differs, re-run
`fdu --install-skill` so this file describes the installed command.
## Compose the Request From Six Axes
Every option belongs to exactly one axis, and any axis composes with any other.
There are no subcommands: the grammar is always “report on a path”.
| Scope | What is scanned and cached? | `PATH`, `--scan-depth N`, `--one-filesystem`, `--gitignore-budget SIZE\|all`, `--gitignore-line-limit SIZE\|all`, `--no-gitignore` |
| Content | Which file bodies are read? | `--analyze none\|lines\|code\|words\|all` |
| Selection | Which entries does this query consider? | `--include`, `--exclude`, `--min-size`, `--modified-since`, `--modified-before`, `--kind`, `--exclude-ignored`, `--only-ignored`, `--depth`, `-n/--limit`, `--sort`, `--reverse`, `--size` |
| View | Which roll-up is reported? | `--view list,summary,tree,families,types,extensions,languages,documents,largest,recent,files`, or `--view full` |
| Format | How is it serialized? | `--format text\|tree\|paths\|long\|json\|jsonl\|yaml`, `--color`, `--progress` |
| Mode | How is work performed? | `--cache auto\|refresh\|read-only\|only\|off`, `--watch`, `--analysis-workers N` |
Scope versus selection is the distinction that matters: scope decides what is scanned
and cached, so one cache serves every query, while selection filters the retained index
at query time. Narrowing a selection never costs a rescan.
Work has three layers.
A single unfiltered `--no-gitignore --view summary PATH` is the one exact composition
that retains only aggregate tallies and no index, under every cache policy except `only`
and `refresh`, whose contracts are about the snapshot itself.
Under the rest a snapshot cannot save the walk that request is already doing, so it
neither reads nor writes one.
Without `--no-gitignore` the summary reads `.gitignore` to report its ignored share,
which needs the index.
Ordinary metadata requests retain the reusable index but never read regular-file
contents. One-shot metadata reports under `auto` skip loading a snapshot when it cannot
avoid the current metadata walk, though a complete indexed report may write one.
Any `--analyze` value other than `none` opts into a separate content sidecar.
fdu reads eligible files whose requested result is absent or stale; a compatible
repeated run reuses unchanged records.
Coverage is scoped to the analyzers too: an unsupported deeper analyzer leaves byte
metadata visible but does not retain a separate lower-level metric record for that file.
## Pick the View, Then Shape It
- `--view list` (default), with `--format tree` for directory roll-ups, `--format paths`
for complete flat matching paths, or `--long` for size, age, and path.
- `--view tree` for the directory hierarchy, in machine output too.
- `--view extensions` for the raw-extension breakdown.
Rows partition the tree and so sum to its total; a derived extension always carries a
leading dot, and names having none are tallied under the literal `(none)`.
- `--view types` for stable detected file types and exact byte shares.
- `--view families` for code, prose, markup, data, binary, and unknown roll-ups.
- `--view languages` for code-family rows and byte shares from path-only detection.
- `--view documents` for prose metrics; it requires any enabled analyzer.
- `--view largest` for the 20 largest regular files, and `--view recent` for the 20 most
recently modified. Both are presets over `files`, not separate machinery: `largest` is
`files --sort size --limit 20` and `recent` is `files --sort mtime --limit 20`, each
restricted to regular files, and `--sort` and `--limit` still override them.
- `--view files` for a complete flat listing: every matching entry, in name order.
One-shot text adds the performance footer described below; use a machine format when
output is consumed programmatically.
- `--view summary` for one aggregate row.
- `--view full` for the bounded digest of every view but List and Files.
- Several views in one run share one scan: `--view summary,types,families`. Text then
labels each block with an all-caps header naming its view; a single-view text report
has no header. Machine formats tag every report with `view` either way.
`--analyze` names a set of analyzers, comma-separated, from `lines`, `code`, and
`words`; `none` and `all` are totals and cannot be combined with anything else.
Anything but `none` may open eligible files and adds content work beyond the metadata
walk. Compatible cached results prevent unchanged bodies from being reread.
Add `--analyze lines` to stream physical, blank, and nonblank lines and raw word counts.
Add `--analyze code` for standard LOC, comment, and code-blank partitions across
supported common languages.
Text labels its code-line percentage denominator while retaining bytes in the first
column. Use `--analyze words` for normalized word volume, paragraphs, aggregate-derived
pages, and reader-visible Markdown that excludes destinations and code.
The `documents` percentage column is document-word share and is also labeled in text.
`--analyze code,words` — or `all` — computes both in one streaming pass.
Requesting analysis without naming a view selects one that displays it: `code` reports
`languages`, `words` reports `documents`, and either both or `lines` alone reports
`families`. Naming `--view` overrides that; a view never enables an analyzer, so a
`--view` that displays no content metric prints a note saying what was read for nothing.
`--view full` includes `documents` only when an analyzer ran, and otherwise names it as
skipped. Use `--analysis-workers` to bound concurrent reads and `--words-per-page` to
control page derivation.
Analysis never truncates a file or excludes it because of size.
Invalid UTF-8, binary data, and unsupported SLOC languages remain visible as normal
coverage outcomes. Only I/O failures, files changed during a read, or stale commits make
analysis operationally partial.
Content analysis is one-shot and cannot be combined with `--watch`.
One-shot text reports end with a compact performance line.
It reports regular files walked and their bytes in the report’s size metric, content
bytes actually read, fresh-analysis file and byte rates, content-sidecar files and
apparent bytes restored from cache, the metadata cache tier, and total report time.
Known binary files can contribute walked bytes but zero read bytes.
Cache-only runs report zero walked files because they never consult the tree.
The line is gray only when color is active and has no ANSI escapes otherwise.
Paths, Long, JSON, JSONL, YAML, skill output, lifecycle output, and watch streams omit
it.
A progress line can appear on stderr for a person at a terminal.
It is never drawn when stderr is not a terminal, `TERM` is `dumb`, or `CI` is set, so
agents need no flag; `fdu --docs` states the full rule.
Common shapes are compositions rather than dedicated flags:
```bash
fdu --view largest -n 100 PATH # the 100 largest files
fdu --kind file --modified-since 2h PATH # files changed in the last two hours
fdu --view files --include '*.{rs,toml}' PATH # by pattern
fdu --view tree --sort mtime PATH # an activity map
```
`--depth` and `--limit` bound only the rendered view; `--scan-depth` bounds what is
scanned and retained, so do not reach for it merely to shorten output.
## Find Stale Environments and Build Outputs
```bash
fdu PATH --kind dir --include .venv --modified-before 7d --long
fdu PATH --kind dir --include node_modules --modified-before 30d --long
fdu PATH --kind dir --include target --modified-before 30d --format paths
fdu PATH --kind dir --include .venv --include venv --include node_modules --include target --modified-before 30d --long --sort mtime --reverse
fdu PATH --kind dir --include .venv --modified-before 30d --format json
```
Kind, basename/relative-path patterns, size, and modification age are filters.
Repeated includes form a union.
Directory bytes sum eligible regular-file contents; recency is the newest root or
eligible descendant mtime, including directories and symlinks.
Empty directories use their own mtime.
This is modification activity, not access or last use.
Directory names such as Cargo’s `target` are conventions, not proof of ownership.
Reported bytes need not be uniquely reclaimable.
Exclusions win throughout selected subtrees before size/age bounds; ignored-only queries
traverse structural ancestors.
Flat output lists matching entries, including nested roots whose sizes overlap.
Aggregate views count the covered union once.
The default is the directory tree; `--tree` makes its format explicit.
Flat lists are complete, size-ranked by default, with global row limits; tree limits
remain per-directory and depth only folds the tree.
Paths escapes control characters only and keeps stdout to paths; bound and rule notices
go to stderr. Long adds size and signed age.
Machine rows retain exact `mtime_ns`, `age_ns`, directory `files`/`dirs`, and the
report’s `age_reference_ns`; unknown ages are null.
Tree/Paths/Long require one compatible list view.
Use automatic Text or machine output for grouped/mixed views and Full.
Largest/recent retain regular-file ranks with Paths or Long.
Explicit Paths/Long overrides the Tree view’s presentation.
Format flags conflict.
Rust/Python callers select format on the query before reading; a detached Report cannot
turn a folded tree into a complete flat inventory.
Request another report from the retained index for that change, without scanning again.
## Read What `.gitignore` Covers
Every report reads the tree’s `.gitignore` files.
Summary, tree, and extension rows end with the part of their size those rules ignore, as
`(128 B ignored)`, left off a row with no ignored file; the performance line counts the
rule files read.
```bash
fdu PATH --exclude-ignored # folders by what the rules leave
fdu PATH --view=files --only-ignored --format=jsonl # every entry the rules cover
fdu PATH --no-gitignore # read no rules, show no share
```
Selecting a side changes sizes, ordering, and `--min-size` together, because they follow
the entries shown.
It filters retained results after the scan; it does not prune metadata
work or content analysis.
`--no-gitignore` with either selection is a usage error.
Only per-directory `.gitignore` files apply, not `core.excludesFile`,
`.git/info/exclude`, or a global ignore file, and matching is case-sensitive.
Unignored does not mean tracked: `.git` is unignored unless a rule names it.
For recent working files, add both `--exclude-ignored` and `--exclude='.git/**'`. An
unreadable `.gitignore` makes the result partial (exit 2), while one past
`--gitignore-budget` or `--gitignore-line-limit` is refused whole and named in a note:
sizes stay exact, the ignored shares under that directory do not.
Under `--watch`, a rule edit that moves an entry into either selection streams the
upsert that draws it and one that moves it out streams the removal, so the stream holds
the entry set the flag names.
An upsert carries `ignored`, and so does a removal a rule edit caused; an ordinary
removal, an invalidation, and every record of a run that read no rules omit it.
Without either flag the stream maintains membership rather than each row’s bit, so
re-read a listing after a rule edit if the bit matters.
## Value Grammars
- Sizes: `512`, `10k`, `10M`, `1.5GiB`. Decimal and binary units, case-insensitive.
- Times: `now`, a compound age (`200ms`, `45s`, `2h`, `1h30m`), an RFC 3339 timestamp
with an offset (`2026-08-10T18:22:31Z`), or `@` epoch seconds.
Calendar units and fractional ages are rejected with the spelling to use instead; a
bare local date-time is rejected because resolving it needs a time-zone database.
- `--modified-since` is inclusive and `--modified-before` is exclusive.
## Use Timestamps as a Sync Watermark
Every report carries `provenance.scan_started_at`. Feeding it back selects exactly what
changed after that scan began, which is what makes incremental follow-up sound:
```bash
fdu --view summary --format json PATH # record provenance.scan_started_at
fdu --view files --kind file --format jsonl --modified-since <that> PATH
```
Use the scan’s *start*, not its end: a file modified mid-scan may have been observed
before the modification, so only the start bound is conservative.
## Validate Every Automated Result
Check the process exit status and these fields:
- `schema` before parsing anything else: a report carries `fdu.report/7`, a `--watch`
stream carries `fdu.stream/2`, and `--cache-status` carries `fdu.cache/2`. Treat an
unrecognized value as a version you cannot parse rather than guessing at the fields.
- Integer fields that exceed 2^53 (fingerprints, option hashes, nanosecond timestamps)
lose precision in IEEE 754 binary64 parsers such as JavaScript `JSON.parse`
- `status.complete`, `status.errors`, and `status.errors_omitted` before trusting totals
- `provenance.freshness` and `provenance.source` before presenting data as current
- `truncated` on a tree node before treating it as exhaustive
- Each requested unit in a metric row’s `coverage` before presenting its metrics as
complete
- `ignored` on a row before calling anything ignored or not: an object, or `true` and
`false` on a file row, where rules were read, and `null` where none were, which never
means nothing is ignored
- `ignore_rules.refused` before trusting an ignored share: below a refused `.gitignore`
the split is not exact, though sizes are
- `detection.sources`, `detection.confidence`, and `detection.flags` before treating a
deep-detected type or origin label as exact
`provenance.source` is `cold_scan`, `warm_revalidate`, or `cache_only`. Only
`--cache only` can return `provenance.freshness: stale`, and it says so rather than
implying currency; it fails outright when no usable snapshot exists rather than silently
scanning.
Exit 0 is accepted success, exit 1 is a fatal failure, and exit 2 is incomplete data or
invalid usage. Do not discard useful stdout from exit 2; inspect the completeness fields
and use `--allow-partial` only when incomplete totals are acceptable.
## Cache Behavior
No ordinary view requires a preexisting cache.
Metadata-only one-shot reports include current sizes or timestamps, so they must inspect
every entry; under `auto` they skip loading a snapshot when it cannot make that
verification cheaper.
A complete indexed scan may still write one.
Content analysis is where ordinary repeated runs benefit most.
The first compatible run reads eligible bodies; a later run restores unchanged records
from the content sidecar and reads only changed or newly eligible files.
A stored analyzer set answers only the same set: a different one, wider or narrower,
reads the files again and replaces it.
The performance footer reports fresh and cached analysis separately.
`--cache=only` is a distinct contract: it never verifies the source tree, labels the
answer stale, and fails unless compatible metadata and any requested content analysis
already exist. `--cache=off` neither reads nor writes fdu cache data.
The snapshot is one file per root under the user cache directory.
`--cache-status` maps a hash-named file back to the tree it describes, and
`--cache-clear` removes it; both run without scanning.
Cache status is its own document, carrying the `fdu.cache/2` schema in every machine
format rather than a report schema.
A current snapshot’s row carries the `identity` of the entry and `.gitignore` tiers it
holds, and every row a `content` object for the sidecar beside it, with its own `state`
and a current sidecar’s `identity`, or `null` when there is none.
Each status row carries a `state`: `current`, `stale` for a snapshot another fdu version
wrote or one this build cannot read, `leftover` for a file fdu left behind, with a
`leftover_kind`, `unrecognized` for a file that is not fdu’s, or `absent`. Clearing
removes current and stale snapshots, so `--cache-clear=all` reclaims what an upgrade
leaves behind; it also reclaims leftovers, and it never removes an unrecognized file.
All shipped metadata views need current per-file attributes because they report sizes or
timestamps. An in-place edit changes no directory timestamp, so a cached directory
fingerprint cannot prove those answers current.
Content reuse still pays because an unchanged file fingerprint avoids the much more
expensive body read and analysis.
Exact names and ordinary extensions remain path-only classifications.
When analysis is enabled, unresolved files and ambiguous `.h` headers may use bounded
shebang, modeline, literal, or signature probes.
Do not collapse their provenance into an unqualified language claim; retain the report’s
source and confidence fields when summarizing or transforming machine output.
Run `fdu --help` for the complete flag, cache, color, scope, and exit contract.