file-engine 2.2.0

Async, cross-platform file operations engine for desktop apps and developer tools: copy, move, sync, watch, and compress files with progress reporting and cancellation.
Documentation
# Changelog

All notable changes to this project are documented here.

This project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [2.2.0]

### Changed

- **`analyze()`'s tree walk is now multithreaded.** Replaced `walkdir`
  with `jwalk` under the `analyze` feature: directory reads and `stat()`
  calls now run across a worker pool instead of one at a time on a
  single thread. Aggregation (the largest-files heap, extension/MIME
  stats, age buckets) still happens serially on the consuming task, so
  no locking was introduced there — only the underlying directory
  traversal is parallelized. On a 330k-file/138k-dir local tree this cut
  wall time from 12.8s (single-threaded, 45% CPU) to 9.0s (194% CPU)
  with a warm disk cache; the gap widens further on cold caches or
  network filesystems, where per-`stat()` latency (not local CPU) is the
  bottleneck. `profiler::scan` (feature `operations`) is unaffected — it
  still uses `walkdir`.

### Added

- **`AnalyzeBuilder::walk_concurrency(n)`** — worker-thread count for the
  parallel walk, defaulting to `available_parallelism()`, matching
  `.hash_concurrency()`.

## [2.1.0]

### Added

- **`FileEngine::analyze()`** — read-only tree inspection: file/directory
  counts, total size, the largest N files, size and count grouped by
  extension, an age histogram (last-modified bucketed into
  under-a-day/week/month/year/older), plus filters (extension,
  glob-based excludes, size range, modified-time range, max depth,
  symlink-following) so a caller narrows what gets counted rather than
  filtering a full report after the fact. Feature `analyze`, on by
  default — it was already reserved in `Cargo.toml` but unimplemented
  until now.

  Follows the same builder shape as every other operation
  (`FileEngine::analyze(path)...start()`), but returns a dedicated
  `AnalysisHandle`/`AnalysisProgress` pair rather than the existing
  `Handle<T>`/`Progress`: those are hard-coded to each other, and both
  live behind the `operations` feature, which `analyze` deliberately
  doesn't require — the same reasoning `WatchHandle` already established
  for `watch`.

  Excluded directories (`.exclude_globs()`, backed by `globset`) prune
  traversal itself rather than being filtered out after a full walk, and
  `max_depth` passes straight through to `walkdir`'s own depth-limiting
  for the same reason. A per-entry error (a permission-denied
  subdirectory, a symlink loop under `.follow_symlinks(true)`) is
  handled per `.on_error(AnalysisErrorStrategy)`, distinct from
  `planner::ErrorStrategy` since analysis never writes anything and
  `Undo` has no meaning here.

- **`.detect_mime_types(bool)`** (feature `analyze`) and
  **`.detect_duplicates(bool)`** (feature `checksum`) on
  `AnalyzeBuilder` — both off by default, since either turns the walk
  from metadata-only into a full extra read per matched file. Duplicate
  detection groups candidates by size first (two files can only be
  byte-identical if they're the same size) and only blake3-hashes files
  that actually collide, with hashing bounded by `.hash_concurrency(n)`
  (same `Arc<Semaphore>` pattern the batching pipeline already uses for
  its own worker pool).

- **`AnalysisReport::errors`/`errors_total`** and (feature `checksum`)
  **`AnalysisReport::duplicates`/`duplicate_groups_total`/
  `duplicate_bytes_wasted`** — the detailed lists are capped
  (`.max_reported_errors()`, `.max_reported_duplicates()`) so a badly
  permissioned or heavily duplicated tree can't grow the report
  unboundedly, but the counts and `duplicate_bytes_wasted` stay uncapped
  and accurate even once the sample is truncated — the wasted-space
  figure in particular is summed over every duplicate group found, not
  just the ones that made it into the capped list.

- **`FE_INVALID_GLOB_PATTERN`** in `errors.toml`, backing
  `Error::InvalidGlobPattern` — an invalid pattern passed to
  `.exclude_globs()`.

### Changed

- The README's feature table no longer describes `analyze` as
  unimplemented.

### Fixed

- Nothing user-visible; 2.1.0 is additive.

## [2.0.0]

### Upgrading from 1.x

Every public output type is now `#[non_exhaustive]` — `Progress`,
`StopReason`, `OperationOutcome`, and `SyncOutcome`. This release absorbs
the churn so that later additions to any of them are not breaking
changes. What that means for your code:

`Progress` gained a variant. If you `match` on it, add a wildcard arm:

```rust
match progress {
    Progress::Started { .. } => { /* ... */ }
    Progress::EntryCompleted { .. } => { /* ... */ }
    // ...
    _ => {}
}
```

`Progress` is now `#[non_exhaustive]`, so this arm is required — and this
is the last time adding a variant will break you.

`OperationOutcome` gained a `duration` field and is now
`#[non_exhaustive]`. Reading it is unaffected, as is
`OperationOutcome::default()`; exhaustive destructuring needs a trailing
`..`. Building one with a struct literal is no longer possible outside
the crate at all — including with `..Default::default()`, which
`#[non_exhaustive]` also blocks. It is an output type, so this should
only affect test fixtures; construct them from `Default::default()` and
assign the fields you need.

`SyncOutcome` is `#[non_exhaustive]` on the same terms — read `copy` and
`delete` as before, but build one via `SyncOutcome::default()` rather
than a struct literal.

`StopReason` is `#[non_exhaustive]`, so a `match` on it needs a `_` arm,
exactly like `Progress`.

Nothing else changed: every builder, `Handle<T>`, `Error`, and every
feature flag keep their 1.x behaviour and signatures.

### Added

- **`EtaEstimator`** — predicts time remaining from the `Progress`
  stream. Feed it every event with `.observe()` and read `.estimate()`,
  which returns `Option<Duration>` (`None` until there is something
  measured to extrapolate from). Purely observational: no I/O, no tasks,
  no reference to the running operation, and no cost if unused.

  It models three cost regimes separately, because a single
  bytes-per-second figure describes none of them: batched small files
  cost per *file*, streamed large files cost per *byte*, and the
  directory pre-pass costs per *directory* and isn't counted in
  `Started`'s `bytes_total` at all.

- **`Handle::elapsed()`** — wall time since the operation was spawned,
  the counterpart to `EtaEstimator::estimate()` for a UI showing elapsed
  next to remaining. Covers the whole run including the directory
  pre-pass, so it won't match a timer started on the first `Progress`
  event. Freezes once the handle has been polled to completion.

- **`OperationOutcome.duration`** — the same figure after the handle is
  gone, since `handle.await` consumes it. `SyncOutcome`'s two outcomes
  are timed per phase and so don't sum to the whole run (the diff
  preceding them is in neither); a delete phase skipped because the copy
  phase stopped early reports `Duration::ZERO`.

- **`Progress::Planned`** — the workload split (directories, small
  files/bytes, large files/bytes, and the threshold that separated them),
  emitted once per phase before the directory pre-pass. Earlier than
  `Started`, which is emitted after that pre-pass and counts only file
  entries. Not emitted by the delete sweeps, which are metadata-only.

- **`Progress::EntryProgress`** — cumulative bytes written for a large
  entry still in flight, sampled from the destination file every 250ms.
  `tokio::fs::copy` is opaque while it runs, so without this a lone large
  file emitted nothing between `EntryStarted` and `EntryCompleted` — its
  transfer rate was unmeasurable for exactly as long as the copy took. A
  copy the filesystem satisfies by copy-on-write finishes before the
  first sample and emits none, which is correct: there was nothing to
  wait for.

### Changed

- **`Progress` is `#[non_exhaustive]`.** See "Upgrading" above.

- **`OperationOutcome`, `SyncOutcome`, and `StopReason` are
  `#[non_exhaustive]`**, so future fields and variants land without a
  major version. See "Upgrading" above.

- **Dispatch order: the smallest large file now runs first.** Previously
  every batch was queued ahead of every stream, so on a workload of many
  small files plus a few large ones, no large file completed until the
  operation was nearly over — and with it, no per-byte transfer rate was
  observable. On a 2.7GB test tree the first stream completed at 95%
  elapsed. Everything behind that first stream is unchanged. This affects
  any consumer of `Progress`, not just time estimation. Applies to
  `copy`/`move`/`sync` and to `compress`'s own pipeline (where archive
  entry order changes as a result — immaterial, since zip entries are
  addressed by name, and order already varied with concurrency).

- **`tokio`'s `time` feature is now enabled.** Required by the
  destination sampling above. If you build a runtime without the time
  driver, sampling is skipped and the copy still succeeds — you simply
  get no `EntryProgress` events.

- `repository` in `Cargo.toml` now points at the real URL.

### Fixed

- Nothing user-visible; 2.0.0 is additive plus the `Progress` break.