safe-chains 0.230.6

Auto-allow safe bash commands in agentic coding tools
Documentation
# Fuzzing: what is specific to safe-chains

General fuzzing method — the two budgets, the seeding ladder, coverage, CI shape, the gotchas — is
in the `rust-fuzzing` skill. This section is only what is specific to safe-chains.

- **Targets** (`fuzz/fuzz_targets/`; each file's header states its properties in full):
  - `parse` — `is_safe_command` on arbitrary bytes: never panics, never hangs, across parse → CST →
    engine → handlers. It **discards the verdict**, so it finds availability bugs only.
  - `equivalence` — a semantics-preserving respelling (flag forms, env-var forms) cannot change the
    verdict.
  - `hook_envelope` — the `targets/*` hook I/O never panics and never emits a grant it was not
    asked for.
  - `explain_render` — the explanation describes the verdict that was enforced, and a command
    cannot forge a marker line or smuggle a control/bidi character into it.
  - `suggest_roundtrip` — every config `--suggest` generates parses back, and merging never drops
    what was there.
  - `level_monotonic` — a stricter level never approves what a looser one refused.
  - `config_load` — a repo `.safe-chains.toml` (the one attacker-placed input) never aborts the
    loader, loads NOTHING when it is not valid TOML, and loads deterministically.
  - `setup_merge` — `--setup` into an arbitrary existing settings file: on refusal the file is
    byte-identical, on success it still parses.
  - `path_admit` — the credential shield outranks every package-content read admit, and those
    admits never grant a write.
  - `gate_prefilter` — a declared path gate never skips a value its own judge would refuse.
- **Not fuzzed:** `docs.rs` (no security property) and `pathctx` (proptests cover it and shrink
  better for a pure function).
- **CI:** `fuzz-replay.yml` replays every target's cached corpus (`-runs=0`) on every push and PR —
  the gate. `fuzz.yml` is a non-gating burst on each push to `main` (and `workflow_dispatch`, which
  takes a larger `max_total_time` and also renders the `parse` coverage report): one job per target
  running `fuzz/burst.sh`: 180s of single-process mutation timed from when the corpus has loaded
  (libFuzzer's `-max_total_time` counts the load, which on a runner can exceed the whole budget),
  then `-merge=1` into the corpus, saved as `fuzz-corpus-<target>-<run>` only when the merge
  completed, which is the prefix the replay restores. The burst's check fails on a crash-,
  timeout- or oom- artifact, or on a burst step that failed without one. `tests/fuzz_targets_wired.rs`
  keeps both matrices equal to `fuzz/Cargo.toml`, holds the cache prefix and those two conditions
  in step, and refuses a `schedule:` trigger. Both fuzz workflows skip docs-only pushes
  (`**.md`, `docs/**`): nothing fuzzed reads either.
- **`fuzz/burst.sh` runs the same locally** (`fuzz/burst.sh <binary> <target> <seconds>`, from the
  repo root). Finds go to `fuzz/new/<target>`, so bursting targets one after another never merges
  one target's finds into the next; the committed `seed-*` inputs keep their names through the
  merge; `fuzz/dict/<target>.dict` is used when present. A timeout reported after its SIGINT is
  libFuzzer stalling on the way out (2026-10-05, `suggest_roundtrip`): the unit is replayed alone
  and dropped if it finishes. `tests/fuzz_burst.rs` holds all of this against a stand-in libFuzzer.
- **No nightly (retired 2026-09-26).** It ran 05:00 UTC: three 5h shards on `parse` plus 1h on each
  property target. Every real find came in a target's first days; after the first week of August it
  ran seven weeks without another, and its later red nights were job timeouts (`config_load`
  overrunning `timeout-minutes`), not findings. The burst fuzzes new code the day it lands, which is
  when finds happen. For a deeper run after a large change, dispatch `fuzz.yml` with a bigger budget.
- **Seeding is registry-derived.** `src/bin/gen_fuzz_corpus.rs` (feature `fuzz-gen`,
  `cargo run --bin gen-fuzz-corpus --features fuzz-gen`) reads `commands/**/*.toml` and emits the
  `examples_safe`/`examples_denied` invocations as seeds and the command/subcommand/flag vocabulary
  as a `-dict` for `parse`. New commands are covered automatically — **no hand-maintained fuzz
  corpus.** The burst's build job regenerates them each run and the `parse` burst merges them in.
  Generated `gen-*` seeds and `fuzz/dict/` are git-ignored. The property targets take other input
  domains (a flag value, an envelope, a config file), so the registry seeds do not apply to them.
- **Regression seeds are the one hand-kept input.** A minimized find is committed as `seed-*` in
  its target's corpus (`.gitignore` lets `seed-*` through for `parse` and `explain_render`), so the
  replay holds it down. `explain_render/seed-unclosed-brace-nest` is the 2026-09-26 timeout: an
  unclosed `{` nest was parsed once as a brace group and again as a command named `{`, doubling per
  level. The parser's guards for that class count work instead of timing it (`cst/budget.rs` holds
  the entry and step budgets; the tests in `cst/parse.rs` assert work linear in nesting and walk
  every committed command seed), so they hold on a loaded machine.
- **Measured coverage** (region, authored source): mutation corpus alone ~26%, registry seeds alone
  ~37%, **combined ~61%**. The two are complementary — seeds unlock the per-command grammars,
  mutation covers parser byte-paths.
- **Two kinds of coverage hole, and what each means:**
  - **Un-exampled commands** — a handler at ~0% (e.g. `glab`, `magick`, `sysctl`) almost always has
    NO `examples_safe`/`examples_denied`, so nothing seeds it and byte mutation never synthesizes
    `glab mr list`. Adding examples is the **cheapest coverage win** and doubles as a
    `toml_examples_match_dispatch` guard. Prefer this before anything fancier.
  - **Out-of-scope layers** — `targets/*` (hook-envelope I/O), `docs.rs`/`registry/docs.rs`, and
    `suggest.rs` sit at ~0% because `is_safe_command` never calls them. They are not blind spots of
    the `parse` target; reaching them needs a **new** target (a hook-envelope target for
    `targets/`). Don't chase these by seeding `parse`.
  - **`cst/explain.rs` + `cst/display.rs` no longer need a target of their own.** The `parse` fuzz
    target still does not reach them — it calls `is_safe_command` — so their FUZZ coverage is
    unchanged and this bullet is not a claim about the corpus. What changed is that the two
    property guards in `handler_property_tests.rs`, which already generate adversarial command
    strings, now also call the renderer.
    `arbitrary_command_strings_never_panic` and `classifier_terminates_on_adversarial_input` now
    run `explain(&line).render()` alongside `command_verdict`, so the renderer is held to the same
    never-panics / never-hangs / deterministic contract. That matters because the layer is not
    merely cosmetic: `--explain` is user-facing AND renders the hook's injected context, where a
    panic is a crash — which for a PreToolUse hook fails OPEN. Extending a corpus beat building a
    target; check that before reaching for new machinery on the layers above.