saccade
saccade checks whether image captures and timings are good enough to support
a claim about a change. compare measures image changes, prove checks exact image identity
or performance claims, and review prepares a decision for a human. Captures
are paired by name and scored with NVIDIA FLIP. Each run writes an offline HTML
report and JSON evidence. Baselines change only when a human approves.

Install
Each release has four archives and a SHA256SUMS file:
saccade-x86_64-unknown-linux-gnu.tar.gz, saccade-aarch64-unknown-linux-gnu.tar.gz,
saccade-aarch64-apple-darwin.tar.gz and saccade-x86_64-pc-windows-msvc.zip.
Download the one for your system from
Releases, check it, extract it,
and put saccade on your PATH:
On macOS use shasum -a 256 -c SHA256SUMS --ignore-missing.
No Rust toolchain is needed. Each archive includes the licences and
third-party notices. See releasing for how releases are
built and checked.
To build from source with Cargo (Rust 1.88 or newer):
From a clone, cargo install --locked --path crates/saccade does the same.
Add --features prechecks for the experimental safety and accessibility checks.
Check the install with saccade --version or saccade doctor --json.
Quickstart (60 seconds)
The demo writes three reports and exits with status 1 on purpose, because it
includes a moved shadow, a changed UI label and a missing capture. view
prints the page to open, saccade-demo/report/index.html. To compare your own
images:
Exit status 0 means no image regression, 1 means at least one image failed,
and 2 means the command could not run. The result is in report/index.html.
saccade does not use the network, a GPU or an account for any of this.
Compare files or directories
Both comparisons exit 1 because the example captures differ. The inspect and export commands exit 0 when they finish. The default metric is mean FLIP, the threshold is 0.01, and the viewing setting is 67 pixels per degree. Passing the threshold only means the configured criteria were met. It does not mean nobody would notice the change, or that the renderer is correct.
Measurement settings, regions, masks, capture requirements and declared changes
go in saccade.toml. Options given on the command line override the project
values. inspect config shows the effective settings and where each came from.
Images are paired by relative name, and new or missing images fail by default. A comparison with no images proves nothing. If captures carry metadata, saccade can refuse to compare images made under different settings; without metadata, comparability is left as unknown. See capture configuration.
Optimization identity
This uses the demo from the quickstart:
Exit 0 establishes exact native decoded-sample equality for the selected images. The demo's PNG files have different bytes but equal samples. The dimensions, sample type and channel interpretation must match, and every selected pair must exist and decode. Perceptual tolerances, masks and region-only acceptance are not accepted as proof of identity.
The proof only covers the captures supplied and the entries selected. Without capture metadata, saccade can still show that the samples are equal, but whether the captures are comparable stays unknown. Identical images on their own do not show that anything got faster. See identity and performance.
Optional model review (bring your own key)
Start with a local preview. These commands do not call a provider:
The preview writes the exact payloads to review-plan/requests.json, with
estimated tokens and, if you configure model rates, an estimated cost.
The request uses the file report because every image in it is paired; a
question cannot be created while required facts are missing. Each request
identifies the exact evidence it refers to. review propose checks answers
against a request and records them as advice. It cannot approve anything.
To call a provider, you configure the endpoints and dedicated credentials in
your user configuration, allow egress for the source roots, and run
review REPORT --run --budget-calls N yourself. Egress is denied for any root
you have not classified. A project's files can pick from approved models and
lower budgets, but cannot grant network, credential, path or approval
authority. Retries and fallbacks all spend from the same attempt budget.
See review for the workflow and limits.
GitHub Action
- uses: Tavrin/saccade@v1
with:
baseline-dir: tests/baseline
capture-dir: artifacts/captures
The v1 tag has not been published yet; the snippet shows the intended
syntax. By default the action installs a release binary and verifies its
checksum. Set install-mode: source to build from source instead. The report
and JUnit results are uploaded even when the comparison fails. Comparisons on
fork pull requests need only read permission and no AI keys.
The verdict, exit code, report URL and immutable artifact ID are separate outputs. Baseline updates run only from a trusted dispatch tied to the selected, reviewed content, and open a pull request for review that is never merged automatically. See CI and Action examples.
For AI agents
Start with the agent guide. It explains the bounded JSON
results (verdict, worst, next_actions), how to page through failures,
and what an agent may not do. Ready-made packs are generated from
one guide: a
Claude Code skill and
Codex instructions.
The MCP server has six tools. It reads only from the --root directories
(read-only) and writes only under --out-root. Provider calls stay off unless a
human starts the server with --allow-provider-calls and a positive
--budget-calls. None of the tools writes baselines. Agents must not relax a
threshold or update a baseline to make a task pass.
Reproducible showcases
9 reproducible cases with commands, expected exits and measured output. Pages gallery.
The datasets are procedural, and simulated timings are labeled illustrative.
Generating them needs Python 3, Pillow and numpy, but no network, GPU or
provider. Validating the showcases needs a build with --features prechecks:
Put the built binary on PATH before running the showcase script. The script
checks exit codes and compares output against the measured EXPECTED.txt
transcripts. It never accepts changed output on its own.
Status
saccade 0.1.2 fixes the case-sensitive MCP Registry namespace.
- Stable:
compare,identity,approve,view,inspect,init,noise,demo,serve, the localreviewpreview and request commands, the MCP server, the HTML and JSON reports, exit codes, and the GitHub Action inputs. The active formats are listed in contracts. Scripts should gate ondoctor --jsoncapability names (list). - Experimental, and their output may change: the
experimentcommands (ablation, sequences, ranking, bisection, and the safety and accessibility prechecks) andreview eval. - AI review accuracy is not yet qualified. The pilot evaluation has fewer labeled cases per question than qualification requires, so we make no claim about how accurate any model's answers are, or which model does better. Treat every model answer as a proposal for a human to check. See evaluation.
Local tests do not cover native installation on each platform or the behaviour of live providers; those need their own checks.
Limits
saccade only measures the captures it is given. It does not run your renderer or schedule captures, it cannot tell when your application is ready to be captured, and it does not treat equal images as proof of correct output. FLIP scores depend on viewing conditions, and HDR display transforms and numerical buffers have to be interpreted explicitly. When repeat noise is missing it is treated as unknown rather than zero. A performance claim needs qualified timing, repeat noise and attribution.
A model cannot establish equality, qualify timing, approve a baseline, or turn a tie into an acceptance. If a model's answers contradict each other across the two blind orderings, or the intent is ambiguous, the item stays unresolved. For an independent blind review, the reviewer must not have the mapping or the implementation context; whoever created the private key cannot review blind.
"Human-final" approval is a policy and an audit record. It does not
authenticate anyone: a shell agent with unrestricted access can still run
saccade approve. Workbench receipts
record a scoped token attestation; CLI receipts record human_attestation: null.
Documentation
- Captures and identity/performance
- Review, agents, CI
- Contracts and generated schema index
- Evaluation, command reference, experimental checks
- Architecture, contributing, changes
License
saccade is licensed under either of MIT or
Apache-2.0, at your option. FLIP comes from the pure-Rust
flip-rs port of NVIDIA FLIP, which is
BSD-3-Clause. See third-party notices; release archives
also carry a generated THIRD_PARTY_NOTICES.md covering every dependency.
MCP Registry name: mcp-name: io.github.Tavrin/saccade