AI coding agents spend most of their input tokens on tool output: build logs, test runs, git status, the same file read again after every edit. All of it goes into the context window raw, and it is re-sent on every turn after that.
sqz compresses tool output before it enters the context. Per-command formatters keep what an agent acts on (the failing test, the assertion, the file:line) and drop what it does not. Anything the model already has in context comes back as a short reference instead of the content. Source code passes through unchanged, and so do stack traces, migrations and credential lines.
Real output from the release binary (assets/demo.tape). Regenerate with vhs assets/demo.tape.
Without sqz: With sqz:
cat auth.py (200 lines): 1,300 tokens 1,300 tokens (source is served in full)
cargo test (1 failure): 366 tokens 185 tokens (failure, assertion, file:line)
sed -n '41,80p' auth.py: 260 tokens 35 tokens (§ref:…:L41-80§ to the read above)
git status, unchanged: 166 tokens 37 tokens (§ref:…§)
───────────────────────────────────── ─────────────
Total: 2,092 tokens 1,557 tokens (26% saved, nothing dropped)
Single Rust binary, deterministic, no LLM calls, works offline. Every compressed result can be recovered byte-exact with sqz expand, except output that mentions credentials, a stack trace or a migration, which sqz never caches.
[!NOTE] Name disambiguation: this repo, ojuschugh1/sqz, is an independent project and is not affiliated with any other similarly named or working tool or other compression projects that shorten "squeeze". If you installed
sqz/sqz-cli/sqz-mcpfrom crates.io, npm, PyPI, or Homebrew, it comes from this repository.
Token Savings
24.7% average reduction across 3,003 real compressions · 13-token refs for output the model already has · 100% of agent-critical facts kept in the quality benchmark · 0% on source code, and on stack traces and credential lines
One developer's week, measured from actual sqz gain output:
$ sqz gain
sqz token savings (last 7 days)
──────────────────────────────────────────────────
04-13 │ │ 2,329 saved
04-14 │ │ 0 saved
04-15 │███ │ 12,954 saved
04-16 │██ │ 9,223 saved
04-17 │████ │ 14,752 saved
04-18 │██████████████████████████████│ 105,569 saved
04-19 │████████ │ 30,882 saved
04-20 │█ │ 4,334 saved
──────────────────────────────────────────────────
Total: 3,003 compressions, 178,442 tokens saved (24.7% avg reduction)
Per-command compression
Single-command compression (measured via cargo test -p sqz-engine benchmarks):
| Content | Before | After | Saved |
|---|---|---|---|
| Repeated log lines | 148 | 39 | 74% |
| Large JSON array | 259 | 111 | 57% |
| JSON API response | 64 | 53 | 17% |
| Git diff | 61 | 54 | 11% |
| Prose/docs | 124 | 122 | 2% |
| Stack trace (safe mode) | 82 | 82 | 0% |
Session-level with dedup
Anything the model already has in its context comes back as a reference. Measured with the release binary against a throwaway database, using the fixtures in sqz/tests/quality_bench.rs and demo/, in cl100k tokens of everything the agent sees (the reference plus sqz's one-line note on stderr):
| Repeat | Without sqz | With sqz | Saved |
|---|---|---|---|
Unchanged git status run again |
166 | 37 | 78% |
| Lines 41-80 of a 200-line file read earlier | 260 | 35 | 87% |
Same cargo test output, nothing changed |
366 | 39 | 89% |
A word on what actually repeats. A user who counted 92 of their own Claude Code sessions (script) found identical whole-file re-reads in 1 of 542 reads, and line-range re-reads in 8.4% of them. So sqz does not lead with "the same file read five times": the repeats that happen in practice are unchanged command output, line ranges of a file already read, and files re-read after a small edit, and each has its own reference type (details). Sessions with noisy command output and repeated commands see the biggest wins.
Install
Prebuilt binaries (no compiler required — works on every platform):
# macOS / Linux
|
# Windows (PowerShell)
|
# Any platform via npm
# macOS / Linux via Homebrew
Build from source via Cargo:
sqz-cli provides the sqz binary; sqz-mcp provides the MCP server. sqz-engine is a library dependency — it compiles automatically and does not need to be installed separately.
Build from source (cargo install sqz-cli) works too, but needs a C toolchain:
- Linux:
build-essential(apt) or equivalent - macOS: Xcode Command Line Tools (
xcode-select --install) - Windows: Visual Studio Build Tools with the "Desktop development with C++" workload. Without these,
cargo installfails withlinker link.exe not found. If you don't already have them, use the PowerShell or npm install above instead.
Then initialize:
# or
--global writes to ~/.claude/settings.json (the user scope per the
Anthropic scope table),
so the sqz hook fires in every Claude Code session on this machine. This is
the common case on first install. Your existing permissions, env,
statusLine, and unrelated hooks in ~/.claude/settings.json are
preserved — sqz merges its entries rather than overwriting.
The entries are its hooks and two allow rules, Bash(set -o pipefail) and
Bash(sqz compress --cmd *), which cover only the parts its rewrite adds.
sqz does not approve commands; your own permission rules still decide.
When Claude Code asks before a rewritten command, or denies it, its message
quotes the rewritten command, such as set -o pipefail; git tag v0 2>&1 for
git tag v0.
Plain sqz init (project scope) is useful when you want sqz active only
inside one repo. If ~/.claude/settings.json already has the sqz hooks, it
leaves them out of the project file, since they already run in every project.
After installing, sqz doctor checks the whole chain — binary, database,
shell hook, which clients are detected vs actually routed through sqz, and
whether anything was compressed recently — and prints the fix for each gap
it finds. sqz doctor --fix applies those fixes. After upgrading sqz, run it
once: it brings hooks and guidance an older version wrote up to date, removes
hook files current sqz no longer uses and a project copy of hooks that
~/.claude/settings.json already has, points hooks whose sqz binary is gone
at the running one, and notes any hook that still runs a different sqz
binary. sqz uninstall removes only sqz's own entries, deletes a file only
when nothing else is left in it, and ends by listing the files it kept.
Without a terminal to answer its prompt it needs --yes, like sqz reset.
See what sqz would have saved, before installing any hook
If you already use Claude Code or Kiro, your transcripts are on disk. Replay them through the sqz engine and get a counterfactual estimate on your own sessions, computed locally with nothing uploaded. From the author's last week of Kiro IDE work, trimmed to the summary:
)
Each output is replayed the way sqz could actually have reached it on that
client. Kiro can't let sqz rewrite a tool call, so the first total is zero
there; on Claude Code the shell rows split into commands the hook rewrites
on its own and ones it skips. The second total assumes the agent pipes shell
output through sqz compress, reads files with the sqz MCP tools (lossless,
so only repeats count) and runs MCP servers behind sqz-mcp proxy, and the
report says how much of that is the proxy cutting long strings and arrays on
first pass. Web fetches, edits and sub-agent results are out of reach on every
client and are counted, not replayed.
It scans ~/.claude/projects and ~/.kiro/sessions (Kiro CLI and IDE), or
point it anywhere with --transcripts PATH. Each transcript replays against a
fresh throwaway cache, so dedup references never pretend to span sessions,
and the report says "estimate" because that's what it is: these outputs were
not compressed at the time.
Only using one agent? Pass --only (or --skip) to limit which
configs are written:
Accepted names: claude, cursor, windsurf, cline, gemini,
kiro, opencode, codex. Aliases (claude-code, gemini-cli, roo,
kiro-cli) also work. --only and --skip can't be combined.
Manual installation (preserve comments in your config)
sqz init round-trips your config file through a JSON parser to merge
the sqz entry, which drops any comments in your opencode.jsonc (and
the analogous JSON-with-comments files other tools accept). If you've
commented your config carefully and want to keep them, install by hand
instead.
OpenCode — two steps:
-
Drop the plugin file in place.
sqzprints the generated TS to stdout so you don't have to hand-write the path-escaping logic: -
Add the MCP entry to your existing
opencode.jsoncyourself. Append this block inside the top-levelmcpobject (create themcpobject if it doesn't exist):"sqz": { "type": "local", "command": ["sqz-mcp", "--transport", "stdio"], "enabled": true }
Comments in the rest of your file stay put. OpenCode auto-discovers
the plugin file; no plugin array entry needed (adding one causes
double-loading, see issue #10).
Other tools: Claude Code and Gemini CLI use plain JSON configs
without comment support, Cursor, Windsurf and Cline get markdown rules
files, and Codex gets a TOML entry in ~/.codex/config.toml that keeps
your comments, so the automated path is non-destructive there. Use
sqz init --only <tool> for those. An existing .gemini/settings.json
gets sqz's hook merged in; if it has comments or doesn't parse, sqz init
leaves it alone and prints the entry to add by hand.
Run sqz doctor afterwards to see what was configured for each client
and which integration tier it is on.
How It Works
On transparent clients (Claude Code, Gemini CLI, Copilot CLI, OpenCode) sqz installs a hook that rewrites shell commands so their output goes through sqz before it reaches the AI tool. On advisory clients the agent is told to pipe output through sqz compress or use the sqz MCP tools. The table below gives each client's tier.
Claude → git status → [sqz hook rewrites] → compressed output
What gets compressed:
- Shell output: per-command formatters (git, cargo, npm/pnpm/yarn/bun, pytest, ruff, mypy, go, docker ps/images, kubectl describe/apply, terraform, gradle, maven, xcodebuild, adb logcat, dotnet, grep/rg, tree, curl, and more)
- JSON: strips nulls and empty values, cuts strings over 500 characters, compact encoding, TOON format
- Logs — collapses repeated lines
- Test output — shows failures only (state-machine parsers for Rust, Go, Python, JS, JVM)
What doesn't get compressed:
- Source code: served in full on first read
- Stack traces, SQL migrations, PEM blocks and credential lines: routed to safe mode and passed through unchanged at any size, and never cached; on the shell hook a per-command formatter may still summarize the rest of a test run around them
- Your prompts and the AI's responses — controlled by the AI tool, not sqz
Supported Tools
| Tool | Integration | Setup |
|---|---|---|
| Claude Code | PreToolUse + PostToolUse hooks + MCP server (transparent) | sqz init |
| Cursor | .cursor/rules/sqz.mdc guidance (advisory); MCP server added by hand |
sqz init |
| Windsurf | .windsurf/rules/sqz.md guidance (advisory); MCP server added by hand |
sqz init |
| Cline | .clinerules guidance (advisory); MCP server added by hand |
sqz init |
| Gemini CLI | BeforeTool hook (transparent) | sqz init |
| Kiro | Steering + MCP server (advisory) | sqz init |
| OpenCode | TypeScript plugin + MCP server (transparent) | sqz init |
| Codex CLI | AGENTS.md guidance + MCP server (advisory) | sqz init |
| Zed | AGENTS.md guidance + MCP server (advisory) | sqz init |
| Copilot CLI | preToolUse hook (transparent), when ~/.copilot exists |
sqz init |
| Copilot coding agent (CI) | Repo-level hook, self-bootstraps in the cloud sandbox | sqz init --ci + commit |
| Any CI pipeline | setup-sqz action — compress logs before an LLM step | uses: ojuschugh1/setup-sqz@v1 |
| Any MCP server | sqz-mcp proxy wraps it, compresses its tool results |
see below |
| VS Code | Extension | Install from Marketplace |
| JetBrains | Plugin | Install from Marketplace |
| Chrome | Browser extension | ChatGPT, Claude.ai, Gemini, Grok, Perplexity |
| Firefox | Browser extension | Same sites |
Transparent clients have a hook that routes shell commands through sqz on
its own. On advisory clients sqz saves tokens only when the agent pipes
output through sqz compress or calls the sqz MCP tools; the client's
built-in shell and file tools bypass it. sqz doctor shows each client's
tier.
Compress Any MCP Server
sqz-mcp proxy sits between your agent and any stdio MCP server and compresses
what flows back: tool results go through the full sqz pipeline (dedup refs on
repeats, safe-mode for stack traces and secrets, error results untouched), and
verbose tool descriptions get compacted so tools/list stops eating your
context window. The proxy injects an sqz_expand tool so the agent can recover
any cached original byte-exact; results that mention credentials, a stack trace or
a migration are never cached.
Wrap a server by prefixing its command in your MCP config:
Flags: --lazy-tools shortens every tool description to one sentence and
injects an sqz_tool_help tool that serves the full original docs on demand
(some MCP servers spend 10-40k tokens on tools/list alone); --no-desc
keeps descriptions verbatim; --no-cache stores nothing, so it sends no
dedup refs and adds no sqz_expand tool. Works with
every MCP client (Claude Code, Cursor, Windsurf, Zed, Codex, Kiro, ...)
because the client just sees a normal MCP server. Full guide, including the
server's own sqz_read_file / sqz_grep / sqz_list_dir tools and per-client
config: MCP context compression.
CLI
Dedup Escape Hatch
When sqz sees the same content twice, it returns a compact §ref:HASH§ token
instead of the full text. Most models handle this fine, but some (e.g., GLM 5.1)
can't parse the ref format and loop. Four ways to work around this:
# 1. Recover original content from a ref
# 2. Compress without dedup (per-invocation)
|
# 3. Disable dedup globally (env var)
# 3b. Or just shorten the ref freshness window (seconds, default 1800).
# Useful on clients without a compaction hook; 0 disables refs.
# 4. MCP passthrough tool (returns input byte-exact, zero transforms)
# Available via tools/list when sqz-mcp is running
Recall: search everything sqz has seen
Everything that flows through sqz is indexed locally (SQLite FTS5, on your machine, nothing leaves it). When your agent compacts its context and loses that error message from an hour ago, search for it instead of re-running the command:
)
Agents get the same thing as the sqz_recall MCP tool.
Track Your Own Savings
Run sqz gain in your shell any time to see your own daily breakdown (see the
Token Savings section above for what the output looks like), and sqz stats
for the full cumulative report:
Add --breakdown to see exactly which commands consume the most tokens:
sqz stats and sqz gain count cl100k tokens of what sqz actually printed,
abbreviation legend and expand hint included. A run of one character class
longer than 2 KB, such as a long separator line, is counted in 2 KB slices,
which can be off by about 1% plus 4 tokens.
Honest accounting: cost and regret
Token reduction is not automatically billed-cost reduction — with provider prompt caching, most context re-transmits at a ~90% discount, and compression that drops something the agent needed costs extra turns instead (arXiv:2607.12161). sqz measures both sides instead of hand-waving:
sqz stats --costestimates dollars saved under a prompt-cached billing model (cache write ×1.25 once, cache read ×0.10 per subsequent turn), with every assumption printed and overridable (--price-in,--reread-turns,--cache-write-mult,--cache-read-mult). This is the conservative estimate: without caching the same savings would bill at full input price on every turn.sqz statsreports regret signals: quick re-runs (the agent re-produced byte-identical output within 2 minutes — the repeat bought nothing; the Claude Code PostToolUse hook never answers with a reference, so its repeats count as ordinary compressions) and ref expands (§ref§tokens recovered to original bytes). Both are proxies for "compression dropped something the model needed." If one command keeps showing up, its formatter needs work — file an issue.
sqz's design already avoids the failure modes that study measured:
compression is deterministic and query-agnostic (never invalidates provider
prompt caches), only new tool output is touched (history is never rewritten),
commands that follow output, run a server or background a job are left alone
(on Claude Code outside bypass mode, the PreToolUse hook leaves pipelines,
loops and heredocs as written and the PostToolUse hook compresses their
output after a successful run), grouped lists and pipelines (Gemini CLI,
Copilot CLI, Claude Code in bypass mode) skip the per-command formatters,
and the 16-token net-win gate skips compressions that wouldn't pay for their
own markers. A command is also left as written when its rewrite could run a
different set of commands than the original. In ls ${DEST:?}; echo next,
piping ls on its own would let echo next run after the failed expansion
stopped the shell, and where a list goes through sqz as one group, a failed
cd in cd build && make; make test would skip the group instead of only
make.
Per-project filtering:
Stats are stored locally in SQLite under ~/.sqz/sessions.db — nothing leaves your machine.
How Compression Works
-
Per-command formatters for the commands below:
Ecosystem Commands Git status, log, diff, show, stash, remote, fetch, push, pull, add, commit Rust cargo build/test/clippy/check/nextest JavaScript npm/pnpm/yarn/bun install/audit/outdated (test and run when they fail), tsc, eslint, biome Python pytest, ruff, mypy, pip Go go test (incl. -jsonstream), go build, go vet, golangci-lintCloud terraform plan/apply/destroy/refresh/init Containers docker/podman ps/images, kubectl describe/apply JVM gradle build/test, maven Mobile xcodebuild (build + test), adb logcat .NET dotnet build/publish/test System grep/rg, tree, curl/wget Other commands, including kubectl get and logs, gh, aws, gcloud, ls and find, go through the generic compression pipeline.
-
Table compactor — aligned-column output from tools without a dedicated formatter (
ps aux,netstat, database CLIs) collapses its padding runs to two-space separators. The detector is strict — indented lines, code, YAML, and JSON never match. -
Structural summaries: the engine has an API that reduces a code file to imports, function signatures and a call graph. The hook and MCP paths do not use it; they serve source in full.
-
Dedup cache: SHA-256 content hash, persistent across sessions. Byte-identical output the model already has = 13-token reference. A line range of a file the model already has in full (
sed -n '40,80p',sqz_read_filewithoffset/limit) =§ref:HASH:L40-80§. A re-read after a small edit = the changed lines with one unchanged line on each side, through the sqz MCP tools (the shell hook sends the edited file again); edits spread through the file come back in full. References are only served while the original is still in context (30-minute window, tunable viaSQZ_REF_TTL_SECS, reset on compaction). -
JSON pipeline: strip nulls → drop bookkeeping fields (
__v,etag, request and correlation ids) → flatten → collapse arrays → TOON encoding (lossless compact format). Fields whose value is an empty string, array or object are dropped at any depth, and strings over 500 characters are cut to their first 500 characters followed by.... -
Safe mode: stack traces (a trace header followed by its frames), SQL DDL at the start of a line, PEM blocks and credential lines (
key=valueorkey: valuewith a secret-looking key), with or without ANSI colors, skip the pipeline and pass through unchanged at any size, with only ANSI codes stripped, and are never cached. On the shell hook a per-command formatter still runs first. Its summary is used only when it keeps every line of each paragraph that holds one of them unchanged, so it may summarize the rest of a test run around a trace; otherwise the whole output passes through. Output that only mentions a password, key, trace or migration outside those structures is compressed as usual but is not cached either, so anything a lossy stage drops from it cannot be expanded.
If a stage panics, sqz compress prints the input unchanged with one line on stderr and keeps the command's exit status, and the Claude Code PostToolUse hook leaves the result as it was.
Measured quality, including what each stage drops and what it never touches: quality benchmark. For the full technical details, see docs/.
Configuration
# ~/.sqz/presets/default.toml
[]
= "default"
= "1.0"
[]
= true
= 3
[]
= true
[]
= 0.70
= 200000
Per-project database
Everything sqz persists (stats, dedup cache, sessions) lives in one SQLite
file, ~/.sqz/sessions.db by default. Set SQZ_DB_PATH to keep it
per-project instead:
# e.g. in the project's .envrc (direnv)
Every surface honors it — the shell hook, sqz stats/gain/expand, and
the MCP server (set it under env in your MCP config). Prefer absolute
paths: relative values resolve against whatever directory the process runs
from. The parent directory is created if missing (mode 0700, like ~/.sqz).
Privacy
- Zero telemetry — no data transmitted, no crash reports
- Fully offline — works in air-gapped environments
- All processing local
Development
License
Elastic License 2.0 (ELv2) — use, fork, modify freely. Two restrictions: no competing hosted service, no removing license notices.
Links
- Documentation index
- How to stop AI coding agents from re-reading the same files
- MCP context compression
- Quality benchmark: compression ratio vs information preservation
- Context rot vs context compression
- Choosing a context compression approach
- White Paper: Pre-Injection Context Compression
- Token Savings Benchmark
- Discord
- Changelog
- MCP Registry name:
mcp-name: io.github.ojuschugh1/sqz
Star History
If sqz saved you tokens today, a star helps other people find it. sqz stats --share prints your own numbers if you want to show them.
Contributors
Thanks to everyone who has contributed code, fixes, and ideas to sqz:
And to everyone who filed the detailed bug reports behind our fixes. Precise repros make this project better with every release. Want to join them? PRs are reviewed fast: see the open issues to get started.
Contributor grid made with contrib.rocks.