vision-squeezer 0.8.1

Fit images to a vision-LLM token budget before you send them. Library, CLI and MCP server for code, agents and pipelines.
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
<p align="center">
  <img src="assets/logo.png" width="300" alt="VisionSqueezer Logo" />
</p>

# VisionSqueezer

<p align="center">
  <a href="https://github.com/eralpozcan/vision-squeezer/actions"><img src="https://img.shields.io/github/actions/workflow/status/eralpozcan/vision-squeezer/ci.yml?label=build" alt="Build"></a>
  <a href="https://crates.io/crates/vision-squeezer"><img src="https://img.shields.io/crates/v/vision-squeezer" alt="crates.io"></a>
  <a href="https://www.npmjs.com/package/vision-squeezer"><img src="https://img.shields.io/npm/v/vision-squeezer" alt="npm"></a>
  <a href="LICENSE"><img src="https://img.shields.io/badge/license-Elastic--2.0-blue" alt="License"></a>
</p>

Fit images to a vision-LLM token budget **before** they reach the model. A library, CLI, and MCP server for code, agents, and pipelines.

> **What it does not do.** VisionSqueezer cannot shrink an image you paste, drag, or `@`-attach into a chat window (Claude Code, Cursor, ChatGPT, and so on). The client attaches it before any tool or hook runs, so it is billed at full size.
>
> **Where it does save tokens:**
> - Images your code sends to a model API: the Rust crate, the Python bindings, the CLI.
> - Images an agent reads from files or gets through tools (screenshots, crawlers, browser automation): the MCP server, or the Claude Code image-read hook for files inside the project.
>
> See [Where it saves tokens]#where-it-saves-tokens for the exact cases.

The MCP server runs in any MCP client (Claude Code, Cursor, Codex, Gemini CLI, and others), but it only sees images the agent passes to it.

---

## Install

### One command per client (recommended)

```bash
npx vision-squeezer install                       # asks for the client and scope
npx vision-squeezer install --client claude --scope user --yes
npx vision-squeezer install --client cursor --yes
```

Every client has exactly one install path, and the installer picks it:

| Client | What `install` does |
|---|---|
| Claude Code | Installs the `vision-squeezer-mcp` plugin: MCP server, `/vision-stats` `/vision-doctor` `/vision-upgrade` skills, and the image-read hook |
| Codex CLI, Qwen Code, Kimi CLI | Runs `<cli> mcp add vision-squeezer -- npx -y vision-squeezer@<version>` |
| Gemini CLI | Same `mcp add`, plus an image-read hook in `settings.json` |
| VS Code | Runs `code --add-mcp` |
| Cursor | Merges one pinned entry into `mcp.json`, plus an image-read hook in `hooks.json` |
| OpenCode | Merges one pinned entry into `opencode.json`, plus an image-read plugin |
| Windsurf, Claude Desktop | Merges one pinned entry into the client's JSON config |

JSON files are merged, never replaced: other servers and hooks are kept, a file that cannot be parsed is left untouched, and re-running replaces our own entry instead of adding a second.

The scope (`user`, `local`, `project`) is asked only when the client supports more than one. `--method` from older versions is accepted and ignored.

### Claude Code

```bash
npx vision-squeezer install --client claude --scope user --yes
```

This runs the two commands below for you (and removes an older `claude mcp add` registration and `settings.json` hook first). To do it by hand:

```bash
claude plugin marketplace add eralpozcan/vision-squeezer
claude plugin install vision-squeezer-mcp@vision-squeezer --scope user
```

Restart Claude Code, or run `/reload-plugins`, to load the MCP server, skills, and hook. Only want the MCP server, with no skills and no hook? `claude mcp add --scope user vision-squeezer -- npx -y vision-squeezer`.

### Claude Desktop

```bash
npx vision-squeezer install --client claude-desktop --yes
```

Or add to `claude_desktop_config.json` (macOS: `~/Library/Application Support/Claude/`, Windows: `%APPDATA%\\Claude\\`):
```json
{
  "mcpServers": {
    "vision-squeezer": {
      "command": "npx",
      "args": ["-y", "vision-squeezer"]
    }
  }
}
```

### Cursor

```bash
npx vision-squeezer install --client cursor --yes
```

Or add to `~/.cursor/mcp.json` (global) or `.cursor/mcp.json` (project):
```json
{
  "mcpServers": {
    "vision-squeezer": {
      "command": "npx",
      "args": ["-y", "vision-squeezer"]
    }
  }
}
```

<details>
<summary><b>Click here to view installation instructions for 10+ other IDEs and Agents (VS Code, JetBrains, Windsurf, Zed, etc.)</b></summary>

### VS Code Copilot

```bash
code --add-mcp '{"name":"vision-squeezer","command":"npx","args":["-y","vision-squeezer"]}'
```

Or add to `.vscode/mcp.json`:
```json
{
  "servers": {
    "vision-squeezer": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "vision-squeezer"]
    }
  }
}
```

### JetBrains (IntelliJ, WebStorm, PyCharm)

Open **Tools → GitHub Copilot → Model Context Protocol (MCP) → Configure**, then add:
```json
{
  "servers": {
    "vision-squeezer": {
      "command": "npx",
      "args": ["-y", "vision-squeezer"]
    }
  }
}
```

### Windsurf

```bash
npx vision-squeezer install --client windsurf --yes
```

Or add to `~/.codeium/windsurf/mcp_config.json`:
```json
{
  "mcpServers": {
    "vision-squeezer": {
      "command": "npx",
      "args": ["-y", "vision-squeezer"]
    }
  }
}
```

### Gemini CLI

```bash
gemini mcp add --scope user vision-squeezer -- npx -y vision-squeezer
```

Or add to `~/.gemini/settings.json` (user) / `.gemini/settings.json` (project):
```json
{
  "mcpServers": {
    "vision-squeezer": {
      "command": "npx",
      "args": ["-y", "vision-squeezer"]
    }
  }
}
```

### Codex CLI

```bash
codex mcp add vision-squeezer -- npx -y vision-squeezer
```

Or add to `~/.codex/config.toml`:
```toml
[mcp_servers.vision-squeezer]
command = "npx"
args = ["-y", "vision-squeezer"]
```

### Qwen Code

```bash
qwen mcp add vision-squeezer -- npx -y vision-squeezer
```

### OpenCode

`opencode mcp add` is interactive-only, so add directly to `~/.config/opencode/opencode.json` (global) or `opencode.json` in the repo root (project):
```json
{
  "mcp": {
    "vision-squeezer": {
      "type": "local",
      "command": ["npx", "-y", "vision-squeezer"],
      "enabled": true
    }
  }
}
```

### Kimi CLI

```bash
kimi mcp add vision-squeezer -- npx -y vision-squeezer
```

### Zed

Add to `~/.config/zed/settings.json`:
```json
{
  "context_servers": {
    "vision-squeezer": {
      "command": "npx",
      "args": ["-y", "vision-squeezer"]
    }
  }
}
```

### Kiro

Add to `.kiro/settings/mcp.json` (workspace) or `~/.kiro/settings/mcp.json` (global):
```json
{
  "mcpServers": {
    "vision-squeezer": {
      "command": "npx",
      "args": ["-y", "vision-squeezer"]
    }
  }
}
```

### Antigravity

MCP-only, no hooks needed. Configure via the Antigravity MCP settings:
```json
{
  "mcpServers": {
    "vision-squeezer": {
      "command": "npx",
      "args": ["-y", "vision-squeezer"]
    }
  }
}
```

</details>

### Manual install (Rust binary)

```bash
# From crates.io
cargo install vision-squeezer

# Or from source
git clone https://github.com/eralpozcan/vision-squeezer && cd vision-squeezer
make install   # builds → ~/.local/bin/
```

Then use the binary path directly in any config above instead of `npx`:
```json
{ "command": "vision-squeezer-mcp" }
```

> **Tip:** Run `npx -y vision-squeezer --setup` to print ready configs with auto-detected paths.

---

## CLI Usage

```bash
vision-squeezer path/to/image.jpg \
  --mode auto|ocr|standard \      # default: auto (= standard, keeps colour; ocr is opt-in)
  --format jpeg|webp|avif \       # default: jpeg
  --quality 85 \                  # output quality 1-100 (default: 75)
  --tile-size 256 \               # patch size in px (default: 512)
  --no-crop \                     # disable padding removal
  --smart-crop \                  # edge-energy crop (vs corner-tolerance)
  --auto-quality 0.95 \           # binary-search quality to hit SSIM target
  --bg-tolerance 25 \             # background detection 0-255 (default: 15)
  --model <provider-alias> \ # model-aware resizing; see the model catalog
  --max-tokens 1600 \            # token budget: downscale until the output fits (0 = off; MCP default 1600)
  --max-tiles 20 \                # hard cap on tile count (model-specific unit)
  --json \                        # machine-readable JSON output
  --dry-run                       # run pipeline, skip disk write
```

### Batch mode

Pass a directory instead of a file:

```bash
vision-squeezer ./screenshots --recursive --output-dir ./optimized
vision-squeezer ./screenshots --recursive --json > report.json
```

---

<details>
<summary><b>The Math: How Vision Models Bill You in 2026</b></summary>

If you send raw images to an LLM, you are leaking tokens. Modern vision models do not care about your file size (MB/KB); they only care about **pixel dimensions**, but each provider calculates costs completely differently. `vision-squeezer` simulates these algorithms to find the mathematical minimum size that drops your token usage without losing visual context.

### 1. OpenAI GPT-6 / GPT-5.6 (32px Patches)
Current OpenAI vision models count **32×32 patches**, fit high-detail images within a 2048px edge and 2500-patch budget, then apply a 1.2 multiplier.
* **The Fix:** `--model gpt6` removes padding, fits the patch budget, and avoids partially used edge patches. A 1024×1024 input is 1024 patches / 1229 tokens.

### 2. Claude 4.7+ (28px Patches)
Claude counts `ceil(Width/28) × ceil(Height/28)` visual tokens. Claude 4.7+ high-resolution vision uses a 2576px edge and 4784-token budget; `claude-standard` covers the earlier 1568px / 1568-token tier.
* **The Fix:** Strip padding, fit the selected tier, and align dimensions to the 28px grid.

### 3. Gemini 3 (Large Tiles)
Gemini uses a massive **768×768 tile** system (if the image is > 384px). Each tile is a flat 258 tokens.
* **The Fix:** An 800×600 image triggers a 2×1 tile grid (516 tokens). `vision-squeezer` snaps it down slightly to fit exactly inside a 768×768 box, dropping the cost to 258 tokens (**50% savings**).

### 4. Llama 3.2 / 3.3 Vision (560px Tiles)
Meta's Mllama vision tiles images on a **560×560** grid, capped at 4 tiles (~1601 tokens each).
* **The Fix:** A 2400×1670 screenshot trimmed to 2400×1200 drops from a 2×2 to a 2×1 canvas: **6,404 → 3,202 tokens (−50%)**. (Llama 4 uses a different vision encoder and is not modeled.)

### 5. Qwen2-VL / 2.5-VL / 3-VL (28px Patch Grid)
Alibaba's Qwen-VL uses a **28px effective grid** (14px patch × 2×2 merge); `tokens = (W/28)·(H/28)` bounded to `[4, 16384]`.
* **The Fix:** The patch is small, so area is the lever — a 1024×1024 image with its border stripped to 896×896 drops **1,369 → 1,024 tokens (−25%)**.

### 6. DeepSeek Flash / DeepSeek-VL2
DeepSeek Flash now accepts images through its OpenAI-compatible API and documents an upper bound of **384 image tokens per image**. Use `--model deepseek` for the API profile; use `--model deepseek-local` for the exact DeepSeek-VL2 open-weight tile math.

### 7. Kimi K2.5 / K2.6 / K3
Kimi supports native image and video input through Moonshot's OpenAI-compatible API. Moonshot does not publish a stable image billing grid, so `--model kimi` gives an advisory native-resolution estimate while still applying exact crop and file-size optimization.

### 8. DeepSeek-VL2 (384px Anyres Tiles)
SigLIP-384 + 2× pixel-shuffle gives 196 tokens/tile on a `(m·384, n·384)` canvas (`m·n ≤ 9`).
* **The Fix:** An 800×768 image snapped to 768×768 drops from 3×2 to 2×2 tiles: **1,415 → 1,023 tokens (−28%)**. (Open weights — the win is local-inference context, not API billing.)

Legacy `gpt4o` and `gpt5` profiles remain available for endpoints that still use 512px high-detail tiles.

> Full provider math and sources: **[visionsqueezer.com/providers]https://visionsqueezer.com/providers/claude**

</details>

## Pipeline

```
Input image
  → crop_padding               remove solid-color borders
  → calculate_optimal_dims     snap to tile boundary (always down)
  → [enforce_max_tokens]       token budget (MCP default 1600) / optional tile cap
  → resize_exact               Lanczos3
  → [binarize]                 OCR mode only: Otsu threshold
  → JPEG/WebP encode           configurable quality & format
```

### 🦀 Why Rust?

VisionSqueezer is a performance-critical middleware. We chose Rust for three uncompromising reasons:

- **Invisible Latency:** Image processing should never be the bottleneck. Rust ensures that snapping, cropping, and encoding happen in milliseconds, making the optimization layer truly invisible to the developer's workflow.
- **Minimal Footprint:** As an MCP server running in the background of your IDE, VisionSqueezer is designed to be ultra-lightweight, consuming near-zero CPU and RAM when idle.
- **Wasm-Ready:** Rust's first-class support for WebAssembly allows us to bring the same high-performance optimization to the browser and the Edge (Cloudflare Workers), enabling client-side squeezing before the image even hits the network.

---

### Benchmark snapshots

Captured with `vision-squeezer <image> [flags] --dry-run --json` on the sample images in `data/`. **The goal is fewer tokens, not fewer megabytes**: providers bill by pixel dimensions, so file size only matters for upload and latency. Without a budget the pipeline snaps to a grid and crops padding, which saves little on large photos because providers already downscale oversized images themselves. A token budget (`--max-tokens`, default **1600** in the MCP server) is what makes the saving real.

### Case Study 1: Standard image (istanbul.jpg)
2400×1670, 0.5 MB.

| Run | Output | File size | Claude 4.7+ tokens | GPT-6 tokens | Gemini tokens |
|---|---|---|---|---|---|
| *no budget* (pre-0.7 default) | 2048×1536 | −29% | 4,674 → 4,070 (−13%) | 2,903 → 2,942 | 3,096 → 1,548 (−50%) |
| `--max-tokens 1600` (MCP default) | 1344×924 | −67% | 4,674 → 1,584 (−66%) | 2,903 → 1,462 (−50%) | 3,096 → 1,032 (−67%) |
| `--max-tokens 1000` | 1064×728 | −79% | 4,674 → 988 (−79%) | 2,903 → 939 (−68%) | 3,096 → 516 (−83%) |
| `--max-tokens 1600 --model gpt6` | 1408×960 | −65% | 4,674 → 1,785 (−62%) | 2,903 → 1,584 (−45%) | 3,096 → 1,032 (−67%) |
| `--max-tokens 1600 --model gemini` | 2304×1536 | −22% | 4,674 → 4,565 (−2%) | 2,903 → 2,880 (−1%) | 3,096 → 1,548 (−50%) |

```bash
vision-squeezer data/istanbul.jpg --max-tokens 1600
```
```text
Input:  2400×1670  (0.5 MB)
Output: 1344×924  (0.2 MB, JPG q75)
File:   67.5% smaller

── Token Estimates ─────────────────────────────────────────
Model          Before    After      Saved
------------------------------------------
Claude 4.7+      4674     1584     3090 (66.1%)
GPT-6            2903     1462     1441 (49.6%)
GPT-4o           1105     1105        0 (0.0%)
GPT-5             910      910        0 (0.0%)
Gemini           3096     1032     2064 (66.7%)
────────────────────────────────────────────────────────────
```

### Case Study 2: 12-megapixel image (istanbul2.jpg)
4096×3072, 2.2 MB.

| Run | Output | File size | Claude 4.7+ tokens | GPT-6 tokens | Gemini tokens |
|---|---|---|---|---|---|
| *no budget* (pre-0.7 default) | 3584×2560 | −40% | 4,661 → 4,698 | 2,942 → 2,924 (−1%) | 6,192 → 5,160 (−17%) |
| `--max-tokens 1600` (MCP default) | 1288×952 | −88% | 4,661 → 1,564 (−66%) | 2,942 → 1,476 (−50%) | 6,192 → 1,032 (−83%) |
| `--max-tokens 1000` | 1036×756 | −92% | 4,661 → 999 (−79%) | 2,942 → 951 (−68%) | 6,192 → 516 (−92%) |
| `--max-tokens 1600 --model gpt6` | 1344×992 | −87% | 4,661 → 1,728 (−63%) | 2,942 → 1,563 (−47%) | 6,192 → 1,032 (−83%) |
| `--max-tokens 1600 --model gemini` | 2304×1536 | −71% | 4,661 → 4,565 (−2%) | 2,942 → 2,880 (−2%) | 6,192 → 1,548 (−75%) |

```bash
vision-squeezer data/istanbul2.jpg --max-tokens 1600
```
```text
Input:  4096×3072  (2.2 MB)
Output: 1288×952  (0.3 MB, JPG q75)
File:   87.9% smaller

── Token Estimates ─────────────────────────────────────────
Model          Before    After      Saved
------------------------------------------
Claude 4.7+      4661     1564     3097 (66.4%)
GPT-6            2942     1476     1466 (49.8%)
GPT-4o            765     1105        0 (0.0%)
GPT-5             630      910        0 (0.0%)
Gemini           6192     1032     5160 (83.3%)
────────────────────────────────────────────────────────────
```

### Reading the numbers

- **Budget off, tokens barely move.** Claude 4.7+ already caps images at 4,784 tokens, so a 4,661-token original can come out at 4,698. The squeezed file is smaller, the bill is not.
- **`--max-tokens 1600` cuts Claude tokens by ~66% and GPT-6 by ~50%** on both photos while keeping composition, colour and large text. Fine detail (distant windows, small signs) is what goes first.
- **The budget is measured with the target model**, Claude when none is set. `--model gemini` with the same 1600 lands on Gemini's 768px tile grid (1,032 tokens), and the file-size column follows.
- **Pick the budget for your content.** 1600 and 1000 kept body text legible in a retina code screenshot; 600 did not. Pass `--max-tokens 0` (or omit it on the CLI) to disable the cap.
- **Colour is never dropped.** `auto` mode behaves like `standard`; black-and-white output only happens with an explicit `--mode ocr`.

---

## Current model sanity checks

| Profile | Input | Estimated tokens |
| --- | --- | ---: |
| GPT-6 / GPT-5.6 | 1024×1024 | 1,229 |
| GPT-6 / GPT-5.6 | 2048×2048 | 3,000 after the 2,500-patch cap |
| Claude 4.7+ | 1000×1000 | 1,296 |
| Legacy GPT-5 / 5.1 | 1024×1024 | 630 |

---

## MCP Tool: `optimize_image`

| Argument | Type | Required | Default |
|----------|------|----------|---------|
| `image_path` | string | one of the two | — (preferred: read locally, optimized copy also written to a temp file) |
| `image_base64` | string | one of the two | — |
| `mode` | `"auto"` \| `"standard"` \| `"ocr"` | — | `"auto"` |
| `output_format` | `"jpeg"` \| `"webp"` | — | `"jpeg"` |
| `quality` | integer 1–100 | — | 75 |
| `tile_size` | integer | — | 512 |
| `crop` | boolean | — | true |
| `bg_tolerance` | integer 0–255 | — | 15 |
| `max_tokens` | integer | — | 1600 (`0` disables the cap) |
| `max_tiles` | integer | — | — |
| `target_model` | string enum | — | Core models plus `glm`, `pixtral`, `gemma`, `internvl`, `minicpm`, `molmo`, `aya`, `phi4`, `granite`, `llava`, `falcon`, `minimax`, `step`, `ling`, `voyage` |

**Response:** an MCP `image` block with the optimized image, then a text block with a small report. The image is never base64 inside the text: a model reads that as text tokens, which would cost more than the image it replaces.
```json
{
  "width": 1344,
  "height": 924,
  "output_path": "/tmp/vision-squeezer/shot-1a2b3c4d.jpg",
  "tokens_before": 4674,
  "tokens_after": 1584,
  "savings_report": {
    "tiles_before": 4674,
    "tiles_after": 1584,
    "tiles_saved": 3090,
    "token_reduction_pct": "66.1",
    "size_reduction_pct": "67.5"
  }
}
```
`output_path` is set when the input was an `image_path`.

### Where it saves tokens

The optimization only helps if the **smaller image is what reaches the model**. What that means in practice:

| How the image gets to the model | Optimized? |
|---|---|
| Claude Code reads an image file inside your project (`Read`) | **Yes**, by the image-read hook that ships with the plugin (`install --client claude`) |
| The agent calls `optimize_image` with an `image_path` | **Yes**, from any folder. Tell your agent to do this in `CLAUDE.md` / `AGENTS.md` |
| You paste or drag an image into the chat | **No.** The client attaches it before any tool or hook runs |
| You attach a file with `@path` | **No.** No tool call happens, so no hook fires |

**The baseline matters.** Claude Code already downsizes large images before it sends them: in our test a 2400×1670 file reached the model as 1999×1392. Against that, the 1600-token budget saves about 56% on `istanbul.jpg` (about 3,600 → 1,584 estimated tokens), not the 66% measured against the original file. All token counts here are estimates from each provider's published rules, not billed usage.

| Client | Image-read hook | Status |
|---|---|---|
| Claude Code | Yes, `PreToolUse` on `Read` (plugin) | Tested in a live session |
| Cursor | Yes, `preToolUse` + `updated_input` on `Read` (`hooks.json`) | Follows Cursor's docs, not tested in the app |
| Gemini CLI | Yes, `BeforeTool` on `read_file` (`settings.json`) | Follows Gemini's docs, not tested in the app. Google replaced Gemini CLI with Antigravity CLI for free and Google One users on 2026-06-18 |
| OpenCode | Yes, a plugin on `tool.execute.before` for `read` | Follows OpenCode's docs, not tested in the app |
| Codex CLI | No. Hooks can rewrite a call, but Codex's docs list no image-read tool path | MCP tool only |
| Windsurf | No. `pre_read_code` can only block | MCP tool only |
| VS Code Copilot | No. The docs show no way to rewrite a tool input | MCP tool only |
| Kimi CLI | No. `PreToolUse` shows the input but the docs describe no rewrite | MCP tool only |
| Qwen Code, Claude Desktop | Not checked / no hook system | MCP tool only |

"MCP tool only" means the agent has to call `optimize_image` itself with an `image_path`; tell it to in `AGENTS.md`. The hooks for Cursor, Gemini CLI, and OpenCode keep their copy in a self-ignoring `.vision-squeezer/` folder inside the project, because those clients restrict reads to the workspace.

Save screenshots into the project (for example `screenshots/`) and reference them by path to get the saving. The hook only rewrites files inside the project: it answers with an allow decision, so it must not open files a plain `Read` would have asked about.

## Config Reference

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `quality` | u8 1–100 | 75 | JPEG/WebP output quality |
| `tile_size` | u32 | 512 | Custom grid size; ignored when `target_model` is set |
| `crop` | bool | true | Remove solid-color padding borders |
| `bg_tolerance` | u8 0–255 | 15 | Max channel delta for background detection |
| `output_format` | jpeg/webp | jpeg | Output encoding. WebP is ~30-50% smaller |
| `max_tiles` | u32 | — | Hard cap on tile count (progressive downscale) |
| `target_model` | string | — | Model-aware profile; use `gpt6` for current OpenAI models |

## Supported Models

## Platform support

Prebuilt MCP binaries cover Linux x86_64/ARM64, macOS Apple Silicon/Intel, and Windows x86_64/ARM64. Python wheels cover Linux x86_64/ARM64, macOS Apple Silicon/Intel, and Windows x86_64. Unsupported Python architectures can install from source with `pip install --no-binary vision-squeezer vision-squeezer`.

| Model | Tile Size | Pre-scaling | Token Formula |
|-------|-----------|-------------|---------------|
| GPT-6 / GPT-5.6 | 32×32 patches | 2048px edge / 2500 patches | ceil(patches × 1.2) |
| Claude 4.7+ high resolution | 28×28 patches | 2576px edge / 4784 patches | 1 token per patch |
| Claude standard | 28×28 patches | 1568px edge / 1568 patches | 1 token per patch |
| GPT-4o / GPT-4.5 (legacy) | 512×512 | fit 2048px → scale short 768px | 85 + tiles × 170 |
| GPT-5 / GPT-5.1 (legacy) | 512×512 | fit 2048px → scale short 768px | 70 + tiles × 140 |
| Gemini 3 | 768×768 | flat tier at ≤384×384 | 258 per tile |
| DeepSeek Flash | provider-managed | up to 384 image tokens | conservative 384-token cap |
| Kimi K2.5/K2.6/K3 | native-resolution | provider-specific | advisory 28px estimate |
| GLM, Pixtral/Mistral, Gemma, InternVL | provider-specific | provider-specific | advisory generic profile |
| MiniCPM, Molmo, Aya, Phi-4, Granite | provider-specific | provider-specific | advisory generic profile |
| LLaVA, Falcon, MiniMax, Step, Ling, Voyage | provider-specific | provider-specific | advisory generic profile |

---

## Advanced Features

### Persistence & Analytics

VisionSqueezer tracks every optimization locally in `~/.vision-squeezer/stats.db`.

```bash
vision-squeezer stats
```
```text
── VisionSqueezer Analytics ────────────────────────────────
Total Optimizations: 42
Total Tokens Saved:  842,500
Total Bytes Saved:   156.40 MB
Estimated USD Saved: $2.11
────────────────────────────────────────────────────────────
```

### Shell Hook & Claude Code Skills

Add to `.zshrc` / `.bashrc`:

```bash
eval "$(vision-squeezer setup-hook)"
```

This installs two things at once:

- **`squeeze` alias** — optimize and capture the output path in one step:
  ```bash
  img_path=$(squeeze data/logo.png --model gemini)
  ```
- **Claude Code skills** — written to `~/.claude/skills/` automatically on first run

#### Available Skills

| Skill | Trigger | What it does |
|-------|---------|--------------|
| `vision-stats` | `/vision-stats` | Show cumulative token & byte savings — reads local stats.db, zero MCP overhead |
| `vision-doctor` | `/vision-doctor` | Check installed version vs latest npm release, show update command if outdated |
| `vision-upgrade` | `/vision-upgrade` | Detect install method (cargo/npm/npx) and run the correct upgrade command |

**Example output:**

```
/vision-stats
── VisionSqueezer Analytics ────────────────────────────────
Total Optimizations: 42
Total Tokens Saved:  842,500
Total Bytes Saved:   156.40 MB
Estimated USD Saved: $2.11
────────────────────────────────────────────────────────────

/vision-doctor
## VisionSqueezer Doctor
- [x] Binary found: /usr/local/bin/vision-squeezer
- [x] Installed version: 0.1.8
- [x] Latest version (npm): 0.1.8
- [x] Status: Up to date
```

**Marketplace install** (alternative to `setup-hook`):
Add to `~/.claude/settings.json`:
```json
{
  "extraKnownMarketplaces": {
    "vision-squeezer": {
      "source": { "source": "github", "repo": "eralpozcan/vision-squeezer" }
    }
  }
}
```
Then run `/plugins add vision-stats@vision-squeezer`, `/plugins add vision-doctor@vision-squeezer`, or `/plugins add vision-upgrade@vision-squeezer` in Claude Code.

### Sandbox Mode: "Think in Code"

Execute atomic operations locally before the image ever reaches an LLM. The agent decides _how_ to process the image; you pay zero tokens for the intermediate steps.

**CLI:**
```bash
vision-squeezer screenshot.png --ops '[{"op":"crop","x":10,"y":20,"width":500,"height":500},{"op":"binarize"}]'
```

**MCP tool — `sandbox_execute`:**

| Argument | Type | Description |
|----------|------|-------------|
| `image_base64` | string | Input image |
| `operations` | array | Ordered list of ops to apply |

Supported ops: `crop`, `grayscale`, `binarize`, `resize`, `contrast`, `brightness`.

### Crawler Integration

Optimize images on the fly in any Playwright/Puppeteer scraping pipeline — zero API token waste on raw screenshots.

```javascript
await page.route('**/*.{png,jpg}', async (route) => {
  const response = await route.fetch();
  const body = await response.body();
  const optimized = await squeeze(body); // pipe through vision-squeezer
  route.fulfill({ body: optimized });
});
```

---

## Contributing

PRs welcome. See [CONTRIBUTING.md](CONTRIBUTING.md) for setup and guidelines.
Please follow our [Code of Conduct](CODE_OF_CONDUCT.md).

## License

Elastic License 2.0 (ELv2) — See [LICENSE](LICENSE) for details.