vision-squeezer 0.8.1

Fit images to a vision-LLM token budget before you send them. Library, CLI and MCP server for code, agents and pipelines.
Documentation
---
title: CLI Options Reference
description: Every VisionSqueezer CLI flag, including --max-tokens token budgets, --model provider profiles, --format, --quality, --smart-crop, --json and --dry-run.
navigation:
  icon: i-lucide-sliders-horizontal
---

## Reference

| Flag | Description |
| --- | --- |
| `--model <name>` | Target model alias; core profiles plus popular families such as `kimi`, `glm`, `pixtral`, `gemma`, `internvl`, `llava`, and `minimax`. |
| `--format jpeg\|webp\|avif` | Output encoding. AVIF is the default for new pipelines. |
| `--quality 1-100` | Output quality (default `75`). |
| `--auto-quality 0.0..1.0` | Binary-search quality in `[40,95]` to hit an SSIM target. |
| `--smart-crop` | Edge-energy (Sobel-lite) crop. Best for photographic content. |
| `--ops '<JSON>'` | Execute [Sandbox]/guides/sandbox operations. |
| `--output <path>` | Custom output destination (single-file mode). |
| `--output-dir <path>` | Output root (batch mode, mirrors structure). |
| `--recursive` | Walk subdirectories in batch mode. |
| `--max-tokens <N>` | Token budget. The image is downscaled until the target model's token estimate fits (Claude when no `--model` is set). `0` disables. The MCP server defaults to `1600`; the CLI has no cap unless you pass this. |
| `--max-tiles <N>` | Hard cap in the target model's own unit (28px patches for Claude, tiles for Gemini/Llama). Prefer `--max-tokens` ([token budget guide]/guides/token-budget). |
| `--mode auto\|standard\|ocr` | `auto` (default) behaves like `standard` and keeps colour. `ocr` is the only mode that converts to black and white. |
| `--json` | Machine-readable JSON output (single-file or batch aggregate). |
| `--dry-run` | Run the full pipeline without writing to disk or updating the stats DB. |

## Quality vs. auto-quality

`--quality` is a fixed encoder setting. `--auto-quality` is smarter: it binary-searches the quality range and lands on the **smallest file that still passes** a perceptual SSIM threshold.

```bash [Terminal]
# Fixed quality
vision-squeezer image.png --quality 80

# Target perceptual fidelity, minimize bytes
vision-squeezer image.png --auto-quality 0.95
```

Use `--auto-quality 0.95` when bandwidth matters and you can tolerate slight perceptual loss.

## Crop strategy

- **Default (corner-tolerance):** strips solid-color padding. Best for screenshots with uniform borders.
- **`--smart-crop`:** keeps the high-information region using gradient energy. Best for photos and saliency-heavy content.

## Estimate before writing

```bash [Terminal]
vision-squeezer image.png --model gpt6 --json --dry-run
```

`--json --dry-run` reports token impact with zero side effects — ideal for pipelines that gate on savings before committing.