supercode-cli 0.4.15

supercode — a lightweight, fully-customizable AI coding agent CLI in Rust. Any model via OpenRouter; natively continues Claude Code and Codex sessions.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
<div align="center">

# supercode

**A lightweight, fully-customizable AI coding agent — SDK + CLI, in Rust.**

Any model via [OpenRouter](https://openrouter.ai) · natively continues real
**Claude Code**, **Codex**, **Gemini CLI**, **Goose**, **Grok**, **OpenCode**, and **Pi** sessions ·
reads and live-drives **Hermes** and **OpenClaw** through their own doors · a single ~8.4 MB binary.

[![CI](https://github.com/volter-ai/supercode/actions/workflows/ci.yml/badge.svg)](https://github.com/volter-ai/supercode/actions/workflows/ci.yml)
[![crates.io](https://img.shields.io/crates/v/supercode-cli.svg)](https://crates.io/crates/supercode-cli)
[![downloads](https://img.shields.io/crates/d/supercode-cli.svg)](https://crates.io/crates/supercode-cli)
[![docs.rs](https://img.shields.io/docsrs/supercode-core)](https://docs.rs/supercode-core)
[![MSRV](https://img.shields.io/badge/MSRV-1.85-blue.svg)](https://www.rust-lang.org)
[![license](https://img.shields.io/badge/license-MIT%20OR%20Apache--2.0-blue.svg)](#license)
[![stars](https://img.shields.io/github/stars/volter-ai/supercode)](https://github.com/volter-ai/supercode/stargazers)
[![last commit](https://img.shields.io/github/last-commit/volter-ai/supercode)](https://github.com/volter-ai/supercode/commits)
[![platforms](https://img.shields.io/badge/platform-linux%20%7C%20macos-informational)](#install)

</div>

<a href="docs/assets/hero.mp4">
<img src="docs/assets/hero.webp" alt="supercode fixing a failing test" width="100%" />
</a>
<sub>Animated poster — click through for the full MP4.</sub>

<p align="center">
<a href="#install">Install</a> ·
<a href="#quickstart">Quickstart</a> ·
<a href="#documentation">Docs</a> ·
<a href="#benchmarks--footprint">Benchmarks</a> ·
<a href="CONTRIBUTING.md">Contributing</a>
</p>

```sh
supercode run "Fix the failing tests, then run pytest -q to confirm."
```

supercode reads your code, edits it, runs the commands, and confirms the result —
streaming every step. It's a native Rust agent loop that talks directly to any
model and is built to be a **superset** of what Claude Code and Codex do.

---

## What supercode is

supercode is the **superset glue tool** for AI coding agents — the connective
tissue *between* harnesses. The purpose of the whole supercode family of
packages is **uniform access to every harness**: one capability-honest way to
discover, read, translate, continue, observe, and drive any harness, through
each harness's own doors. It is explicitly **not** meant to be your primary
coder. Rescuing a rate-limited session and slashing continuation cost is the
purpose of one package (`crates/reduce`), not of the family.

Four capabilities deliver that purpose (numbered as the estate cites them, not
ranked):

1. **Translate between session formats.** A universal converter across the
   declared format set — **Claude Code, Codex, Gemini CLI, Goose, Grok, opencode, and pi** — load a session in any harness's
   on-disk format and faithfully emit it in any other, backed by a measured
   per-pair translation-fidelity matrix and lossless `A→B→A` round-trips.
2. **Emulate-to-continue any harness losslessly.** Pick up a real session from
   harness *X*, keep running it while emulating *X*'s semantics, and emit it
   back in *X*'s native format so *X*'s **own** tool can resume it with no loss.
3. **Continue losslessly *with token reduction*** — the `reduce` package's purpose.
   Keep running a session at drastically reduced token cost with **zero semantic
   loss** and full reversibility. The **rate-limit rescue** (hit a limit →
   reduce + switch provider at minimum cost → continue → export back) is the
   flagship instance.
4. **Universal compatibility & UI layer over harnesses.** One capability-honest
   surface — discovery, live observation, and control — over every harness on
   the machine, reached through each harness's **own doors** (ACP, gateway
   APIs, stock CLIs, session stores). Support is tiered per harness
   (**observed / driven / translated / written-back**), every tier
   receipt-verified. supercode renders and supervises what harnesses do —
   including orchestrator harnesses' cron, fleets, and channels — but never
   *becomes* the orchestrator, the coder, or the daemon.

> **Status:** Native discovery, import, export, passive follow, and
> emulate-to-continue are implemented for all seven declared formats. The
> committed 7×7 fixture matrix measures every format pair and restores exact
> canonical semantics in all 49 `A→B→A` runs. Native-only
> metadata and cross-format raw-byte residue remain named separately; semantic
> losslessness is not a claim of byte-identical foreign formats. Live runtime
> operations are only called verified when a current executable receipt proves
> them; Grok's ACP start/input/events/interrupt/respond/load-session and both
> Supercode/stock continuation paths are covered by `npm run probe:live:grok`.
> The [stock-resume matrix]docs/interop/STOCK-RESUME-MATRIX.md
> additionally proves real stock Claude Code, Codex, OpenCode, Pi, and Grok
> recall facts from both sides of a GLM 5.2/OpenRouter continuation while the
> original source stores remain unchanged.
> ACP load-session resumes persisted history in a new protocol process; it does
> not attach to an arbitrary already-running TUI. That separate capability is
> tracked explicitly as `runtime.attach_existing_process` and remains unsupported
> where the upstream harness exposes no reachable endpoint. Supercode's own
> SDK runtime now provides such an endpoint: `resume --serve` hosts one canonical
> continuation in an attachable tmux session by default, while ACP, HTTP, and
> terminal frontends join the same authenticated live process without creating
> duplicate agents. Tmux is only the local process supervisor; the SDK registry
> and persisted session family remain authoritative.
> See [ROADMAP]ROADMAP.md.

Rust consumers can take only the layer they need: `supercode-interchange` for
session discovery and translation, `supercode-reduce` for reversible token
reduction, `supercode-runtime` for provider/runtime primitives, and
`supercode-harness` for Supercode's complete native Agent and tool loop.
`supercode-core` is the backwards-compatible facade and contains no second
implementation. See [package boundaries](docs/architecture/package-boundaries.md).

---

## Why supercode

- 🪶 **Light & fast.** A single ~8.4 MB static binary (`cargo build --release`,
  stripped + thin-LTO, Linux x86_64 — see [`Cargo.toml`]Cargo.toml release
  profile); the agent loop runs in
  **~13–19 MB of RAM** — measured **~7× lighter than Codex and ~27× lighter than
  Claude Code** on identical tasks ([benchmarks]#benchmarks--footprint).
- 🔁 **A superset.** It natively loads, continues, and translates real
  **Claude Code, Codex, Gemini CLI, Grok, opencode, and pi** sessions. Resume a session against
  another provider, then export it back to the source harness.
- 🎛️ **Yours to shape.** Every prompt, every tool description, and every tool's
  on/off state is configurable — from CLI flags or the `Config` builder.
- 🌐 **Any model.** One OpenAI-compatible provider reaches Claude, GPT, Gemini,
  Llama, DeepSeek — anything OpenRouter routes to, or any endpoint you point it at.

## Install

```sh
curl -fsSL https://raw.githubusercontent.com/volter-ai/supercode/main/scripts/install.sh | sh
```

```powershell
irm https://raw.githubusercontent.com/volter-ai/supercode/main/scripts/install.ps1 | iex
```

The install scripts pull checksum-verified prebuilt binaries from GitHub
Releases. Building from source with `cargo install supercode-cli` remains
available on platforms without a prebuilt asset.

<details><summary>Other channels (cargo, binstall, Homebrew, source)</summary>

| Channel | Command |
|---|---|
| Install script | `curl -fsSL .../scripts/install.sh \| sh` (prebuilt, checksum-verified) — Windows: `irm .../scripts/install.ps1 \| iex` |
| cargo | `cargo install supercode-cli` |
| cargo-binstall | `cargo binstall supercode-cli` (prebuilt, no compile) |
| Homebrew | `brew install volter-ai/tap/supercode` (once the tap is published) |
| from source | `git clone … && cargo install --path crates/cli` |

Prebuilt channels resolve to GitHub Releases. Building from source needs Rust
1.85+.
</details>

<details><summary>Uninstall</summary>

```sh
curl -fsSL https://raw.githubusercontent.com/volter-ai/supercode/main/scripts/uninstall.sh | sh
```

```powershell
irm https://raw.githubusercontent.com/volter-ai/supercode/main/scripts/uninstall.ps1 | iex
```

By default this removes only the binary (and the PATH line `install.sh` may
have added to your shell rc) — your config, credentials, MCP registry, and
session history under `~/.config/supercode` are left alone. Add `--purge` to
remove those too, and `--dry-run` to preview either mode without deleting
anything:

```sh
sh scripts/uninstall.sh --dry-run          # preview: binary + PATH line only
sh scripts/uninstall.sh --dry-run --purge  # preview: + config/credentials/sessions
sh scripts/uninstall.sh --purge            # prompts, then removes everything
```

Windows (`scripts/uninstall.ps1`, same flags as `-DryRun` / `-Purge` / `-Yes`):

```powershell
.\uninstall.ps1 -DryRun            # preview: binary + user-PATH entry only
.\uninstall.ps1 -DryRun -Purge     # preview: + config/credentials/sessions
.\uninstall.ps1 -Purge             # prompts, then removes everything
```

Installed via `cargo install`? The uninstaller detects this and runs
`cargo uninstall supercode-cli` for you. Homebrew or npm installs are managed
by their own package manager (`brew uninstall supercode` /
`npm uninstall -g supercode`) — the script tells you which applies rather than
reaching into those directories itself. Full details: [Install & setup →
Uninstall](docs/setup.md#uninstall).
</details>

**Platform support** (from the release build matrix, `.github/workflows/release.yml`):

| Platform | Target | Status |
|---|---|---|
| Linux x86_64 | `x86_64-unknown-linux-gnu` | ✓ prebuilt release binary |
| Linux aarch64 | `aarch64-unknown-linux-gnu` | ✓ prebuilt release binary |
| macOS Apple Silicon | `aarch64-apple-darwin` | ✓ prebuilt release binary |
| macOS Intel | `x86_64-apple-darwin` | ✓ prebuilt release binary |
| Windows x86_64 | `x86_64-pc-windows-msvc` | ✓ prebuilt release binary |
| Windows aarch64 | `aarch64-pc-windows-msvc` | build from source (`cargo install supercode-cli`) |
| Anything else with a Rust toolchain || `cargo install supercode-cli` / build from source |

<details><summary>Tell your coding agent to install it for you (copy-paste)</summary>

```text
Install supercode, a Rust CLI coding agent, and verify it works:

1. Install it with one of:
   - curl -fsSL https://raw.githubusercontent.com/volter-ai/supercode/main/scripts/install.sh | sh
   - cargo install supercode-cli   (needs Rust 1.85+; use this on a platform
     without a prebuilt release asset)
2. Confirm it's on PATH: `supercode --version`
3. Set an API key: `export OPENROUTER_API_KEY=sk-or-...` (or run `supercode login`
   to save it interactively), then run `supercode doctor` and confirm every line
   is green, especially `provider: reachable (200)`.
4. Try it: `supercode run "list the top-level files in this repo"`.

If `supercode doctor` reports a missing key or an unreachable provider, stop and
tell me the exact output rather than guessing at a fix.
```

</details>

## Quickstart

```sh
supercode login          # paste your OpenRouter key → ~/.config/supercode
supercode doctor         # ✓ config, key, and live provider reachability
supercode run "summarize this repo's module layout"
```

No key yet? Any of: `supercode login`, `export OPENROUTER_API_KEY=sk-or-…`, or
`--api-key`. On a genuinely fresh install with no config yet, an interactive
`run`/`chat`/`resume` walks you through this automatically (detects an
existing key/env var, lets you pick a default model, and finishes with a
`doctor` check) — it runs at most once and never fires on a non-interactive
or scripted invocation, so the copy-paste block above still works unattended.

<table>
<tr>
<td width="50%" valign="top">

**Health check** — `supercode doctor`

<a href="docs/assets/doctor.mp4">
<img src="docs/assets/doctor.webp" alt="doctor" width="100%" />
</a>
<sub>Animated poster — click through for the full MP4.</sub>

</td>
<td width="50%" valign="top">

**Corpus audit** — `supercode audit`

<a href="docs/assets/audit.mp4">
<img src="docs/assets/audit.webp" alt="audit" width="100%" />
</a>
<sub>Animated poster — click through for the full MP4.</sub>

</td>
</tr>
</table>

## The superpower: continue any session, in any model

supercode treats sessions like an image editor treats files — one canonical
in-memory model, importers and exporters for each format.

<a href="docs/assets/session-interop.mp4">
<img src="docs/assets/session-interop.webp" alt="inspect and convert a session" width="100%" />
</a>
<sub>Animated poster — click through for the full MP4.</sub>

```sh
# Continue a Claude Code transcript with GPT-5:
supercode --model gpt-5 resume ~/.claude/projects/<proj>/<id>.jsonl "finish the refactor"

# Continue a Codex rollout with Claude Opus:
supercode --model opus resume ~/.codex/sessions/.../rollout-*.jsonl

# Inspect or convert any session — no API key needed:
supercode inspect <session>.jsonl
supercode convert <claude-session>.jsonl --to codex -o as-codex.jsonl

# Follow normalized messages while another harness writes the session:
supercode watch <session>.jsonl
```

`resume` auto-detects the format and normalizes it to a provider-neutral history,
so a conversation that started in one tool continues seamlessly in another model.
`watch` passively emits snapshots and appended messages as NDJSON; it observes
persisted session state and does not attach to or control the harness's terminal.
→ [Sessions & interop](docs/sessions.md) · [Resume & format-detection reference](docs/resume.md)

<details><summary>A static still, if the GIF above didn't load</summary>

![converting a Claude Code session to Codex format, then inspecting it](docs/assets/session-interop-still.png)

</details>

## Use it as a library

```rust
use supercode::{Agent, Config};

let config = Config::builder()
    .model("anthropic/claude-opus-4-8")
    .system_prompt("You are a terse, expert pair programmer.")
    .disable_tool("bash")                          // toggle tools off…
    .tool_description("search", "Grep the repo.")  // …or re-describe them
    .build();

let mut agent = Agent::new(config)?;
let reply = agent.send("Find every TODO and group them by file.").await?;
println!("{reply}");
```

Register your own tools by implementing `Tool`. → [SDK guide](docs/sdk.md)

## Built-in tools

`read_file`, `write_file`, `edit_file`, `list_dir`, `glob`, `search` (regex,
gitignore-aware), `apply_patch`, `bash`, `shell` (persistent), `update_plan` —
each disableable and re-describable. Writes are sandboxed (the CLI defaults to
workspace-write; the library defaults to full access unless you set
`.sandbox(...)`; on macOS the OS seatbelt confines `bash` too). Full breakdown
of what each sandbox/approval mode actually enforces: → [Safety](docs/safety.md).

## Privacy

supercode sends **no telemetry, ever** — no analytics, no usage pings, no
phone-home, on any command. The only outbound network calls a `run`/`chat`/
`doctor`/`resume` makes are to the model provider *you* configured
(`--base-url`, default OpenRouter) and, if you've wired one up, an MCP server
you configured yourself. `inspect`/`convert`/`watch` make no network calls at all.
This is a stronger claim than an opt-out — there is no telemetry sink to opt
out of. See [Install & setup → Privacy](docs/setup.md#privacy) for the
grep-able verification.

## Package map

Every crate and SDK package, generated from the manifests by
`scripts/docs-package-map.mjs`; the merge gate fails when it is stale.

<!-- package-map:start -->
| Kind | Path | Name | Description |
|---|---|---|---|
| crate | `crates/bench` | `supercode-bench` | Task-based benchmark harness for the supercode agent. |
| crate | `crates/cli` | `supercode-cli` | supercode — a lightweight, fully-customizable AI coding agent CLI in Rust. Any model via OpenRouter; natively continues Claude Code and Codex sessions. |
| crate | `crates/codex-frontend` | `supercode-codex-frontend` | Version-pinned stock Codex app-server compatibility adapter for Supercode SDK runtimes. |
| crate | `crates/core` | `supercode-core` | Compatibility facade composing Supercode interchange, reduction, runtime, and harness packages |
| crate | `crates/frontend-model` | `supercode-frontend-model` | Donor-neutral client state and semantic intents for Supercode frontends. |
| crate | `crates/frontend-tui` | `supercode-frontend-tui` | Attachable terminal frontend primitives for Supercode SDK runtimes. |
| crate | `crates/harness` | `supercode-harness` | The optional native Supercode agent and tool harness |
| crate | `crates/interchange` | `supercode-interchange` | Canonical, provider-neutral session interchange primitives for Supercode |
| crate | `crates/opencode-frontend` | `supercode-opencode-frontend` | Version-pinned stock OpenCode HTTP compatibility adapter for Supercode SDK runtimes. |
| crate | `crates/reduce` | `supercode-reduce` | Optional lossless, reversible session reduction for Supercode |
| crate | `crates/runtime` | `supercode-runtime` | Optional native model and tool runtime for Supercode |
| sdk, npm | `sdk/client` | `@volter-ai-dev/supercode-client` | Headless state and control layer for Supercode-powered coding-agent frontends |
| sdk, npm | `sdk/frontend` | `@volter-ai-dev/supercode-frontend` | Generated language-neutral attachment client for Supercode SDK runtimes |
| sdk, npm | `sdk/frontend-browser` | `@volter-ai-dev/supercode-frontend-browser` | Reference browser observer and capability-gated controller for Supercode SDK runtimes |
| sdk, npm | `sdk/frontend-pi` | `@volter-ai-dev/supercode-frontend-pi` | Optional Pi-derived terminal frontend for an existing Supercode SDK runtime |
| sdk | `sdk/orchestrator` | `@volter-ai-dev/supercode-orchestrator` | The orchestrator runtime over the supercode ontology: one typed operational model whose folder is its serialization, read and written through the harness world doors (docs/ORCHESTRATOR-IR.md) |
| sdk, npm | `sdk/playwright-shim` | `@volter-ai-dev/supercode-playwright-shim` | Playwright-shaped in-page DOM runtime for Supercode browser providers |
| sdk, npm | `sdk/remote-access` | `@volter-ai-dev/supercode-remote-access` | Optional remote-access lifecycle and tunnel-provider adapter for Supercode frontends |
| sdk, npm | `sdk/terminal` | `@volter-ai-dev/supercode-terminal` | Optional leak-safe terminal and tmux host adapter for Supercode frontends |
| sdk, npm | `sdk/typescript` | `@volter-ai-dev/supercode-harness-sdk` | TypeScript client for Supercode's harness.v1 session and runtime service |
| sdk, npm | `sdk/ui` | `@volter-ai-dev/supercode-ui` | Composable default UI kit for Supercode-powered coding-agent experiences |
<!-- package-map:end -->

## Documentation

Files are filed by kind. At the root, `SPEC.md` is the normative reference
for the reduction engine and is cited from source; `ROADMAP.md` holds open
work and the vision; `ARCHITECTURE-REVIEW.md` and `UNIVERSAL-LAYER-BACKLOG.md` are
plans that open tickets under `.volter/tracker` cite; `COOKBOOK.md` is a
how-to. Architecture explanations live in [`docs/architecture`](docs/architecture/),
research and proposals in [`docs/plans`](docs/plans/), and evidence in
[`docs/evidence`](docs/evidence/).
[`docs/README.md`](docs/README.md) is the map.

| | |
|---|---|
| 🚀 [Using the agent]docs/agent.md | run, chat, approvals, models |
| 🔁 [Sessions & interop]docs/sessions.md | resume / inspect / convert / watch |
| ↩️ [Resume & format detection]docs/resume.md | auto-detection internals, drop/round-trip semantics |
| 🔍 [Corpus audit]docs/audit.md | typed-schema coverage report |
| 🛠️ [SDK guide]docs/sdk.md | embed, custom tools, sandbox, customize |
| 🔌 [Integrations]docs/integrations.md | MCP, pipelines, any endpoint |
| 🌐 [Browser capabilities]docs/browser-capabilities.md | shared CLI/SDK/MCP contract, active-page providers, and the in-page Playwright shim (`sdk/playwright-shim`) |
| 🧩 [Repository agent packages]docs/agent-packages.md | agent folders, capability config, and portable frontend contributions |
| 📦 [Install & setup]docs/setup.md | channels, login, doctor, config |
| 🛡️ [Safety]docs/safety.md | sandbox policies, approval policies, defaults |
| 🩺 [doctor guide]docs/doctor.md | what each health check means, how to fix red |
| 📖 [Cookbook]COOKBOOK.md | every use case, runnable |
| 🖥️ [Terminal capabilities]docs/terminal-capabilities.md | color tiers, spinner/title/picker gating across `NO_COLOR`/`TERM=dumb`/piped/tty |

## Related repositories

The browser execution runtime that once lived under `packages/browser-*` here
is its own MIT repository, [volter-ai/browser-substrate](https://github.com/volter-ai/browser-substrate):
one durable filesystem, process host, virtual ports, and sandboxed previews in
a tab, with Node, Python, and WALI-compiled native programs as lanes. supercode
runs there unchanged as a pinned WALI program bundle; the recipe is that
repository's `programs/supercode`. Nothing in this repository imports it.

## Benchmarks & footprint

"Light and fast" is measured, not asserted — supercode's own footprint is
measured on every bench run; the cross-harness comparison below was measured
once, see the caveat after the table for what that means for reproducibility.
[`bench/`](bench/) runs the agent against a live model and profiles the
harness's own footprint (peak RSS, CPU, startup, binary size). Same model
(DeepSeek V4 Flash via OpenRouter), same tasks, same verifier, same per-process
sampler — only the harness changes:

| harness | own-process RSS | vs supercode | runtime |
|---|--:|--:|---|
| **supercode** | **~13–19 MB** || single 8.4 MB Rust binary |
| codex | ~89 MB | ~7× | Rust core (+ Node launcher) |
| claude code | ~346 MB | ~27× | Node / TypeScript |

The codex/claude columns above are a one-time measurement on the author's
machine, off-repo — rerunning `compare_harnesses.py` needs the `codex` and
`claude` binaries installed locally plus a live OpenRouter API key, so the
comparison is not reproducible from a clean checkout. The committed evidence
is [`bench/suites/harness-comparison.json`](bench/suites/harness-comparison.json)
and the methodology write-up at
[`bench/suites/harness-comparison.md`](bench/suites/harness-comparison.md).
The supercode own-process numbers, by contrast, are measured on every bench
run via `getrusage`.

All three solved the same tasks; supercode also used the least CPU. On a real
subset of the [Aider polyglot benchmark](https://github.com/Aider-AI/polyglot-benchmark)
(12 Exercism exercises across Python/Rust/Go) supercode driving DeepSeek V4 Flash
scores **12/12**. Full methodology and per-task data: [`bench/`](bench/); full
comparison methodology and caveats:
[`bench/suites/harness-comparison.md`](bench/suites/harness-comparison.md).

**Startup latency** — jcode's README also reports time-to-first-frame /
time-to-first-input against its competitors; here's the same axis for
supercode. Unlike the RSS table above, this needs **no API key** — it times
process spawn to first PTY output on each tool's `--help` (network-free, so
it's reproducible from a clean checkout with just the binaries installed):

| harness | first-output latency (median of 9) | version |
|---|--:|---|
| **supercode** | **~4 ms** | 0.1.0 |
| codex | ~35 ms | codex-cli 0.143.0 |
| claude code | ~183 ms | 2.1.205 |

This is a **cold-start proxy** (`--help`, not a live chat frame): supercode's
`chat` command gates all rendering behind `require_api_key()`, so a true
first-chat-frame or first-input-echo number would need live credentials (or a
mock completion endpoint) for all three tools — not set up here, and the
table above should not be read as chat responsiveness. Raw data
[`bench/suites/startup-latency.json`](bench/suites/startup-latency.json),
methodology and the full caveat:
[`bench/suites/startup-latency.md`](bench/suites/startup-latency.md).

```sh
export OPENROUTER_API_KEY=sk-or-...
cargo run -p supercode-bench -- --tasks bench/suites/polyglot --sandbox danger-full-access
python3 bench/tools/compare_harnesses.py     # supercode vs claude vs codex (RSS, needs API key)
python3 bench/tools/startup_latency.py       # supercode vs claude vs codex (startup latency, no API key needed)
```

Or reproduce a scorecard end to end with one command —
[`scripts/run-benchmarks.sh`](scripts/run-benchmarks.sh) prints toolchain/git-SHA provenance,
always measures the real release-binary size (no API key needed), and (with
`OPENROUTER_API_KEY` set) runs a suite and writes a timestamped scorecard under
`bench/results/`; see [`bench/README.md`](bench/README.md#reproducing-a-scorecard-end-to-end).

## Status

Implemented and tested: OpenRouter / OpenAI-compatible streaming provider with
tool calls · built-in tool suite + per-tool customization · agent loop with
streaming events and an iteration budget · native Claude Code **and** Codex
session loading + continuation · saving / converting sessions back to both
formats. Written files have been **manually** verified resumable by the stock
`claude` and `codex` binaries (a one-off check, not part of CI).

Not yet built (natural next steps): a lossless native session format, richer
prompt-cache controls, and Linux (landlock/seccomp) sandboxing — the OS process
sandbox is macOS-only today. See [ROADMAP](ROADMAP.md).

## License

MIT OR Apache-2.0.