nanocodex 0.2.0

A small, high-performance, headless Rust agents SDK
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
<div align="center">

<h1>Nanocodex</h1>

<p><strong>Blazing-fast, minimal, library-first reimplementation of Codex.</strong></p>

[![CI](https://img.shields.io/github/actions/workflow/status/gakonst/nanocodex/ci.yml?branch=master)][ci]
[![Crates.io](https://img.shields.io/crates/v/nanocodex.svg)][crates]
[![Docs.rs](https://img.shields.io/docsrs/nanocodex)][docs]
[![License](https://img.shields.io/badge/license-MIT%20OR%20Apache--2.0-blue.svg)][license]

**[Install]#installation** | **[Thesis]#model-and-harness-co-design** | **[Why Code Mode?]#why-code-mode** | **[API]#api** | **[Examples]examples** | **[Benchmarks]#how-fast**

[ci]: https://github.com/gakonst/nanocodex/actions/workflows/ci.yml
[crates]: https://crates.io/crates/nanocodex
[docs]: https://docs.rs/nanocodex
[license]: LICENSE-MIT

</div>

---

Nanocodex is a Code Mode-first Rust agents SDK. It provides typed turns, tools,
events, steering, cancellation, queueing, durable session snapshots, and fast
historical forks over the OpenAI Responses API. It supports persistent WebSocket
and HTTPS/SSE transports while keeping the complete coding-agent conversation
inside your process, without requiring an app server or durable control plane.

## Installation

Install the daily-driver CLI on macOS or Linux:

```sh
curl -fsSL https://nanocodex.paradigm.xyz | bash
```

The installer tracks stable releases. Switch an installed CLI to the rolling
nightly channel with:

```sh
nanocodex update --nightly
```

Multi-architecture Linux images are published to GHCR as
`ghcr.io/gakonst/nanocodex:latest` and `ghcr.io/gakonst/nanocodex:nightly`.
Immutable version, commit, and `nightly-<commit>` tags are also available.

Add the library to a Rust project:

```sh
cargo add nanocodex
```

Or add it directly to `Cargo.toml`:

```toml
[dependencies]
nanocodex = "0.1"
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }
```

Node.js 12.22 or newer must be available on `PATH` for Code Mode.

## Model and harness co-design

Nanocodex starts from a simple thesis: **the model and its harness are one
system**. A coding model is not an interchangeable completion engine floating
above an arbitrary tool loop. Its effective capabilities depend on the exact
instructions, tool contracts, response shapes, history ordering, continuation
semantics, cache identity, streaming behavior, and failure recovery that
surround it. Change the harness and you change the agent.

That is why Nanocodex begins with how Codex actually works. We inspect the
concrete Codex implementation, identify the model-facing invariants and
operational behavior it relies on, and adapt them rather than designing a
provider-neutral agent abstraction from first principles. Typed Responses
items, stable prompt prefixes, incremental conversation continuation, complete
history replay, tool-call ordering, WebSocket lifecycle, cancellation, and
subprocess cleanup are parts of the behavioral contract, not incidental
plumbing.

The adaptation is as important as the fidelity. Nanocodex keeps the pieces that
shape model behavior, then recasts them for a headless, library-first product:
one owned driver instead of an app server, typed results and optional events
instead of a durable rollout control plane, caller-defined tools instead of a
product-wide integration catalog, and generated Code Mode orchestration instead
of a generic workflow or multi-agent scheduler. Codex is the evidence and the
behavioral reference; Nanocodex chooses the smaller API boundary.

This also means agent quality and performance belong to the **model–harness
pair**. We evaluate the complete path—prompt, tools, transport, caching,
execution, and recovery—because a model-only comparison would miss the system
we are actually building.

## Why Code Mode?

Most tool-calling agents expose a flat catalog and return to the model between
operations. Nanocodex instead presents caller-defined Rust tools and MCP tools
as typed JavaScript functions behind one Code Mode entrypoint. The model can
write a small program that uses loops, conditions, data transformations, and
`Promise.all`, so related tool work can be composed inside one cell without a
model round trip between every operation.

This keeps the model-facing surface small while the application keeps normal
Rust ownership. Tool implementations, credentials, retry policy, and mutable
state stay outside the generated program. Code Mode runs cells on a prewarmed,
session-persistent Node host and gives the model explicit controls for yielding,
resuming long work, and bounding returned output.

Subagents make that composition especially useful. An application can expose
`AgentHandle::spawn()` as a clean-room worker, `AgentHandle::fork()` as a worker
with the latest safe conversation context, and another tool for follow-up turns
on a retained child. Code Mode can then generate the orchestration topology for
the task instead of selecting a hard-coded workflow DAG:

```js
const [independent, contextual] = await Promise.all([
  tools.spawn_agent({
    role: "reviewer",
    task: "Find assumptions the parent may have missed."
  }),
  tools.fork_agent({
    role: "investigator",
    task: "Trace the suspected regression using our existing context."
  })
]);

const followUp = await tools.prompt_agent({
  agent_id: independent.agent_id,
  task: `Challenge this conclusion:\n\n${contextual.report}`
});

text({ independent, contextual, followUp });
```

That allows dynamic fan-out, fan-in, independent checks, contextual branches,
and targeted follow-ups while Nanocodex core remains one owned agent lifecycle
rather than a generic multi-agent scheduler. The bundled CLI exposes this
example surface with `nanocodex --subagents true`; library consumers define the
tools and policy themselves. See [`subagents.rs`](examples/subagents.rs) for the
complete implementation.

## API

```rust
let (agent, _events) = Nanocodex::new(api_key)?;

// Accepted now.
let turn = agent.prompt("Inspect this repository.").await?;
// Optional cloneable control.
let _control = turn.control();
// Steer the same active turn.
turn.steer("Focus on the failing tests.").await?;
// Completed result and checkpoint.
let checkpoint = turn.result().await?;

let _follow_on = agent.prompt("Now propose a fix.").await?.result().await?;

let turn = agent.prompt("Run a long investigation.").await?;
// Cancel queued or active work.
turn.cancel().await?;
// Returns Err(TurnCancelled).
let _cancelled = turn.result().await;

// Fork from the latest safe model/tool boundary.
let (latest, _events) = agent.fork().await?;
// Fork from the exact older state.
let (historical, _events) = agent.fork_from(&checkpoint).await?;

// The application owns snapshot storage and retention.
let stored = serde_json::to_vec(&checkpoint.snapshot())?;
drop(agent);
let snapshot = serde_json::from_slice(&stored)?;
let (resumed, _events) = Nanocodex::builder(api_key)
    .resume(snapshot)
    .build()?;
```

`prompt().await` means accepted, not completed. The agent retains conversation
history, tools, cache identity, response chain, and its WebSocket automatically.
Snapshots contain the complete unredacted conversation and must be protected
like the prompts and tool data they preserve. Resuming creates fresh runtime
resources while retaining authoritative typed history and cache lineage. They
do not persist provider response IDs, so the first resumed request replays that
history before subsequent turns follow the configured incremental or full-replay
policy. Use `RolloutConfig` for the existing Codex-compatible on-disk recording
and `codex resume` handoff;
snapshots are the library-level checkpoint for application-owned restoration.
Both project the same immutable committed session boundary: snapshots encode
the complete native restore state, while rollouts append the Codex-specific
turn envelope and only the newly committed history.

Creating and retaining that committed boundary is constant-time: forks,
`TurnResult`, and rollout projection share the immutable typed-history prefix.
Calling `TurnResult::snapshot()` is deliberately different. It copies the
complete history into an independently serializable value, so its work and
memory scale with total retained history; JSON encoding adds another
content-sized pass. Treat it as a durability boundary rather than something to
call after every turn unless every boundary must be retained. Rollout recording
remains incremental during ordinary turns, and the first request after
restoring a serialized snapshot performs the required full-history replay.

GPT-5.6 Pro is selected independently from reasoning effort; it does not use a
different model slug. All six effort levels are available in either mode:

```rust
let (agent, _events) = Nanocodex::builder(api_key)
    .reasoning_mode(ReasoningMode::Pro)
    .thinking(Thinking::Xhigh) // None, Low, Medium, High, Xhigh, or Max
    .fast_mode(true) // Requests priority service.
    .build()?;

// Later turns can switch priority service without replacing the session.
agent.set_fast_mode(false).await?;
```

The CLI equivalents are `--reasoning-mode pro --thinking xhigh`, with matching
`OPENAI_REASONING_MODE` and `OPENAI_REASONING_EFFORT` environment variables.

Independent agents that use the same immutable instructions and tool definitions
may deliberately share only their provider-side prompt cache while retaining
separate conversations, response chains, tools, and workspaces:

```rust
let (agent, _events) = Nanocodex::builder(api_key)
    .session_id(attempt_id)
    .prompt_cache_key(agent_recipe_id)
    .shared_prompt_cache()
    .build()?;
```

Builders cloned after `shared_prompt_cache()` singleflight the first warmup for
each exact prefix. Later agents skip that redundant request and send their first
complete generation with the shared provider cache key. The cache key is
inherited by forks and clean spawned agents. Without an explicit key, every
clean root or spawned agent retains its own cache identity.

### Lifecycle and dataflow

```text
NanocodexBuilder
 private agent driver ───────────────► AgentEvents
       ▲                                  side channel
 Nanocodex
 cloneable conversation handle
       ├── prompt(...) ──► Turn
       │                    ├── steer(...)
       │                    ├── cancel(...)
       │                    ├── control() ──► cloneable TurnControl
       │                    └── result() ──► TurnResult
       │                                         │
       ├── fork()                                │ checkpoint
       └── fork_from(&TurnResult) ◄──────────────┘

AgentHandle, supplied to tools_factory(...)
       ├── spawn()  clean child
       └── fork()   child from latest safe model/tool boundary
```

That is the complete ownership model. See the runnable
[`lifecycle.rs`](examples/lifecycle.rs) for all of it in one file.

### Authentication

Native applications can use either an `OpenAI` API key or a `ChatGPT`
subscription. Existing API-key construction remains unchanged:

```rust
let (agent, events) = Nanocodex::new(api_key)?;
```

For a `ChatGPT` subscription, the bundled CLI performs an authorization-code
OAuth login with PKCE and reuses Codex's credential file at
`$CODEX_HOME/auth.json`, or `~/.codex/auth.json` by default. If Codex is already
logged in, no separate Nanocodex login is required:

```sh
nanocodex auth login
nanocodex auth status
nanocodex
```

Plain `nanocodex` and `nanocodex run` prefer an existing Codex subscription
session in `$CODEX_HOME/auth.json` (normally `~/.codex/auth.json`). If that file
is absent, the CLI falls back to `OPENAI_API_KEY`; direct binary runs load the
key from the nearest `.env` automatically. An explicit `--api-key` overrides
both automatic sources, while `NANOCODEX_AUTH_FILE` or `--auth-file` selects a
specific subscription file. A present but invalid auth file fails visibly
instead of silently incurring API-key charges. `nanocodex auth logout` removes
the shared file and therefore logs both Codex and Nanocodex out.

The default Responses policy depends on the selected authorization:

| Authorization | Default transport | Default `store` | Default history |
| --- | --- | ---: | --- |
| ChatGPT subscription | WebSocket | `false` | connection-local incremental |
| API key | WebSocket | `true` | stored incremental |

Selecting HTTPS keeps API-key sessions stored and incremental. ChatGPT HTTPS
remains `store: false` and automatically switches to complete client-history
replay because it has no connection-local checkpoint. Explicit
`--store-responses` and `--responses-history` values override these defaults
when the combination is supported.

These defaults optimize the normal long-lived interactive session, where the
retained WebSocket delivered warm first events in 80-151 ms versus 277-369 ms
over HTTPS. HTTPS is the better explicit choice for cold or fork-heavy
one-shot work: cold first events were 374-405 ms versus 587-679 ms for
WebSocket, and fresh HTTPS forks reached their first event in 373-436 ms versus
about 1.0-1.1 seconds for a new WebSocket. With API-key authentication,
`store: true` also reduced three historical-fork requests by roughly 97%.
Completion latency was backend-noisy, and every policy had identical token,
cache, and estimated-cost behavior in this workload. See the full
[transport benchmark](docs/RESPONSE_TRANSPORT_BENCH.md).

Library consumers own their login UX and can reuse the same managed session:

```rust
use nanocodex::{Nanocodex, load_chatgpt_auth};

let auth = load_chatgpt_auth("/path/to/.codex/auth.json")?;
let (agent, events) = Nanocodex::new(auth)?;
```

The shared authorization handle selects the matching Responses endpoints,
attaches the account and FedRAMP routing headers, refreshes shortly before JWT
expiry, reloads credentials rotated by another process, and retries one rejected
request after a serialized refresh. Forks and built-in HTTP tools share that
same rotating session. Browser/WASM embeddings continue to receive an
already-authorized WebSocket from the host and never own refresh tokens.

## Lifecycle details

`Nanocodex::new` installs the standard instructions, medium thinking, built-in
tools, persistent WebSocket, and retry/reconnect policy. The builder can instead
fix HTTPS/SSE, response storage, and full-replay policy for the agent and all of
its forks. Dropping the event receiver is supported; events then become a no-op.

Callers never pass transcripts, response IDs, tool outputs, or turn IDs back to
the agent. An incremental transport sends only the new delta with
`previous_response_id`; full-replay policy transparently sends the authoritative
typed history. A replacement ephemeral socket and HTTPS with `store: false`
also replay that history.

The Responses path encodes each wire request once and rejects common
non-metadata frames before attempting metadata decoding. Every streamed attempt
records request and response bytes, encode/send and socket-wait time, parsing,
public-event emission, typed decoding, and time to first event/output as
structural tracing fields.

### Queue, steer, cancel

Every `prompt` is an ordinary queued turn. Steering is explicit because it has
different semantics: it joins one already-active turn and is sampled only
between complete model responses and tool outputs. It does not create a second
turn or terminal event.

Cancellation targets the same opaque unfinished turn. Cancelling queued work
removes it before it reaches the model. Cancelling active work waits for model
work, Code Mode cells, and shell process groups to stop, then resolves the turn
as `NanocodexError::TurnCancelled`. Partial model or tool work is never
committed; surviving queued prompts resume from the last completed checkpoint.

Call methods directly when one task owns the turn:

```rust
let turn = agent.prompt("Investigate the failing tests.").await?;
turn.steer("Prioritize deterministic failures.").await?;
let result = turn.result().await?;
```

Use `TurnControl` only when result and control ownership need to split:

```rust
use nanocodex::NanocodexError;

let turn = agent.prompt("Run a long investigation.").await?;
let control = turn.control();
let result_task = tokio::spawn(async move { turn.result().await });

control.steer("Check the integration tests first.").await?;
control.cancel().await?;
assert!(matches!(result_task.await?, Err(NanocodexError::TurnCancelled)));
```

### Continue and fork conversations

Follow-on prompts reuse retained context automatically:

```rust
let first = agent
    .prompt("Choose one word for this project.")
    .await?
    .result()
    .await?;

let second = agent
    .prompt("Return the word you chose in uppercase.")
    .await?
    .result()
    .await?;
```

Each completed result is also an opaque historical checkpoint. The mainline can
keep advancing while multiple branches start from different points:

```rust
let turn_2 = agent
    .prompt("Record design decision A.")
    .await?
    .result()
    .await?;

agent
    .prompt("Record later decision B.")
    .await?
    .result()
    .await?;

// The mainline may continue while both new agents are being constructed.
let mainline = agent.prompt("Continue the primary analysis.").await?;
let ((historical, _), (latest, _)) = tokio::try_join!(
    agent.fork_from(&turn_2),
    agent.fork(),
)?;

let historical_turn = historical.prompt("Explore an alternative to A.").await?;
let latest_turn = latest.prompt("Challenge our newest assumptions.").await?;
let (mainline, historical, latest) = tokio::try_join!(
    mainline.result(),
    historical_turn.result(),
    latest_turn.result(),
)?;
```

#### Why checkpoint forks are efficient

With the stored incremental policy, used by default with API-key
authentication, each Nanocodex `response.create` request sets `store: true`.
Once a response completes, the API can retain it as a checkpoint; Nanocodex
keeps its response ID private inside the completed `TurnResult`. The next
healthy model call sends that ID as `previous_response_id` plus only the new
delta—the user message, steer, or tool output added since the stored
response—not the transcript again.

The same mechanism makes a historical fork cheap:

```text
prewarm       input: stable instructions/tools, store: true   → prefix response
root turn A   previous_response_id: prefix, input: A delta    → response A
root turn B   previous_response_id: A, input: B delta         → response B
fork from A   previous_response_id: A, input: branch delta    → branch response
```

The root remains attached to response B and continues independently. The fork
gets its own driver, WebSocket, response chain, service stack, and tool runtime,
but its first request references response A and uploads only the branch delta.
Locally, immutable typed-history segments and stable cache lineage are shared,
so constructing the branch does not copy the retained conversation either.
The API still evaluates the complete logical context—and token usage reflects
that context—even though the request payload carries only the delta.

The local fork remains constant-time under every transport and storage policy.
What changes is its first wire request: a fresh `store: false` fork cannot use a
durable provider checkpoint and replays its shared committed history. HTTPS
with `store: false` also replays on every turn; a persistent WebSocket may use
connection-local incremental continuation until it is replaced. Stable cache
lineage is preserved in all cases.

Stored responses are an optimization, not the source of truth. Nanocodex keeps
complete client-owned typed history for every checkpoint. If the provider no
longer has a response ID, the retry drops `previous_response_id`, replays that
committed history once, and then resumes delta-only requests from the new stored
response. Partial or failed responses are never committed or used as fork
points.

In the retained checkpoint benchmark, three concurrent branch requests sent
2,175 bytes instead of an equivalent 84,612-byte full replay—a 97.4% payload
reduction. See [`benchmarks/fork_results.md`](benchmarks/fork_results.md) for the
live methodology, cache observations, and raw trials.

See [`fork_conversations.rs`](examples/fork_conversations.rs) for a complete
ten-checkpoint example with parallel historical forks and a caller-defined
Tower stack.

### Configure only what your application owns

The common paths remain short; factories appear only when lifecycle isolation
requires them.

```rust
use std::time::Duration;

use nanocodex::{Nanocodex, Responses, Thinking};
use tower::timeout::TimeoutLayer;

let responses = Responses::builder()
    .layer(TimeoutLayer::new(Duration::from_secs(120)))
    .build();

let (agent, events) = Nanocodex::builder(api_key)
    .codex_home("/home/me/.codex")
    .instructions("You are a concise repository maintenance agent.")
    .thinking(Thinking::High)
    .workspace("/work/project")
    .tools(tools)
    .responses(responses)
    .build()?;
```

Configuring a Codex home loads its `AGENTS.override.md` or `AGENTS.md` before
workspace-scoped project instructions. Native applications can also opt into a
Codex-compatible committed-history rollout; rollout configuration implies the
same Codex home. The agent session ID is also the UUID accepted by
`codex resume`:

```rust
use nanocodex::{Nanocodex, RolloutConfig};

let (agent, events) = Nanocodex::builder(api_key)
    .rollout(RolloutConfig::new("/home/me/.codex"))
    .build()?;

println!("codex resume {}", agent.session_id());
println!("rollout: {}", agent.rollout().unwrap().path().display());
agent.flush_rollout().await?;
```

Recording appends Codex's model-context response items and its legacy turn and
message events after each completed or cancelled turn, so both resumed model
context and the visible Codex transcript are restored. Compaction uses explicit
replacement-history records, and failed partial output is excluded.
`nanocodex resume <thread-id>` discovers the same rollout beneath
`$CODEX_HOME/sessions` (or `archived_sessions`) that `codex resume` uses and
materializes its latest replacement-history boundary plus subsequent response
items. No Nanocodex-specific sidecar is required, so Codex-created threads can
be continued directly. Pass `--prompt "..."` to submit a follow-on prompt as
soon as the TUI opens.
The rollout writer projects the same committed session boundary exposed by
`TurnResult::snapshot()` rather than maintaining a second conversation state.
`flush_rollout()` retries pending writes and provides a durability barrier.
Treat a resume as a single-writer handoff: release the Nanocodex session before
continuing the same rollout in Codex.

Use `tools_factory` when a tool must spawn or fork the agent that invoked it.
The factory receives a weak `AgentHandle`, not credentials:

```rust
let (agent, events) = Nanocodex::builder(api_key)
    .tools_factory(|handle| build_agent_tools(handle))
    .build()?;
```

`handle.spawn()` creates a clean child; `handle.fork()` creates a contextual
child from the invoking agent. Both privately reuse the builder's credentials
and policy. The application can expose these as Code Mode tools and let the
model generate loops, fan-out, follow-up prompts, and synthesis without encoding
a DAG in the SDK. See [`subagents.rs`](examples/subagents.rs) for that complete
orchestration pattern.

One Tower call is one complete streamed Responses attempt. Nanocodex owns retry
and reconnect policy; caller middleware can own deadlines, load shedding,
tracing, metrics, and circuit breaking without creating a second retry loop.
See [`docs/RESPONSES_TOWER.md`](docs/RESPONSES_TOWER.md) for the boundary and
ordering rules.

### Tools, MCP, events, and errors

`#[tool]` turns an async Rust function into a typed tool and derives its input
schema. `Tools::builder()` accepts generated or manual `Tool` implementations;
`Mcp::builder()` adds deferred Streamable HTTP or stdio MCP providers. The model
normally sees only Code Mode and its wait operation, then composes nested tools
with generated JavaScript, including loops, conditionals, and `Promise.all`.

Code Mode prewarms one persistent Node host alongside the first model call and
reuses it for the session. Cells receive one shared owned history snapshot;
resumed waits do not copy history they cannot read. A nested shell request can
extend the default outer-cell yield deadline while an explicit `@exec` deadline
still wins. Live shell session IDs remain visible for later `write_stdin`
calls, and stdout/stderr drains share one bounded completion deadline.

`AgentEvents` is an optional ordered stream independent of `TurnResult`. A TUI,
server, notebook, or binding can consume all events, select a subset, or drop
the receiver without changing prompt/result behavior. Libraries emit diagnostic
`tracing` spans but never install a global subscriber. Nested tools that finish
after a yielded cell retain their original Code Mode and model-call lineage, so
the public event stream and trace hierarchy agree.

Lifecycle failures are direct `NanocodexError` variants. Common control flow can
match `TurnCancelled` or `TurnNotSteerable`; transport and API details remain
available through `responses_error()` and the standard `Error::source` chain.

Runnable API tours:

```sh
cargo run -p nanocodex-examples --bin minimal
cargo run -p nanocodex-examples --bin lifecycle
cargo run -p nanocodex-examples --bin follow-on
cargo run -p nanocodex-examples --bin custom-tool
cargo run -p nanocodex-examples --bin mcp
cargo run -p nanocodex-examples --bin fork-conversations
cargo run -p nanocodex-examples --bin subagents
```

## CLI and repository

Install the daily-driver CLI and start it in the workspace the agent should
edit:

```sh
curl -fsSL https://nanocodex.paradigm.xyz | bash

# Reuse an existing Codex login, or fall back to OPENAI_API_KEY when absent.
nanocodex

# Create or explicitly select the same subscription store Codex uses.
nanocodex auth login
nanocodex --auth-file "${CODEX_HOME:-$HOME/.codex}/auth.json"
```

The CLI merges enabled MCP servers from
`${CODEX_HOME:-$HOME/.codex}/config.toml` with its deferred
`openaiDeveloperDocs`, `tempo`, and `cloudflare` defaults. Explicit `--mcp` and
`--mcp-stdio` entries take precedence over Codex config, and Codex entries take
precedence over same-named defaults. Their catalogs stay out of the stable
model prompt until Code Mode calls `tool_search`. Disable the built-in set with
`--mcp-defaults false` and Codex config loading with
`--mcp-codex-config false`; the environment equivalents are
`NANOCODEX_MCP_DEFAULTS=false` and `NANOCODEX_MCP_CODEX_CONFIG=false`.
Streamable HTTP servers also reuse Codex's file-backed MCP OAuth credentials.
Expired access tokens refresh through the server's OAuth metadata and rotated
credentials are persisted back to `.credentials.json` under the same
cross-process lock as Codex.

Provider prewarming is scheduled before agent construction returns, so the TUI
does not wait for HTTP client setup, DNS/TLS, MCP handshakes, `tools/list`, or
BM25 indexing. Discovery and index construction run in parallel in the
background while the terminal is idle. Activating a deferred tool changes only
the nested Code Mode runtime; the model-visible `exec`/`tool_search` prefix stays
byte-stable for prompt-cache reuse.

The CLI records Codex-compatible rollouts beneath
`${CODEX_HOME:-$HOME/.codex}/sessions` by default. The `request_id` in headless
JSONL is the resumable UUID. Restore its workspace, committed typed history, and
rollout identity in Nanocodex with:

```sh
nanocodex resume <request_id>
nanocodex resume <request_id> --prompt "Continue where we left off"
```

Or hand the interoperable rollout to Codex:

```sh
nanocodex run "Remember this thread"
codex resume <request_id>
```

Use `nanocodex --rollouts false ...` (or `NANOCODEX_ROLLOUTS=false`) when a CLI
consumer does not want local session recording.

The TUI retains one session across prompts. Enter submits, Tab explicitly queues
a follow-up while work is active, and `/cancel` stops the focused turn. At any
safe model/tool boundary, `/btw <question>` opens a fast fork in a vertical pane
while the mainline continues. The fork inherits the last completed response ID
plus complete tool results and applied steers after that response; partial model
output and unmatched tool calls remain excluded. With local telemetry running,
`/trace` opens Jaeger filtered to every turn in the focused main or `/btw`
session. `/mcp login <server>` opens browser OAuth and hot-reloads that server's
tools into the running session after the callback; `/mcp reload <server>` retries
discovery without restarting the TUI. The complete keybinding reference,
retained Amp and Codex research, and
prioritized Ratatui backlog live in
[`docs/TUI_NOTES.md`](docs/TUI_NOTES.md). The headless `nanocodex run` adapter
emits flushed JSONL for scripts and Harbor.

`just run-otel` exports compact per-turn streaming summaries alongside the full
agent trace. When diagnosing a jagged stream, opt into the individual correlated
API-delta, TUI-application, and presented-frame records:

```sh
just run-otel-detail
```

The TUI log at `.nanocodex/logs/tui.log` correlates each request and event across
socket receipt, agent emission, TUI receipt, state application, frame
coalescing, Ratatui changed cells, terminal output bytes, and final flush.
Detailed mode is intentionally opt-in because long responses can create
thousands of records. Use `just bench-stream` for the focused event-delivery,
transcript-update, and steady-frame regression gate. See
[`docs/OBSERVABILITY.md`](docs/OBSERVABILITY.md) for the full trace contract.

The workspace also contains publishable [JavaScript](js) and [Python](py)
libraries plus thin [Node](examples/node), [browser Worker](examples/react-vite),
and [website](web) consumers.
Architecture and current work are tracked in [`PLAN.md`](PLAN.md); benchmark
runner research lives in [`docs/HARBOR_RS_LOG.md`](docs/HARBOR_RS_LOG.md).

```sh
# Install pinned host dependencies.
just bootstrap
# Run the native smoke test.
just run
# Build and cache benchmark inputs.
just prepare-evals
# Run the pinned Terminal-Bench suite.
just eval
# Inspect retained Harbor jobs.
just view
```

## Nanocodex versus Codex

Use Nanocodex when the agent is a component of your Rust application. Use Codex
when you want the complete product: durable threads, approval UX, broad built-in
integrations, managed subagents, and a mature TUI and IDE ecosystem.

| | Nanocodex | Codex |
| --- | --- | --- |
| Product boundary | Rust library in your process | Application and durable agent runtime |
| State | In-memory authority; optional Codex-compatible rollout | Persisted threads and rollouts |
| Follow-on turns | Fixed WebSocket or HTTPS policy; automatic delta or full replay | Full Codex session lifecycle |
| Historical forks | Exact completed checkpoint; parent keeps running | Durable thread reconstruction |
| Tools | Code Mode over Rust tools and MCP | Broad built-in tool and integration surface |
| Middleware | Your concrete Tower stack | Codex-owned runtime policy |
| Results and events | Typed `TurnResult` plus optional ordered `AgentEvents` | Product-wide rollout/event lifecycle |
| Orchestration | Model composes application-defined agent tools | Managed agents, task identities, mailboxes, and budgets |

The smaller boundary is the feature. A caller builds an agent, receives
`(Nanocodex, AgentEvents)`, sends prompts through a cheap cloneable handle, and
awaits independently owned `TurnResult`s. The CLI, Harbor adapter, Python
binding, and Rust/WASM binding all consume that same API.

### Codex request parity

The deterministic differential drives stock Codex and Nanocodex through basic
tools, continued PTY shells, bounded output, failures, images, WebSocket
reconnect/replay, automatic compaction, and SIGINT cleanup:

```bash
cargo build -p nanocodex-bin
python3 benchmarks/codex_request_parity.py \
  --codex-bin ~/github/openai/codex/codex-rs/target/debug/codex \
  --nanocodex-bin target/debug/nanocodex \
  --output /tmp/nanocodex-request-parity.json
```

The default command succeeds when both implementations complete every scenario,
preserve stable request identity, replay or compact complete history, and clean
up cancelled descendants. Add `--check` when exact normalized request and
transport equality is required; it currently reports the documented
WebSocket-to-HTTPS fallback. App-owned internal turn metadata is an explicit
normalization exclusion. Use repeated `--scenario <name>` flags to run only
selected cases.

Nanocodex deliberately defaults to `Thinking::High`; stock Codex currently
defaults the bundled Sol model to `low`. The differential passes the same
explicit effort to both binaries so it tests harness behavior rather than that
product-policy choice.

### How Fast?

Focused local profiling also covers the ordinary streaming and tool path. On an
M1 Max, prewarming the retained Node host reduced first-cell host latency from
301 ms to 1.7–2.7 ms and complete top-level Code Mode time from 312 ms to 11–13
ms across three trials. On a retained 41-task workload, model generation and
caller-requested subprocesses accounted for 99.864% of summed run time; the
unattributed local remainder was 0.136%. These are diagnostic measurements, not
model-service speed guarantees. See
[`single_prompt_profile_2026-07-20.md`](benchmarks/single_prompt_profile_2026-07-20.md)
and
[`long_prompt_profile_2026-07-20.md`](benchmarks/long_prompt_profile_2026-07-20.md)
for methodology and reproduction commands.

The committed-history benchmark on the current session/rollout implementation
kept an active checkpoint plus parent append at 0.369–0.394 µs from 100 through
10,000 retained items. Iterating only the next request suffix stayed near 26
ns. The full-history flattening component required by an explicit durable
snapshot scaled from 8.2 µs at 100 items to 92 µs at 1,000 and 992 µs at
10,000. Those synthetic items contain about 512 bytes of text each; the
flattening numbers exclude JSON encoding. See the
[transport benchmark](docs/RESPONSE_TRANSPORT_BENCH.md#committed-session-projection-cost)
for the table and interpretation.

Our live checkpoint benchmark uses `gpt-5.6-sol`, a deterministic 600-fact
prefix, ten sequential turns, and concurrent historical forks. Three runs on
2026-07-20 compared Nanocodex `210ac85` with stock Codex CLI
`0.145.0-alpha.18`:

| Measurement | Nanocodex | Stock Codex | Difference |
| --- | ---: | ---: | ---: |
| Ten sequential turns, median total | 14.78 s | 24.99 s | **1.69x faster** |
| Warm turn p50, turns 3–10 | 1.304 s | 1.532 s | **1.18x faster** |
| Historical fork to first answer, p50 | 1.570 s | 6.530 s | **4.16x faster** |
| Historical fork model time, p50 | 1.291 s | 5.862 s | **4.54x faster** |

**Takeaway: Nanocodex was 1.69x faster across ten turns and 4.16x faster to
the first historical-fork answer in this checkpoint benchmark.**

A Nanocodex fork sent about 725 bytes of new request data from its stored
checkpoint. Replaying the same history would send 27–29 KB: a 97.4% reduction.
On a separate 41-task coding gate, Nanocodex completed 38 tasks with 92.23% of
input tokens cached, zero Responses retries, and zero WebSocket reconnects.

These are checkpoint-path measurements, not a normalized full-agent quality
comparison. The Nanocodex arm used a minimal benchmark developer message and no
production tool definitions; the Codex arm ran the complete stock app-server
agent. See [`benchmarks/fork_results.md`](benchmarks/fork_results.md) for the
methodology, cache observations, raw trials, and reproduction commands.

### The tradeoff

Nanocodex currently supports one model family (`gpt-5.6-sol`), one Responses
WebSocket transport, and caller-defined tools. Sessions and branches live only
as long as your process. Your application owns sandboxing, permissions,
durability, and recursive cancellation policy for application-defined child
agents. Code Mode requires Node.js 12.22 or newer on `PATH`.

That is substantially less product than Codex. It is also much less machinery
between your code and an agent turn.