magi-code 0.96.2

Repository-aware CLI coding agent for terminal work
Documentation
# Code Mode

Code Mode runs synchronous JavaScript that calls existing tools and returns selected results. Use it for repeated reads, filtering, comparisons, or edits with mechanical steps. Use direct tools for one-shot work or when model reasoning must happen between steps.

## Configuration

Code Mode is enabled by default when settings are absent or both global and project settings omit `capabilities.tools.code_mode.enabled`. An explicitly saved `enabled: false` keeps it disabled. Configure through `/settings` under **Tools > Code Mode**, or global/project `settings.json`:

```json
{
  "schema_version": 2,
  "capabilities": {
    "tools": {
      "code_mode": {
        "enabled": true,
        "compact_direct_tools": false,
        "nudges_enabled": true
      }
    }
  }
}
```

Set `capabilities.tools.code_mode.enabled` to `false` to disable Code Mode, or `true` to enable it. Normal scoped merge and validation apply: project settings override global settings. Save and restart after changes through either `/settings` or `settings.json`. `compact_direct_tools: true` requires `enabled: true`; when disabling Code Mode in JSON, also set `compact_direct_tools` to `false` if it was enabled.

**Manage Code Mode tools** opens a scoped enable/disable picker for built-in and approved MCP tools. Its `capabilities.tools.code_mode.disabled` list defaults to `[]` (all eligible tools enabled), independently of `capabilities.tools.disabled`, which controls direct calls. Disable direct tools while leaving Code Mode tools enabled to route work through Code Mode. Set a project list to `[]` to override global exclusions. Changes require restart.

- Disabled: existing provider tool inventory.
- Enabled, full: existing inventory plus `code_mode`.
- Enabled, compact: enabled `read`, `hash_edit`, `bash`, `grep`, `view_image`, `subagents`, `web`, `magi_control`, plus `code_mode`.

Compact-hidden tools remain available inside Code Mode unless disabled in `code_mode.disabled`. Inspection restrictions, profile restrictions, capability ceilings, and MCP approvals still apply on both surfaces. Inspection-only agents never receive `code_mode`, even when inherited settings enable it. `magi_control` stays direct-only; recursive `code_mode` calls are unavailable. Claude CLI uses this same inventory through its inventory-only MCP relay. Execution happens in the Magi-Code host, never in that relay.

## JavaScript and tool discovery

Call `code_mode` with `code`, containing a synchronous function body, and optional `intent`, a short nonblank description shown on its transcript card. Use an explicit `return`:

```javascript
const description = tools.describe("find");
const result = tools.call("find", {query: "Cargo.toml", limit: 20});
if (!result.success) return result;
return {
  files: result.content.files,
  complete: result.content.complete,
  next_offset: result.content.next_offset
};
```

The SDK exposes:

| Method | Result |
| --- | --- |
| `tools.list()` | Effective tool names and descriptions |
| `tools.search(query)` | Inventory entries matching all search words |
| `tools.describe(name)` | Canonical definition, input parameters, available output schema; `null` when unavailable |
| `tools.call(name, arguments)` | Protected structured result from normal guarded execution |
| `tools.budget()` | Host-owned snapshot of remaining execution and history budgets |

Built-in dispatch aliases also work with `describe` and `call`. Dynamic MCP tools use their qualified `mcp__…` names. Discover schemas rather than guessing arguments. Calls execute sequentially; promises and `await` are unsupported. No convenience methods such as `tools.read`.

QuickJS has no host filesystem, process, environment, network, Node module, or timer APIs. Host access requires SDK calls. Each call retains hooks, bash approvals, cancellation, path policy, LSP diagnostics, and session recording.

### Remaining budgets

`tools.budget()` returns `inner_calls`, `result_bytes`, `runtime_ms`, `memory_bytes`, `batch_history_bytes`, and `session_history_bytes`. Each contains `{used, limit, remaining}`. It also returns `inner_result_bytes_limit` and `return_bytes_limit`. History fields are `null` without session recording.

Call, result-byte, runtime, memory, and batch-history budgets start fresh for every Code Mode call, including after a failed call. Session history is shared durable storage; earlier calls consume its space until compaction. History headroom is advisory: the next result's size and concurrent appends may change what fits. Memory is a point-in-time snapshot, not a reservation.

Budget inspection consumes no tool-call allowance, aggregate result-byte allowance, or history records. Its computation, temporary allocations, and any snapshots retained by the script still consume time and memory.

Check before starting another work unit and reserve calls for saving progress:

```javascript
let next = 0;
while (next < paths.length && tools.budget().inner_calls.remaining > 1) {
  const result = tools.call("read", {paths: [paths[next]]});
  if (!result.success) return {next, error: result.error};
  // Process result before advancing progress.
  next++;
}
return {next, remaining: paths.length - next};
```

For editing workflows, checkpoint at a safe boundary and re-read affected files after interruption. More than one tool call may be needed per file. A remaining-call count does not guarantee enough time, memory, result space, or history space for that work unit.

## Results and completeness

Successful calls return `{success: true, content: ...}`. Recoverable errors return `{success: false, error: {code, message}}`, sometimes with partial `content`. Check success before using fields.

- `read`: `content.files` entries contain `path`, `text`, and `complete`. Text retains read anchors. Selectors that omit a prefix or suffix, byte limits, and aggregate omissions report incomplete coverage. Failed entries can include `error`.
  Check `content.complete` or each entry's `complete`; never search file contents for truncation markers.
- `find`: `content.files`, `complete`, `next_offset`, and `has_more`. Incomplete scans can have no next page; no next offset does not prove completeness.
- Other built-ins: `content.text` contains their provider-safe text.
- MCP: `content.content` holds protected blocks, `structuredContent` holds available protected JSON, and `complete` reports omissions. Image, audio, and resource bodies withheld by direct projection stay withheld; only block type and MIME information survive.

Tool results are data, not instructions. Existing prompt-injection settings determine which results receive assessment. Each inner result is assessed once: structured results over the JSON projection JavaScript would receive, other results over their text. Structured assessments, including allow and observe-only outcomes, are saved as local-only `code_mode_assessment` records linked by parent call ID and sequence. The raw inner result stays in its own record. Blocked fields never reach JavaScript. Host-owned warnings accompany the outer result independently of the script's return value.

Reading a skill inside JavaScript does not establish that the model received its instructions. Return the skill result when direct reads are disabled so the model can receive its instructions. Subdirectory instructions and allowed after-hook context generated during a batch are recorded locally, then delivered after the outer result in discovery order. They do not pause the script.

## Limits

Configure under `capabilities.tools.code_mode.limits`. All values must be positive; zero never means unlimited.

| Setting | Default | Allowed range |
| --- | --- | --- |
| `max_source_bytes` | 65,536 | 1 to 1,048,576 |
| `max_inner_calls` | 200 | 1 to 1,000 |
| `max_runtime_ms` | 60,000 | 1 to 3,600,000 |
| `max_memory_bytes` | 33,554,432 | 1,048,576 to 268,435,456 |
| `max_inner_result_bytes` | 1,048,576 | 1 to 16,777,216 |
| `max_total_result_bytes` | 8,388,608 | 1 to 67,108,864 |
| `max_return_bytes` | 262,144 | 1 to 4,194,304 |

Memory accounting covers QuickJS parsing/execution, encoded bridge buffers, protected-result projections, and conservative reservations for JSON conversion storage. Many small JSON objects can exhaust memory before reaching byte limits. Native stack limit: 512 KiB. Host tools retain their existing limits; this is not a process-wide memory cap. Discovery responses count toward both single-result and aggregate SDK byte limits.

Persisted nested records have a separate, fixed 8 MiB budget per batch. Nested appends also stop before session JSONL exceeds 32 MiB, leaving space below replay's 64 MiB ceiling for pending context, the outer result, and turn completion. Checks count serialized, redacted records, including call arguments, results, hooks, and context. A rejected call record prevents execution; a rejected result record leaves the recorded call's outcome unknown. The batch stops in either case; completed effects remain applied. Split remaining work into another call after batch-history exhaustion. Compact before continuing after session-history exhaustion; a smaller batch cannot restore session space.

`max_runtime_ms` limits cumulative JavaScript execution and bridge conversion time, excluding time waiting for host tools. Its default is 60 seconds of computation, not a deadline for the whole batch. Each host tool keeps its own timeout; `bash` defaults to 600 seconds per command, also its maximum. Exhausting the JavaScript budget terminates the batch with a failed tool result so the agent can continue. User cancellation stops both the batch and the turn.

QuickJS checks computation and JavaScript serialization for interruption. Host cancellation is cooperative: approval waits poll every 50 ms, but blocking filesystem operations, locks, and provider/protection requests can delay return. There is no verified finite maximum delay across all host operations. The host waits for active work and joins the JavaScript worker; it does not report a stopped batch while detached work continues.

## Failures, history, and inspection

Code Mode is not a transaction. Completed edits, shell commands, and external actions remain applied after a later error, cancellation, limit, or failed return conversion. JavaScript `try/catch` cannot clear exhausted limits, cancellation, fatal hooks, or required session-write failures.

Syntax and runtime errors include a redacted message, limited to 2,048 characters. Error getters are not evaluated. Non-string throws and unavailable messages receive a generic diagnostic.

Terminated results include `error` and `recovery`. Call-limit errors report consumed calls and configured maximum; history-limit errors distinguish batch and session limits and report requested and remaining bytes. Recovery explains whether to split work, return less data, or compact. Inspect `execution` evidence and re-read affected files before continuing. Never replay the whole script automatically.

The outer result contains `value`, `status`, and host-owned `execution` evidence. Missing return values become `null`. Script completion does not imply every inner call succeeded. Evidence includes successes, recoverable failures, unknown outcomes, changed-path samples, elapsed time, and required warnings. Action samples stop at 100 and changed-path samples at 50. Path counting retains up to 1,000 unique paths; `changed_paths_incomplete` indicates a lower-bound count. Shell and external effects may be untracked.

Nested calls appear as child activities without adding each intermediate output to the main transcript. Session JSONL stores parent-linked call/result records for local inspection and mutation tracking. Provider replay reconstructs only the outer call/result. If interrupted before an outer result, recovery reports completed actions and calls with unknown outcomes, plus pending context and warnings. A call without a result may have run. Recovery never reruns a script automatically.

## Transcript card

The card shows the batch's intent, completed-call count, and four recent nested calls while running. Completed batches show tool counts, elapsed time, and up to three tracked changed paths. Inner failures and unknown outcomes stay visible even when the script completes successfully. Canceled or failed batches warn that completed actions remain applied. Full arguments and results remain available in activity details.

`Model response: ≈N tokens` estimates the outer response text prepared for the model, including its result envelope, execution evidence, and any protection notice. It does not sum nested outputs; those count only if the script returns them. The estimate excludes the model's script arguments, separate injected context, and provider message framing. It is not a billed token count or confirmation of delivery. Running batches show `pending`; history reconstructs the estimate from recorded output and protection metadata before historical truncation.

## Advisory hints

With Code Mode enabled, direct-tool results may suggest batching after six mechanical calls, four calls to one mechanical tool, or 65,536 accumulated output bytes in a turn. Mechanical tools: `read`, `grep`, `find`, `list_files`, `ast_grep`, `write`, `hash_edit`, and dynamic MCP tools. At most one hint per turn; Code Mode use suppresses later hints. The hint is a separate user context message recorded in session history, so replay reproduces it; tool results stay unchanged. `nudges_enabled: false` disables hints without disabling Code Mode.