magi-code 0.95.2

Repository-aware CLI coding agent for terminal work
Documentation
# Jev and internal tooling

[Feature docs index](README.md) · [Configuration](configuration.md) · [Tools and safety model](tools-and-safety.md)

magi-code uses [TypeSafe Jev](https://docs.typesafe.ai/) for bounded semantic judgments. Jev receives state plus a typed question and returns structured probabilities. It does not replace the main conversation model.

All Jev features are optional and disabled by default.

## Set up Jev

1. Create a TypeSafe API key.
2. Export it before starting magi-code:

   ```sh
   export TYPESAFE_API_KEY="your-key"
   magi-code
   ```

3. Open `/settings` and select **Internal Tooling**.
4. Enable **Jev tooling**, then enable only the features you need.
5. Save settings and restart magi-code.

Do not put the API key in `settings.json`. Enabling Jev without a non-empty `TYPESAFE_API_KEY` causes settings validation to fail.

## Bash protection

**Bash protection** scores overall operational risk and four independent hazards: destructive effects, credential exposure, privilege or persistence, and untrusted code execution. Any dimension reaching the selected risk threshold or falling below its confidence threshold requires approval. State includes command text, working directory, workspace, and explicit unknowns; scripts and environment values are not inspected.

Protection levels, from least to most restrictive:

- `permissive`
- `relaxed` (default)
- `balanced`
- `cautious`
- `strict`

Mission Control shows Jev's score, confidence, probabilities, cache status, and decision on the tool card. Flagged commands open an approval prompt before execution.

Metadata records all dimensions and the returned model. The displayed score is the highest dimension score. Commands over 32 KiB cannot be assessed and use the configured failure policy.

**API failure policy** controls what happens when Jev cannot assess a command:

- `auto_approve` allows execution. This is the default.
- `auto_deny` blocks execution.

## Bash file tracking

**Track Bash file activity** adds inferred reads, writes, and edits to **Session files**. Enable it under `/settings` → **Internal Tooling**, or set `internal_tooling.jev.bash_file_tracking.enabled` to `true` with Jev enabled. Restart after saving.

The tracker extracts literal path candidates without executing them. Plain `cat`, `head`, and `tail` calls with literal filenames are handled locally. Other supported commands go to Jev in one batch. Entries carry an **(inferred)** label; a later dedicated-tool operation replaces the inferred label for that operation. Inferences survive session reloads but never grant `hash_edit` anchors, load directory instructions, or count as confirmed mutations.

Bash transcript cards show a separate **Jev: file tracking** row with inferred file counts, R/W/E counts, and minimum confidence among accepted classifications. Local-only reads say **local**. Empty results, skipped assessments, failures, and timeouts get an outcome instead of file counts. Disabled tracking adds no row.

Limits and omissions:

- At most 16 KiB of command text and 32 candidate paths per call. Jev must return both confidence and selected probability of at least `0.9`. This starting threshold has not been calibrated against shell workloads.
- One background assessment at a time across sessions and subagents. It runs alongside Bash; short commands may wait for the remainder of a two-second assessment budget. Busy workers, cancellation, API errors, and late answers omit tracking without changing command results; the completed card reports the outcome.
- Only commands exiting with code `0` contribute entries. Paths must identify regular files after execution. Deleted files, directories, final symlinks, computed paths, glob expansion, working-directory changes, and unsupported control flow are omitted. A successful exit still does not prove every inferred operation happened.
- No repository scan or file-content upload. The bounded command, including inline scripts or heredocs, and literal candidate paths may be sent to TypeSafe. Commands matching secret-redaction patterns are skipped. Redaction cannot detect every secret; leave tracking disabled for commands you cannot send to TypeSafe.

Tracking assessments bypass the shared Jev response cache and do not retry. Their results are saved as local tool-result metadata, not added to provider context.

## Prose protection

**Prose protection** checks added or replacement documentation and code comments through `write` or `hash_edit`. It compares prepared before/after content under file mutation locks, before any writes. Deletion-only edits and unchanged files require no assessment. Prose presence and quality run together; quality matters only when prose is present. A style rejection tells the agent to load `$writer-humanizer`, revise, and retry.

Settings:

- **Prose detection threshold** decides when an edit counts as comments or documentation. Default: `0.7`.
- **Humanized threshold** sets the probability required to allow prose. Default: `0.65`.
- **API failure policy** either allows or denies the edit when assessment fails. Default: `auto_deny`.

Both thresholds must be between `0` and `1`.

Changed passages have a 32 KiB evidence budget, with bounded surrounding context. Budget, file-read, and API failures use the failure policy and are reported as unavailable assessments, not poor prose. Quoted examples, code samples, generated content, and necessary technical terms are excluded by the rubric; these judgments are not guaranteed.

## Prompt injection protection

**Prompt injection protection** assesses selected untrusted tool outputs before they return to provider context. It checks for instructions aimed at an agent, injection likelihood, and potential impact.

Assessment defaults to observe-only. Enable **Enforce decisions** to apply the result:

- `annotate` retains risky output with a warning.
- `quarantine` withholds risky output from provider context. This is the default failure policy.

Choose a protection level from `permissive`, `relaxed`, `balanced` (default), `cautious`, or `strict`. **Maximum assessed bytes** bounds content sent for one assessment; default is 32 KiB. **Protected tools** defaults to `web`.

`write`, `hash_edit`, and `magi_control` results are trusted control output and bypass this assessment.

With enforcement enabled, output exceeding the assessment budget is withheld in full, without an API call. Observe-only mode preserves it unchanged. The `annotate` failure policy can forward unassessed output with a warning after an API failure; it does not override the oversized-output rule. `escalate`, like `quarantine`, withholds output; it does not open an approval prompt.

Raw results and execution effects are saved before assessment, with normal session redaction. Assessment records link to a unique result id; intervening background events do not change that link. Live delivery checks the assessed bytes and SHA-256; replay and compaction check full assessment coverage plus the redacted stored content's bytes and SHA-256. Enforced pending or incomplete records remain withheld. Output without recorded protection metadata replays unchanged with a visible warning, not an implied assessment. Start a new session if that history is untrusted. Existing summaries are not reassessed.

Only configured tools are covered. Semantic detection can miss adversarial instructions; permissions and source trust still apply independently.

## Completion verification

**Completion verification** evaluates a proposed terminal response against bounded prior context and current-turn evidence. It checks request satisfaction, accurate disclosure of blockers, support for material claims, and scope. Each dimension selects `pass`, `fail`, `insufficient_evidence`, or `not_applicable`.

Assessment defaults to observe-only. With **Enforce completion decisions** enabled, supported failures trigger bounded corrective continuation or a limitation message. Low confidence or insufficient evidence produces `unknown`, accepts the response with a warning, and does not invent unfinished work. Service or evidence-budget failures produce `unavailable`. Only positive assessments report `passed`; diagnostics retain judgments, coverage, model, and cache status.

Settings:

- **Completion threshold** applies to both selected probability and confidence for each dimension. Default: `0.7`.
- **Maximum evidence bytes** bounds state sent to Jev. Default: 32 KiB.
- **Maximum continuations** caps verifier-triggered continuations per user turn. Default: `1`.
- **Verifier failure policy** accepts with a warning, or rejects with a warning when enforcement is enabled and assessment is unavailable. Default: `accept_with_warning`.
- **Show Jev debug messages** displays successful completion-check diagnostics. Default: off. Warnings remain visible.

Request and proposed response remain intact or assessment fails explicitly. History and tool receipts retain complete JSON records, recent outcomes, and omission counts; long fields use marked excerpts. Missing evidence is uncertainty, not proof of failure. Confidence measures answer-distribution concentration, not task correctness. Thresholds require evaluation on representative workflows; synthetic tests verify application behavior, not Jev accuracy.

## Skill suggestion

**Skill suggestion** checks each prompt you type against your installed skills. When one clearly fits, it adds a short note after your prompt naming that skill. The agent can ignore the note if the skill does not fit what you asked for.

It is off by default. When off, requests are unchanged. The note is added as a user message, so the system prompt and prompt caching are not affected. It is saved with the session and restored when the session is reloaded.

No note is added for steering messages, automatic compaction, subagents, prompts that already name a skill with `$name`, or projects with fewer than two skills. If Jev fails, the prompt is sent without a note.

Settings:

- **Gate threshold** decides whether the prompt needs a skill at all. Default: `0.3`.
- **Fit threshold** sets how well the best skill must fit before it is suggested. Default: `0.3`.

Each typed prompt waits for up to two Jev requests while this is enabled.

## Compaction timing

**Compaction timing** can compact a session before the normal limit when you start unrelated work. Between the soft threshold and automatic compaction limit, Jev compares the incoming prompt against recent user/assistant history, excluding that prompt. It asks whether the request depends on prior task details; ambiguous follow-ups should not count as independent work.

It is off by default. The automatic compaction limit still applies, and a failed Jev request leaves compaction unchanged.

Settings:

- **Soft threshold percent** sets the share of the model's context window where checks begin. Range: `1` to `99`. Default: `60`.
- **Threshold** sets the probability of unrelated work required to compact early. Default: `0.7`.

Compaction timing requires Jev to be enabled. It does not verify generated summaries.

## Cache and data handling

Jev requests send the state needed for an enabled assessment to TypeSafe. This may include commands, proposed file edits, selected tool output, or bounded conversation evidence. Do not enable a feature for data you cannot send to TypeSafe.

Responses are cached by request hash under the magi-code cache directory in `jev/`; cache files use owner-only permissions on Unix. Open `/settings` and choose **Reset Jev cache** to delete them. A repeated request may use its cached response instead of calling TypeSafe again.

Requests pin `jev-1.13.0`; the request hash includes that model and the questions. Assessment metadata retains the returned model where available. Cache writes are best effort; network requests do not hold the cache lock. Reset prevents in-flight requests in the current process from restoring deleted entries. Cached clients allow one retry for HTTP 408, 429, or server errors within the original 15-second deadline. Background inference keeps its own deadline without retries. Response bodies are limited to 2 MiB.

## Troubleshooting

- `TYPESAFE_API_KEY is required when Jev tooling is enabled`: export the key in the environment that launches magi-code, then restart.
- Unexpected blocks after an API error: check the feature's failure policy.
- Prompt injection findings do not change behavior: enable **Enforce decisions**; assessment alone is observe-only.
- Completion findings do not trigger correction: enable **Enforce completion decisions**.
- Stale-looking assessment: reset the Jev cache and retry.