# The containerised test environment
JupyterLab in a browser tab, a terminal inside it, and `mindfork` built from the
current working tree talking to a CPU-only llama.cpp stack. One command up, one
command down, nothing installed on the host but Docker.
Why it exists, what was measured and which alternatives were rejected:
[docs/research/docker-jupyter-env.md](../docs/research/docker-jupyter-env.md).
## 1. Start
```bash
cd docker && cp .env.example .env && docker compose up --build
```
Then open **<http://localhost:8888/lab?token=mindfork>** → *Other → Terminal* →
type `mindfork`.
The first run downloads ~6.2 GiB of GGUF weights and compiles the crate, so it
takes a while; both are cached afterwards. Later starts are `docker compose up`
and are up in seconds, plus however long llama.cpp needs to load the weights.
Stop with `docker compose down`. The models and the app's data survive it — they
live in named volumes. `docker compose down -v` is the reset button that removes
both.
## 2. What runs where
| `models` | one-shot download into the `models` volume, then exits | — |
| `chat` | `llama-server` with Gemma 4 E2B-it Q8_0 (+ the vision projector) | `chat:8000` inside, `localhost:8000` outside |
| `embed` | `llama-server --embeddings` with bge-m3 Q8_0 | `embed:8001` inside, `localhost:8001` outside |
| `lab` | JupyterLab 4 + the `mindfork` binary | `localhost:8888` |
Inside the lab container:
- `~/mindfork/data` — the app's data root (`settings.json`, `chats/*.json`,
`data.db`, `logs/`). A named volume; it survives image rebuilds.
- `~/work` — the host's `docker/work/`, so dropping a file there from Windows
makes it reachable by `/file attach` or `/rag add`.
- `mindfork` is on `PATH`; the real binary and the spellcheck dictionaries sit in
`/opt/mindfork`, mirroring the layout of the Linux packages.
The app is pre-configured through a seeded `settings.json` — engine and
embeddings in **external** mode pointing at the two servers. Everything stays
editable in the settings screen and every change persists, including switching
the engine to a cloud provider; that is why the environment does *not* use
`MINDFORK_ENGINE_URL`, which would pin the mode to external on every launch.
## 3. After a code change
```bash
docker compose up -d --build lab
```
Only the app is rebuilt; the models stay loaded in the two server containers.
The cargo registry and the `target/` directory are BuildKit cache mounts, so this
is an incremental compile, not a fresh one.
## 4. Running the live smokes from the host
The two server ports are published, so the mandatory `#[ignore]` gate
(AGENTS.md §3) can run from Windows against this stack with no GPU:
```powershell
$env:MINDFORK_ENGINE_URL = "http://127.0.0.1:8000/v1"
$env:MINDFORK_EMBED_URL = "http://127.0.0.1:8001/v1"
cargo test -- --ignored --nocapture --test-threads=1
```
Measured on the first run (2026-08-26), the protocol group
`cargo test ignored_smoke -- --ignored` gives **22 passed, 1 failed in 243 s**:
streaming, anti-self-termination on EOS text, tool-call parsing, thoughts, the
sampling extensions, the typed non-transient 400 on an oversized prompt, and the
vision path all pass. The failure is `control_tools_are_callable` — the
2B-effective model did not call `rewrite_current_message`.
**Which is exactly why this is not the gate of record.** Several smokes assert on
model *behaviour* and were calibrated on Gemma 4 31B / Qwen 3.6 27B; a small
model fails some of them for reasons that are not defects. Use this for protocol,
streaming, tool-call and RAG plumbing, and keep `python tools/e2e_hf.py run` for
a verdict worth writing "Smoke — GO" about. The full `cargo test -- --ignored`
also runs here, but at CPU speed it is an hours-long affair.
## 5. Knobs
Everything is in `.env` (copied from `.env.example`, which documents each one).
The ones that come up most:
- `CHAT_GGUF` — swap the quant. `Q4_K_M` (~2.7 GiB) roughly halves the memory and
doubles the speed; the download follows the file name automatically.
- `CHAT_CTX` — the chat server's context window, `16384` by default. It is the
KV cache, so lowering it back to `8192` is the first thing to try on a
memory-tight box.
- `CHAT_EXTRA_ARGS` / `EMBED_EXTRA_ARGS` — appended verbatim to the server
command lines. Where `--mmproj` lives, and where `-t <threads>` goes.
- `MODELS_DIR` — a bare name is the named volume, a path is a bind mount of a
host directory of GGUFs you already have.
- `LAB_LANG` — `ru` or `en`, applied while the data volume is still empty.
- `LAB_THEME` — `auto` (default) / `light` / `dark`. `auto` asks the terminal
for its background at start-up and matches to it, and JupyterLab's terminal
answers — so switching the *lab* between its light and dark themes is enough,
and nothing needs pinning here. The question is asked once at start-up, so
restart `mindfork` after switching the lab theme. `light`/`dark` pin a
polarity regardless of what the terminal says.
- `LAB_SSH_PUBKEY` / `LAB_SSH_PUBKEY_FILE` — a **public** key turns on an sshd
inside the container; empty (the default) means none runs. See §6.1.
- `LAB_SSH_PORT` — host-side port for it, `2222` by default.
- `OPENAI_API_KEY` / `GEMINI_API_KEY` / `ANTHROPIC_API_KEY` / `XAI_API_KEY` —
passed through to the app, and the seeded settings already name them, so
switching the engine to a cloud provider in the settings screen just works.
## 6. Things worth knowing
### 6.1. SSH in, when the browser terminal is not the point
The terminal in JupyterLab is xterm.js: one emulator, with one set of habits.
Anything terminal-facing behaves differently elsewhere: background-colour
reporting, clipboard escapes, keyboard protocols. Measuring that needs hosts
other than the browser's — and one of them is always "plain SSH into a Linux
box". This is that box, without leaving the compose file.
Put a **public** key in `.env` and bring the lab up:
```bash
LAB_SSH_PUBKEY="ssh-ed25519 AAAAC3... you@host"
```
```bash
ssh -p 2222 jovyan@127.0.0.1
```
`mindfork` and `tmux` are both on `PATH` there — and so is the rest of the
container's environment. That is not free: an SSH session does not inherit
Docker's `ENV`, and the usual fallback (PAM reading `/etc/environment`) needs a
**root** daemon, which this deliberately is not. Left alone, a session lands on
the bare system `PATH` — no `conda` (a "command not found" on every login, from
the base image's own `.bashrc`), no `node`, and no `mcp-server-filesystem`, so
`mindfork` started over SSH would have a broken plugin host while the same
binary in the browser terminal works. The hook therefore writes `SetEnv` into
the config, taking `PATH` and `LANG` from its own environment so they follow the
image. With no key set, no daemon runs at all.
What it is, precisely: **sshd as uid 1000**, on port 2222, keys only, `jovyan`
only, published on `127.0.0.1` alone. A session lands exactly where
`docker exec` already lands, so this adds a transport rather than a privilege —
which is also why it needs no root and gets none. The host key lives on the data
volume (`.lab-ssh/`), so recreating the container does not greet you with
REMOTE HOST IDENTIFICATION HAS CHANGED. `StrictModes` is off because the base
image makes `$HOME` group-writable and sshd refuses to serve such a home; with a
single user behind a loopback-only port there is nothing left for that check to
protect.
It is started by `before-notebook.d/20-sshd`, which reports on the container's
log either way — a daemon that quietly failed to start looks exactly like a
wrong port from the outside.
- **RAM.** Weights 4.63 GiB + KV cache (`CHAT_CTX`, 16384 by default) + the
projector ~0.94 GiB + the embedder
0.6 GiB + JupyterLab: budget ~9–10 GiB for the Docker VM. Below that the chat
server is OOM-killed while loading and it looks like a hang —
`docker compose logs chat` says so plainly. On Windows the limit is
`%UserProfile%\.wslconfig` (`[wsl2] memory=…`).
- **Speed.** CPU-only inference of a Q8_0 model is single-digit-to-low-teens
tokens per second. Fine for exercising the UI, tool loops and RAG; slow for
long generations.
- **Stored cloud keys do not survive an image rebuild.** The Linux key scheme
derives from `/etc/machine-id`, which is generated per build. Name the
environment variable instead (§5) — that is the documented alternative and it
is stable here.
- **TERM is left exactly as JupyterLab sets it.** Reproducing that terminal is
the point; normalising it would hide the very differences this environment
exists to expose.
- **`docker compose exec lab bash`** gives the same binary in a normal terminal —
useful for telling "broken" apart from "broken *in a browser terminal*".
- **One instance per machine — and a closed browser tab does not close the app.**
JupyterLab's terminal is a server-side pty: closing the tab leaves `mindfork`
running, and the next launch is refused with *"mindfork is already running on
this machine"* (spec §1.4). Reattach from the *Running Terminals* panel in the
left sidebar, or `docker compose exec lab pkill mindfork`.
- **An `up` that failed while *starting* leaves a container that then starts
without its ports.** Compose does not recreate a container whose configuration
has not changed — it starts the one already there, and if that container's
first start failed (a port taken by something else, an OOM) the record can be
left half-initialised: the port bindings are recorded, but the container joins
no network and publishes nothing. The next `up` reports it *healthy* — that
check runs inside the container — while `localhost:8888` refuses the
connection and the app cannot resolve `chat`/`embed` either. `docker ps` tells
the two apart: a published port reads `0.0.0.0:8888->8888/tcp`, an unpublished
one is a bare `8888/tcp`. The fix is
`docker compose up -d --force-recreate lab`, and it costs nothing — the data
lives on the volume.
## 7. MCP plugin tools
The image carries **Node 26** and npm, so any `npx` server from the Model
Context Protocol ecosystem runs here (docs/install.md §4.2). The reference
filesystem server is installed at build time and pre-configured, scoped to the
mounted `~/work` directory:
```jsonc
"mcp": {
"enabled": false, // ← the master switch, deliberately off
"servers": [{ "id": "fs",
"command": "mcp-server-filesystem",
"args": ["/home/jovyan/work"],
"enabled": true }]
}
```
To use it: `Ctrl+P` → **Plugins** → turn the master switch on, then enable the
tools on the profile (the host is double opt-in by design — an MCP server is an
arbitrary user-privileged program, so a convenience seed is not allowed to be
what turns it on).
Verified in the container — `secure-filesystem-server 0.2.0`, protocol
`2025-06-18`, 14 tools — and end to end through the app: asked whether any Python
files were reachable, Gemma 4 E2B-it called `mcp__fs__list_allowed_directories`,
got back the one allowed directory, chained
`mcp__fs__search_files(pattern=*.py)` and listed what it found.
Addressed by its binary rather than `npx @modelcontextprotocol/server-filesystem`
on purpose — `npx` would reach the registry on every launch. For any other
server, the `npx` spelling in the install docs works as written.
An environment created before this existed keeps its own `settings.json` (the
start-up hook never overwrites one), so its Plugins section is empty. Either
`docker compose down -v` for a clean slate, or add just the server entry by hand
in `Ctrl+P` → Plugins → `Ctrl+N`.
## 8. What does not work in here, and why that is fine
- **The system clipboard.** There is no X server, so `F5` cannot reach one — and
OSC 52 does not survive JupyterLab either. That is not a container defect, it
is the exact situation `/export` exists for
([docs/history/chat-export-file.md](../docs/history/chat-export-file.md)); the
file lands in the folder the file browser is already showing.
- **Speech (`/tts`).** No audio device. The app says "audio unavailable" and
carries on, which is the designed behaviour.
- **The Wasmer Python sandbox** is not provisioned — `python_exec` is seeded in
*local* mode instead, against the container's own Python 3.13. To test the
sandbox itself: `docker compose exec lab mindfork sandbox setup` (~300 MB into
the data volume).