ferrox-server
OpenAI-compatible HTTP server for Ferrox (ferrox-server binary).
Serves chat/completions against a GGUF (or Kimi safetensors dir) via
FERROX_MODEL_PATH. See docs/API.md.
Parallel serving (Metal)
Multiple concurrent clients are supported via continuous batching (llama.cpp slots + one batched decode worker). On Metal, this is on by default when KV pool and prefix cache are not configured.
See docs/plans/metal-parallel-concurrency.md.
0.15.3 serving receipt (Host B, Llama-3.2-3B Q4_K_M, CB on): concurrency
1→8 all OK, ~24 aggregate tok/s at 8 clients, mean TTFT ~118 ms sequential /
~957 ms at concurrency 8. Receipts under
benchmarks/receipts/serving/.
The web UI is a separate app
This binary serves the HTTP API and nothing else. GET / is a 404 like
any other unknown path. Ferrox Studio, the chat / models / activity /
connect frontend, lives at ui/ in the repository root and
reaches this server over the same public API an editor would use.
&& &&
npm run dev proxies /v1, /admin, /health, /metrics and
forwards to the server address printed at startup.