ferrox-server 0.15.2

OpenAI-compatible HTTP server for the Ferrox inference engine
Documentation

ferrox-server

OpenAI-compatible HTTP server for Ferrox (ferrox-server binary).

Serves chat/completions against a GGUF (or Kimi safetensors dir) via FERROX_MODEL_PATH. See docs/API.md.

Parallel serving (Metal)

Multiple concurrent clients are supported via continuous batching (llama.cpp slots + one batched decode worker). On Metal, this is on by default when KV pool and prefix cache are not configured.

ferrox serve -m model.gguf -dev metal -ngl all              # CB auto-on
ferrox serve -m model.gguf -dev metal -cb -np 4             # explicit slot cap
ferrox serve -m model.gguf -dev metal --no-cont-batching      # private path (serialized on Metal)

See docs/plans/metal-parallel-concurrency.md.

The web UI is a separate app

This binary serves the HTTP API and nothing else. GET / is a 404 like any other unknown path. Ferrox Studio, the chat / models / activity / connect frontend, lives at ui/ in the repository root and reaches this server over the same public API an editor would use.

cargo run -p ferrox-server -- -m model.gguf     # terminal 1
cd ui && npm install && npm run dev             # terminal 2 -> :5173

npm run dev proxies /v1, /admin, /health, /metrics and forwards to the server address printed at startup.