# ferrum-infer-rs
[](https://crates.io/crates/ferrum-cli)
[](https://github.com/sizzlecar/ferrum-infer-rs/blob/main/LICENSE)
> Rust-native LLM inference for OpenAI-compatible local and private serving.
**One binary. No Python runtime. Apple Silicon Metal and NVIDIA CUDA acceleration.**
[中文说明](README_zh.md)
## Quick Start
Install Ferrum:
```bash
brew tap sizzlecar/ferrum
# macOS Apple Silicon
brew install ferrum
# Linux x86_64, NVIDIA CUDA sm89
brew install ferrum-cuda
```
Run a model directly:
```bash
# macOS Metal (GGUF)
ferrum run qwen3.5:4b-q4_k_m
# Linux CUDA (safetensors)
ferrum run qwen3.5:4b
```
Serve the same model through an OpenAI-compatible API:
```bash
# macOS Metal
ferrum serve --model qwen3.5:4b-q4_k_m --served-model-name ferrum --port 8000
# Linux CUDA
ferrum serve --model qwen3.5:4b --served-model-name ferrum --port 8000
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"ferrum","messages":[{"role":"user","content":"Hello"}]}'
```
## Features
- `ferrum run` and `ferrum serve` in one Rust binary.
- OpenAI-compatible chat completions, streaming, tools, and structured output.
- Apple Silicon Metal and NVIDIA CUDA from the same runtime.
- Continuous batching, paged KV cache, prefix cache, and typed admission control.
- GGUF on Metal and GPTQ/safetensors on CUDA.
- v0.8 covers language-model inference only. Release scope: Qwen3.5 4B, Qwen3.5 35B-A3B,
Qwen3 30B-A3B, and Llama 3.1 8B dense. [Support matrix](docs/release/runtime-vnext/0.8.0/SUPPORT_MATRIX.md).
## Performance Snapshot
Latest R2 development `ferrum serve` checkpoint. Metal uses random 64-token
input / 128-token output; CUDA uses 256 / 128. Values are mean tok/s with the
95% confidence-interval half-width across three repeats.
| Qwen3.5 4B | c=16 · 61.9 ± 0.1 | c=32 · 241.3 ± 0.6 |
| Qwen3.5 35B-A3B | c=4 · 26.1 ± 0.2 | c=16 · 174.1 ± 1.0 |
| Qwen3 30B-A3B | c=16 · 39.6 ± 1.2 | c=32 · 214.9 ± 2.7 |
`c` is active server concurrency. Every row completed 100 requests × 3 repeats
with zero errors. [Measurement details](docs/release/runtime-vnext/0.8.0/PERFORMANCE_REPORT.md).
## OpenAI-Compatible API
Ferrum supports:
- chat completions and streaming usage
- function tools with `auto`, `none`, `required`, or a named function
- `json_object` and strict `json_schema` structured output
- multi-turn sessions, prefix cache, and session cache
- typed concurrency, memory, and scheduler controls
See [OpenAI API compatibility](docs/openai-api-compatibility.md) for the exact
request contract and [cache product controls](docs/cache-product.md) for prefix
and session caching.
## Installation
Homebrew:
```bash
brew tap sizzlecar/ferrum
# macOS Apple Silicon Metal
brew install ferrum
# Linux x86_64 CUDA sm89
brew install ferrum-cuda
```
Prebuilt release tarballs:
```bash
# Linux x86_64 CUDA sm89
curl -L https://github.com/sizzlecar/ferrum-infer-rs/releases/latest/download/ferrum-linux-x86_64-cuda-sm89.tar.gz | tar xz
LD_LIBRARY_PATH=/usr/local/cuda/lib64:${LD_LIBRARY_PATH:-} ./ferrum --version
# macOS Apple Silicon Metal
```
Install from crates.io:
```bash
# macOS Apple Silicon Metal
cargo install ferrum-cli --version 0.8.0 --locked --features metal
# NVIDIA CUDA
cargo install ferrum-cli --version 0.8.0 --locked \
--features cuda,vllm-moe-marlin,vllm-paged-attn-v2
```
The official prebuilt CUDA asset targets `sm89`. CUDA installation requires a
compatible NVIDIA driver, CUDA runtime, and NCCL runtime on the target host.
## Architecture
- Contracts: `ferrum-types`, `ferrum-interfaces`
- Execution: `ferrum-engine`, `ferrum-scheduler`, `ferrum-kv`, `ferrum-sampler`
- Models and compute: `ferrum-models`, `ferrum-kernels`, `ferrum-native-ops`, `ferrum-quantization`
- Product surface: `ferrum-cli`, `ferrum-server`, `ferrum-tokenizer`
- Validation: `ferrum-bench-core`, `ferrum-testkit`
## License
MIT