openkind-server
The
openkinddproduction inference daemon.
openkind-server provides the openkindd daemon binary that powers self-hosted openkind deployments. It sets up structured tracing, initializes the EngineRegistry, mounts both HTTP (axum) and gRPC (tonic) listeners, registers Prometheus metrics, and manages graceful shutdown on Ctrl-C (SIGINT, plus SIGTERM on Unix).
Installation & Build
Homebrew (macOS & Linux)
Cargo
Build from source
# Binary produced at target/release/openkindd
Command-Line Usage
Options
| Flag | Environment Variable | Default | Description |
|---|---|---|---|
--http-addr <ADDR> |
OPENKIND_HTTP_ADDR |
127.0.0.1:8080 |
Address to bind HTTP server on |
--grpc-addr <ADDR> |
OPENKIND_GRPC_ADDR |
127.0.0.1:9090 |
Address to bind gRPC server on (port 0 or literal 0/off/none/disabled disables gRPC) |
--models <ALIASES> |
OPENKIND_MODELS |
mock,jev-latest |
Comma-separated list of model aliases to register |
--qwen35-aliases <ALIASES> |
OPENKIND_QWEN35_ALIASES |
qwen35-native |
Aliases in --models that use the native Qwen engine |
--qwen35-bundle-root <PATH> |
OPENKIND_QWEN35_BUNDLE_ROOT |
(none) | Offline selected-profile bundle root required by native aliases |
--qwen35-checkpoint-root <PATH> |
OPENKIND_QWEN35_CHECKPOINT_ROOT |
(none) | Offline pinned checkpoint root required by native aliases |
--qwen35-tokenizer <PATH> |
OPENKIND_QWEN35_TOKENIZER |
(none) | Digest-locked tokenizer JSON required by native aliases |
--qwen35-concurrency <N> |
OPENKIND_QWEN35_CONCURRENCY |
1 |
Maximum concurrent native evaluations |
--qwen35-queue <N> |
OPENKIND_QWEN35_QUEUE |
2 |
Additional native requests allowed to wait |
--qwen35-max-tensor-bytes <N> |
OPENKIND_QWEN35_MAX_TENSOR_BYTES |
(none) | Optional continuation tensor-payload ceiling per request |
--qwen35-max-process-bytes <N> |
OPENKIND_QWEN35_MAX_PROCESS_BYTES |
(none) | Optional process-memory admission ceiling |
--api-key <TOKEN> |
OPENKIND_API_KEY |
(none) | Optional bearer token required for /v1/* and gRPC |
--log-filter <FILTER> |
RUST_LOG |
info |
Tracing filter (e.g. info, debug, openkind=trace) |
Optional API keys
Authentication is off when no key is configured. --api-key <TOKEN> or
OPENKIND_API_KEY enables checking on /v1/* and gRPC. /health and
/metrics stay public. Clients must send the configured key as a bearer token.
Generate a key with the CLI and supply it before starting the daemon:
Configured keys must be nonempty visible ASCII without whitespace. Invalid values fail startup without printing the secret. Keys are not persisted. The CLI guide covers generation and PowerShell; the environment reference covers aliases and conflicting settings.
Example
# Start openkindd locally with mock backends
# Query health
# Evaluate a request
|
Native registration is opt-in and fail-closed. Add a Qwen alias to --models
and supply all three offline artifact paths. Choice requests must reserve the
literal __none__ option for native semantic-none mass. Choice and Score
confidence use normalized entropy; internal review policy is not serialized as
semantic none. When a process-memory ceiling is configured, the engine refreshes
the loaded process high-water mark before each evaluation and conservatively
splits remaining headroom across the configured concurrency limit.
Testing