offgrid-server 0.2.1

OpenAI-compatible HTTP sidecar for offgrid
docs.rs failed to build offgrid-server-0.2.1
Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.

offgrid-server

OpenAI-compatible HTTP server for offgrid: local, offline chat, tool calling, structured output, embeddings, and image and audio input on top of llama.cpp (Vulkan or CPU).

cargo install offgrid-server
offgrid-server --model qwen3-4b.gguf [--embedding-model embed.gguf] [--mmproj mmproj.gguf] [--port 8080] [--api-key KEY]

Point any OpenAI client at http://127.0.0.1:8080/v1. Endpoints: POST /v1/chat/completions (with SSE streaming, reasoning_content and tool_calls deltas), POST /v1/embeddings, GET /v1/models, GET /health.

Building compiles llama.cpp from source, which needs CMake, a C/C++ compiler and, with the default vulkan feature, the Vulkan headers and glslc. Use --no-default-features for a CPU-only build. Prebuilt binaries are attached to the GitHub releases.

OFFGRID_LOG=debug shows llama.cpp's own logs; offgrid-server --help lists all options.