Expand description
Windows LLM via a subprocess llama-cli (mirrors the sd-cli
pattern) so Windows workers reach in-process-llama parity.
llama-cpp-2 (the in-process backend) doesn’t link on Windows MSVC
(a static-vs-dynamic CRT clash, documented in Cargo.toml), so a
Windows release worker would otherwise fall back to the synthetic
LLM. Instead we auto-provision the official llama.cpp Windows
Vulkan release binary into <models_root>/bin/ on first use and run
llama-cli per request, exactly like the image engine runs
sd-cli.
The module is always compiled (so its pure argv / response logic and
the provisioner are unit-tested on every platform), but only
registered on Windows — Linux/macOS keep the faster in-process
llama-cpp-2 backend.
Functions§
- asset_
name - The Windows-Vulkan release asset name for
build. - build_
argv - Build the
llama-cliargument vector for a chat completion. Pure so the flag set is unit-tested without a binary.--no-display-promptkeeps the echoed prompt out of stdout so only the completion is captured;-st(single-turn) +-no-cnvruns one non-interactive turn and exits. - prompt_
from_ params - The last non-empty line of
prompt_for-style chat input: the user turn we feed llama-cli. (llama-cli takes a single-pstring; a full chat template is a future enhancement.) - provision
- Ensure
llama-cliis present under<models_root>/bin/, downloading + extracting the pinned release on first use. Excluded from coverage: the happy path needs a real multi-hundred-MB download; the URL/asset/argv/response logic is unit-tested and the extractor is fixture-tested. - resolve_
url - The full download URL: a
STUDIO_WORKER_LLAMA_URLoverride wins (air-gapped mirror / tests), else the pinned GitHub release asset. - select_
build - Resolve the release build tag:
STUDIO_WORKER_LLAMA_BUILDwins, else the pinned default. - wrap_
response - Wrap raw
llama-clistdout in an OpenAIchat.completionobject so the local API / studio consumers parse it uniformly with the synthetic + in-process engines. Pure.