Skip to main content

Module llama_subprocess

Module llama_subprocess 

Source
Expand description

Windows LLM via a subprocess llama-cli (mirrors the sd-cli pattern) so Windows workers reach in-process-llama parity.

llama-cpp-2 (the in-process backend) doesn’t link on Windows MSVC (a static-vs-dynamic CRT clash, documented in Cargo.toml), so a Windows release worker would otherwise fall back to the synthetic LLM. Instead we auto-provision the official llama.cpp Windows Vulkan release binary into <models_root>/bin/ on first use and run llama-cli per request, exactly like the image engine runs sd-cli.

The module is always compiled (so its pure argv / response logic and the provisioner are unit-tested on every platform), but only registered on Windows — Linux/macOS keep the faster in-process llama-cpp-2 backend.

Functions§

asset_name
The Windows-Vulkan release asset name for build.
build_argv
Build the llama-cli argument vector for a chat completion. Pure so the flag set is unit-tested without a binary. --no-display-prompt keeps the echoed prompt out of stdout so only the completion is captured; -st (single-turn) + -no-cnv runs one non-interactive turn and exits.
prompt_from_params
The last non-empty line of prompt_for-style chat input: the user turn we feed llama-cli. (llama-cli takes a single -p string; a full chat template is a future enhancement.)
provision
Ensure llama-cli is present under <models_root>/bin/, downloading + extracting the pinned release on first use. Excluded from coverage: the happy path needs a real multi-hundred-MB download; the URL/asset/argv/response logic is unit-tested and the extractor is fixture-tested.
resolve_url
The full download URL: a STUDIO_WORKER_LLAMA_URL override wins (air-gapped mirror / tests), else the pinned GitHub release asset.
select_build
Resolve the release build tag: STUDIO_WORKER_LLAMA_BUILD wins, else the pinned default.
wrap_response
Wrap raw llama-cli stdout in an OpenAI chat.completion object so the local API / studio consumers parse it uniformly with the synthetic + in-process engines. Pure.