Expand description
Unified readiness pre-flight for local inference providers.
Local servers (Ollama, LM Studio, llama.cpp) are frequently stopped or have no model loaded. Previously generation simply failed with a raw connection or 404 error. This module centralizes the “is the server up?” and “is the requested model available?” checks so every local provider can return a single, actionable error (with the exact fix command) instead of a cryptic one. It also resolves a placeholder/default request model to the single loaded model when appropriate.
Design notes:
- Cloud Ollama models (
:cloud/-cloud) are remote and bypass readiness. - Results are cached per-process for a short TTL so we do not probe the
local server on every single generation, while still allowing a
ServerDownerror to clear quickly after the user starts the server. - This module is intentionally provider-agnostic: it only consumes the
LocalProviderenum and the per-providerfetch_*_modelshelpers.
Enums§
- Local
Readiness Error - Failure modes surfaced before a generation attempt.