Skip to main content

Module local_readiness

Module local_readiness 

Source
Expand description

Unified readiness pre-flight for local inference providers.

Local servers (Ollama, LM Studio, llama.cpp) are frequently stopped or have no model loaded. Previously generation simply failed with a raw connection or 404 error. This module centralizes the “is the server up?” and “is the requested model available?” checks so every local provider can return a single, actionable error (with the exact fix command) instead of a cryptic one. It also resolves a placeholder/default request model to the single loaded model when appropriate.

Design notes:

  • Cloud Ollama models (:cloud / -cloud) are remote and bypass readiness.
  • Results are cached per-process for a short TTL so we do not probe the local server on every single generation, while still allowing a ServerDown error to clear quickly after the user starts the server.
  • This module is intentionally provider-agnostic: it only consumes the LocalProvider enum and the per-provider fetch_*_models helpers.

Enums§

LocalReadinessError
Failure modes surfaced before a generation attempt.