Expand description
The model host: owns loaded models, drives each model’s lifecycle,
persists residency and enforces admission and exclusive groups
(see docs/runtime/model-lifecycle.md).
Synchronous by design (the local API is a thread-pool server): loads
and unloads run on their own threads and report back through the
lifecycle. Lock order is always entries before residency.
Structs§
- Lane
- The serving path of one loaded model: one request at a time, and a cancel flag an unload raises so a long request (a stream) can end.
- Model
Host - Cheap to clone; every clone is the same host.
- Model
Status - One model’s observable status.
Enums§
Constants§
- DRAIN_
TIMEOUT - How long an unload waits for the request in flight to notice it was cancelled before freeing anyway (the request keeps the weights alive until it returns). A streaming chunk takes well under a second. Safe range 1..=60 s.
- SWAP_
TIMEOUT - How long a swap waits for the outgoing group member to unload before
loading anyway. Covers
DRAIN_TIMEOUTplus freeing. Safe rangeDRAIN_TIMEOUT..=120 s.
Traits§
- Chat
Model - A loaded model that answers chat completions.
- Loaded
Model - A model whose weights are in memory. Engines downcast it back to
their own type through
as_any. - Model
Runtime - Loads catalogue models into memory. Freed by dropping the result.
- Streaming
Model - A loaded streaming speech model. Each
openis an independent utterance state over the shared weights.