Skip to main content

Module host

Module host 

Source
Expand description

The model host: owns loaded models, drives each model’s lifecycle, persists residency and enforces admission and exclusive groups (see docs/runtime/model-lifecycle.md).

Synchronous by design (the local API is a thread-pool server): loads and unloads run on their own threads and report back through the lifecycle. Lock order is always entries before residency.

Structs§

Lane
The serving path of one loaded model: one request at a time, and a cancel flag an unload raises so a long request (a stream) can end.
ModelHost
Cheap to clone; every clone is the same host.
ModelStatus
One model’s observable status.

Enums§

HostError

Constants§

DRAIN_TIMEOUT
How long an unload waits for the request in flight to notice it was cancelled before freeing anyway (the request keeps the weights alive until it returns). A streaming chunk takes well under a second. Safe range 1..=60 s.
SWAP_TIMEOUT
How long a swap waits for the outgoing group member to unload before loading anyway. Covers DRAIN_TIMEOUT plus freeing. Safe range DRAIN_TIMEOUT..=120 s.

Traits§

ChatModel
A loaded model that answers chat completions.
LoadedModel
A model whose weights are in memory. Engines downcast it back to their own type through as_any.
ModelRuntime
Loads catalogue models into memory. Freed by dropping the result.
StreamingModel
A loaded streaming speech model. Each open is an independent utterance state over the shared weights.