pub struct InferencePools { /* private fields */ }Expand description
The world’s live per-model inference pools. Cheap to clone-share behind an
Arc; semaphores are created lazily the first time a model is seen.
Implementations§
Source§impl InferencePools
impl InferencePools
Sourcepub fn new(config: InferencePoolConfig) -> Self
pub fn new(config: InferencePoolConfig) -> Self
Build the pools from a configuration.
Sourcepub fn with_wake(self, wake: Arc<Notify>) -> Self
pub fn with_wake(self, wake: Arc<Notify>) -> Self
Attach the tick-loop wake handle, so that releasing a permit wakes the driver.
This is load-bearing, not a nicety. dispatch_inference leaves a
slot-starved agent ReadyToInfer to be retried “on a later tick”, and
the loop is event-driven: a later tick only happens when something wakes
it. So every path that frees a permit owes the loop a wake, or the freed
slot is invisible and everything queued behind it stays parked (issue
#189).
Hanging the wake off the permit’s Drop rather than off each release
site is the point: the obligation can’t be forgotten by a new call site,
and it covers the paths that don’t report an outcome at all - notably a
cancelled job, which frees its permit and returns with nothing to send.
Sourcepub async fn acquire(&self, model: &str) -> InferencePermit
pub async fn acquire(&self, model: &str) -> InferencePermit
Acquire a permit for model, waiting for a free slot if the pool is
full. The returned InferencePermit releases the slot when dropped -
so the caller holds it for exactly the duration of the inference request.
Sourcepub fn try_acquire(&self, model: &str) -> Option<InferencePermit>
pub fn try_acquire(&self, model: &str) -> Option<InferencePermit>
Try to take a permit for model without waiting. Returns None if
the pool is currently full.
This is what the synchronous ECS inference-dispatch system calls: a
system can’t .await, so instead of blocking on a full pool it leaves
the agent ReadyToInfer and retries on a later tick.
Sourcepub fn occupancy(&self) -> Vec<PoolOccupancy>
pub fn occupancy(&self) -> Vec<PoolOccupancy>
How many slots are in use per model, against each model’s cap. None as
the cap means the model is unbounded. Only models that have actually been
used appear - the semaphores are created lazily.
This is what makes “the pool is full and has been for hours” observable rather than inferred.
Trait Implementations§
Auto Trait Implementations§
impl !Freeze for InferencePools
impl RefUnwindSafe for InferencePools
impl Send for InferencePools
impl Sync for InferencePools
impl Unpin for InferencePools
impl UnsafeUnpin for InferencePools
impl UnwindSafe for InferencePools
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
impl<T> ConditionalSend for Twhere
T: Send,
Source§impl<T> Downcast for Twhere
T: Any,
impl<T> Downcast for Twhere
T: Any,
Source§fn into_any(self: Box<T>) -> Box<dyn Any>
fn into_any(self: Box<T>) -> Box<dyn Any>
Box<dyn Trait> (where Trait: Downcast) to Box<dyn Any>, which can then be
downcast into Box<dyn ConcreteType> where ConcreteType implements Trait.Source§fn into_any_rc(self: Rc<T>) -> Rc<dyn Any>
fn into_any_rc(self: Rc<T>) -> Rc<dyn Any>
Rc<Trait> (where Trait: Downcast) to Rc<Any>, which can then be further
downcast into Rc<ConcreteType> where ConcreteType implements Trait.Source§fn as_any(&self) -> &(dyn Any + 'static)
fn as_any(&self) -> &(dyn Any + 'static)
&Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot
generate &Any’s vtable from &Trait’s.Source§fn as_any_mut(&mut self) -> &mut (dyn Any + 'static)
fn as_any_mut(&mut self) -> &mut (dyn Any + 'static)
&mut Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot
generate &mut Any’s vtable from &mut Trait’s.