Skip to main content

GpuRuntime

Struct GpuRuntime 

Source
pub struct GpuRuntime {
    pub device: GpuDeviceInfo,
    pub devices: Vec<GpuDeviceInfo>,
    pub policy: GpuDispatchPolicy,
    pub memory_budget_bytes: usize,
}

Fields§

§device: GpuDeviceInfo

Highest-scoring probed CUDA device. Existing dispatch code routes one-shot kernels through this device.

§devices: Vec<GpuDeviceInfo>

All usable CUDA devices discovered at probe time, ordered by score.

§policy: GpuDispatchPolicy§memory_budget_bytes: usize

Implementations§

Source§

impl GpuRuntime

Source

pub fn probe() -> Result<GpuAvailability, GpuError>

Source

pub fn availability() -> Result<GpuAvailabilityRef<'static>, GpuError>

Return the cached probe outcome without collapsing faults into absence.

Source

pub fn resolve(policy: GpuPolicy) -> Result<Option<&'static Self>, GpuError>

Resolve CUDA under an explicit policy. Ok(None) is reserved for a genuine absence under Auto/Off; probe faults always remain Err, and Required converts absence into RequiredDeviceUnavailable.

Source

pub fn require() -> Result<&'static Self, GpuError>

Resolve CUDA under Required semantics and return the device handle.

Source

pub fn resolution_call_count() -> u64

Number of times Self::availability has been entered process-wide.

Test-facing instrumentation for the laziness contract: a size-gated caller that returns before resolving availability leaves this unchanged, so a test can assert a CPU-sized decision path created no CUDA context. This is a monotone call counter, NOT a probe-success flag.

Source

pub fn resolve_if_dense_work_exceeds_floor( policy: GpuPolicy, work_flops: u128, ) -> Result<Option<&'static Self>, GpuError>

Size-gated Self::resolve: resolve the process-wide runtime only when the estimated dense arithmetic work_flops clears the GPU-dispatch flop floor.

This is the ordering fix for the CUDA startup tax. For a CPU-sized problem (work_flops below the floor) it returns Ok(None) without calling Self::resolve, so the device probe — and the cuDevicePrimaryCtxRetain primary-context creation it performs on every GPU — never runs. The problem-size decision therefore strictly precedes any driver contact, and a CPU-sized fit pays ZERO CUDA cost.

The floor is GpuDispatchPolicy::MIN_CALIBRATABLE_GEMM_FLOPS — the smallest gemm_min_flops ANY reachable policy (default seed or device-calibrated) can carry, known WITHOUT a device — so the gate never needs a probe to decide it should not probe, and refusing below it can never block work that any policy would have dispatched. Work at or above the floor falls through to the identical lossless resolution path (where the real, possibly calibrated policy still gates each op), so device behaviour for genuinely GPU-sized problems is unchanged.

Source

pub fn resolve_if_fused_batch_exceeds_floor( policy: GpuPolicy, rows: usize, ) -> Result<Option<&'static Self>, GpuError>

Size-gated Self::resolve for independent fused row kernels.

Batches below GpuDispatchPolicy::MIN_CALIBRATABLE_FUSED_KERNEL_N cannot be admitted by either the default or any device-calibrated policy. Refuse them before availability resolution so a CPU-sized first call does not create CUDA contexts and run calibration merely to learn that it should stay on the CPU. At and above the universal floor, the concrete runtime’s calibrated policy remains authoritative.

Source

pub fn policy(&self) -> &GpuDispatchPolicy

Source

pub fn selected_device(&self) -> &GpuDeviceInfo

Source§

impl GpuRuntime

Source

pub fn device_ordinals(&self) -> Vec<usize>

Ordinals of all usable devices, highest-score first.

self.devices is already score-sorted at probe time, so this simply projects out the ordinals. Empty only if the probe somehow produced no devices (the public probe() guarantees at least one on Ok(Some(_))).

Source

pub fn device_count(&self) -> usize

Number of usable devices in the pool.

Source

pub fn memory_budget_for(&self, ordinal: usize) -> usize

Per-device byte budget: free memory capped at half of total, matching the primary-device budget computed in device_runtime::probe. Falls back to the primary memory_budget_bytes when the ordinal is not in the pool so a caller that passes a stale ordinal still gets a usable (conservative) budget rather than zero.

Trait Implementations§

Source§

impl Clone for GpuRuntime

Source§

fn clone(&self) -> GpuRuntime

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Debug for GpuRuntime

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> ByRef<T> for T

Source§

fn by_ref(&self) -> &T

Source§

impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> DistributionExt for T
where T: ?Sized,

Source§

fn rand<T>(&self, rng: &mut (impl Rng + ?Sized)) -> T
where Self: Distribution<T>,

Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T, U> Imply<T> for U
where T: ?Sized, U: ?Sized,

Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> IntoEither for T

Source§

fn into_either(self, into_left: bool) -> Either<Self, Self>

Converts self into a Left variant of Either<Self, Self> if into_left is true. Converts self into a Right variant of Either<Self, Self> otherwise. Read more
Source§

fn into_either_with<F>(self, into_left: F) -> Either<Self, Self>
where F: FnOnce(&Self) -> bool,

Converts self into a Left variant of Either<Self, Self> if into_left(&self) returns true. Converts self into a Right variant of Either<Self, Self> otherwise. Read more
Source§

impl<T> Pointable for T

Source§

const ALIGN: usize

The alignment of pointer.
Source§

type Init = T

The type for initializers.
Source§

unsafe fn init(init: <T as Pointable>::Init) -> usize

Initializes a with the given initializer. Read more
Source§

unsafe fn deref<'a>(ptr: usize) -> &'a T

Dereferences the given pointer. Read more
Source§

unsafe fn deref_mut<'a>(ptr: usize) -> &'a mut T

Mutably dereferences the given pointer. Read more
Source§

unsafe fn drop(ptr: usize)

Drops the object pointed to by the given pointer. Read more
Source§

impl<T> Read<Exclusive, BecauseExclusive> for T
where T: ?Sized,

Source§

impl<T> Same for T

Source§

type Output = T

Should always be Self
Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = Infallible

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, <T as TryFrom<U>>::Error>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<V, T> VZip<V> for T
where V: MultiLane<T>,

Source§

fn vzip(self) -> V