pub struct WgpuServer<C: WgpuCompiler> {
pub compilation_options: WgpuCompilationOptions,
/* private fields */
}Expand description
Wgpu compute server.
Fields§
§compilation_options: WgpuCompilationOptionsImplementations§
Source§impl<C: WgpuCompiler> WgpuServer<C>
impl<C: WgpuCompiler> WgpuServer<C>
Sourcepub fn load_cached_pipeline(
&mut self,
kernel_id: &KernelId,
bindings: &KernelArguments,
mode: ExecutionMode,
) -> Result<Option<Result<PipelineEntry, (u64, KernelCacheKey)>>, CompilationError>
pub fn load_cached_pipeline( &mut self, kernel_id: &KernelId, bindings: &KernelArguments, mode: ExecutionMode, ) -> Result<Option<Result<PipelineEntry, (u64, KernelCacheKey)>>, CompilationError>
Loads a cached kernel if present and creates the pipeline for it.
Returns None if the cache isn’t enabled, Some(Ok(pipeline)) if a cache entry was found,
and Some(Err(cache_key)) if the cache is enabled but doesn’t contain this kernel.
pub fn create_module( &self, entrypoint_name: &str, cube_dim: CubeDim, source: ModuleSource<'_>, mode: ExecutionMode, ) -> Result<ShaderModule, CompilationError>
pub fn create_pipeline( &self, entrypoint_name: &str, repr: Option<AutoRepresentationRef<'_>>, module: ShaderModule, bindings: &KernelArguments, ) -> Arc<ComputePipeline> ⓘ
Source§impl<C: WgpuCompiler> WgpuServer<C>
impl<C: WgpuCompiler> WgpuServer<C>
Sourcepub fn new(
memory_properties: MemoryDeviceProperties,
memory_config: MemoryConfiguration,
compilation_options: WgpuCompilationOptions,
device: Device,
queue: Queue,
tasks_max: usize,
backend: Backend,
timing_method: TimingMethod,
utilities: ServerUtilities,
) -> Self
pub fn new( memory_properties: MemoryDeviceProperties, memory_config: MemoryConfiguration, compilation_options: WgpuCompilationOptions, device: Device, queue: Queue, tasks_max: usize, backend: Backend, timing_method: TimingMethod, utilities: ServerUtilities, ) -> Self
Create a new server.
Trait Implementations§
Source§impl<C: Debug + WgpuCompiler> Debug for WgpuServer<C>
impl<C: Debug + WgpuCompiler> Debug for WgpuServer<C>
Source§impl<C: WgpuCompiler> DeviceService for WgpuServer<C>
impl<C: WgpuCompiler> DeviceService for WgpuServer<C>
Source§impl<C: WgpuCompiler> Server for WgpuServer<C>
impl<C: WgpuCompiler> Server for WgpuServer<C>
Source§fn sync(
&mut self,
handles: Vec<BufferBinding>,
stream_id: StreamId,
) -> DynFut<Result<(), ServerError>>
fn sync( &mut self, handles: Vec<BufferBinding>, stream_id: StreamId, ) -> DynFut<Result<(), ServerError>>
Returns the total time of GPU work this sync completes.
Source§fn staging(
&mut self,
_sizes: &[usize],
_stream_id: StreamId,
) -> Result<Vec<Bytes>, ServerError>
fn staging( &mut self, _sizes: &[usize], _stream_id: StreamId, ) -> Result<Vec<Bytes>, ServerError>
Reserves N Bytes of the provided sizes to be used as staging to load data.
Source§fn initialize_memory(
&mut self,
memory: ManagedMemoryHandle,
size: u64,
stream_id: StreamId,
)
fn initialize_memory( &mut self, memory: ManagedMemoryHandle, size: u64, stream_id: StreamId, )
Source§fn read(
&mut self,
descriptors: Vec<CopyDescriptor>,
stream_id: StreamId,
) -> DynFut<Result<Vec<Bytes>, ServerError>>
fn read( &mut self, descriptors: Vec<CopyDescriptor>, stream_id: StreamId, ) -> DynFut<Result<Vec<Bytes>, ServerError>>
Given bindings, returns the owned resources as bytes. Read more
Source§fn write(
&mut self,
descriptors: Vec<(CopyDescriptor, Bytes)>,
stream_id: StreamId,
)
fn write( &mut self, descriptors: Vec<(CopyDescriptor, Bytes)>, stream_id: StreamId, )
Writes the specified bytes into the buffers given
Source§fn check(
&mut self,
handles: Vec<BufferBinding>,
_stream_id: StreamId,
) -> Result<(), ServerError>
fn check( &mut self, handles: Vec<BufferBinding>, _stream_id: StreamId, ) -> Result<(), ServerError>
Whether the bytes the handles name can be trusted, right now and with
no barrier: the claim check a read makes, without the read. Instant —
enqueue-time failures only. A device fault needs
sync,
which drains first.Source§unsafe fn launch(
&mut self,
kernel: Box<dyn CubeKernel>,
count: CubeCount,
args: KernelArguments,
stream_id: StreamId,
launch_mode: LaunchMode,
)
unsafe fn launch( &mut self, kernel: Box<dyn CubeKernel>, count: CubeCount, args: KernelArguments, stream_id: StreamId, launch_mode: LaunchMode, )
Source§fn flush(&mut self, stream_id: StreamId) -> Result<(), ServerError>
fn flush(&mut self, stream_id: StreamId) -> Result<(), ServerError>
Flush all outstanding tasks in the server. Read more
Source§fn start_profile(
&mut self,
stream_id: StreamId,
) -> Result<ProfilingToken, ServerError>
fn start_profile( &mut self, stream_id: StreamId, ) -> Result<ProfilingToken, ServerError>
Enable collecting timestamps.
Source§fn end_profile(
&mut self,
stream_id: StreamId,
token: ProfilingToken,
) -> Result<ProfileDuration, ProfileError>
fn end_profile( &mut self, stream_id: StreamId, token: ProfilingToken, ) -> Result<ProfileDuration, ProfileError>
Disable collecting timestamps.
Source§fn abandon_profile(&mut self, stream_id: StreamId, token: ProfilingToken)
fn abandon_profile(&mut self, stream_id: StreamId, token: ProfilingToken)
Drop the window
token opened without measuring it, for a caller that
will never close it with end_profile. Read moreSource§fn memory_usage(&mut self, stream_id: StreamId) -> MemoryUsage
fn memory_usage(&mut self, stream_id: StreamId) -> MemoryUsage
Memory usage of the given stream.
Source§fn memory_report(&mut self, stream_id: StreamId) -> MemoryReport
fn memory_report(&mut self, stream_id: StreamId) -> MemoryReport
Structured per-pool report of the given stream’s main GPU memory:
each pool’s shape, usage, and high-water marks, in allocation-routing
order. The read side of a measured memory plan — see
MemoryManagement::memory_report in cubecl-server.Source§fn stream_ids(&self) -> Vec<StreamId>
fn stream_ids(&self) -> Vec<StreamId>
Stream ids the client should iterate to aggregate across the device. Read more
Source§fn memory_cleanup(&mut self, stream_id: StreamId)
fn memory_cleanup(&mut self, stream_id: StreamId)
Ask the server to release memory that it can release.
Source§fn allocation_mode(&mut self, mode: MemoryAllocationMode, stream_id: StreamId)
fn allocation_mode(&mut self, mode: MemoryAllocationMode, stream_id: StreamId)
Update the memory mode of allocation in the server.
Source§fn install_memory_pools(
&mut self,
config: MemoryConfiguration,
stream_id: StreamId,
) -> Result<(), InstallMemoryPoolsError>
fn install_memory_pools( &mut self, config: MemoryConfiguration, stream_id: StreamId, ) -> Result<(), InstallMemoryPoolsError>
Install a new dynamic-pool layout for the device’s main GPU memory. Read more
Source§fn graph_prepare(&mut self, stream_id: StreamId) -> Result<(), ServerError>
fn graph_prepare(&mut self, stream_id: StreamId) -> Result<(), ServerError>
Prepare
stream_id for an upcoming graph capture: route allocations
into a stable pool and snapshot it, so every buffer allocated between
here and end_capture can be pinned for
the graph’s lifetime. Call this before the warmup run so the capture
window reuses the slices warmup left in the pool rather than allocating
its own — which a hardware-graph backend cannot do at all (a device
malloc inside the capture is illegal there), and which on any backend
would grow the memory a graph pins beyond what it replays against. Read moreSource§fn begin_capture(&mut self, stream_id: StreamId) -> Result<(), ServerError>
fn begin_capture(&mut self, stream_id: StreamId) -> Result<(), ServerError>
Begin recording the launches issued on
stream_id into a graph instead
of executing them, so the sequence can later be
replayed without paying the launch path again.
Call graph_prepare and warm up first. Read moreSource§fn end_capture(&mut self, stream_id: StreamId) -> Result<GraphId, ServerError>
fn end_capture(&mut self, stream_id: StreamId) -> Result<GraphId, ServerError>
Stop recording (see
begin_capture),
store the captured graph in the backend’s registry, and return its
GraphId, ready to replay.Source§fn replay(
&mut self,
graph: GraphId,
stream_id: StreamId,
) -> Result<(), ServerError>
fn replay( &mut self, graph: GraphId, stream_id: StreamId, ) -> Result<(), ServerError>
Replay the graph identified by
graph on stream_id, re-running the
whole recorded launch sequence against its original buffers. A hardware
graph replays as a single dispatch; a software graph re-encodes the
recorded dispatches, which is still far cheaper than the launch path but
stays O(n) in recorded launches. Read moreSource§fn graph_destroy(&mut self, graph: GraphId, stream_id: StreamId)
fn graph_destroy(&mut self, graph: GraphId, stream_id: StreamId)
Release the graph identified by
graph, destroying whatever it recorded
and unpinning the buffers it retained. Replay returns at enqueue time,
so the backend must guarantee no in-flight replay can still read those
buffers once they return to the pool — by syncing stream_id where
nothing weaker will do (CUDA, HIP), or by relying on the queue ordering
that already places a submitted replay ahead of any later write (wgpu).
A no-op by default and for an unknown id.Source§impl<C: WgpuCompiler> ServerCommunication for WgpuServer<C>
impl<C: WgpuCompiler> ServerCommunication for WgpuServer<C>
Source§fn sync_collective(&mut self, stream_id: StreamId) -> Result<(), ServerError>
fn sync_collective(&mut self, stream_id: StreamId) -> Result<(), ServerError>
Ensure that all queued collective operations have been executed. Read more
Source§fn comm_init(&mut self, device_ids: Vec<DeviceId>) -> Result<(), ServerError>
fn comm_init(&mut self, device_ids: Vec<DeviceId>) -> Result<(), ServerError>
Initialize the communication between the devices in
device_ids. Read moreSource§fn all_reduce(
&mut self,
src: BufferBinding,
dst: BufferBinding,
dtype: ElemType,
stream_id: StreamId,
op: ReduceOperation,
device_ids: Vec<DeviceId>,
) -> Result<(), ServerError>
fn all_reduce( &mut self, src: BufferBinding, dst: BufferBinding, dtype: ElemType, stream_id: StreamId, op: ReduceOperation, device_ids: Vec<DeviceId>, ) -> Result<(), ServerError>
Performs an
all_reduce operation on the input data and writes it to the output buffer.
see https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/usage/collectives.html#allreduce Read moreSource§impl<C: WgpuCompiler> ServerStorage for WgpuServer<C>
impl<C: WgpuCompiler> ServerStorage for WgpuServer<C>
Source§type Storage = WgpuStorage
type Storage = WgpuStorage
The storage type defines how data is stored and accessed.
Source§fn get_resource(
&mut self,
binding: BufferBinding,
stream_id: StreamId,
) -> Result<ManagedResource<WgpuResource>, ServerError>
fn get_resource( &mut self, binding: BufferBinding, stream_id: StreamId, ) -> Result<ManagedResource<WgpuResource>, ServerError>
Given a resource handle, returns the storage resource. Read more
Source§impl<C: WgpuCompiler> WriteScoped for WgpuServer<C>
impl<C: WgpuCompiler> WriteScoped for WgpuServer<C>
Source§type Streams = SchedulerMultiStream<ScheduledWgpuBackend>
type Streams = SchedulerMultiStream<ScheduledWgpuBackend>
The multi-stream driver the server keeps.
Source§fn write_streams(&mut self) -> &mut Self::Streams
fn write_streams(&mut self) -> &mut Self::Streams
The driver, split-borrowed from the server so the scope can claim on
the way in and settle on the way out while
body holds the rest.Source§fn on_failure(&mut self, stream: StreamId, error: &ServerError)
fn on_failure(&mut self, stream: StreamId, error: &ServerError)
Told about every failure a scope settles, so a measurement in flight
on
stream is invalidated wherever the failure happened. Read moreAuto Trait Implementations§
impl<C> !RefUnwindSafe for WgpuServer<C>
impl<C> !Sync for WgpuServer<C>
impl<C> !UnwindSafe for WgpuServer<C>
impl<C> Freeze for WgpuServer<C>where
PhantomData<C>: Freeze,
impl<C> Send for WgpuServer<C>where
PhantomData<C>: Send,
impl<C> Unpin for WgpuServer<C>where
PhantomData<C>: Unpin,
impl<C> UnsafeUnpin for WgpuServer<C>where
PhantomData<C>: UnsafeUnpin,
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
Mutably borrows from an owned value. Read more
impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
Source§impl<T> Downcast for Twhere
T: Any,
impl<T> Downcast for Twhere
T: Any,
Source§fn into_any(self: Box<T>) -> Box<dyn Any>
fn into_any(self: Box<T>) -> Box<dyn Any>
Converts
Box<dyn Trait> (where Trait: Downcast) to Box<dyn Any>, which can then be
downcast into Box<dyn ConcreteType> where ConcreteType implements Trait.Source§fn into_any_rc(self: Rc<T>) -> Rc<dyn Any>
fn into_any_rc(self: Rc<T>) -> Rc<dyn Any>
Converts
Rc<Trait> (where Trait: Downcast) to Rc<Any>, which can then be further
downcast into Rc<ConcreteType> where ConcreteType implements Trait.Source§fn as_any(&self) -> &(dyn Any + 'static)
fn as_any(&self) -> &(dyn Any + 'static)
Converts
&Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot
generate &Any’s vtable from &Trait’s.Source§fn as_any_mut(&mut self) -> &mut (dyn Any + 'static)
fn as_any_mut(&mut self) -> &mut (dyn Any + 'static)
Converts
&mut Trait (where Trait: Downcast) to &Any. This is needed since Rust cannot
generate &mut Any’s vtable from &mut Trait’s.Source§impl<T> DowncastSend for T
impl<T> DowncastSend for T
impl<T> ErasedDestructor for Twhere
T: 'static,
Source§impl<T> Instrument for T
impl<T> Instrument for T
Source§fn instrument(self, span: Span) -> Instrumented<Self> ⓘ
fn instrument(self, span: Span) -> Instrumented<Self> ⓘ
Source§fn in_current_span(self) -> Instrumented<Self> ⓘ
fn in_current_span(self) -> Instrumented<Self> ⓘ
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
Converts
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
Converts
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more