pub struct Stream { /* private fields */ }Expand description
A CUDA stream (GPU command queue).
Streams provide ordered, asynchronous execution of GPU commands. Commands enqueued on the same stream execute sequentially, while commands on different streams may execute concurrently.
The stream holds an Arc<Context> to ensure the parent context
outlives the stream.
§A Stream is a shared handle, not a unique owner
Cloning yields a second handle to the same queue — Stream::raw
returns the same CUstream — and the queue is destroyed once the last
handle drops. That is what lets two subsystems which each want to hold
“their” stream be collapsed onto one queue: oxicuda-dnn’s DnnHandle and
the BlasHandle nested inside it now share one, so a convolution’s output
is ordered before a GEMM that reads it by stream semantics alone — no
event choreography, no host rendezvous, and a capture of the pair is a
linear chain rather than a fork/join.
Clone is written out rather than derived so the doc comment can say what
it means: this is another reference to one queue, not a copy of it.
Implementations§
Source§impl Stream
impl Stream
Sourcepub fn new(ctx: &Arc<Context>) -> CudaResult<Self>
pub fn new(ctx: &Arc<Context>) -> CudaResult<Self>
Creates a new stream with CU_STREAM_NON_BLOCKING flag.
Non-blocking streams do not implicitly synchronise with the default (NULL) stream, allowing maximum concurrency.
§Errors
Returns a CudaError if the driver
call fails (e.g. invalid context, out of resources).
Sourcepub fn with_priority(ctx: &Arc<Context>, priority: i32) -> CudaResult<Self>
pub fn with_priority(ctx: &Arc<Context>, priority: i32) -> CudaResult<Self>
Creates a new stream with the specified priority and
CU_STREAM_NON_BLOCKING flag.
Lower numerical values indicate higher priority. The valid range
can be queried via cuCtxGetStreamPriorityRange.
§Errors
Returns a CudaError if the priority
is out of range or the driver call otherwise fails.
Sourcepub fn synchronize(&self) -> CudaResult<()>
pub fn synchronize(&self) -> CudaResult<()>
Sourcepub fn wait_event(&self, event: &Event) -> CudaResult<()>
pub fn wait_event(&self, event: &Event) -> CudaResult<()>
Makes all future work submitted to this stream wait until the given event has been recorded and completed.
This is the primary mechanism for inter-stream synchronisation:
record an Event on one stream, then call wait_event on
another stream to establish an ordering dependency.
§Errors
Returns a CudaError if the driver
call fails (e.g. invalid event handle).
Sourcepub fn is_same_queue(&self, other: &Self) -> bool
pub fn is_same_queue(&self, other: &Self) -> bool
Whether self and other are handles to the same driver queue.
The question a caller asks before deciding that stream order alone
sequences two pieces of work: on one queue it does, on two it does not
and an event is required. Compares the driver handle rather than the
Arc, so a queue reached through two independently-built handles (were
that ever possible) still answers truthfully.