pub struct DeviceBuffer<T: DeviceRepr> { /* private fields */ }Expand description
Owned, typed allocation of device memory (Runtime API).
Implementations§
Source§impl<T: DeviceRepr> DeviceBuffer<T>
impl<T: DeviceRepr> DeviceBuffer<T>
Sourcepub fn new(len: usize) -> Result<Self>
pub fn new(len: usize) -> Result<Self>
Allocate an uninitialized buffer of len elements on the current device.
Sourcepub fn from_slice(src: &[T]) -> Result<Self>
pub fn from_slice(src: &[T]) -> Result<Self>
Allocate and synchronously copy src from host memory.
Sourcepub fn copy_from_host(&self, src: &[T]) -> Result<()>
pub fn copy_from_host(&self, src: &[T]) -> Result<()>
Synchronous H2D copy.
Sourcepub fn copy_to_host(&self, dst: &mut [T]) -> Result<()>
pub fn copy_to_host(&self, dst: &mut [T]) -> Result<()>
Synchronous D2H copy.
Sourcepub fn copy_from_host_async(&self, src: &[T], stream: &Stream) -> Result<()>
pub fn copy_from_host_async(&self, src: &[T], stream: &Stream) -> Result<()>
Asynchronous H2D copy on stream.
Sourcepub fn copy_to_host_async(&self, dst: &mut [T], stream: &Stream) -> Result<()>
pub fn copy_to_host_async(&self, dst: &mut [T], stream: &Stream) -> Result<()>
Asynchronous D2H copy on stream.
Sourcepub fn as_device_ptr(&self) -> u64
pub fn as_device_ptr(&self) -> u64
Raw device pointer as the u64 value kernels expect. Convenience
wrapper around as_raw.
Source§impl<T: DeviceRepr> DeviceBuffer<T>
impl<T: DeviceRepr> DeviceBuffer<T>
Sourcepub fn new_async(len: usize, stream: &Stream) -> Result<Self>
pub fn new_async(len: usize, stream: &Stream) -> Result<Self>
Asynchronously allocate len elements on stream from the device’s
default memory pool (CUDA 11.2+).
The buffer retains stream (a cheap Arc clone), so it also
frees stream-ordered: Drop enqueues cudaFreeAsync on
stream, ordered by the driver after every operation already
submitted to it — including a kernel still reading the buffer. That
makes “launch on the stream, then let the buffer drop” safe by
construction, with no host synchronize and no retention pool: the
per-op stream.synchronize() that today keeps scratch / output
buffers alive can be dropped. Freed blocks return to the device’s
stream-ordered memory pool for reuse across a chain.
Precondition (single-stream): this ordering guarantee holds only
for work submitted to this same stream. If the buffer is used on a
different stream, the Drop free on the origin stream is not
automatically ordered after that work — record an event on the other
stream and have the origin stream wait on it before the buffer drops,
or free explicitly with free_async on a
suitably-ordered stream. Using the buffer only on its origin stream
needs no such care.
free_async is still available if you want to
free at an explicit point rather than at scope exit.
Sourcepub fn zeros_async(len: usize, stream: &Stream) -> Result<Self>
pub fn zeros_async(len: usize, stream: &Stream) -> Result<Self>
Stream-ordered allocate-and-zero: cudaMallocAsync on stream
followed by cudaMemsetAsync on the same stream, so the buffer is
zeroed in stream order before any subsequent op observes it.
This is the async counterpart of zeros, and the
constructor to use for output buffers that should free
stream-ordered: like new_async the buffer
retains stream and reclaims via cudaFreeAsync on Drop, so
an output written by a pipelined kernel can be evicted without a
host synchronize or an executor-side lifetime guard.
Requires CUDA 11.2+.
Sourcepub fn free_async(self, stream: &Stream) -> Result<()>
pub fn free_async(self, stream: &Stream) -> Result<()>
Free this buffer asynchronously on stream. Consumes self so
the sync Drop does not also free.