Expand description
Safe Rust wrappers for the CUDA Runtime API.
The Runtime API is “higher level” than the Driver API: contexts are
implicit (each device has a primary context the runtime uses
automatically), kernels are typically linked at build time by nvcc,
and most operations dispatch to the current thread’s current device.
baracuda-runtime mirrors the Driver-side types where it makes sense
(Device, Stream, Event, DeviceBuffer) and uses the
CUDA 12.0+ library API (Library, Kernel) for loading PTX at
runtime — the Driver-API equivalent of Module::load_ptx +
Module::get_function.
§Driver ↔ Runtime interop
CUstream and cudaStream_t are the same C type. With the
driver-interop feature, Stream::as_raw_driver() and
Event::as_raw_driver() return views usable by baracuda-driver
APIs. See [interop].
§Examples
Device query — discover the visible GPUs and inspect compute capability + SM count.
use baracuda_runtime::Device;
let count = Device::count()?;
for d in Device::all()? {
let (major, minor) = d.compute_capability()?;
println!("device {}: cc {major}.{minor}, {} SMs", d.ordinal(),
d.multiprocessor_count()?);
}Async memory copy — overlap H2D upload with later kernel launches
by issuing on a non-blocking Stream.
use baracuda_runtime::{Device, DeviceBuffer, Stream};
Device::from_ordinal(0).set_current()?;
let stream = Stream::non_blocking()?;
let host: Vec<f32> = (0..4096).map(|i| i as f32).collect();
let device: DeviceBuffer<f32> = DeviceBuffer::new(host.len())?;
device.copy_from_host_async(&host, &stream)?;
let mut back = vec![0.0f32; host.len()];
device.copy_to_host_async(&mut back, &stream)?;
stream.synchronize()?;
assert_eq!(host, back);Event timing — measure the elapsed device time between two
Event::record calls.
use baracuda_runtime::{Device, DeviceBuffer, Event, Stream};
Device::from_ordinal(0).set_current()?;
let stream = Stream::new()?;
let start = Event::new()?;
let end = Event::new()?;
// Record START -> issue some work -> record END.
start.record(&stream)?;
let buf: DeviceBuffer<f32> = DeviceBuffer::zeros(1 << 20)?;
end.record(&stream)?;
end.synchronize()?;
let ms = Event::elapsed_time_ms(&start, &end)?;
println!("device-side elapsed: {ms} ms");Re-exports§
pub use device::Device;pub use error::Error;pub use error::Result;pub use event::Event;pub use graph::CaptureMode;pub use graph::Graph;pub use graph::GraphExec;pub use graph::GraphNode;pub use graph::UpdateResult;pub use init::device_synchronize;pub use init::driver_version;pub use init::get_device_flags;pub use init::last_error;pub use init::peek_last_error;pub use init::runtime_version;pub use init::set_device_flags;pub use launch::Dim3;pub use launch::LaunchBuilder;pub use memory::DeviceBuffer;pub use module::Kernel;pub use module::Library;pub use stream::Stream;
Modules§
- array
- CUDA arrays + texture / surface objects (Runtime API).
- device
- Device enumeration + queries via the Runtime API.
- driver_
entry - Runtime-to-Driver entry-point bridge —
cudaGetDriverEntryPoint. - error
- Error type for
baracuda-runtime. - event
- Runtime-API events.
- external
- External memory / semaphore interop via the Runtime API.
- graph
- CUDA Graphs (Runtime API).
- graphics
- Graphics-API interop (Runtime API).
- green
- Green contexts (Runtime API, CUDA 13.1+) — lightweight sub-contexts that share the primary context’s memory but carve out an SM subset.
- init
- Runtime-API initialization helpers.
- ipc
- Runtime-API IPC — share events and device allocations between
processes. Linux-primary; Windows returns
NOT_SUPPORTEDor similar on most paths. - launch
- Kernel launch builder for the Runtime API.
- launch_
attr - Extended launch attributes +
cudaLaunchKernelEx(cluster launches, programmatic stream serialization, preferred shmem carveout). - memcpy2d
- Runtime-API 2-D memory copies + pitched device allocations.
- memcpy3d
- 3D memcpy +
cudaMalloc3Dpitched 3D buffers. - memory
- Runtime-API device memory.
- mempool
- Stream-ordered memory pools (Runtime API, CUDA 11.2+).
- module
- Runtime-API library + kernel loading.
- multicast
- Multicast objects (CUDA 12.0+).
- profiler
- Runtime-side profiler start / stop — bracket the region you want Nsight Systems / nvprof to capture.
- query
- Runtime-API queries: pointer attributes, device properties, kernel
attributes. Typed wrappers around the
cuda*GetAttributes/cudaGetDevicePropertiesfamily. - stream
- Runtime-API streams.
- user_
object - Runtime-API graph user objects (CUDA 12.0+).
- vmm
- Virtual memory management (Runtime API).