oxicuda
Pure Rust CUDA replacement for the COOLJAPAN ecosystem.
Part of the OxiCUDA project.
Overview
Version: 0.5.5 — 2026-08-13 — see Status below for test counts
oxicuda is the umbrella crate that re-exports all OxiCUDA sub-crates behind
feature flags. It provides a single dependency entry point for applications that
need GPU compute capabilities without installing the CUDA Toolkit -- libcuda.so
(or nvcuda.dll) is loaded dynamically at runtime.
The core crates (driver, memory, launch) are enabled by default. Higher-level
libraries -- BLAS, DNN, FFT, sparse, solver, and random number generation -- are
opt-in via feature flags. Enable full to get everything.
A prelude module and init() function provide convenient imports and
one-call CUDA driver initialization. Additional built-in modules cover
profiling, multi-GPU device pools, collective communication (NCCL equivalent),
pipeline parallelism, and multi-node distributed training. The backend and
compute modules (always compiled in -- see Feature Flags)
provide a portable ComputeBackend trait plus one-call backend auto-selection
that works even where the CUDA driver does not exist at all, such as macOS --
see Using oxicuda on macOS below.
Architecture
oxicuda (umbrella)
+---------+---------+---------+---------+
| | | | |
driver memory launch ptx autotune
| | | | |
+----+----+---------+---------+---------+
|
+------+------+------+------+------+
| | | | | |
blas dnn fft sparse solver rand
Quick Start
use *;
Feature Flags
| Feature | Description | Default |
|---|---|---|
driver |
CUDA driver API wrapper | Yes |
memory |
GPU memory management | Yes |
launch |
Kernel launch infrastructure | Yes |
nvrtc |
NVRTC runtime JIT compiler (CUDA-C to PTX) | No |
ptx |
PTX code generation DSL | No |
autotune |
Autotuner engine (implies ptx) |
No |
blas |
cuBLAS equivalent | No |
dnn |
cuDNN equivalent (implies blas) |
No |
fft |
cuFFT equivalent | No |
sparse |
cuSPARSE equivalent | No |
solver |
cuSOLVER equivalent | No |
rand |
cuRAND equivalent | No |
pool |
Stream-ordered memory pool | No |
backend |
No-op flag, not in the default list -- but oxicuda-backend is a mandatory (non-optional) dependency and the backend/compute modules carry no #[cfg(feature = ...)] gate, so the ComputeBackend trait, CpuBackend, and auto-selection are compiled in regardless of this flag. It exists only so dependants that name backend explicitly keep compiling. |
No (functionality present either way) |
primitives |
CUB-equivalent parallel GPU primitives | No |
vulkan |
Vulkan compute backend (cross-vendor) | No |
metal |
Apple Metal compute backend (macOS/iOS) | No |
webgpu |
WebGPU compute backend (via wgpu) | No |
rocm |
AMD ROCm/HIP backend (Linux + AMD GPU) | No |
level-zero |
Intel Level Zero backend | No |
onnx-backend |
ONNX operator runtime and graph executor | No |
tensor-backend |
ToRSh GPU tensor backend with autograd | No |
transformer-backend |
TrustformeRS transformer inference backend | No |
wasm-backend |
WASM + WebGPU backend for browser environments (implies webgpu, so compute::default_backend also registers it) |
No |
gpu-tests |
Enables on-device GPU test suites in the crates that have them | No |
full |
Enable all optional features | No |
Sub-crates
| Crate | Volume | Description |
|---|---|---|
oxicuda-driver |
Vol.1 | CUDA driver API bindings |
oxicuda-memory |
Vol.1 | Device, pinned, unified memory |
oxicuda-launch |
Vol.1 | Kernel launch and grid configuration |
oxicuda-nvrtc |
Vol.1 | Runtime CUDA-C to PTX JIT (NVRTC), dlopen'd |
oxicuda-ptx |
Vol.2 | PTX code generation DSL |
oxicuda-autotune |
Vol.2 | Autotuner for kernel parameters |
oxicuda-blas |
Vol.3 | Dense linear algebra (GEMM, etc.) |
oxicuda-dnn |
Vol.4 | Deep learning primitives |
oxicuda-fft |
Vol.5 | Fast Fourier Transform |
oxicuda-sparse |
Vol.5 | Sparse matrix operations |
oxicuda-solver |
Vol.5 | Matrix decompositions and solvers |
oxicuda-rand |
Vol.5 | Random number generation |
oxicuda-primitives |
Vol.5 | CUB-equivalent warp/block/device primitives |
oxicuda-backend |
— | Abstract compute backend trait |
oxicuda-vulkan |
— | Vulkan compute backend |
oxicuda-metal |
— | Apple Metal compute backend |
oxicuda-webgpu |
— | WebGPU compute backend |
oxicuda-rocm |
— | AMD ROCm/HIP backend |
oxicuda-levelzero |
— | Intel Level Zero backend |
Using oxicuda on macOS
There is no NVIDIA driver on macOS, so the CUDA-driver Quick Start example
above does not apply there -- oxicuda::init() and the driver/memory/
launch APIs return Err(CudaError::NotInitialized) at runtime (the crate
still compiles). The supported way to compute on a Mac's GPU is the portable
ComputeBackend trait, which this crate always compiles in, backed by either
the metal or webgpu feature:
[]
= { = "0.5", = ["metal"] }
let backend = default_backend?;
println!;
default_backend() probes every backend compiled into the build and returns
the best one already initialised -- Metal on a Mac with the metal feature
enabled, otherwise the always-available CpuBackend, never an error. Full
op-coverage detail (exactly which ComputeBackend methods run on the GPU
today) lives in the root README and
the oxicuda-metal /
oxicuda-webgpu crate READMEs.
Status
- Version: 0.5.5 (2026-08-13)
- Tests: 164 passing with
--features metal,webgpu,nvrtc(the macOS-relevant feature set; measured 2026-08-13 on real Apple M3 hardware). This crate's full--all-featurescount (covering every CUDA-only subsystem too) is tracked in the root README's crate table rather than duplicated here, to avoid two figures drifting apart.
License
Apache-2.0 -- (C) 2026 COOLJAPAN OU (Team KitaSan)