dlpark
A pure Rust implementation of dmlc/dlpack.
This implementation focuses on transferring tensors between Rust and Python, and between Rust tensor/array libraries, without copying.
Installation
dlpark ships no default features — enable the interop backends you need:
The cpu-all feature group enables every CPU-testable backend (candle, half, image, ndarray, pyo3) in one go. The crate targets Rust edition 2024.
Mental model
A producer wraps its data into a [Builder], which holds the owning context plus scalar tensor fields. Building the builder produces a [ManagedBox<M>] — an RAII handle over a raw DLPack managed tensor pointer that calls the DLPack deleter on drop. M selects the ABI:
- [
legacy::Dlpack] =ManagedBox<DLManagedTensor>— the pre-v0.8"dltensor"capsule. - [
versioned::Dlpack] =ManagedBox<DLManagedTensorVersioned>— the current"dltensor_versioned"capsule, carrying version and flags.
A consumer receives a ManagedBox (from Python, another Rust library, or a raw pointer) and reads its metadata and data through the accessors below. The flow is always: owning value → Builder → ManagedBox → borrowed views/slices, never the reverse.
What is DLPack?
DLPack is a common in-memory tensor structure that enables sharing tensor data between different deep learning frameworks. It provides a standardized way to exchange tensor data without copying, making it efficient for framework interoperability.
Key features of DLPack:
- Zero-copy tensor sharing between frameworks
- Support for various data types and devices (CPU, GPU, etc.)
- Memory management through deleter functions
- Versioned ABI for compatibility
Implementation Details
Versioning
The library implements both legacy and versioned DLPack structures:
legacy::Dlpack: Legacy managed tensor capsule supportversioned::Dlpack: Versioned managed tensor capsule support with:- Major version: 1
- Minor version: 3
- Additional flags for tensor properties (read-only, copied, sub-byte type padding)
Safe Abstractions
The library provides a Rust ownership wrapper over the C-style DLPack structures, ManagedBox<M>:
- RAII wrapper around a raw managed
DLPacktensor pointer - Automatic cleanup through the DLPack deleter function on drop
legacy::Dlpackandversioned::Dlpackare convenience aliases forManagedBox<DLManagedTensor>andManagedBox<DLManagedTensorVersioned>— the two concrete forms you'll actually use
Other key features:
- Memory safety through Rust's ownership system
- Support for
imagebuffers, ndarray, and candle tensors, plus raw DLPack tensor layouts - Python interoperability through PyO3
- Optional DLPack 1.3 C Exchange API fast path when a producer type exposes
__dlpack_c_exchange_api__
Choosing Builder Metadata
Builder<C, L> uses the metadata type L to select its allocation strategy:
CopiedArray<S, T, N>— fixed rank known at compile time. Shape and strides are copied into the managed tensor allocation.buildandbuild_raware infallible.CopiedSlice<S, T>— dynamic rank. The containers may be borrowed slices or owned values such asVec<i64>. Usetry_buildortry_build_raw.GenericArray<S, T, A, B, N>— fixed-rank metadata with elements that implementTryInto<i64>. Values are converted directly into the managed tensor allocation; usetry_buildortry_build_raw.GenericSlice<S, T, A, B>— dynamic-rank metadata with elements that implementTryInto<i64>. It avoids allocating temporaryVec<i64>values.BorrowedArray<N>— fixed-rank, zero-copy metadata. Itsbuildmethods are unsafe because the arrays must outlive the managed tensor.BorrowedSlice— dynamic-rank, zero-copy metadata. Itstry_buildmethods are unsafe for the same lifetime reason.
Copied metadata and the managed tensor share one allocation. Borrowed metadata only allocates the managed tensor header.
Use GenericArray or GenericSlice when a framework exposes dimensions or
strides as an integer type other than i64:
use ;
use GenericArray;
let shape = ;
let strides = ;
let mut data = vec!;
let data_ptr = data.as_mut_ptr.cast;
let builder = new;
// SAFETY: the boxed Vec owns six initialized f32 values at data_ptr.
let dlpack: Dlpack = unsafe
.try_build
.unwrap;
The generic path converts each value directly into the trailing i64
metadata storage. In the included length-64 microbenchmark, direct conversion
takes approximately 11.7 ns, while allocating temporary Vec<i64> storage
and then calling copy_nonoverlapping takes approximately 54.9 ns on the
development machine. Reproduce it with:
The ndarray exporter uses the same direct path for its usize shape and
isize strides, so exporting an owned array does not allocate temporary
Vec<i64> metadata.
Python Exchange Paths
The pyo3 feature supports the standard Python DLPack capsule protocol:
legacy::Dlpackconsumes or produces legacy"dltensor"capsules.versioned::Dlpackconsumes or produces"dltensor_versioned"capsules.python::dlpack_device(obj)calls and validatesobj.__dlpack_device__(), returning a RustDLDevice.- When extracting a versioned tensor from a Python object, dlpark first checks the object's type for a
__dlpack_c_exchange_api__PyCapsule named"dlpack_exchange_api". If present, it uses the DLPack C Exchange API no-sync function table. Otherwise it callsobj.__dlpack__(max_version=(1, 3)). Producers that only implement the legacy no-argument protocol must be extracted aslegacy::Dlpack, because they return the incompatible"dltensor"capsule ABI. - Consumers can call
versioned::Dlpack::extract_with_options(obj, stream, copy)to pass optional stream and tri-state copy requests to__dlpack__;extract_with_streamis the typed convenience path for GPU consumers. Thecudarcfeature implements stream mapping forCudaStream; other backends can implement the unsafepython::DlpackStreamtrait.
The C Exchange API is intended for extension/library use where the consumer can borrow tensors and coordinate work on the producer's current stream. It is not a replacement for the normal __dlpack__ ingestion path.
Reading tensor data
Once you hold a ManagedBox, its consumer-side accessors read metadata and CPU data without unsafe:
use DlpackElement;
let shape = dlpack.shape?; // &[i64]
let strides = dlpack.strides?; // Option<&[i64]> (None = compact)
let n = dlpack.num_elements?;
let bytes = dlpack.num_bytes?; // sub-byte-packing aware
let data = dlpack.?; // compact CPU data, dtype-checked
cpu_data_slice validates device (CPU only), dtype match, and compact row-major layout before forming the slice. Accessors directly on raw DLTensor are unsafe, because DLPack metadata cannot prove that its public pointers are readable or within an allocation. Prefer the corresponding safe ManagedBox accessors. Low-level consumers of non-compact layouts may use DLTensor::cpu_data_ptr / cpu_data_ptr_bytes after proving the raw tensor pointer contract.
Mutable access and the IS_COPIED flag. Writing into a DLPack tensor is gated by two versioned flags, because exclusive ownership cannot be proven from a &mut ManagedBox alone — the producer may hold aliases:
DlpackFlags::IS_COPIEDasserts the export owns an unaliased copy.cpu_data_slice_mutrequires it and needs nounsafe.DlpackFlags::READ_ONLYforbids mutation; both mut accessors reject it.
Without IS_COPIED, use the unsafe ..._mut_unchecked accessors and prove exclusivity yourself. Legacy DLManagedTensor has no flags field, so it can never satisfy IS_COPIED — mutation of a legacy tensor always goes through the _unchecked path. Interop adapters mirror this: ArrayViewMutD::try_from(&mut dlpack) is the safe, IS_COPIED-gated path; array_view_from_dlpack_mut_unchecked is the escape hatch.
ManagedBox::flags() / version() read the versioned fields; flags_mut is unsafe because setting IS_COPIED or clearing READ_ONLY asserts the corresponding ownership/mutability guarantee.
Features
No features are enabled by default — enable the backends you need (see Installation).
| Feature | Description | Status |
|---|---|---|
pyo3 |
Python interop via pyo3 (capsule protocol + DLPack C Exchange API fast path) | ✅ |
image |
Zero-copy conversion with image buffers | ✅ |
ndarray |
Zero-copy conversion with ndarray arrays/views | ✅ |
half |
f16/bf16 element type support (via half) |
✅ |
candle |
Conversion with candle Tensor — CPU only; candle's CUDA backend needs separate integration work |
✅ |
cudarc |
Zero-copy conversion with cudarc CudaSlice<T> — no automated tests here, needs a CUDA-capable device to exercise |
✅ |
Quick Start
Two runnable examples:
examples/dlparkimg— a Python extension module (viapyo3) transferringimage::RgbImageto/from Python (e.g.torch.Tensor).examples/ndarray-candle— a plain binary round-tripping data through DLPack:ndarray::Array2→versioned::Dlpack→candle::Tensor→versioned::Dlpack→ndarrayview, run withcargo run -p ndarray-candle.
Usage Examples
Converting between Rust and Python
use ;
use *;
// Rust to Python
// Python to Rust
Image Processing
use ;
use ;
let img = from_vec?;
let tensor: Dlpack = from.build;
let img2 = try_from?;
ndarray
use ;
use ;
let array = arr2;
let tensor: Dlpack = from.try_build?;
let view = try_from?;
assert_eq!;
let dynamic: = arr2.into_dyn;
let dynamic_tensor: Dlpack = from.try_build?;
candle
Zero-copy from candle::Tensor to DLPack; the reverse direction (DLPack to candle::Tensor) always copies, since candle has no borrowed CPU tensor type.
use Tensor;
use ;
let tensor = new?;
let dlpack: Dlpack = try_from?.try_build?;
let tensor2 = try_from?;
assert_eq!;
let tensor = new?;
let builder = try_from?;
let dlpack = builder
.insert_flags?
.?;
cudarc
Zero-copy in both directions between a cudarc CudaSlice<T> and a DLPack tensor. Builder::try_from returns a builder with IS_COPIED and a contiguous 1-D default layout (shape = [len], strides = [1]); replace its metadata for higher-rank tensors. The reverse direction consumes the managed tensor through TryFrom<ManagedBox<M>> for BorrowedCudaSlice<M, T>, keeping it alive for as long as the CUDA view exists.
use ;
let dlpack: Dlpack = try_from?
.metadata
.try_build?;
let borrowed = try_from?;