ruTENSOR
English | 简体中文 | 日本語 | Deutsch | Русский
Tensor linear algebra for Ruda device tensors: contractions and einsum, reductions, physical permutations, and elementwise operations.
- Cargo package:
ruTENSOR - Rust crate:
rutensor - English user guide
- 中文使用文档
ruTENSOR User Guide
Compute libraries · Tensor framework · 中文
ruTENSOR provides named-axis tensor contractions, reductions, physical permutations, and elementwise operations. Inputs use RudaTensor<R>; the application selects a device Runtime. This library is distinct from the higher-level ruda-tensor framework.
1. Configure dependencies
The Cargo package is ruTENSOR; the Rust import name is rutensor. Defaults enable std and device tensor computation without selecting a driver. The following configuration places the application directory alongside the RUDA source directory:
[]
= { = "ruTENSOR", = "../RUDA/ruTENSOR" }
= { = "../RUDA/ruda-core", = false, = ["std", "tensor-host-data"] }
= { = "../RUDA/ruda-kernel", = false, = ["frontend-std", "device-tensor"] }
= { = "../RUDA/ruda-driver-cuda", = false, = ["std"] }
2. Contraction, reduction, and permutation
Save the following as the application's src/main.rs:
use TensorData;
use ;
use ;
use ;
einsum("ik,kj->ij", ...) sums over k and returns shape [2, 2]. reduce retains mode 0 and reduces mode 1, returning shape [2]. permute returns a newly allocated [3, 2] tensor, not a view sharing the input storage.
3. Einsum expressions
einsum(expression, inputs) accepts one or more inputs. Case-sensitive letters identify axes; the right side of the arrow selects and orders output axes.
| Expression | Operation |
|---|---|
ik,kj->ij |
Matrix multiplication |
...ik,...kj->...ij |
Matrix multiplication with broadcast batch axes |
abc,cde->abde |
Multidimensional tensor contraction |
ij,jk,kl->il |
Three-input contraction |
i,j->ij |
Outer product |
ii->i |
Diagonal |
ii-> |
Trace, returning a rank-zero scalar |
ijk->ki |
Reduce j and reorder the remaining axes |
...i->i |
Reduce all axes represented by the ellipsis |
- Matching modes across inputs must have equal extents or an extent of 1. Ellipsis axes broadcast with right alignment.
- Repeated labels within an input select a diagonal; those axes must have exactly equal extents.
- Output labels must be unique and present in the inputs.
- Without
->, output axes start with the ellipsis axes, followed by alphabetically sorted labels that occur exactly once. - A scalar input uses an empty label segment; for example,
,ij->ijmultiplies the first scalar input into the matrix. - Inputs can be noncontiguous. Computation does not read tensor values back to the host.
For fixed expressions, shapes, strides, and dtypes, construct EinsumPlan::new(expression, descriptors) and reuse execute(inputs). To specify output and arithmetic precision, use EinsumPlan::with_options or einsum_with_options.
4. Descriptors and execution plans
A Mode is an i32 label. The same label identifies the same logical axis across tensors, regardless of its physical position. TensorDescriptor stores extents, element strides, and storage dtype. OperandDescriptor adds labels and an input unary transform.
These functions create and execute a plan for D = alpha * A @ B + beta * C:
use ;
use ;
Plans do not retain input buffers. Subsequent executions can use different tensors with matching shapes, strides, and dtypes. All inputs must reside on the same device.
| Constructor | Execution input order | Scalar order |
|---|---|---|
contraction |
A, B; optional C | alpha; beta when C is present |
sum_product |
All product inputs; optional C | alpha; beta when C is present |
reduction |
A; optional C | alpha; beta when C is present |
permutation |
A | alpha |
elementwise_binary |
A, B | alpha, beta |
elementwise_trinary |
A, B, C | alpha, beta, gamma |
Plan::execute allocates the output. Plan::execute_into accepts and returns an output tensor whose shape, strides, and dtype match the output descriptor. Its buffer must be exclusively owned, without sharing with inputs or other views.
Use TensorDescriptor::new(extents, strides, dtype) for padded or reordered output layouts; output axes must not overlap. Explicit operation descriptors can introduce additional broadcast axes through the output shape.
5. Reduction and elementwise operations
reduce(input, input_modes, output_modes, operation) reduces labels absent from the output; reduced axes are removed:
| ReductionOp | Operation | Empty reduction |
|---|---|---|
Sum |
Sum | 0 |
Product |
Product | 1 |
Min |
Minimum | Positive infinity |
Max |
Maximum | Negative infinity |
elementwise_binary and elementwise_trinary align and broadcast named axes, using BinaryOp::{Add, Mul, Min, Max}. Trinary operations evaluate (alpha * op(A) op_ab beta * op(B)) op_abc gamma * op(C).
Select Identity, Negate, Abs, Sqrt, Exp, Log, Sin, Cos, Tanh, Relu, Reciprocal, or Conjugate through OperandDescriptor::with_unary. Log is the natural logarithm; Conjugate equals Identity for real values. Min and Max propagate NaNs.
6. Precision, storage, and errors
- Storage types are F16, BF16, F32, and F64. Quantized, integer, and complex storage are not accepted.
ComputeType::F32orF64controls input conversion, unary transforms, products, reductions, and scalar arithmetic. F64 inputs require F64 computation.- Defaults use F64 arithmetic if any input is F64, otherwise F32. Identical input storage types preserve that output type; mixed inputs produce F64 if any input is F64, otherwise F32.
- Request an explicit output dtype for higher-precision storage; increasing storage precision cannot recover precision already lost in an input.
- Output conversion happens at the final write. Inputs with a different dtype from the compute type require device conversion buffers; matching inputs are read in their original layout.
- Reductions use parallel workgroup tree merging. Floating-point addition and multiplication order can differ from serial evaluation.
- Empty outputs submit no computation; rank-zero outputs contain one scalar. Shape, stride, and address calculations use platform
usize; dispatches also obey backend resource limits. - Returned
Resultvalues report expression, descriptor, shape, device, and dispatch-range errors. Device submission follows the Runtime's error handling; asynchronous execution errors surface at synchronization or readback. - General contractions directly traverse reduction coordinates. They do not search multi-input contraction paths or automatically lower to Tensor Core matrix multiplication. Storage conversion requires additional device memory.
ruTENSOR exposes a Ruda Rust API, not the NVIDIA cuTENSOR C ABI. The device must support the selected storage and compute types.