ruCCL 0.21.26

Ruda collective communication algorithms and orchestration.
Documentation

ruCCL

English | 简体中文 | 日本語 | Deutsch | Русский

This crate is part of the RUDA workspace. Run the commands below from the RUDA monorepo root.

Ruda's collective communication library. The public tensor interface reuses Ruda tensor backends and compute libraries; rank and in_process provide device-independent communication protocols, scheduling, and device-adapter contracts.

CUDA tensor example

cargo run -p ruCCL --features cuda --example all_reduce

Requires a working NVIDIA driver and CUDA Toolkit. The example creates four logical ranks on GPU 0, performs tensor computation and Ring AllReduce, reads back results, and closes the session. It checks Sum/Mean over 257 FP32 elements and verifies that the original inputs remain unchanged.

Features

Feature Scope
Default General communication core and tensor-backend interfaces; does not automatically enable a GPU backend
cuda CUDA tensor backend, retaining its default fusion and tuning configuration
test-cuda Selects the existing CUDA test backend in addition to cuda
test-wgpu / test-metal / test-vulkan Existing WGPU test entry points; run separately from the CUDA test features
tracing Existing cross-layer tracing integration

The registered tensor API includes register, all_reduce, reduce, broadcast, and finish_collective. All ranks must call matching collective operations in the same order. These low-level calls use the inner backend; differentiable explicit-rank operations are provided by ruda_autodiff::collective, and the optimizer layer handles parameter-gradient synchronization.

ruCCL User Guide

Compute libraries · Tensors and frameworks · 中文

1. Layers and entry points

The Cargo package is ruCCL and the Rust crate is ruccl.

ruCCL includes tensor Backend collectives, a rank core, and in-process implementations. ruda-communication provides communication infrastructure. The orchestrator feature enables orchestration entry points.

2. Tensor collective API

Function Behavior
register<B> Registers a peer, device, and CollectiveConfig
all_reduce<B> Returns the reduced result to participants
broadcast<B> Sender passes Some(tensor); receivers pass None
reduce<B> Reduces to a specified root; non-root participants receive None
finish_collective<B> Ends the peer's collective session
reset_collective<B> Resets the local collective service and discards registrations and in-progress operation state

These registered interfaces use B: ruda_tensor::Backend and B::FloatTensorPrimitive. Register the inner backend for low-level calls; use ruda_autodiff::collective for the explicit-rank forward/backward graph rules described below.

3. Registration and call contracts

Create configuration with CollectiveConfig::default(). Use with_num_devices for the number of local participating devices. Configure strategies and multinode addresses through their configuration methods.

Participants must agree on device counts, use unique peer IDs, and call matching collectives in the same order. Shape, reduction operation, root, and other parameters must agree. Each broadcast must have exactly one sender.

For multinode execution, configure node counts, global and local addresses, and data-service ports together.

4. Errors and lifecycle

CollectiveError covers duplicate or missing registrations, shape mismatches, inconsistent reduction operations or roots, and invalid broadcast sender counts.

Use finish_collective for normal completion. reset_collective forgets in-progress state; it does not complete an operation, checkpoint a device task, or provide lossless recovery.

5. CUDA example

The cuda feature enables the CUDA tensor backend. Run cargo run --locked -p ruCCL --features cuda --example all_reduce to execute Ring AllReduce with four logical ranks on GPU 0. It checks Sum/Mean over 257 FP32 elements, input preservation, and session exit.

Device adapters are in tensor_device. For the optimizer interface, see explicit-rank gradient reduction. Transfers include a host-staged path, not zero-copy P2P.

Source: collective API, configuration, rank, and in-process implementation.

6. Collective training

Enable collective in ruda-optim. With an explicitly owned rank communicator, convert backward gradients into GradientsParams, call grads.all_reduce_with::<InnerBackend>(&communicator, ReduceOperation::Mean)?, then pass the returned gradients to optimizer.step. Parameter IDs, gradient shapes, dtypes, and call order must match across ranks. For autodiff training, InnerBackend is the backend without the Autodiff wrapper.

Run the two-rank training example from the source tree:

cargo run --locked -p ruda-optim --features collective,cuda --example collective_training -- run ./collective-training-state
cargo run --locked -p ruda-optim --features collective,cuda --example collective_training -- resume ./collective-training-state

run requires a directory that does not yet exist. It saves each rank's model and optimizer after the first update, then executes the second update. resume restores that directory and executes the second update. With CUDA enabled, both logical ranks in this example use the same default device.

See the collective training example for the complete call sequence. To also save scheduler state and pending accumulated gradients, use TrainingRecord from Training and saving state.

7. Explicit-rank tensor collectives

RankCommunicator<TensorDevice<B>> exposes floating broadcast_float, all_reduce_float, all_gather_float and reduce_scatter_float, plus the corresponding I32/I64 operations. Gather concatenates equal axis-zero shards in rank order; scatter requires axis zero to divide evenly by world size. Integer tensors retain their storage width without floating-point conversion. Tensor payloads use the existing host-staged transport, not native NCCL or zero-copy P2P.

For tracked tensors, use ruda_autodiff::collective::{all_gather, all_gather_dim, reduce_scatter_sum, reduce_scatter_sum_dim, reduce_scatter_mean, reduce_scatter_mean_dim, all_reduce_sum, all_reduce_mean, broadcast}. All ranks must enter matching forward and backward operations, with matching dtypes, shapes, gradient tracking and root. See differentiable rank collectives for the backward rules.

ruda_optim::data_parallel::DataParallel<B, C> reuses replica validation, token/sample weighting, frozen/tied parameters and FP32 gradient output with a DataParallelCommunicator<B::InnerBackend>. The default is this ruCCL transport; rust-ascend supplies a native HCCL implementation with TCP used only for metadata.