ruda-optim
Optimizer updates, gradient handling, clipping, learning-rate schedules, and training-state records for Ruda models. Device-fused optimizer paths and collective gradient synchronization are explicit features.
Interfaces
- Crate-root optimizer exports provide optimizer configurations and gradient updates.
grad_clippingandlr_schedulerhandle clipping and schedules.trainingcombines model, optimizer, scheduler, and accumulated-gradient records.Fp32MasterOptimizer::new(existing_optimizer).init()keeps authoritative FP32 parameters and the wrapped optimizer's state while returning parameters in their original F32/F16/BF16 storage dtype.with_gradient_scaleexplicitly unscales in FP32 before optional per-parameterwith_grad_clipping; neither option enables dynamic scaling or automatic step skipping.GradientsAccumulator::accumulate_with_dtype(&model, gradients, FloatDType::F32)converts incoming and pending gradients before addition without changing parameter storage or loss normalization.TrainingRecord::capture_with_dtypesrecords per-parameter floating storage metadata with caller state; load the resulting record type and userestore_with_dtypesto restore mixed storage, FP32 masters and pending gradients together. Use full-precision recorder settings for FP32 state.- With
collective,data_parallel::DataParallel::initializevalidates replica paths, shapes, dtypes and tied aliases, then broadcasts floating parameters from an explicit root. Local parameter IDs are preserved. DataParallel::initialize_with_buffersadditionally broadcasts I32/I64 and Bool parameter buffers once. Buffers retain IDs, widths and aliases and do not enter gradient updates; this does not enable per-forward buffer synchronization.DataParallel::reducesynchronizes gradients of local loss sums and divides by the total effective token/sample count.reduce_fp32retains FP32 output for half-storage parameters and accepts FP32 accumulation; all ranks must select the same reduction mode. Accumulate locally before reducing;MissingGradientPolicyexplicitly selects rejection or zero contribution for unused parameters. Globally unused parameters remain absent. The default transport is ruCCL host-staged;DataParallel<B, C>accepts aDataParallelCommunicator<B::InnerBackend>such as rust-ascend's native HCCL adapter. This is replicated data parallelism, not TP/PP/FSDP. Save each rank's model/optimizer/continuation records together.fused_adamwsupplies opt-in AdamW/AMSGrad implementations.fused_adamw::storagesupplies preallocated native-adapter kernels underfused-adamw-device;stats_plan::StatsPlanplans bounded hierarchical gradient reductions. The out-of-placeadamw_stepAPI and default Rust optimizer remain separate. See the fused AdamW guide and native PyTorch training.
Usage
Cargo package: ruda-optim. Rust import: ruda_optim.
[]
= "0.21"
Features
Default features: std, ruda-model/default.
| Feature | Purpose |
|---|---|
fused-adamw |
Enable host fused-optimizer interfaces. |
fused-adamw-device |
Enable the runtime-generic device implementation. |
fused-adamw-cuda |
Enable CUDA fused-optimizer integration. |
collective |
Enable ruCCL gradient synchronization. |
gradient-guard |
Enable the gradient-guard integration. |