ruda-optim
Optimizer updates, gradient handling, clipping, learning-rate schedules, and training-state records for Ruda models. Device-fused optimizer paths and collective gradient synchronization are explicit features.
Interfaces
- Crate-root optimizer exports provide optimizer configurations and gradient updates.
grad_clippingandlr_schedulerhandle clipping and schedules.trainingcombines model, optimizer, scheduler, and accumulated-gradient records.- With
collective,data_parallel::DataParallel::initializevalidates replica paths, shapes, dtypes and tied aliases, then broadcasts floating parameters from an explicit root. Local parameter IDs are preserved. DataParallel::reducesynchronizes gradients of local loss sums and divides by the total effective token/sample count. Accumulate locally before reducing;MissingGradientPolicyexplicitly selects rejection or zero contribution for unused parameters. Globally unused parameters remain absent. Transport is ruCCL host-staged; this is replicated data parallelism, not TP/PP/FSDP. Save each rank's model/optimizer/continuation records together.fused_adamwsupplies opt-in AdamW/AMSGrad implementations.fused_adamw::storagesupplies preallocated native-adapter kernels underfused-adamw-device;stats_plan::StatsPlanplans bounded hierarchical gradient reductions. The out-of-placeadamw_stepAPI and default Rust optimizer remain separate. See the fused AdamW guide and native PyTorch training.
Usage
Cargo package: ruda-optim. Rust import: ruda_optim.
[]
= "0.21"
Features
Default features: std, ruda-model/default.
| Feature | Purpose |
|---|---|
fused-adamw |
Enable host fused-optimizer interfaces. |
fused-adamw-device |
Enable the runtime-generic device implementation. |
fused-adamw-cuda |
Enable CUDA fused-optimizer integration. |
collective |
Enable ruCCL gradient synchronization. |
gradient-guard |
Enable the gradient-guard integration. |