ruda-optim
Optimizer updates, gradient handling, clipping, learning-rate schedules, and training-state records for Ruda models. Device-fused optimizer paths and collective gradient synchronization are explicit features.
Interfaces
- Crate-root optimizer exports provide optimizer configurations and gradient updates.
grad_clippingandlr_schedulerhandle clipping and schedules.trainingcombines model, optimizer, scheduler, and accumulated-gradient records.fused_adamwsupplies opt-in AdamW/AMSGrad implementations.
Usage
Cargo package: ruda-optim. Rust import: ruda_optim.
[]
= "0.21"
Features
Default features: std, ruda-model/default.
| Feature | Purpose |
|---|---|
fused-adamw |
Enable host fused-optimizer interfaces. |
fused-adamw-device |
Enable the runtime-generic device implementation. |
fused-adamw-cuda |
Enable CUDA fused-optimizer integration. |
collective |
Enable ruCCL gradient synchronization. |
gradient-guard |
Enable the gradient-guard integration. |