docs.rs failed to build ruda-optim-0.21.38
Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.
Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.
Visit the last successful build:
ruda-optim-0.21.35
ruda-optim
Optimizer updates, gradient handling, clipping, learning-rate schedules, and training-state records for Ruda models. Device-fused optimizer paths and collective gradient synchronization are explicit features.
Interfaces
- Crate-root optimizer exports provide optimizer configurations and gradient updates.
grad_clippingandlr_schedulerhandle clipping and schedules.trainingcombines model, optimizer, scheduler, and accumulated-gradient records.Fp32MasterOptimizer::new(existing_optimizer).init()keeps authoritative FP32 parameters and the wrapped optimizer's state while returning parameters in their original F32/F16/BF16 storage dtype.with_gradient_scaleexplicitly unscales in FP32 before optional per-parameterwith_grad_clipping; neither option enables dynamic scaling or automatic step skipping.GradientsAccumulator::accumulate_with_dtype(&model, gradients, FloatDType::F32)converts incoming and pending gradients before addition without changing parameter storage or loss normalization.TrainingRecord::capture_with_dtypesrecords per-parameter floating storage metadata with caller state; load the resulting record type and userestore_with_dtypesto restore mixed storage, FP32 masters and pending gradients together. Use full-precision recorder settings for FP32 state.- With
collective,data_parallel::DataParallel::initializevalidates replica paths, shapes, dtypes and tied aliases, then broadcasts floating parameters from an explicit root. Local parameter IDs are preserved. DataParallel::initialize_with_buffersadditionally broadcasts I32/I64 and Bool parameter buffers once. Buffers retain IDs, widths and aliases and do not enter gradient updates; this does not enable per-forward buffer synchronization.DataParallel::reducesynchronizes gradients of local loss sums and divides by the total effective token/sample count.reduce_fp32retains FP32 output for half-storage parameters and accepts FP32 accumulation; all ranks must select the same reduction mode. Accumulate locally before reducing;MissingGradientPolicyexplicitly selects rejection or zero contribution for unused parameters. Globally unused parameters remain absent. The default transport is ruCCL host-staged;DataParallel<B, C>accepts aDataParallelCommunicator<B::InnerBackend>such as rust-ascend's native HCCL adapter. This is replicated data parallelism, not TP/PP/FSDP. Save each rank's model/optimizer/continuation records together.fused_adamwsupplies opt-in AdamW/AMSGrad implementations.fused_adamw::storagesupplies preallocated native-adapter kernels underfused-adamw-device;stats_plan::StatsPlanplans bounded hierarchical gradient reductions. The out-of-placeadamw_stepAPI and default Rust optimizer remain separate. See the fused AdamW guide and native PyTorch training.
Usage
Cargo package: ruda-optim. Rust import: ruda_optim.
[]
= "0.21"
Features
Default features: std, ruda-model/default.
| Feature | Purpose |
|---|---|
fused-adamw |
Enable host fused-optimizer interfaces. |
fused-adamw-device |
Enable the runtime-generic device implementation. |
fused-adamw-cuda |
Enable CUDA fused-optimizer integration. |
collective |
Enable ruCCL gradient synchronization. |
gradient-guard |
Enable the gradient-guard integration. |