Expand description
Ruda optimizer updates, gradient state, clipping and learning-rate schedules.
Modules§
- adaptor
- Adaptor module for optimizers.
- decay
- Weight decay module for optimizers.
- grad_
clipping - Gradient clipping module.
- lr_
scheduler std - Learning rate scheduler module.
- momentum
- Momentum module for optimizers.
- record
- Record module for optimizers.
- training
std - Combined model, optimizer, scheduler and accumulated-gradient records.
Structs§
- AdaGrad
- AdaGrad optimizer
- AdaGrad
Config - AdaGrad configuration.
- AdaGrad
State - AdaGrad state.
- AdaGrad
State Item - The record item type for the module.
- Adam
- Adam optimizer.
- Adam
Config - Adam configuration.
- Adam
State - Adam state.
- Adam
State Item - The record item type for the module.
- AdamW
- AdamW optimizer.
- AdamW
Config AdamWConfiguration.- AdamW
State - AdamW state.
- AdamW
State Item - The record item type for the module.
- Adan
- Adan optimizer.
- Adan
Config AdanConfiguration.- Adan
State - Adan state.
- Adan
State Item - The record item type for the module.
- Adaptive
Momentum State - Adaptive momentum state.
- Adaptive
Momentum State Item - The record item type for the module.
- Adaptive
Nesterov Momentum State - Adaptive Nesterov momentum state.
- Adaptive
Nesterov Momentum State Item - The record item type for the module.
- Centered
State - CenteredState is to store and pass optimizer step params.
- Centered
State Item - The record item type for the module.
- Gradients
Accumulator - Accumulate gradients into a single Gradients object.
- Gradients
Params - Data type that contains gradients for parameters.
- Gradients
Params Record - Host-owned gradient state, keyed by the original model parameter IDs.
- LBFGS
- L-BFGS optimizer.
- LBFGS
Config - LBFGS Configuration.
- LBFGS
State - L-BFGS optimizer state
- LBFGS
State Item - The record item type for the module.
- LrDecay
State - Learning rate decay state (also includes sum state).
- LrDecay
State Item - The record item type for the module.
- Multi
Gradients Params - Exposes multiple gradients for each parameter.
- Muon
- Muon optimizer.
- Muon
AdamW - Mixed optimizer with fixed, explicit parameter identities.
Missing gradients skip both momentum and decay for that parameter.
A tied parameter is updated once, by its
ParamId. - Muon
AdamW Config - Muon for explicitly selected hidden matrices and AdamW for the remainder.
- Muon
AdamW Record - Versioned optimizer record. Save the model record with it so ParamIds survive. Configuration and routing are checked on load; lower-precision record settings can round momentum, so use full-precision records for continuation comparisons.
- Muon
Config - Muon configuration.
- Muon
State - Muon state.
- Muon
State Item - The record item type for the module.
- RmsProp
- Optimizer that implements stochastic gradient descent with momentum. The optimizer can be configured with RmsPropConfig.
- RmsProp
Config - Configuration to create the RmsProp optimizer.
- RmsProp
Momentum - RmsPropMomentum is to store config status for optimizer.
(, which is stored in optimizer itself and not passed in during
step()calculation) - RmsProp
Momentum State - RmsPropMomentumState is to store and pass optimizer step params.
- RmsProp
Momentum State Item - The record item type for the module.
- RmsProp
State - State of RmsProp
- RmsProp
State Item - The record item type for the module.
- Sgd
- Optimizer that implements stochastic gradient descent with momentum.
- SgdConfig
- Configuration to create the Sgd optimizer.
- SgdState
- State of Sgd.
- SgdState
Item - The record item type for the module.
- Square
AvgState - SquareAvgState is to store and pass optimizer step params.
- Square
AvgState Item - The record item type for the module.
Enums§
- Adjust
LrFn - Learning rate adjustment method for Muon optimizer.
- Line
Search Fn - Strategy for the line search optimization phase
- Muon
Error - Configuration, metadata or explicit parameter-group error for Muon. Device execution faults remain backend errors; no host fallback is attempted.
- Muon
Matrix Layout - Logical orientation used ONLY for shape-based learning-rate scaling. The tensor itself is not reshaped and its update keeps the same geometry.
- Muon
Momentum Mode - Momentum convention. Checkpoint buffers are NOT interchangeable between modes.
Traits§
- Optimizer
- General trait to optimize module.
- Simple
Optimizer - Simple optimizer is an opinionated trait to simplify the process of implementing an optimizer.
Type Aliases§
- Learning
Rate - Type alias for the learning rate.