pub struct Fp32MasterOptimizer<O> { /* private fields */ }Expand description
Opt-in FP32 master parameters for an existing simple optimizer.
The wrapped optimizer receives FP32 parameters and gradients; its algorithm, hyperparameters and per-parameter state are unchanged. Returned parameters retain their incoming storage dtype. Existing optimizers do not opt in.
Implementations§
Source§impl<O> Fp32MasterOptimizer<O>
impl<O> Fp32MasterOptimizer<O>
Sourcepub fn new(optimizer: O) -> Self
pub fn new(optimizer: O) -> Self
Wrap an existing optimizer without modifying its configuration.
Sourcepub fn with_grad_clipping(self, clipping: GradientClipping) -> Self
pub fn with_grad_clipping(self, clipping: GradientClipping) -> Self
Explicitly apply RUDA’s existing per-parameter clipping after FP32 conversion.
Sourcepub fn with_gradient_scale(self, scale: f32) -> Self
pub fn with_gradient_scale(self, scale: f32) -> Self
Explicitly divide gradients by this loss scale in FP32 before clipping.
The caller scales the loss and uses one scale for the accumulation window. No dynamic scaling, non-finite policy or optimizer-step skipping is enabled.
Sourcepub fn init<B, M>(self) -> OptimizerAdaptor<Self, M, B>
pub fn init<B, M>(self) -> OptimizerAdaptor<Self, M, B>
Build RUDA’s original module adaptor, preserving parameter IDs and records. Clipping is disabled unless explicitly configured on the wrapper or adaptor.