pub struct NAdamW {
pub lr: f32,
pub beta1: f32,
pub beta2: f32,
pub eps: f32,
pub weight_decay: f32,
/* private fields */
}Expand description
Nesterov AdamW. Per-tensor state: two f32 buffers.
Fields§
§lr: f32Learning rate.
beta1: f32First-moment EMA decay β₁. Default 0.9.
beta2: f32Second-moment EMA decay β₂. Default 0.999.
eps: f32Denominator stability constant. Default 1e-8.
weight_decay: f32Decoupled weight-decay coefficient λ. Default 0.01.
Implementations§
Trait Implementations§
Source§impl Optimizer for NAdamW
impl Optimizer for NAdamW
Source§fn set_lr(&mut self, lr: f32)
fn set_lr(&mut self, lr: f32)
Set the base learning rate (for LR schedules / warmup). Default is a
no-op for algorithms without a scalar
lr (e.g. Adafactor’s relative
step size); every algorithm in this crate that has an lr field
overrides this to update it.fn step( &mut self, name: &str, _shape: &[usize], param: &mut [f32], grad: &[f32], )
Source§fn end_iteration(&mut self)
fn end_iteration(&mut self)
Advance the global step counter. Most algorithms increment per
call to
step, so most implementations leave this a no-op.Source§fn step_batch(&mut self, items: &mut [OptItem<'_>])
fn step_batch(&mut self, items: &mut [OptItem<'_>])
Batched step over ALL parameters in one call. Default: sequential
step per item — bit-identical to the per-parameter loop. Optimizers
whose parameter groups are independent (e.g. Muon on the 2-D weight
matrices vs AdamW on the embeddings/biases/norms) can override this to
run the groups on separate threads; because the groups touch disjoint
parameters and disjoint optimizer state, the result is bit-for-bit the
same as the serial loop — only the wall-clock (the CPU-side optimizer
bubble) shrinks toward max(group_times) instead of their sum.Source§fn lr_scale(&self, _name: &str) -> f32
fn lr_scale(&self, _name: &str) -> f32
Per-tensor multiplier on the effective learning rate. Default
is
1.0 for every name. Override when wrapping this crate to
support per-name LR schedules (e.g. embedding-vs-attention
splits, or the Gaussian-splat attribute-typed LR setup). The
CPU impls in this crate currently honor this only when the
caller passes a pre-scaled lr for the relevant call —
backends are encouraged to consult it inside their fused
kernel.Source§fn state_dict(&self) -> Option<OptimizerState>
fn state_dict(&self) -> Option<OptimizerState>
Snapshot the optimizer’s state for a checkpoint. Read more
Source§fn load_state_dict(&mut self, _state: &OptimizerState) -> bool
fn load_state_dict(&mut self, _state: &OptimizerState) -> bool
Restore a snapshot. Returns false when unsupported or when the state does
not belong to this algorithm.
Auto Trait Implementations§
impl Freeze for NAdamW
impl RefUnwindSafe for NAdamW
impl Send for NAdamW
impl Sync for NAdamW
impl Unpin for NAdamW
impl UnsafeUnpin for NAdamW
impl UnwindSafe for NAdamW
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
Mutably borrows from an owned value. Read more
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
Converts
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
Converts
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more