pub struct Muon<B: Backend> { /* private fields */ }Expand description
Muon optimizer.
Muon internally runs standard SGD-momentum, and then performs an orthogonalization post-processing step using a finite quintic Newton-Schulz iteration. This is an approximate spectral transformation, NOT an exact polar decomposition or a guarantee that U*U^T is identity. This implementation computes in the tensor’s dtype; it does not silently cast to BF16 or create FP32 master parameters.
§Important Notes
-
Only for nonempty 2D parameters: Muon is designed for hidden weight matrices. Use AdamW or SGD for biases, embeddings, and layer norms.
-
Learning rate adjustment: Muon automatically adjusts the learning rate based on parameter shape. See
AdjustLrFnfor details. -
Weight decay timing: Unlike typical optimizers, Muon applies weight decay AFTER orthogonalization but uses the original (unadjusted) learning rate for it.
Implementations§
Source§impl<B: Backend> Muon<B>
impl<B: Backend> Muon<B>
Sourcepub fn validate_step<const D: usize>(
&self,
lr: LearningRate,
tensor: &Tensor<B, D>,
grad: &Tensor<B, D>,
state: Option<&MuonState<B, D>>,
) -> Result<(), MuonError>
pub fn validate_step<const D: usize>( &self, lr: LearningRate, tensor: &Tensor<B, D>, grad: &Tensor<B, D>, state: Option<&MuonState<B, D>>, ) -> Result<(), MuonError>
Check tensor metadata before submitting kernels. Does not scan values for NaN/Inf. Unscale and validate gradients before calling (on every rank).
Sourcepub fn try_step<const D: usize>(
&self,
lr: LearningRate,
tensor: Tensor<B, D>,
grad: Tensor<B, D>,
state: Option<MuonState<B, D>>,
) -> Result<(Tensor<B, D>, Option<MuonState<B, D>>), MuonError>
pub fn try_step<const D: usize>( &self, lr: LearningRate, tensor: Tensor<B, D>, grad: Tensor<B, D>, state: Option<MuonState<B, D>>, ) -> Result<(Tensor<B, D>, Option<MuonState<B, D>>), MuonError>
Fallible metadata-checked update. Output and state may still be executing asynchronously. An accepted submission is not proof of device completion.
Uses a complete 2D matrix, never a flattened mixed-parameter buffer or an arbitrary shard. Missing gradients must be skipped by the caller.
Trait Implementations§
Source§impl<B: Backend> SimpleOptimizer<B> for Muon<B>
impl<B: Backend> SimpleOptimizer<B> for Muon<B>
Source§fn step<const D: usize>(
&self,
lr: LearningRate,
tensor: Tensor<B, D>,
grad: Tensor<B, D>,
state: Option<Self::State<D>>,
) -> (Tensor<B, D>, Option<Self::State<D>>)
fn step<const D: usize>( &self, lr: LearningRate, tensor: Tensor<B, D>, grad: Tensor<B, D>, state: Option<Self::State<D>>, ) -> (Tensor<B, D>, Option<Self::State<D>>)
Perform a single Muon optimization step.
§Algorithm
- Apply momentum to gradient
- Orthogonalize update via Newton-Schulz
- Adjust learning rate based on parameter shape
- Apply weight decay (using original lr)
- Update parameter (using adjusted lr)
§Notes
Unlike typical optimizers, the weight decay and parameter update use different learning rates:
- Weight decay uses the original
lr - Parameter update uses the shape-adjusted
lr
§Panics
This function will panic if the input tensors are not 2D.