pub struct MuonAdamWConfig { /* private fields */ }Expand description
Muon for explicitly selected hidden matrices and AdamW for the remainder.
Optimizer::step(lr, ...) uses lr for Muon and lr * adamw_lr_ratio
for AdamW. try_step_with_lrs accepts independent rates instead.
This is a high-level tensor optimizer, not the experimental fused AdamW API.
Implementations§
Source§impl MuonAdamWConfig
impl MuonAdamWConfig
Sourcepub fn new() -> Self
pub fn new() -> Self
Create a new instance of the config.
§Arguments
§Default Arguments
§muon
Muon settings, with legacy numerical defaults preserved.
- Defaults to
"MuonConfig::new()"
§adamw
Settings for all parameters not explicitly selected for Muon.
- Defaults to
"AdamWConfig::new()"
§adamw_lr_ratio
AdamW/Muon learning-rate ratio. 0.015 maps 0.02 to 0.0003.
- Defaults to
0.015
Source§impl MuonAdamWConfig
impl MuonAdamWConfig
Sourcepub fn with_muon(self, muon: MuonConfig) -> Self
pub fn with_muon(self, muon: MuonConfig) -> Self
Sets the value for the field muon.
Muon settings, with legacy numerical defaults preserved.
- Defaults to
"MuonConfig::new()"
Sourcepub fn with_adamw(self, adamw: AdamWConfig) -> Self
pub fn with_adamw(self, adamw: AdamWConfig) -> Self
Sets the value for the field adamw.
Settings for all parameters not explicitly selected for Muon.
- Defaults to
"AdamWConfig::new()"
Sourcepub fn with_adamw_lr_ratio(self, adamw_lr_ratio: f64) -> Self
pub fn with_adamw_lr_ratio(self, adamw_lr_ratio: f64) -> Self
Sets the value for the field adamw_lr_ratio.
AdamW/Muon learning-rate ratio. 0.015 maps 0.02 to 0.0003.
- Defaults to
0.015
Source§impl MuonAdamWConfig
impl MuonAdamWConfig
Sourcepub fn init<B: AutodiffBackend, M: AutodiffModule<B>>(
&self,
module: &M,
muon_parameters: &[ParamId],
) -> Result<MuonAdamW<M, B>, MuonError>
pub fn init<B: AutodiffBackend, M: AutodiffModule<B>>( &self, module: &M, muon_parameters: &[ParamId], ) -> Result<MuonAdamW<M, B>, MuonError>
Validate configuration and resolve selected parameter IDs against a model. Do not select embeddings, output heads, normalization gains or biases.
Trait Implementations§
Source§impl Clone for MuonAdamWConfig
impl Clone for MuonAdamWConfig
Source§impl Config for MuonAdamWConfig
impl Config for MuonAdamWConfig
Source§fn save<P>(&self, file: P) -> Result<(), Error>
fn save<P>(&self, file: P) -> Result<(), Error>
std only.Source§fn load<P>(file: P) -> Result<Self, ConfigError>
fn load<P>(file: P) -> Result<Self, ConfigError>
std only.