pub struct BakeHyper {Show 24 fields
pub steps_a: usize,
pub steps_b: usize,
pub l1_init: f64,
pub l1_step: f64,
pub eval_every: usize,
pub lr_a: f64,
pub lr_b: f64,
pub tau: f32,
pub fcd_layers: usize,
pub batch: usize,
pub fcd_batch: usize,
pub seed: u64,
pub target_sparsity: f64,
pub l1_mult: f64,
pub mask_init: f32,
pub softplus_l1: bool,
pub checkpoint_accuracy: bool,
pub checkpoint_min_accuracy: Option<f64>,
pub checkpoint_min_balanced_accuracy: Option<f64>,
pub checkpoint_raw_priority: bool,
pub align: usize,
pub uniform_inter: bool,
pub focus_tokens: Vec<u32>,
pub focus_follow_tokens: Vec<u32>,
}Expand description
Hyper-parameters — defaults are the certified recipe.
Fields§
§steps_a: usize§steps_b: usize§l1_init: f64§l1_step: f64§eval_every: usize§lr_a: f64§lr_b: f64§tau: f32§fcd_layers: usize§batch: usizeIndependent fixed-length records per optimizer step. The loss is normalized over all focused targets in the batch.
fcd_batch: usizeFocused records per cached final-FFN optimizer step. This is separate
from batch because Phase A stores every layer’s activations while a
one-layer FCD cache stores only two hidden vectors per record.
seed: u64§target_sparsity: f64Target sparsity (0..1). When >0, the best checkpoint must have at least this fraction of neurons pruned; if none qualifies the highest-sparsity checkpoint is used.
l1_mult: f64L1 aggression multiplier: scales both l1_init and l1_step.
1.0 = harder pruning push, <1.0 = softer.
mask_init: f32Effective unlooped mask logit at step zero. The per-visit value is solved so that the product over loop visits equals sigmoid(init). 2.0 preserves the native recipe; 4.0 reproduces the older DTG-MA trading notebooks’ near-identity start.
softplus_l1: boolPenalize softplus(logit), whose derivative is sigmoid(logit), instead of penalizing sigmoid(logit) itself. This reproduces the older DTG-MA recipe and avoids an extra (1-sigmoid) attenuation near an open gate.
checkpoint_accuracy: boolSelect Phase-A checkpoints by held-out hard balanced accuracy when focused class tokens are configured. Otherwise held-out PPL remains the checkpoint metric.
checkpoint_min_accuracy: Option<f64>Optional strict lower bound for a focused checkpoint’s raw accuracy. A checkpoint is eligible only when its measured accuracy is greater than this value. This lets callers impose a natural-distribution majority guard instead of selecting from a balanced holdout.
checkpoint_min_balanced_accuracy: Option<f64>Optional strict lower bound for focused balanced accuracy. Combined
with checkpoint_min_accuracy, this prevents a majority-only mask
from becoming the shipped specialist.
checkpoint_raw_priority: boolWhen a joint accuracy guard is configured, rank eligible checkpoints by raw accuracy first, then balanced accuracy and PPL. The historical balanced-first selector remains the default for compatibility.
align: usizeRound each layer’s kept-neuron count UP to a multiple of this (0/1 = off). 32 keeps the defragged FFN on grouped codecs (in % 32 == 0) and SIMD kernels off their scalar tails.
uniform_inter: boolForce one FFN width across all layers (the max aligned count) — the whole-token GPU graphs require a uniform intermediate size.
focus_tokens: Vec<u32>When non-empty, LM loss is accumulated only where the next token is one of these ids. The whole chunk is still forwarded as context. This is useful for supervised corpora with a long input and a one-token answer, where ordinary all-token LM loss would drown the task signal in prompt reconstruction.
focus_follow_tokens: Vec<u32>Optional token(s) that must immediately follow a focused target.
Supervised ChatML uses the one-token label followed by <|im_end|>;
this prevents label names mentioned inside the user instruction from
being mistaken for answer positions.