pub struct RefitAcc {
pub support: Vec<u32>,
pub gss: Vec<f32>,
pub ya: Vec<f32>,
pub hidden: usize,
pub tokens: u64,
pub buf_g: Vec<f32>,
pub buf_o: Vec<f32>,
pub buf_t: usize,
}Expand description
Online accumulators for the AWNP refit of a narrowed FFN.
The refit needs Gss = A_SᵀA_S and YA = YᵀA_S per layer, where A_S
are the calibration activations of the KEPT neurons and Y the full
FFN output. Both are small enough to hold; the thing that is not is
the activations they are built from — a 27B layer would dump a
gigabyte per thousand tokens. So they are accumulated as the
calibration runs and written once at the end.
CMF_FFN_REFIT=<dir> holds support.<L>.u32 (a u32 count then the
kept indices) for every layer to accumulate; CMF_FFN_REFIT_FROM/TO
bound the layer span so the accumulators fit in RAM.
Fields§
§support: Vec<u32>§gss: Vec<f32>§ya: Vec<f32>§tokens: u64§buf_g: Vec<f32>Activations staged transposed ([ns, t] and [hidden, t]) until the
batch is worth a GEMM. The product costs ns² to move and add
REGARDLESS of how many tokens went into it, so folding 16 chunks
into one call cuts that cost 16× — it was 15 TB of traffic per
calibration pass at one call per 256 tokens.
buf_o: Vec<f32>§buf_t: usize