Skip to main content

dequant_mxfp4_row

Function dequant_mxfp4_row 

Source
pub fn dequant_mxfp4_row(
    packed: &[u8],
    scales: &[u8],
) -> Result<Vec<f32>, QuantError>
Expand description

Dequantizes one row of Kimi K3’s MXFP4-packed expert weights. Unlike every other kernel in this module, MXFP4 here is NOT a single interleaved byte stream – Kimi K3’s real safetensors checkpoint stores the packed 4-bit codes and the per-group E8M0 scales as two separate tensors (*.weight_packed, *.weight_scale; confirmed directly against a real shard header’s tensor shapes, not ggml’s own combined-block GGUF convention), so this takes both buffers directly rather than one combined block stream. packed is in_dim/2 bytes (2 nibble-packed E2M1 codes per byte, low-nibble-first-half / high-nibble-second-half within each 32-element group – same convention as this module’s other nibble-packed formats); scales is in_dim/MXFP4_GROUP_SIZE bytes (one E8M0 scale byte per group).