pub fn dequant_mxfp4_row(
packed: &[u8],
scales: &[u8],
) -> Result<Vec<f32>, QuantError>Expand description
Dequantizes one row of Kimi K3’s MXFP4-packed expert weights. Unlike
every other kernel in this module, MXFP4 here is NOT a single
interleaved byte stream – Kimi K3’s real safetensors checkpoint
stores the packed 4-bit codes and the per-group E8M0 scales as two
separate tensors (*.weight_packed, *.weight_scale; confirmed
directly against a real shard header’s tensor shapes, not ggml’s own
combined-block GGUF convention), so this takes both buffers directly
rather than one combined block stream. packed is in_dim/2 bytes
(2 nibble-packed E2M1 codes per byte, low-nibble-first-half /
high-nibble-second-half within each 32-element group – same
convention as this module’s other nibble-packed formats); scales is
in_dim/MXFP4_GROUP_SIZE bytes (one E8M0 scale byte per group).