pub enum QDType {
Show 21 variants
Q4_0,
Q4_1,
Q5_0,
Q5_1,
Q8_0,
Q8_1,
Q2_K,
Q3_K,
Q4_K,
Q5_K,
Q6_K,
Q8_K,
IQ2_XXS,
IQ2_XS,
IQ2_S,
IQ3_XXS,
IQ3_S,
IQ1_S,
IQ1_M,
IQ4_NL,
IQ4_XS,
}Expand description
Quantized data type of a GGUF super-block.
The numeric codes match the ggml_type enum (ggml.h): the same codes
Tensor::load_gguf matches on. Every variant
loads as raw [num_blocks, block_bytes] DType::U8 via load_gguf and
decodes to dense floats via
Tensor::dequantize.
Variants§
Q4_0
4-bit, 32 elements in 18 bytes (block_q4_0).
Q4_1
4-bit, 32 elements in 20 bytes (block_q4_1).
Q5_0
5-bit, 32 elements in 22 bytes (block_q5_0).
Q5_1
5-bit, 32 elements in 24 bytes (block_q5_1).
Q8_0
8-bit, 32 elements in 34 bytes (block_q8_0).
Q8_1
8-bit, 32 elements in 36 bytes (block_q8_1).
Q2_K
2-bit super-block, 256 elements in 84 bytes (block_q2_K).
Q3_K
3-bit super-block, 256 elements in 110 bytes (block_q3_K).
Q4_K
4-bit super-block, 256 elements in 144 bytes (block_q4_K).
Q5_K
5-bit super-block, 256 elements in 176 bytes (block_q5_K).
Q6_K
6-bit super-block, 256 elements in 210 bytes (block_q6_K).
Q8_K
8-bit super-block, 256 elements in 292 bytes (block_q8_K).
IQ2_XXS
Improved 2-bit, 256 elements in 66 bytes (block_iq2_xxs).
IQ2_XS
Improved 2-bit, 256 elements in 74 bytes (block_iq2_xs).
IQ2_S
Improved 2-bit, 256 elements in 82 bytes (block_iq2_s).
IQ3_XXS
Improved 3-bit, 256 elements in 98 bytes (block_iq3_xxs).
IQ3_S
Improved 3-bit, 256 elements in 110 bytes (block_iq3_s).
IQ1_S
Improved 1-bit, 256 elements in 50 bytes (block_iq1_s).
IQ1_M
Improved 1-bit, 256 elements in 56 bytes (block_iq1_m).
IQ4_NL
Non-linear 4-bit, 32 elements in 18 bytes (block_iq4_nl).
IQ4_XS
Non-linear 4-bit, 256 elements in 136 bytes (block_iq4_xs).
Implementations§
Source§impl QDType
impl QDType
Sourcepub const fn elems_per_block(self) -> i64
pub const fn elems_per_block(self) -> i64
Number of elements one super-block decodes to.
Sourcepub const fn block_bytes(self) -> i64
pub const fn block_bytes(self) -> i64
Size in bytes of one super-block on disk.