pub enum Encoding {
Plain = 0,
PlainDictionary = 2,
Rle = 3,
BitPacked = 4,
DeltaBinaryPacked = 5,
DeltaLengthByteArray = 6,
DeltaByteArray = 7,
RleDictionary = 8,
ByteStreamSplit = 9,
}Expand description
Encodings supported by Parquet. Not all encodings are valid for all types. These enums are also used to specify the encoding of definition and repetition levels. See the accompanying doc for the details of the more complicated encodings.
Variants§
Plain = 0
Default encoding. BOOLEAN - 1 bit per value. 0 is false; 1 is true. INT32 - 4 bytes per value. Stored as little-endian. INT64 - 8 bytes per value. Stored as little-endian. FLOAT - 4 bytes per value. IEEE. Stored as little-endian. DOUBLE - 8 bytes per value. IEEE. Stored as little-endian. BYTE_ARRAY - 4 byte length stored as little endian, followed by bytes. FIXED_LEN_BYTE_ARRAY - Just the bytes.
PlainDictionary = 2
Deprecated: Dictionary encoding. The values in the dictionary are encoded in the plain type. in a data page use RLE_DICTIONARY instead. in a Dictionary page use PLAIN instead
Rle = 3
Group packed run length encoding. Usable for definition/repetition levels encoding and Booleans (on one bit: 0 is false; 1 is true.)
BitPacked = 4
Bit packed encoding. This can only be used if the data has a known max width. Usable for definition/repetition levels encoding.
DeltaBinaryPacked = 5
Delta encoding for integers. This can be used for int columns and works best on sorted data
DeltaLengthByteArray = 6
Encoding for byte arrays to separate the length values and the data. The lengths are encoded using DELTA_BINARY_PACKED
DeltaByteArray = 7
Incremental-encoded byte array. Prefix lengths are encoded using DELTA_BINARY_PACKED. Suffixes are stored as delta length byte arrays.
RleDictionary = 8
Dictionary encoding: the ids are encoded using the RLE encoding
ByteStreamSplit = 9
Encoding for floating-point data. K byte-streams are created where K is the size in bytes of the data type. The individual bytes of an FP value are scattered to the corresponding stream and the streams are concatenated. This itself does not reduce the size of the data but can lead to better compression afterwards.