Skip to main content

Module encoding

Module encoding 

Source
Expand description

section payload encoding (raw / zstd / float16 / int8).

Two orthogonal axes:

  • Wire encoding of a section payload (SECTION_ENCODING_* in layout): how the bytes are stored on disk. raw and zstd apply to any non-embedding section. float16 and int8 only apply to the embeddings section.

  • Logical dtype of an embeddings section (Manifest::dtype): how to interpret the bytes after wire decoding. float32 is the recall-max baseline; float16 and int8 are smaller-but-lossy variants that always accumulate in f32 at search time.

Section checksums (SectionEntry::checksum) hash the physical bytes as stored. content_hash hashes the decoded bytes so two files with the same logical content but different wire encoding still produce the same content_hash for non-quantized sections.

Structs§

Deduped
the result of a dedup pass: the first-seen unique texts (in first-seen order, deterministic) and a per-chunk back-reference into that pool.
Int4EmbeddingsView
Decoded view over an int4 embeddings payload. Slices borrow the input bytes (no copy); accessors decode scales/codes on demand.
Int8EmbeddingsView
Decoded view over an int8 embeddings payload. The slices borrow from the input bytes (no copy).
IntpackReader
random-access reader over an intpack payload. parse validates the header and directory once; get then reaches any element in O(1) block lookups, decoding only the touched block.
TxtStreams
a parsed txt_streams payload. parse validates the header and offset table once; stream/text then reach any single chunk in O(1) without touching the others (the O(1)-reopen enabler). decode_payload rebuilds the full canonical section by concatenating all of them.

Constants§

DEDUP_MAP_V1
dedup-map payload version byte (leads the back-reference array).
DEFAULT_ZSTD_LEVEL
Default zstd compression level. 19 is in the “high” tier - slow to encode but a one-time cost and yields ~30% smaller text payloads than the default level 3.
INT4_BLOCK
Block size for the per-group absmax scale. dim must be a multiple.
INT4_PAYLOAD_VERSION
INT4_PREFIX_SIZE
INT4_SCALE_KIND_PER_GROUP
INT8_PAYLOAD_VERSION
INT8_PREFIX_SIZE
INT8_SCALE_KIND_PER_VECTOR
INTPACK_BLOCK
values per frame-of-reference block.
MAX_DICT_BYTES
max trained-dictionary size in bytes (~110 KiB). the ZDICT trainer returns a SMALLER dict when the sample set does not justify the cap, so this is an upper bound, not a fixed size. chosen in the 64-112 KiB band the work order pins; large enough to capture pt-br boilerplate, small enough to stay a negligible slice of the file.
TXT_STREAMS_V1
leading kind/version byte. v1 = plain per-stream zstd. the dict (V2) and fsst (V3) variants live in zstd_dict.rs / fsst.rs and claim the next values here, reusing this container without a new encoding id beyond their section-entry codec id.
TXT_STREAMS_V2
kind/version byte for the dict-framed variant. distinct from TXT_STREAMS_V1 so a reader dispatches the right per-stream codec without a new encoding id beyond zstd_dict (5) on the section entry.
TXT_STREAMS_V3
kind/version byte for the fsst-framed variant.

Functions§

decode_dedup_map
parse a dedup-map payload back to the back-reference array, bounds-checked.
decode_fsst_payload
reconstruct the canonical chunks_canonical payload from an fsst-framed txt_streams V3 payload. byte-identical to sections::encode_chunks_canonical, so content_hash is preserved.
decode_payload
Decode a section payload from its on-disk encoding to the logical bytes a reader consumes, via the wire-codec registry. For raw this is a borrow; for zstd an owned decompressed buffer. The zstd_dict (id 5) codec needs the shared dictionary (section 0x0A) and is decoded via decode_payload_with_dict, so it is rejected here. Unknown or reserved-but-unimplemented encodings are rejected.
decode_payload_with_dict
Decode a chunks_canonical payload that MAY be dict-framed (zstd_dict, id 5), supplying the shared dictionary from section 0x0A. all other encodings ignore the dict and route through decode_payload. the dict variant decodes BYTE-IDENTICALLY to the raw chunks_canonical payload, so content_hash and citations are unchanged.
decode_txt_streams_payload
reconstruct the canonical (raw-encoding) chunks_canonical payload from a full txt_streams payload (including the leading kind byte). the output is byte-identical to sections::encode_chunks_canonical, so content_hash is preserved. dispatched by encoding::decode_payload.
decode_zstd_dict_payload
reconstruct the canonical chunks_canonical payload from a dict-framed txt_streams V2 payload using the shared dict. byte-identical to sections::encode_chunks_canonical, so content_hash is preserved.
dedup
run the first-seen dedup pass over texts (the decompressed canonical strings, in chunk order). deterministic: a HashMap keyed by the text only decides membership, while first-seen ORDER is driven by the input sequence, so two builds over the same corpus match exactly.
encode_dedup_map
serialize the back-reference array: a version byte then an intpack (encoding-id-4 primitive, reused) packing of the u32 refs as u64. a pure function of the refs, so two builds are byte-identical.
encode_fsst
encode texts as per-chunk fsst frames behind the shared txt_streams offset table, with one corpus-wide symbol table embedded after the container header. a pure function of the inputs, so two builds match.
encode_int4_embeddings
Encode the int4 embeddings section payload. embeddings is n * dim row-major f32; dim must be divisible by INT4_BLOCK.
encode_int8_embeddings
Encode the int8 embeddings section payload. Layout matches INT8_PAYLOAD_VERSION / INT8_SCALE_KIND_PER_VECTOR.
encode_smallest
Cost-driven encoder: try every candidate wire encoding and return the (encoding_id, bytes) of the SMALLEST result, so the writer can record the chosen id in the section entry. Ties break toward the EARLIEST candidate (cheaper-to-decode wins an equal-size race). This only auto- picks among non-embedding encodings; existing presets that name an encoding explicitly are untouched, so the frozen output stays byte-identical.
encode_txt_streams
encode texts (the canonical strings, in chunk order) as per-chunk independent zstd streams behind an intpack offset table. the layout is a pure function of the inputs, so two builds are byte-identical.
encode_zstd_dict
encode texts as per-chunk dict-framed zstd streams behind the shared txt_streams offset table. byte-identical output for identical inputs + identical dict (the dict itself is deterministic), so two builds match.
expand_dedup
re-expand the unique pool through the back-references to the full ordered list of canonical texts, byte-identical to the original. every ref must index a real unique entry; a hostile map errors, never panics.
expected_embeddings_size
Expected size of the embeddings section for a given dtype. None for an unknown dtype AND when n * dim overflows: both come from an untrusted header, and a wrapped product would let a hostile file claim a huge corpus behind a tiny section (mutation-fuzz finding).
f16_bytes_to_f32
Decode a float16 byte slice into f32 values, accumulating in f32 for downstream dot-product accuracy.
f32_to_f16_bytes
Convert an L2-normalized float32 embedding to float16, returning the little-endian byte representation. Not a checked NaN/Inf path - callers must validate inputs upstream.
int4_blocks_per_row
Number of 64-dim blocks per row. dim must be divisible by INT4_BLOCK.
nibble_to_i4
Sign-extend a 4-bit nibble (low 4 bits of b) to an i8 in [-8, 7].
pack_nibbles
Pack signed 4-bit codes ([-7, 7]) into bytes, two nibbles per byte, low nibble first. The nibble is the two’s-complement low 4 bits, so -7..=7 maps to 0x9..=0x7 and unpacks back exactly via sign extension.
pack_u64s
pack a slice of u64 with per-128-block frame-of-reference bitpacking. the inverse is unpack_u64s (and random access via IntpackReader).
quantize_f32_to_i4
Quantize one L2-normalized f32 row to int4 with per-64-block absmax scales. Returns (scales, codes) where scales[g] is the f16 absmax scale of block g and codes[j] in [-7, 7] reconstructs as codes[j] as f32 * scales[j / 64]. Mirrors quantize_f32_to_i8 but per 64-dim group; a zero block maps to all-zero codes with scale 1.
quantize_f32_to_i8
Quantize an L2-normalized float32 embedding to int8 with a per-vector scale. Returns (scale, i8_bytes) where f32_value ≈ i8 * scale.
train_dict
train ONE deterministic zstd dictionary over the canonical texts.
unpack_u64s
decode an entire intpack payload back to its u64 values.
zstd_encode
Compress with zstd at DEFAULT_ZSTD_LEVEL. Returns the compressed bytes ready to write as the section payload.