Expand description
section payload encoding (raw / zstd / float16 / int8).
Two orthogonal axes:
-
Wire encoding of a section payload (
SECTION_ENCODING_*inlayout): how the bytes are stored on disk.rawandzstdapply to any non-embedding section.float16andint8only apply to the embeddings section. -
Logical dtype of an embeddings section (
Manifest::dtype): how to interpret the bytes after wire decoding.float32is the recall-max baseline;float16andint8are smaller-but-lossy variants that always accumulate inf32at search time.
Section checksums (SectionEntry::checksum) hash the physical
bytes as stored. content_hash hashes the decoded bytes so two
files with the same logical content but different wire encoding
still produce the same content_hash for non-quantized sections.
Structs§
- Deduped
- the result of a dedup pass: the first-seen unique texts (in first-seen order, deterministic) and a per-chunk back-reference into that pool.
- Int4
Embeddings View - Decoded view over an int4 embeddings payload. Slices borrow the input bytes (no copy); accessors decode scales/codes on demand.
- Int8
Embeddings View - Decoded view over an int8 embeddings payload. The slices borrow from the input bytes (no copy).
- Intpack
Reader - random-access reader over an intpack payload.
parsevalidates the header and directory once;getthen reaches any element in O(1) block lookups, decoding only the touched block. - TxtStreams
- a parsed
txt_streamspayload.parsevalidates the header and offset table once;stream/textthen reach any single chunk in O(1) without touching the others (the O(1)-reopen enabler).decode_payloadrebuilds the full canonical section by concatenating all of them.
Constants§
- DEDUP_
MAP_ V1 - dedup-map payload version byte (leads the back-reference array).
- DEFAULT_
ZSTD_ LEVEL - Default zstd compression level. 19 is in the “high” tier - slow to encode but a one-time cost and yields ~30% smaller text payloads than the default level 3.
- INT4_
BLOCK - Block size for the per-group absmax scale.
dimmust be a multiple. - INT4_
PAYLOAD_ VERSION - INT4_
PREFIX_ SIZE - INT4_
SCALE_ KIND_ PER_ GROUP - INT8_
PAYLOAD_ VERSION - INT8_
PREFIX_ SIZE - INT8_
SCALE_ KIND_ PER_ VECTOR - INTPACK_
BLOCK - values per frame-of-reference block.
- MAX_
DICT_ BYTES - max trained-dictionary size in bytes (~110 KiB). the ZDICT trainer returns a SMALLER dict when the sample set does not justify the cap, so this is an upper bound, not a fixed size. chosen in the 64-112 KiB band the work order pins; large enough to capture pt-br boilerplate, small enough to stay a negligible slice of the file.
- TXT_
STREAMS_ V1 - leading kind/version byte. v1 = plain per-stream zstd. the dict (V2) and
fsst (V3) variants live in
zstd_dict.rs/fsst.rsand claim the next values here, reusing this container without a new encoding id beyond their section-entry codec id. - TXT_
STREAMS_ V2 - kind/version byte for the dict-framed variant. distinct from
TXT_STREAMS_V1so a reader dispatches the right per-stream codec without a new encoding id beyond zstd_dict (5) on the section entry. - TXT_
STREAMS_ V3 - kind/version byte for the fsst-framed variant.
Functions§
- decode_
dedup_ map - parse a dedup-map payload back to the back-reference array, bounds-checked.
- decode_
fsst_ payload - reconstruct the canonical
chunks_canonicalpayload from an fsst-framedtxt_streamsV3 payload. byte-identical tosections::encode_chunks_canonical, socontent_hashis preserved. - decode_
payload - Decode a section payload from its on-disk encoding to the logical bytes
a reader consumes, via the wire-codec registry. For
rawthis is a borrow; forzstdan owned decompressed buffer. Thezstd_dict(id 5) codec needs the shared dictionary (section 0x0A) and is decoded viadecode_payload_with_dict, so it is rejected here. Unknown or reserved-but-unimplemented encodings are rejected. - decode_
payload_ with_ dict - Decode a chunks_canonical payload that MAY be dict-framed (
zstd_dict, id 5), supplying the shared dictionary from section 0x0A. all other encodings ignore the dict and route throughdecode_payload. the dict variant decodes BYTE-IDENTICALLY to the raw chunks_canonical payload, so content_hash and citations are unchanged. - decode_
txt_ streams_ payload - reconstruct the canonical (raw-encoding)
chunks_canonicalpayload from a fulltxt_streamspayload (including the leading kind byte). the output is byte-identical tosections::encode_chunks_canonical, socontent_hashis preserved. dispatched byencoding::decode_payload. - decode_
zstd_ dict_ payload - reconstruct the canonical
chunks_canonicalpayload from a dict-framedtxt_streamsV2 payload using the shareddict. byte-identical tosections::encode_chunks_canonical, socontent_hashis preserved. - dedup
- run the first-seen dedup pass over
texts(the decompressed canonical strings, in chunk order). deterministic: aHashMapkeyed by the text only decides membership, while first-seen ORDER is driven by the input sequence, so two builds over the same corpus match exactly. - encode_
dedup_ map - serialize the back-reference array: a version byte then an intpack (encoding-id-4 primitive, reused) packing of the u32 refs as u64. a pure function of the refs, so two builds are byte-identical.
- encode_
fsst - encode
textsas per-chunk fsst frames behind the shared txt_streams offset table, with one corpus-wide symbol table embedded after the container header. a pure function of the inputs, so two builds match. - encode_
int4_ embeddings - Encode the int4 embeddings section payload.
embeddingsisn * dimrow-major f32;dimmust be divisible byINT4_BLOCK. - encode_
int8_ embeddings - Encode the int8 embeddings section payload. Layout matches
INT8_PAYLOAD_VERSION/INT8_SCALE_KIND_PER_VECTOR. - encode_
smallest - Cost-driven encoder: try every candidate wire encoding and return the
(encoding_id, bytes)of the SMALLEST result, so the writer can record the chosen id in the section entry. Ties break toward the EARLIEST candidate (cheaper-to-decode wins an equal-size race). This only auto- picks among non-embedding encodings; existing presets that name an encoding explicitly are untouched, so the frozen output stays byte-identical. - encode_
txt_ streams - encode
texts(the canonical strings, in chunk order) as per-chunk independent zstd streams behind an intpack offset table. the layout is a pure function of the inputs, so two builds are byte-identical. - encode_
zstd_ dict - encode
textsas per-chunk dict-framed zstd streams behind the shared txt_streams offset table. byte-identical output for identical inputs + identicaldict(the dict itself is deterministic), so two builds match. - expand_
dedup - re-expand the unique pool through the back-references to the full ordered list of canonical texts, byte-identical to the original. every ref must index a real unique entry; a hostile map errors, never panics.
- expected_
embeddings_ size - Expected size of the embeddings section for a given dtype.
Nonefor an unknown dtype AND whenn * dimoverflows: both come from an untrusted header, and a wrapped product would let a hostile file claim a huge corpus behind a tiny section (mutation-fuzz finding). - f16_
bytes_ to_ f32 - Decode a float16 byte slice into f32 values, accumulating in f32 for downstream dot-product accuracy.
- f32_
to_ f16_ bytes - Convert an L2-normalized float32 embedding to float16, returning the little-endian byte representation. Not a checked NaN/Inf path - callers must validate inputs upstream.
- int4_
blocks_ per_ row - Number of 64-dim blocks per row.
dimmust be divisible byINT4_BLOCK. - nibble_
to_ i4 - Sign-extend a 4-bit nibble (low 4 bits of
b) to ani8in[-8, 7]. - pack_
nibbles - Pack signed 4-bit codes (
[-7, 7]) into bytes, two nibbles per byte, low nibble first. The nibble is the two’s-complement low 4 bits, so-7..=7maps to0x9..=0x7and unpacks back exactly via sign extension. - pack_
u64s - pack a slice of u64 with per-128-block frame-of-reference bitpacking.
the inverse is
unpack_u64s(and random access viaIntpackReader). - quantize_
f32_ to_ i4 - Quantize one L2-normalized f32 row to int4 with per-64-block absmax
scales. Returns
(scales, codes)wherescales[g]is the f16 absmax scale of blockgandcodes[j]in[-7, 7]reconstructs ascodes[j] as f32 * scales[j / 64]. Mirrorsquantize_f32_to_i8but per 64-dim group; a zero block maps to all-zero codes with scale 1. - quantize_
f32_ to_ i8 - Quantize an L2-normalized float32 embedding to int8 with a per-vector
scale. Returns
(scale, i8_bytes)wheref32_value ≈ i8 * scale. - train_
dict - train ONE deterministic zstd dictionary over the canonical texts.
- unpack_
u64s - decode an entire intpack payload back to its
u64values. - zstd_
encode - Compress with zstd at
DEFAULT_ZSTD_LEVEL. Returns the compressed bytes ready to write as the section payload.