Skip to main content

Module codec

Module codec 

Source
Expand description

Residual quantization codec for PLAID.

Once k-means has produced a set of coarse centroids, every token embedding can be represented as:

token ≈ centroid[centroid_id] + decode(residual_codes)

The residual is the element-wise difference between the token and its nearest centroid. Each residual dimension is then placed into one of 2^nbits buckets according to a precomputed set of cutoffs, and the bucket index (0…2ⁿ-1) is what we store on disk. At read time the bucket index is mapped back to a reconstruction value via bucket_weights and added to the centroid, yielding an approximate copy of the original token.

This module exposes the codec state and the encode/decode operations assuming the codec has already been trained. Bucket cutoffs and weights are learned from a sample of residuals via train_quantizer.

Storage layout: residual codes are LSB-first bit-packed at nbits bits each. Supported widths are {1, 2, 4, 8} — enough to cover every value the ColBERTv2/PLAID papers use in practice. For a 128-d embedding at 2 bits, this is 32 bytes per token (vs. 128 bytes unpacked), matching the paper’s §4.5 packed-index layout.

Structs§

DecodeTable
Precomputed 256-entry lookup table mapping every possible packed byte to the sequence of bucket_weights values it decodes to.
EncodedVector
A single encoded token: a centroid reference plus a bit-packed buffer of per-dim bucket codes.
ResidualCodec
A trained residual-quantization codec.

Functions§

packed_bytes_per_vector
Number of bytes required to pack dim codes at nbits bits each.
read_code
Read the code at logical position i from a packed buffer.
train_quantizer
Learn bucket cutoffs and reconstruction weights from a sample of residual values.