# khive-quant
SQ8 scalar quantization codecs for approximate distance computation in ANN
indexes. Two codecs, chosen by which distance metric the index needs:
`Sq8Codec` (per-dimension affine scale, for dot product / cosine) and
`GsSq8Codec` (global shared scale, for L2 — the Vamana acquisition path).
## Usage
```rust
use khive_quant::Sq8Codec;
let corpus: Vec<Vec<f32>> = vec![
vec![0.1, 0.9, 0.4],
vec![0.8, 0.2, 0.6],
];
let codec = Sq8Codec::train(&corpus);
let encoded: Vec<_> = corpus.iter().map(|v| codec.encode(v)).collect();
let dot = codec.approx_dot(&encoded[0], &encoded[1]);
let cosine_dist = codec.approx_cosine_dist(&encoded[0], &encoded[1]);
```
`train` / `train_flat` compute per-dimension `min`/`max` from the corpus and
derive `scale_i = (max_i - min_i) / 255`; `encode` maps each `f32` dimension to
a `u8` code via `round((x - min_i) / scale_i)`. `encode_par` / `encode_flat_par`
parallelize encoding across a batch with `rayon`. `approx_dot`,
`approx_cosine_dist`, and `approx_l2_sq` reconstruct original-scale distances
from `u8` codes with per-dimension weighted sums accumulated in `f64`. This
preserves small dimensions when another dimension has a much wider range.
## GsSq8Codec — the Vamana acquisition path
`GsSq8Codec` uses one shared scale `gs = max_range_across_dims / 255` for every
dimension (per-dim `min_i` offsets are still subtracted before quantizing).
This makes squared L2 in code space `gs² * sum((a_i - b_i)^2)` algebraically
exact after the lossy `f32` -> `u8` encode — the offset terms cancel and `gs²`
factorizes out, so `GsSq8Codec::l2_sq` needs no residual pass or anisotropy
gate. The trade-off is honest, not hidden: narrow-range dimensions get fewer
`u8` levels and contribute proportionally less L2 signal than the wide-range
dimensions that set `gs`. `GsSq8Codec::is_in_distribution` flags query vectors
whose components fall outside the trained range so a caller can fall back to
exact `f32` distance for out-of-distribution queries — see
`VamanaIndex::search`.
## Hot-loop kernels
`GsSq8Codec` uses `u8_l2sq_u32` for squared L2: NEON `vabdq_u8` +
`vmull_u8` on aarch64, or a chunked portable fallback elsewhere.
`Sq8Codec` computes dot product and squared L2 with per-dimension weighted
`f64` sums. The `u8_dot_u32` integer helper remains test-only because a shared
scale loses narrow dimensions.
## Where this sits
Built on `rayon` only — no khive-* dependencies. Consumed today by
[khive-vamana](https://crates.io/crates/khive-vamana) for its SQ8-quantized
acquisition path. Governed by
[ADR-052](https://github.com/ohdearquant/khive/blob/main/docs/adr/ADR-052-ann-production-lifecycle.md),
which documents why the predecessor per-dimension L2 codec (with an
anisotropy gate calibrated on a synthetic corpus) silently fell back to a full
residual pass on real transformer embeddings, and why the global-scale design
eliminates the gate entirely.
## License
Apache-2.0.