Expand description
SIMD-accelerated dot products for the search hot path.
Three dtype paths, three SIMD targets:
AVX2 (x86_64) NEON (aarch64) Scalar
f32 · f32 8 lanes (f32x8) 4 lanes (f32x4) autovec
f32 · f16 8 lanes (load+cvt) 4 lanes (load+cvt) autovec
f32 · i8 16 lanes (i8 -> i32) 16 lanes (i8 -> i32) autovec
f32 · i4 32-nib unpack/step 32-nib unpack/step autovecThe int4 kernel is block-64 (per-group f16 absmax scale). SIMD vectorizes the nibble unpack but reduces per group identically to the scalar path, so all three backends agree bit-for-bit.
Accumulators are always f32. The query is f32 (L2-normalized), the database is f32 / f16 / i8 (i8 with a per-vector scale). Final score is the real cosine.
Detection happens once at module load via OnceLock. The dispatch
function is a function pointer chosen at first call, so the per-query
cost is one indirect call, not a CPUID check per vector.
Safety boundary. Every unsafe kernel below reads through raw pointers
sized by q.len() / dim, so the safe dispatchers in this module are
the ONLY place the length invariants are checked, and they check with
assert! (kept in release builds), never debug_assert!. The reader
validates section sizes at open time too, but a defence three layers
away from the pointer is not a defence; the cost is one integer compare
per row, invisible next to the dot product itself.
Enums§
- Simd
Backend - What backend is the runtime using right now? Useful for
urna stats/ benchmarks so the user can see whether SIMD is active.
Functions§
- detect_
backend - The SIMD backend selected at runtime. Cached after the first call.
- dot_
f32_ bytes - Dot product between an f32 query and an f32 row stored as little-endian bytes (the way embeddings live in mmap).
- dot_
f32_ f16_ bytes - Dot product between an f32 query and an f16 row stored as little-endian bytes. Accumulates in f32. The query stays f32 (it is normalized once per call, no need to drop precision there).
- dot_
f32_ f16_ scalar - dot_
f32_ i8 - Dot product between an f32 query and a single i8 row, multiplied by
the row’s f32 scale.
qstays f32; the i8 row is widened to i32 in the inner loop, multiplied by f32 lanes ofq, accumulated in f32. - dot_
f32_ i4_ blocked - Fused dequant + dot for an int4 block-
blockrow against an f32 query.codesisdim/2packed nibble bytes (low nibble first),group_scalesis one f32 perblock-dim group. The SIMD backends vectorize the nibble unpack but reduce per-group identically to scalar, so the result is bit-for-bit equal across all three backends (float add is not associative; a lane-parallel reduction would diverge in the last ulp). - dot_
f32_ i4_ blocked_ scalar - Fused dequant + dot for int4 block-
blockcodes against an f32 query.codesisdim/2packed bytes (two nibbles each, low nibble first);group_scalesis one f32 perblock-dim group. Each component contributesq[j] * code[j] * group_scales[j / block], accumulated in f32 per group so the per-group scale multiplies the group’s partial sum once (matching the SIMD backends bit-for-bit). - dot_
f32_ i8_ scalar - dot_
f32_ scalar - score_
int8_ section - Score every row of an int8 embeddings section against
q.out[i]is the cosine score; the runtime sorts these.