Skip to main content

Module simd

Module simd 

Source
Expand description

SIMD-accelerated dot products for the search hot path.

Three dtype paths, three SIMD targets:

               AVX2 (x86_64)         NEON (aarch64)        Scalar
  f32 · f32    8 lanes (f32x8)       4 lanes (f32x4)       autovec
  f32 · f16    8 lanes (load+cvt)    4 lanes (load+cvt)    autovec
  f32 · i8     16 lanes (i8 -> i32)  16 lanes (i8 -> i32)  autovec
  f32 · i4     32-nib unpack/step    32-nib unpack/step    autovec

The int4 kernel is block-64 (per-group f16 absmax scale). SIMD vectorizes the nibble unpack but reduces per group identically to the scalar path, so all three backends agree bit-for-bit.

Accumulators are always f32. The query is f32 (L2-normalized), the database is f32 / f16 / i8 (i8 with a per-vector scale). Final score is the real cosine.

Detection happens once at module load via OnceLock. The dispatch function is a function pointer chosen at first call, so the per-query cost is one indirect call, not a CPUID check per vector.

Safety boundary. Every unsafe kernel below reads through raw pointers sized by q.len() / dim, so the safe dispatchers in this module are the ONLY place the length invariants are checked, and they check with assert! (kept in release builds), never debug_assert!. The reader validates section sizes at open time too, but a defence three layers away from the pointer is not a defence; the cost is one integer compare per row, invisible next to the dot product itself.

Enums§

SimdBackend
What backend is the runtime using right now? Useful for urna stats / benchmarks so the user can see whether SIMD is active.

Functions§

detect_backend
The SIMD backend selected at runtime. Cached after the first call.
dot_f32_bytes
Dot product between an f32 query and an f32 row stored as little-endian bytes (the way embeddings live in mmap).
dot_f32_f16_bytes
Dot product between an f32 query and an f16 row stored as little-endian bytes. Accumulates in f32. The query stays f32 (it is normalized once per call, no need to drop precision there).
dot_f32_f16_scalar
dot_f32_i8
Dot product between an f32 query and a single i8 row, multiplied by the row’s f32 scale. q stays f32; the i8 row is widened to i32 in the inner loop, multiplied by f32 lanes of q, accumulated in f32.
dot_f32_i4_blocked
Fused dequant + dot for an int4 block-block row against an f32 query. codes is dim/2 packed nibble bytes (low nibble first), group_scales is one f32 per block-dim group. The SIMD backends vectorize the nibble unpack but reduce per-group identically to scalar, so the result is bit-for-bit equal across all three backends (float add is not associative; a lane-parallel reduction would diverge in the last ulp).
dot_f32_i4_blocked_scalar
Fused dequant + dot for int4 block-block codes against an f32 query. codes is dim/2 packed bytes (two nibbles each, low nibble first); group_scales is one f32 per block-dim group. Each component contributes q[j] * code[j] * group_scales[j / block], accumulated in f32 per group so the per-group scale multiplies the group’s partial sum once (matching the SIMD backends bit-for-bit).
dot_f32_i8_scalar
dot_f32_scalar
score_int8_section
Score every row of an int8 embeddings section against q. out[i] is the cosine score; the runtime sorts these.