1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
//! Asymmetric int8-query × fused-LUT MaxSim over packed residual codes.
//!
//! This is the stage-2 scoring kernel of a ColBERT / PLAID-style late
//! interaction engine, extracted so any engine can call it: score one
//! candidate document's *stored* residual codes against a query, without
//! decompressing the document to floats first.
//!
//! ```text
//! q · token = q · centroid[cid] (optional, supplied by the host)
//! + Σ_d q_d · bucket_weight[code_d] (int8 query × int8 table, integer MACs)
//! score = Σ_q max_t (q · token_t) · inv_norm_t (optional per-token normalisation)
//! ```
//!
//! The document side is a table that turns each packed residual *byte*
//! straight into its `8/nbits` int8 bucket weights ([`Lut`]). The query side
//! is int8 codes with one f32 scale per row ([`PreparedQuery`]). A
//! [`Scorer`] binds the two, plus the host's optional centroid term, and
//! scores [`DocView`]s: borrowed slices pointing at wherever the host keeps
//! its bytes (heap, mmap, cell-contiguous).
//!
//! # What is inside, what is outside
//!
//! Inside: the loop order (doc-token-outer, so each token expands once and is
//! reused across every query row), the SIMD (NEON `sdot` and `smmla`, AVX2,
//! AVX-VNNI and AVX-512 VNNI), the runtime dispatch, and a scalar reference
//! every SIMD path must match **bit-for-bit**.
//!
//! Where an architecture offers more than one int8 dot instruction, the
//! faster one depends on the core rather than on the feature bits: Arm's
//! Neoverse N2 doubles its throughput with `smmla` while Apple's M4 loses
//! with it. Dispatch therefore measures the candidates once per process and
//! caches the winner ([`supported_kernels`], [`Lut::pin_kernel`]). Because
//! every kernel is bit-identical, that choice can only change speed.
//!
//! Outside: candidate generation, IVF, storage, threads. The crate has no
//! dependencies and owns no thread pool; [`Lut`], [`PreparedQuery`] and
//! [`Scorer`] are `Send + Sync`, so the host parallelises across queries or
//! across candidate chunks however it already does.
//!
//! # Example
//!
//! ```
//! use maxsim_lut::{ColbertPacking, Codes, DocView, Lut, PreparedQuery, Scorer};
//!
//! let dim = 128;
//! let nbits = 4;
//! // Bucket weights come from the host's residual quantiser (2^nbits of them).
//! let weights: Vec<f32> = (0..16).map(|i| -0.3 + 0.04 * i as f32).collect();
//! let lut = Lut::new(&ColbertPacking::new(nbits).unwrap(), &weights).unwrap();
//!
//! // One query: 32 tokens of dim 128, row-major f32.
//! let query = vec![0.01f32; 32 * dim];
//! let q = PreparedQuery::new(&lut, &query, 32, dim).unwrap();
//!
//! // Stage-1 product the host already has: [num_centroids, n_query_tokens].
//! let num_centroids = 1024;
//! let cdot = vec![0.0f32; num_centroids * 32];
//! let scorer = Scorer::new(&lut, &q).with_centroid_term(&cdot, num_centroids).unwrap();
//!
//! // A candidate document: 200 tokens, 64 packed bytes each, one centroid id per token.
//! let packed = vec![0u8; 200 * 64];
//! let codes = vec![7u32; 200];
//! let doc = DocView::new(&packed, 200, 64).codes(Codes::U32(&codes));
//! let score: f32 = scorer.score(doc);
//! assert!(score.is_finite());
//! println!("kernel in use: {}", lut.kernel(dim));
//! ```
//!
//! # Preconditions the host must know
//!
//! * `dim · nbits` must be a multiple of 8 (whole packed bytes) and `dim ≤ 256`.
//! * The SIMD paths need `dim % 8 == 0`; other dims score correctly on the
//! scalar path. [`Lut::kernel`] tells you which path will run.
//! * A document's packed rows must be contiguous, one row per token, at a
//! fixed `row_stride` of at least `dim / (8/nbits)` bytes.
//! * The win depends on that contiguity. A host that scatters a document's
//! tokens across cells will see the kernel run and the speedup vanish.
//!
//! # Provenance
//!
//! The kernels are extracted from next-plaid's `residual_lut.rs`
//! (Apache-2.0, <https://github.com/lightonai/next-plaid>), with the codec
//! coupling replaced by the [`Packing`] trait and the ndarray types by slices.
pub use ;
pub use ;
pub use ;
pub use PreparedQuery;
pub use ;
/// Errors from building tables and queries or from shape validation.
/// Row stride, in lanes, that the SIMD kernels read query rows at: `dim`
/// rounded up to 64 lanes, a multiple of the NEON (16), AVX2 (32) and
/// AVX-512 (64) chunk widths. Padding lanes are zero and contribute
/// `anything · 0 = 0`.
pub