Expand description
Incremental updates: mutate an existing Index without retraining.
A freshly built index “freezes” its codec — the centroids and
residual cutoffs/weights were learned from the token distribution
at that point in time. In docbert’s sync loop the vast majority of
documents are unchanged from one sync to the next; only a handful
are added, updated, or deleted. Retraining the codec every sync is
wasteful: k-means and quantizer training dominate build_index.
apply_update takes a small mutation plan and produces a new
index with deleted documents removed, upserted documents re-encoded
against the existing codec, and the centroid → tokens inverted
file rebuilt from the final encoded token list. No k-means, no
quantizer retraining.
Callers are expected to trigger a full rebuild (via
build_index) periodically to combat codec drift as the corpus
evolves — the existing codec only stays well-calibrated while the
underlying token distribution doesn’t change dramatically.
Structs§
- Index
Update - Mutation plan consumed by
apply_update.
Functions§
- apply_
update - Produce a new
Indexreflectingupdateapplied toindex, reusing the existing codec and centroids.