Skip to main content

Module update

Module update 

Source
Expand description

Incremental updates: mutate an existing Index without retraining.

A freshly built index “freezes” its codec — the centroids and residual cutoffs/weights were learned from the token distribution at that point in time. In docbert’s sync loop the vast majority of documents are unchanged from one sync to the next; only a handful are added, updated, or deleted. Retraining the codec every sync is wasteful: k-means and quantizer training dominate build_index.

apply_update takes a small mutation plan and produces a new index with deleted documents removed, upserted documents re-encoded against the existing codec, and the centroid → tokens inverted file rebuilt from the final encoded token list. No k-means, no quantizer retraining.

Callers are expected to trigger a full rebuild (via build_index) periodically to combat codec drift as the corpus evolves — the existing codec only stays well-calibrated while the underlying token distribution doesn’t change dramatically.

Structs§

IndexUpdate
Mutation plan consumed by apply_update.

Functions§

apply_update
Produce a new Index reflecting update applied to index, reusing the existing codec and centroids.