Skip to main content

Module bulk_embed

Module bulk_embed 

Source
Expand description

inillucent embed: fill a vector column for every row of a table, on the processor or a graphics card.

Invariant: a row is either embedded by the same model and the same text that embed(prefix || text) would use, or it is reported. The command reads the rows whose vector column is NULL, embeds each text with the prefix put in front, and writes the vector in the layout embed() returns. A row whose text is NULL or empty is skipped and counted. A row cut at the model’s token limit is embedded from the tokens that fit and is counted and named, so nothing is silently shortened. The device is the one that was asked for: a card that will not start is an error that names the installer command and never a run on the processor.

§Why it is a command and not UPDATE t SET v = embed(text)

embed() embeds one row per call on the processor, which for the 558,429 chunks of the study would take about 12 hours. This command sorts a slice of rows by length, groups them under a memory ceiling, runs several sessions at once, and commits every --commit-every rows, so a stopped run resumes at the first row still NULL.

The model is opened here from inillucent_core because the SQL engine sits below the retrieval engine and cannot reach it; this crate already links both.

Constants§

DEFAULT_COMMIT_EVERY
The rows written in one transaction when --commit-every is not given.
MAX_NAMED_TRUNCATED
How many rowids of truncated rows the report names.

Functions§

embed_table
embed: embeds a text column into a vector column.