Expand description
Local embedding generation (LLM-only, one-shot per invocation). Embedding generation for the GraphRAG memory.
v1.0.76: the default build is LLM-only — the binary does NOT bundle
fastembed / ort / ndarray / tokenizers. All embeddings are produced
by the OpenRouter REST embeddings API and stored as a BLOB in
memory_embeddings(memory_id, embedding, source). Vector similarity is
computed in pure Rust at query time.
§Workload classification (G42/S3, BLOCK 1 — MANDATORY)
LLM embedding is I/O-bound: each call waits on a network round-trip to the OpenRouter REST API while the local CPU stays idle. Concurrency therefore uses tokio (async I/O concurrency) and NEVER rayon (reserved for CPU-bound work).
§Permit formula (G42/S3, BLOCO 2)
permits = clamp(--llm-parallelism, 1, 32)
.min(available_parallelism())
.min(available_ram_mb * 0.5 / LLM_WORKER_RSS_MB)LLM_WORKER_RSS_MB = 350 (crate::constants): the historical
per-worker RSS budget, retained as the RAM bound on the permit
formula.
Structs§
- Embed
Cache Stats - G56: stats snapshot returned by
embed_entity_texts_cached.
Enums§
- Embedding
Error Kind - GAP-004 (v1.0.88): typed classifier for embedding error messages.
- Fallback
Reason - G58/S1: reason an embedding call could not be completed and the caller must fall back to a non-vector retrieval path (FTS5 prefix + LIKE).
- LlmBackend
Kind - LLM backend kind for the fallback chain. Mirrors the CLI
--llm-backendenum so users can pass the same value to--llm-fallbackwithout translation.
Constants§
- CHUNK_
EMBED_ BATCH_ SIZE - Calibration base: chunk (long-text) batch size per LLM call at the
calibration dimensionality (G42/S2). Use
chunk_embed_batch_sizefor the dim-adaptive value (G44). - EMBED_
BATCH_ CALIBRATION_ DIM - Dimensionality the batch bases above were calibrated against (G44).
- ENTITY_
EMBED_ BATCH_ SIZE - Calibration base: entity-name (short-text) batch size per LLM call at
the calibration dimensionality (G42/S2). Use
entity_embed_batch_sizefor the dim-adaptive value (G44).
Functions§
- bytes_
to_ f32 - Bytes to f 32.
- chunk_
embed_ batch_ size - Dim-adaptive batch size for chunk (long-text) embedding calls (G44).
- classify_
embedding_ error - Classify an embedding
AppErrorinto a typedFallbackReason. - effective_
permits - G42/S3 BLOCO 2: effective permit count.
- embed_
entity_ texts_ cached - G56: embeds entity-name texts through a process-wide cache.
- embed_
passage_ or_ skip - v1.0.89 (BUG-SKIP-EMBED + GAP-EMBED-PROPAGATION): embed a passage
honouring both
--llm-backendand--skip-embedding-on-failure. - embed_
passage_ with_ choice - Embed a single passage using the LLM backend selected by the user via
--llm-backend. Routes toembed_with_fallbackso failures fall through to the next backend in the chain before giving up. - embed_
passage_ with_ embedding_ choice - v1.0.93: embedding with
EmbeddingBackendChoiceawareness. - embed_
passages_ parallel_ shared - Embeds many passages with
EmbeddingBackendChoiceawareness (GAP-SG-147). - embed_
via_ backend - Embeds a single text via the given backend. Used by
embed_with_fallbackand exposed to allow direct one-shot selection without a chain. Embeds a single text via the given backend. Used byembed_with_fallbackand exposed to allow direct one-shot selection without a chain. - embed_
via_ backend_ legacy - Legacy one-shot wrapper around
embed_via_backendthat discards the resolved backend. Kept for call sites that only care about the vector and ignore the executed-backend signal. New code should preferembed_via_backenddirectly. - embed_
via_ backend_ strict - Embed via backend strict.
- embed_
with_ fallback - Tries each LLM backend in
chainin order, returning the first successful embedding. On failure, the diagnostic tail of the last error is preserved in the returnedAppError::Embeddingso the operator can see WHY every backend failed. - embedding_
dim - Returns the dimensionality of the embedding space. Used to validate LLM responses and to size the in-memory cache.
- entity_
embed_ batch_ size - Dim-adaptive batch size for entity-name (short-text) embedding calls (G44).
- f32_
to_ bytes - F 32 to bytes.
- get_
openrouter_ chat_ client - v1.0.95 (ADR-0054): initialises the process-wide OpenRouter chat client on
first use and returns it.
modelis the text model the enrich JUDGE will call (no default; the caller validates presence upfront). - get_
openrouter_ embedder - Initialises the process-wide OpenRouter embedding client on first use and returns it.
- is_
openrouter_ initialized - Returns true when the process-wide OpenRouter embed client is ready.
- openrouter_
chat_ client - v1.0.95: returns the process-wide OpenRouter chat client if it has already
been initialised via
get_openrouter_chat_client. Used by the enrich JUDGE dispatch, which initialises the singleton once at startup and then fetches it per item without re-threading the API key. - should_
skip_ embedding_ on_ failure - v1.0.89 (BUG-SKIP-EMBED): reads
--skip-embedding-on-failure/ runtime_config (flag > XDG; product env is not read). Returnstruewhen the user opted to persist with NULL embedding on failure. - try_
embed_ query_ with_ choice - failure, returns a structured
FallbackReasonso the caller can surfacevec_degradedinstead of a hard exit 11. - try_
embed_ query_ with_ deterministic_ fallback - G58 / ADR-0043 (v1.0.85): deterministic fallback for
recallandhybrid-search. - try_
embed_ query_ with_ embedding_ choice - v1.0.93 (GAP-OR-INGEST): query embedding with
EmbeddingBackendChoiceawareness. Mirrorstry_embed_query_with_choicebut routes throughembed_passage_with_embedding_choiceso OpenRouter API is used when configured. - try_
embed_ query_ with_ fallback - G58/S1: try to embed a query, mapping any failure to a structured
FallbackReasonso callers can route to FTS5 + LIKE fallback instead of returning exit 11 to the user.