Skip to main content

Crate onnx_embedding_plugin

Crate onnx_embedding_plugin 

Source
Expand description

In-process ONNX embedding provider for the graph-storage gear.

ADR-0005 makes this the default: a MiniLM-class sentence-embedding model run through ONNX Runtime, in the gear’s own process, so a small deployment needs no inference service to use vector search at all.

§Artifacts are supplied, never fetched

The model and tokenizer are read from paths the operator configures. The crate downloads nothing. That is not caution about the network: the embedding-space identity has to be verifiable, and the only identity a downloader can offer is the name it asked for. Reading a file lets the identity be the SHA-256 of the bytes actually loaded, so two deployments claiming one space either agree on that hash or are visibly different.

§The runtime is loaded, not linked

ort is pinned with load-dynamic, so ONNX Runtime is resolved by dlopen at first use through ORT_DYLIB_PATH. Building this crate needs no runtime headers; running it needs the shared library. See OnnxEmbeddingProvider::load for what happens when that path is wrong, which is worse than an error.

Structs§

OnnxEmbeddingProvider
A MiniLM-class sentence-embedding model, in this process.
OnnxProviderConfig
What a deployment declares about its model.

Enums§

OnnxLoadError
Pooling
How the model turns a sequence of token vectors into one sentence vector.