pub struct EmbeddingModel { /* private fields */ }Expand description
A loaded embedding model: tokenizer + encoder + the checkpoint’s own pooling rule.
Implementations§
Source§impl EmbeddingModel
impl EmbeddingModel
Sourcepub fn from_gguf_path(path: impl AsRef<Path>) -> Result<Self, EmbedError>
pub fn from_gguf_path(path: impl AsRef<Path>) -> Result<Self, EmbedError>
Opens path and builds whichever embedding stack its
general.architecture names, or refuses naming what is missing.
pub fn architecture(&self) -> &str
Sourcepub fn name(&self) -> &str
pub fn name(&self) -> &str
The checkpoint’s general.name, or its architecture when the
file carries none. What /v1/embeddings reports as model.
pub fn n_embd(&self) -> usize
pub fn n_ctx_train(&self) -> usize
pub fn pooling_type(&self) -> PoolingType
Sourcepub fn token_ids(&self, text: &str) -> Vec<u32>
pub fn token_ids(&self, text: &str) -> Vec<u32>
The exact ids the encoder will see for text: the tokenizer’s
pieces wrapped in the model’s own special tokens. Public because
/v1/embeddings has to report usage.prompt_tokens, and that
number is this length — llama.cpp counts the specials too.
Sourcepub fn decode_tokens(&self, ids: &[u32]) -> String
pub fn decode_tokens(&self, ids: &[u32]) -> String
Text for ids, through this checkpoint’s own vocabulary.
The counterpart to Self::token_ids, so /v1/detokenize
answers for an encoder rather than refusing. An embedding
model’s whole contract is the vector it returns for a string,
and when that vector is surprising the first question is what
tokens it actually saw. Without this the only way to ask was to
load the checkpoint in a second tool.
Not wrap_special’s inverse: it decodes exactly the ids given,
including specials if the caller passes them, because a caller
checking a tokenization wants to see what it sent.
Sourcepub fn embed(&self, text: &str, normalize: bool) -> Result<Vec<f32>, EmbedError>
pub fn embed(&self, text: &str, normalize: bool) -> Result<Vec<f32>, EmbedError>
Pooled embedding for text. normalize applies L2 normalization,
which is what an OpenAI-compatible /v1/embeddings response is
expected to carry and what llama.cpp’s server does by default;
the raw pooled vector is what the graph produced.
Un-pooled n_tokens × n_embd hidden states, for a caller that
wants to pool differently (or not at all).
Sourcepub fn rank_head(&self) -> Option<&RankHead>
pub fn rank_head(&self) -> Option<&RankHead>
The checkpoint’s reranker classification head, or None for a
plain embedding model. What /v1/rerank checks before it
promises a caller a relevance score.
Sourcepub fn rerank_input(
&self,
query: &str,
document: &str,
) -> Result<PairSequence, EmbedError>
pub fn rerank_input( &self, query: &str, document: &str, ) -> Result<PairSequence, EmbedError>
The exact input Self::rerank_score will see for one
(query, document) pair: [CLS] query [SEP] document [SEP],
with the segment id of every position.
Separate from the scoring call for the same reason
Self::token_ids is separate from Self::embed — a route
has to report usage.prompt_tokens, and that number is
tokens.len().
Sourcepub fn rerank_score(&self, pair: &PairSequence) -> Result<f32, EmbedError>
pub fn rerank_score(&self, pair: &PairSequence) -> Result<f32, EmbedError>
The head’s relevance score for a pair sequence built by
Self::rerank_input.
This is upstream’s RANK path in full: encode, take the CLS
row, run the classification head, report output 0
(send_rerank’s embd[0]). The CLS row is taken here regardless
of what {arch}.pooling_type says, because the head was trained
on that position — pooling_type = RANK is the checkpoint
declaring this path, not naming a pooling rule, which is why
crate::pooling::pool still refuses RANK and must keep
refusing it.
No L2 normalization and no sigmoid: upstream reports the raw
logit, so a score is comparable only against other scores from
the same head, and this must not quietly squash it into 0..1.
Un-pooled n_tokens × n_embd hidden states for a pair built by
Self::rerank_input — Self::hidden_states’s counterpart
for the cross-encoder input, and the one graph call
Self::rerank_score itself makes.
Public for the same reason Self::hidden_states is: when a
relevance score is surprising, the first questions are what
tokens the model saw and what came out before the head, and
without this the only way to ask was to load the checkpoint a
second time — which for a reranker does not even work, because
crate::load_bert_encoder_from_path alone leaves cls.*
unconsumed and refuses.
The pair’s own segments are honoured, so passing a
PairSequence whose segments are all zero reproduces the
segment-blind graph exactly, without a second copy of it to
drift.
Auto Trait Implementations§
impl !RefUnwindSafe for EmbeddingModel
impl !UnwindSafe for EmbeddingModel
impl Freeze for EmbeddingModel
impl Send for EmbeddingModel
impl Sync for EmbeddingModel
impl Unpin for EmbeddingModel
impl UnsafeUnpin for EmbeddingModel
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more