Skip to main content

Module pooling

Module pooling 

Source
Expand description

{arch}.pooling_type: how a sequence of hidden states becomes one embedding vector.

This is llama.cpp’s enum llama_pooling_type and the switch in llm_graph_context::build_pooling (src/llama-graph.cpp), transcribed. The key is written by the HF converter (gguf_writer.add_pooling_type) as a uint32 holding that enum’s value, so the wire numbers below are load-bearing and are not ferrox’s own invention.

Its own module rather than a section of crate::bert_encoder because pooling is not a BERT fact: llama-embed and gemma-embedding are decoders that carry the same key, and /v1/embeddings needs to honour it for whatever produced the hidden states.

§What is implemented, and what refuses

NONE, MEAN, CLS and LAST are here. RANK is not: it is not a pooling rule at all but a classification head — upstream runs cls/cls_out matrices, a tanh, and an optional head norm over the pooled row, and reports the result through /v1/rerank against classifier.output_labels. None of that exists in ferrox yet, so PoolingType::Rank parses (so the refusal can name it) and pool returns PoolingError::Unimplemented rather than quietly handing back a CLS row that means something else.

Enums§

PoolingError
PoolingType
enum llama_pooling_type, by its wire values.

Functions§

l2_normalize
L2-normalizes in place, which is what every BGE/E5/GTE consumer expects of an embedding and what common_embd_normalize’s default (p == 2) does. A zero vector is left alone, exactly as upstream leaves it (it divides by norm > 0 ? 1/norm : 0, i.e. it zeroes — and a zero vector is already zero).
pool
Pools hidden (n_tokens rows of n_embd floats, in row order) down to one vector — except for PoolingType::None, which returns every row unchanged.