pub enum LayerCompressor {
None,
Csa,
Hca,
}Expand description
Which compressor one layer runs, and everything that follows from it.
The ratio is not a tuning knob with a smooth range: it selects one of three different mechanisms, and the parameters below are not interpolations between them. Deriving them here, once, is what stops a single scalar ratio from building an indexer on an HCA layer or none on a CSA layer – both of which run, produce numbers, and are wrong.
Variants§
None
Ratio 0: no compressed tier at all. Attention sees the raw sliding window and nothing else.
A real entry in the shipped schedule rather than a disabled
state – (0, 0, 4, 128, 4, 128, 4, 0) opens with two of them
and closes with one – so it must be executable, not skipped.
Csa
Ratio 4: Compressed Sparse Attention. Overlapping blocks, a compressor projection twice as wide, and the Lightning Indexer restricting which compressed entries a query may see.
Hca
Ratio 128: Heavily Compressed Attention. Non-overlapping blocks, a single-width compressor, and dense visibility over every compressed entry – no indexer.
Implementations§
Source§impl LayerCompressor
impl LayerCompressor
Sourcepub fn from_ratio(ratio: u32) -> Option<Self>
pub fn from_ratio(ratio: u32) -> Option<Self>
Reads one layer’s ratio.
Returns None for a ratio that is not 0, 4 or 128, rather than
approximating it to the nearest mechanism: there is no nearest
mechanism, and picking one would give that layer the wrong
compressor width and the wrong indexer, silently.
Sourcepub fn has_indexer(self) -> bool
pub fn has_indexer(self) -> bool
Whether this layer instantiates the Lightning Indexer.
CSA only. On an HCA layer every compressed entry is visible, so there is nothing for a top-k selector to select; building one there costs its own compressor, its own keys and its own tier of device memory to answer a question with a fixed answer.
Sourcepub fn projection_width_multiple(self) -> usize
pub fn projection_width_multiple(self) -> usize
How many times wider this layer’s raw per-token compressor projection is than one head.
2 for CSA, and this is the detail a single scalar ratio gets
wrong on half the stack. Each raw token is projected twice:
once for its role as the tail of the block ending at it, and
once as the head of the next, overlapping block – two different
learned projections of the same token, not one reused twice
(llama.cpp load_arch_tensors’ coff = ratio == 4 ? 2 : 1, and
build_overlap_compressed_kv_from_state’s
GGML_ASSERT(kv_state->ne[0] == 2*n_embd_head)).
Sourcepub fn overlapping(self) -> bool
pub fn overlapping(self) -> bool
Whether consecutive compression blocks share raw positions.
Sourcepub fn visible_compressed(self, pos: usize) -> usize
pub fn visible_compressed(self, pos: usize) -> usize
How many compressed entries a query at position pos may see.
(pos + 1) / ratio: ratio-derived, like everything else here,
and zero on a layer with no compressor. The +1 is because
pos is an index and the count of tokens through it is one
more – without it the query at the last position of a block
cannot see the block it just completed, which is off by exactly
one entry for the whole of the sequence.
Trait Implementations§
Source§impl Clone for LayerCompressor
impl Clone for LayerCompressor
Source§fn clone(&self) -> LayerCompressor
fn clone(&self) -> LayerCompressor
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read moreimpl Copy for LayerCompressor
Source§impl Debug for LayerCompressor
impl Debug for LayerCompressor
impl Eq for LayerCompressor
Source§impl PartialEq for LayerCompressor
impl PartialEq for LayerCompressor
impl StructuralPartialEq for LayerCompressor
Auto Trait Implementations§
impl Freeze for LayerCompressor
impl RefUnwindSafe for LayerCompressor
impl Send for LayerCompressor
impl Sync for LayerCompressor
impl Unpin for LayerCompressor
impl UnsafeUnpin for LayerCompressor
impl UnwindSafe for LayerCompressor
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more