pub struct KvCache {
pub layers: Vec<LayerKvCache>,
pub max_seq_len: usize,
pub policy: EvictionPolicy,
}Expand description
Full KV cache for all layers.
Fields§
§layers: Vec<LayerKvCache>§max_seq_len: usize§policy: EvictionPolicyImplementations§
Source§impl KvCache
impl KvCache
pub fn new( num_layers: usize, num_kv_heads: usize, head_dim: usize, max_seq_len: usize, ) -> Self
pub fn clear(&mut self)
pub fn total_memory_bytes(&self) -> usize
Sourcepub fn recurrent_state_bytes(&self) -> usize
pub fn recurrent_state_bytes(&self) -> usize
Bytes owned by linear-core recurrent state (including the tentative speculative scratch). This is reported separately from attention KV so a serving slot’s O(1) capacity can be compared with its context cache without guessing from model geometry.
Sourcepub fn attention_state_bytes(&self) -> usize
pub fn attention_state_bytes(&self) -> usize
Attention KV (or sealed O(1) attention state) bytes, excluding the
linear recurrent vectors returned by [recurrent_state_bytes].
Sourcepub fn seq_len(&self) -> usize
pub fn seq_len(&self) -> usize
Current sequence length (max across layers — dead layers may lag):
the absolute depth, which a trimmed sliding layer keeps in
pos_len() while it stores only its tail.
Sourcepub fn bounded_state_bytes(&self) -> usize
pub fn bounded_state_bytes(&self) -> usize
Bytes owned by bounded-anchor rings (constant in context).
Sourcepub fn needs_eviction(&self) -> bool
pub fn needs_eviction(&self) -> bool
True when some layer with PER-POSITION storage reached the cap.
Bounded anchors hold nothing per position and never need it: a
model whose every layer is O(1) has no eviction cliff at all. A
sliding tail (trim_window) bounds itself and is not counted.