pub struct KvCache {
pub layers: Vec<LayerKvCache>,
pub max_seq_len: usize,
pub policy: EvictionPolicy,
}Expand description
Full KV cache for all layers.
Fields§
§layers: Vec<LayerKvCache>§max_seq_len: usize§policy: EvictionPolicyImplementations§
Source§impl KvCache
impl KvCache
pub fn new( num_layers: usize, num_kv_heads: usize, head_dim: usize, max_seq_len: usize, ) -> Self
pub fn clear(&mut self)
pub fn total_memory_bytes(&self) -> usize
Sourcepub fn recurrent_state_bytes(&self) -> usize
pub fn recurrent_state_bytes(&self) -> usize
Bytes owned by linear-core recurrent state (including the tentative speculative scratch). This is reported separately from attention KV so a serving slot’s O(1) capacity can be compared with its context cache without guessing from model geometry.
Sourcepub fn attention_state_bytes(&self) -> usize
pub fn attention_state_bytes(&self) -> usize
Attention KV (or sealed O(1) attention state) bytes, excluding the
linear recurrent vectors returned by [recurrent_state_bytes].
Sourcepub fn seq_len(&self) -> usize
pub fn seq_len(&self) -> usize
Current sequence length (max across layers — dead layers may lag).
Sourcepub fn bounded_state_bytes(&self) -> usize
pub fn bounded_state_bytes(&self) -> usize
Bytes owned by bounded-anchor rings (constant in context).
Sourcepub fn needs_eviction(&self) -> bool
pub fn needs_eviction(&self) -> bool
True when some layer with PER-POSITION storage reached the cap. Bounded anchors hold nothing per position and never need it: a model whose every layer is O(1) has no eviction cliff at all.
Trait Implementations§
Auto Trait Implementations§
impl Freeze for KvCache
impl RefUnwindSafe for KvCache
impl Send for KvCache
impl Sync for KvCache
impl Unpin for KvCache
impl UnsafeUnpin for KvCache
impl UnwindSafe for KvCache
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
Mutably borrows from an owned value. Read more