pub struct LayerKvCache {
pub mode: KvMode,
pub seq_len: usize,
pub num_kv_heads: usize,
pub head_dim: usize,
pub linear_state: Vec<f32>,
pub linear_scratch: Vec<f32>,
pub o1: Option<O1State>,
/* private fields */
}Expand description
KV cache for a single layer, head-major.
Fields§
§mode: KvMode§seq_len: usizePositions appended so far (grows once per token, dead heads included).
num_kv_heads: usize§head_dim: usize§linear_state: Vec<f32>Linear-core condensate S (vmf_phase), f64; empty on full layers.
linear_scratch: Vec<f32>Tentative lane-2 state during speculative verify.
o1: Option<O1State>O(1) Nyström override (None = plain cache attention).
Implementations§
Source§impl LayerKvCache
impl LayerKvCache
pub fn new(num_kv_heads: usize, head_dim: usize) -> Self
Sourcepub fn o1_begin(&mut self, m: usize, w: usize, sink: usize, rect: O1Rect)
pub fn o1_begin(&mut self, m: usize, w: usize, sink: usize, rect: O1Rect)
Arm query collection for a fresh prompt pass (a cleared cache).
Sourcepub fn o1_push_q(&mut self, q_all: &[f32])
pub fn o1_push_q(&mut self, q_all: &[f32])
Record one position’s rotated queries ([num_heads × head_dim])
during the exact prompt pass. No-op unless collecting — the hook
sits inside qwen_attention so every prefill flavor (sequential,
batched) feeds the same trace.
pub fn o1_sealed(&self) -> bool
Sourcepub fn o1_seal(&mut self, num_heads: usize) -> bool
pub fn o1_seal(&mut self, num_heads: usize) -> bool
Freeze the prompt into per-KV-group Nyström states and drop this
layer’s full KV. Returns false (layer stays exact, KV kept) when
the preconditions fail: the seal needs f32 KV rows (CMF_KV=q8
stores int8), every group densely stored, a full q trace, and a
GQA fan-out that actually divides.
Sourcepub fn o1_step(
&mut self,
q_all: &[f32],
k_new: &[f32],
v_new: &[f32],
num_heads: usize,
) -> Vec<f32>
pub fn o1_step( &mut self, q_all: &[f32], k_new: &[f32], v_new: &[f32], num_heads: usize, ) -> Vec<f32>
One decode step on a sealed layer: per KV group, insert the
group’s fresh (k, v) ONCE and read every Q head’s attention
output. Returns [num_heads × head_dim]. Head h belongs to group
h/hpk, so a group’s Q heads are contiguous in q_all/out —
same math as the shared KV row the exact path appends once.
Sourcepub fn o1_memory_bytes(&self) -> usize
pub fn o1_memory_bytes(&self) -> usize
Bytes held by the O(1) override (query trace while collecting, per-KV-group states once sealed).
Sourcepub fn append(&mut self, k_new: &[f32], v_new: &[f32], alive: &[bool])
pub fn append(&mut self, k_new: &[f32], v_new: &[f32], alive: &[bool])
Append K/V for one position. k_new/v_new are
[num_kv_heads × head_dim]; heads with alive[h] == false are
skipped (their slices stay empty).
Sourcepub fn attend(&self, q: &[f32], kv_head: usize) -> (Vec<f32>, Vec<f32>)
pub fn attend(&self, q: &[f32], kv_head: usize) -> (Vec<f32>, Vec<f32>)
Per-head attention over its own storage: the f32 branch is bit-for-bit equal to attention_head() over slices; the q8 branch computes score = s_k·⟨q⊙col_k, k_q⟩ and the weighted sum of V in i8 with f32 accumulation. Returns (output [head_dim], probs [stored]).
Sourcepub fn attend_group(
&self,
q_group: &[f32],
kv_head: usize,
out: &mut [f32],
imp_acc: &mut [f32],
scale: f32,
first: usize,
)
pub fn attend_group( &self, q_group: &[f32], kv_head: usize, out: &mut [f32], imp_acc: &mut [f32], scale: f32, first: usize, )
Grouped GQA attention: all Q-heads of one KV group in a single
pass over the stored K rows and a single pass over the V rows
(per-head attend re-read the shared group storage
heads_per_kv times — roadmap §3 P1). Per-head score order,
softmax and V accumulation are IDENTICAL to attend, so each
head’s output is bit-for-bit the same.
q_group: [n_heads_in_group × head_dim] (global head order);
out: same shape; imp_acc[0..stored] accumulates the probs of
every head (Born importance), matching the caller’s former loop.
scale is the score scale (1/√hd unless the arch overrides);
first is the earliest visible position — sliding-window layers
pass stored − window so older rows get zero probability.
Sourcepub fn truncate_last(&mut self, n_drop: usize)
pub fn truncate_last(&mut self, n_drop: usize)
Roll back the last n_drop positions (speculative-decode reject).
Sourcepub fn accumulate_imp(&mut self, probs: &[f32])
pub fn accumulate_imp(&mut self, probs: &[f32])
Accumulate attention mass per stored position (summed over heads).
Sourcepub fn head_keys(&self, kv_head: usize) -> &[f32]
pub fn head_keys(&self, kv_head: usize) -> &[f32]
Contiguous keys of one head: [stored_len × head_dim].
pub fn head_values(&self, kv_head: usize) -> &[f32]
Sourcepub fn head_len(&self, kv_head: usize) -> usize
pub fn head_len(&self, kv_head: usize) -> usize
Number of positions actually stored for a head (0 for dead heads).
Sourcepub fn memory_bytes(&self) -> usize
pub fn memory_bytes(&self) -> usize
Memory usage in bytes.
Trait Implementations§
Source§impl Clone for LayerKvCache
impl Clone for LayerKvCache
Source§fn clone(&self) -> LayerKvCache
fn clone(&self) -> LayerKvCache
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read more