pub struct LayerKvCache {
pub mode: KvMode,
pub seq_len: usize,
pub num_kv_heads: usize,
pub head_dim: usize,
pub linear_state: Vec<f32>,
pub linear_scratch: Vec<f32>,
pub o1: Option<O1State>,
/* private fields */
}Expand description
KV cache for a single layer, head-major.
Fields§
§mode: KvMode§seq_len: usizePositions appended so far (grows once per token, dead heads included).
num_kv_heads: usize§head_dim: usize§linear_state: Vec<f32>Linear-core condensate S (vmf_phase), f64; empty on full layers.
linear_scratch: Vec<f32>Tentative lane-2 state during speculative verify.
o1: Option<O1State>O(1) Nyström override (None = plain cache attention).
Implementations§
Source§impl LayerKvCache
impl LayerKvCache
pub fn new(num_kv_heads: usize, head_dim: usize) -> Self
Sourcepub fn k_heads(&self) -> &[Vec<f32>]
pub fn k_heads(&self) -> &[Vec<f32>]
Per-KV-head stored keys [seq_len × head_dim] (GPU token graph sync).
Sourcepub fn o1_begin(&mut self, m: usize, w: usize, sink: usize, rect: O1Rect)
pub fn o1_begin(&mut self, m: usize, w: usize, sink: usize, rect: O1Rect)
Arm query collection for a fresh prompt pass (a cleared cache).
Sourcepub fn o1_push_q(&mut self, q_all: &[f32])
pub fn o1_push_q(&mut self, q_all: &[f32])
Record one position’s rotated queries ([num_heads × head_dim])
during the exact prompt pass. No-op unless collecting — the hook
sits inside qwen_attention so every prefill flavor (sequential,
batched) feeds the same trace.
pub fn o1_sealed(&self) -> bool
Sourcepub fn o1_seal(&mut self, num_heads: usize) -> bool
pub fn o1_seal(&mut self, num_heads: usize) -> bool
Freeze the prompt into per-KV-group Nyström states and drop this
layer’s full KV. Returns false (layer stays exact, KV kept) when
the preconditions fail: the seal needs f32 KV rows (CMF_KV=q8
stores int8), every group densely stored, a full q trace, and a
GQA fan-out that actually divides.
Sourcepub fn o1_views(&self) -> Option<Vec<O1DeviceView<'_>>>
pub fn o1_views(&self) -> Option<Vec<O1DeviceView<'_>>>
One decode step on a sealed layer: per KV group, insert the
group’s fresh (k, v) ONCE and read every Q head’s attention
output. Returns [num_heads × head_dim]. Head h belongs to group
h/hpk, so a group’s Q heads are contiguous in q_all/out —
same math as the shared KV row the exact path appends once.
Device views of the sealed o1 groups, or None when o1 is not
sealed on this layer (or any group is in the degenerate
exact-only mode the GPU path does not carry).
pub fn o1_step( &mut self, q_all: &[f32], k_new: &[f32], v_new: &[f32], num_heads: usize, ) -> Vec<f32>
Sourcepub fn o1_memory_bytes(&self) -> usize
pub fn o1_memory_bytes(&self) -> usize
Bytes held by the O(1) override (query trace while collecting, per-KV-group states once sealed).
Sourcepub fn append(&mut self, k_new: &[f32], v_new: &[f32], alive: &[bool])
pub fn append(&mut self, k_new: &[f32], v_new: &[f32], alive: &[bool])
Append K/V for one position. k_new/v_new are
[num_kv_heads × head_dim]; heads with alive[h] == false are
skipped (their slices stay empty).
Sourcepub fn attend(&self, q: &[f32], kv_head: usize) -> (Vec<f32>, Vec<f32>)
pub fn attend(&self, q: &[f32], kv_head: usize) -> (Vec<f32>, Vec<f32>)
Per-head attention over its own storage: the f32 branch is bit-for-bit equal to attention_head() over slices; the q8 branch computes score = s_k·⟨q⊙col_k, k_q⟩ and the weighted sum of V in i8 with f32 accumulation. Returns (output [head_dim], probs [stored]).
Sourcepub fn attend_group(
&self,
q_group: &[f32],
kv_head: usize,
out: &mut [f32],
imp_acc: &mut [f32],
scale: f32,
first: usize,
softcap: f32,
)
pub fn attend_group( &self, q_group: &[f32], kv_head: usize, out: &mut [f32], imp_acc: &mut [f32], scale: f32, first: usize, softcap: f32, )
Grouped GQA attention: all Q-heads of one KV group in a single
pass over the stored K rows and a single pass over the V rows
(per-head attend re-read the shared group storage
heads_per_kv times — roadmap §3 P1). Per-head score order,
softmax and V accumulation are IDENTICAL to attend, so each
head’s output is bit-for-bit the same.
q_group: [n_heads_in_group × head_dim] (global head order);
out: same shape; imp_acc[0..stored] accumulates the probs of
every head (Born importance), matching the caller’s former loop.
scale is the score scale (1/√hd unless the arch overrides);
first is the earliest visible position — sliding-window layers
pass stored − window so older rows get zero probability.
Sourcepub fn truncate_last(&mut self, n_drop: usize)
pub fn truncate_last(&mut self, n_drop: usize)
Roll back the last n_drop positions (speculative-decode reject).
Sourcepub fn accumulate_imp(&mut self, probs: &[f32])
pub fn accumulate_imp(&mut self, probs: &[f32])
Accumulate attention mass per stored position (summed over heads).
Sourcepub fn head_keys(&self, kv_head: usize) -> &[f32]
pub fn head_keys(&self, kv_head: usize) -> &[f32]
Contiguous keys of one head: [stored_len × head_dim].
pub fn head_values(&self, kv_head: usize) -> &[f32]
Sourcepub fn head_len(&self, kv_head: usize) -> usize
pub fn head_len(&self, kv_head: usize) -> usize
Number of positions actually stored for a head (0 for dead heads).
Sourcepub fn memory_bytes(&self) -> usize
pub fn memory_bytes(&self) -> usize
Memory usage in bytes.
Trait Implementations§
Source§impl Clone for LayerKvCache
impl Clone for LayerKvCache
Source§fn clone(&self) -> LayerKvCache
fn clone(&self) -> LayerKvCache
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read more