Skip to main content

LayerKvCache

Struct LayerKvCache 

Source
pub struct LayerKvCache {
    pub mode: KvMode,
    pub seq_len: usize,
    pub num_kv_heads: usize,
    pub head_dim: usize,
    pub linear_state: Vec<f32>,
    pub linear_scratch: Vec<f32>,
    pub o1: Option<O1State>,
    /* private fields */
}
Expand description

KV cache for a single layer, head-major.

Fields§

§mode: KvMode§seq_len: usize

Positions appended so far (grows once per token, dead heads included).

§num_kv_heads: usize§head_dim: usize§linear_state: Vec<f32>

Linear-core condensate S (vmf_phase), f64; empty on full layers.

§linear_scratch: Vec<f32>

Tentative lane-2 state during speculative verify.

§o1: Option<O1State>

O(1) Nyström override (None = plain cache attention).

Implementations§

Source§

impl LayerKvCache

Source

pub fn new(num_kv_heads: usize, head_dim: usize) -> Self

Source

pub fn k_heads(&self) -> &[Vec<f32>]

Per-KV-head stored keys [seq_len × head_dim] (GPU token graph sync).

Source

pub fn v_heads(&self) -> &[Vec<f32>]

Per-KV-head stored values [seq_len × head_dim].

Source

pub fn o1_begin(&mut self, m: usize, w: usize, sink: usize, rect: O1Rect)

Arm query collection for a fresh prompt pass (a cleared cache).

Source

pub fn o1_push_q(&mut self, q_all: &[f32])

Record one position’s rotated queries ([num_heads × head_dim]) during the exact prompt pass. No-op unless collecting — the hook sits inside qwen_attention so every prefill flavor (sequential, batched) feeds the same trace.

Source

pub fn o1_sealed(&self) -> bool

Source

pub fn o1_seal(&mut self, num_heads: usize) -> bool

Freeze the prompt into per-KV-group Nyström states and drop this layer’s full KV. Returns false (layer stays exact, KV kept) when the preconditions fail: the seal needs f32 KV rows (CMF_KV=q8 stores int8), every group densely stored, a full q trace, and a GQA fan-out that actually divides.

Source

pub fn o1_views(&self) -> Option<Vec<O1DeviceView<'_>>>

One decode step on a sealed layer: per KV group, insert the group’s fresh (k, v) ONCE and read every Q head’s attention output. Returns [num_heads × head_dim]. Head h belongs to group h/hpk, so a group’s Q heads are contiguous in q_all/out — same math as the shared KV row the exact path appends once. Device views of the sealed o1 groups, or None when o1 is not sealed on this layer (or any group is in the degenerate exact-only mode the GPU path does not carry).

Source

pub fn o1_step( &mut self, q_all: &[f32], k_new: &[f32], v_new: &[f32], num_heads: usize, ) -> Vec<f32>

Source

pub fn o1_memory_bytes(&self) -> usize

Bytes held by the O(1) override (query trace while collecting, per-KV-group states once sealed).

Source

pub fn append(&mut self, k_new: &[f32], v_new: &[f32], alive: &[bool])

Append K/V for one position. k_new/v_new are [num_kv_heads × head_dim]; heads with alive[h] == false are skipped (their slices stay empty).

Source

pub fn attend(&self, q: &[f32], kv_head: usize) -> (Vec<f32>, Vec<f32>)

Per-head attention over its own storage: the f32 branch is bit-for-bit equal to attention_head() over slices; the q8 branch computes score = s_k·⟨q⊙col_k, k_q⟩ and the weighted sum of V in i8 with f32 accumulation. Returns (output [head_dim], probs [stored]).

Source

pub fn attend_group( &self, q_group: &[f32], kv_head: usize, out: &mut [f32], imp_acc: &mut [f32], scale: f32, first: usize, softcap: f32, )

Grouped GQA attention: all Q-heads of one KV group in a single pass over the stored K rows and a single pass over the V rows (per-head attend re-read the shared group storage heads_per_kv times — roadmap §3 P1). Per-head score order, softmax and V accumulation are IDENTICAL to attend, so each head’s output is bit-for-bit the same.

q_group: [n_heads_in_group × head_dim] (global head order); out: same shape; imp_acc[0..stored] accumulates the probs of every head (Born importance), matching the caller’s former loop. scale is the score scale (1/√hd unless the arch overrides); first is the earliest visible position — sliding-window layers pass stored − window so older rows get zero probability.

Source

pub fn truncate_last(&mut self, n_drop: usize)

Roll back the last n_drop positions (speculative-decode reject).

Source

pub fn accumulate_imp(&mut self, probs: &[f32])

Accumulate attention mass per stored position (summed over heads).

Source

pub fn head_keys(&self, kv_head: usize) -> &[f32]

Contiguous keys of one head: [stored_len × head_dim].

Source

pub fn head_values(&self, kv_head: usize) -> &[f32]

Source

pub fn head_len(&self, kv_head: usize) -> usize

Number of positions actually stored for a head (0 for dead heads).

Source

pub fn clear(&mut self)

Clear cache (e.g. on new conversation or task switch).

Source

pub fn export_wire(&self, f16: bool) -> Result<Vec<u8>, String>

Memory usage in bytes. Serialize this layer’s state for the wire: f16 halves it and is the caller’s explicit choice, exactly like the hidden-state wire.

REFUSES rather than travelling half-complete. A cache carrying frozen columns, accumulated importance, a Nyström overlay or q8 storage holds state this format does not describe, and shipping the rest would land a plausible-looking cache that answers differently — the failure mode this whole format exists to avoid.

Source

pub fn import_wire(&mut self, buf: &[u8]) -> Result<(), String>

Install a peer’s state over this layer. The geometry must match the model both sides hold — it is checked, not assumed.

Source

pub fn memory_bytes(&self) -> usize

Trait Implementations§

Source§

impl Clone for LayerKvCache

Source§

fn clone(&self) -> LayerKvCache

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Debug for LayerKvCache

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self>

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self>

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = Infallible

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, <T as TryFrom<U>>::Error>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self>
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self>

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more