Skip to main content

LayerKvCache

Struct LayerKvCache 

Source
pub struct LayerKvCache {
    pub mode: KvMode,
    pub seq_len: usize,
    pub num_kv_heads: usize,
    pub head_dim: usize,
    pub linear_state: Vec<f32>,
    pub linear_scratch: Vec<f32>,
    pub o1: Option<O1State>,
    pub sinks: Option<Vec<f32>>,
    pub bounded: Option<BoundedState>,
    pub wire_kind: WireKind,
    pub wire_layer: u32,
    pub wire_identity: u64,
    /* private fields */
}
Expand description

KV cache for a single layer, head-major.

Fields§

§mode: KvMode§seq_len: usize

Positions appended so far (grows once per token, dead heads included).

§num_kv_heads: usize§head_dim: usize§linear_state: Vec<f32>

Linear-core recurrent state S (vmf_phase), f64; empty on full layers.

§linear_scratch: Vec<f32>

Tentative lane-2 state during speculative verify.

§o1: Option<O1State>

O(1) Nyström override (None = plain cache attention).

§sinks: Option<Vec<f32>>

Learned per-Q-head attention-sink logits of this layer (gpt-oss / MiMo-V2 self_attn.sinks, one f32 per Q head). The sink is an extra softmax column with no value: it joins the max and the denominator of every head’s softmax and so lets a head attend to “nothing”. These are WEIGHTS, not sequence state — clear() and the wire import keep them. None = an ordinary softmax.

§bounded: Option<BoundedState>

Natively bounded anchor (swa_sink_v1): a fixed-size ring installed from the header at load, never per prompt. A layer that carries it stores NOTHING per position (k/v stay empty).

§wire_kind: WireKind

Which state record this layer exchanges on the wire (v2).

§wire_layer: u32

This layer’s index in the stack (the wire header names it).

§wire_identity: u64

hash64 of the model’s operator identity (ModelArch::linear_core_identity JSON); 0 = no operator record.

Implementations§

Source§

impl LayerKvCache

Source

pub fn new(num_kv_heads: usize, head_dim: usize) -> Self

Source

pub fn install_bounded(&mut self, window: usize)

Give this layer its fixed-size ring ([kvh][window][hd] K and V). Called once at load from the header; the record is zeroed on clear() and never reallocated.

Source

pub fn bounded_step( &mut self, q: &[f32], k: &[f32], v: &[f32], w: &BoundedWeights, rope: &BoundedRope, scale: f32, num_heads: usize, out: &mut [f32], )

One position of the bounded operator: insert the raw k, v ([kvh][hd]) into slot t mod W, then attend every Q head of q ([nh][hd], raw) over sinks ∪ window into out ([nh][hd]). Nothing is appended per position.

Source

pub fn bounded_state_bytes(&self) -> usize

Bytes of the bounded ring (0 on other layers).

Source

pub fn bounded_snapshot(&self) -> Option<BoundedSnapshot>

Bit-for-bit copy of the ring for speculation (None on other layers).

Source

pub fn bounded_restore(&mut self, s: &BoundedSnapshot)

Restore a ring snapshot taken on this layer; seq_len follows the restored insert counter.

Source

pub fn set_linear_wire_allowed(&mut self, allowed: bool)

Bind this layer to the legacy cache-wire policy. Delta layers must refuse untagged state exchange rather than risk a plausible additive interpretation on the peer.

Source

pub fn discard_linear_scratch(&mut self)

Discard tentative recurrent state after a speculative rejection or any other path that abandons the lane-2 result.

Source

pub fn k_heads(&self) -> &[Vec<f32>]

Per-KV-head stored keys [seq_len × head_dim] (GPU token graph sync).

Source

pub fn v_heads(&self) -> &[Vec<f32>]

Per-KV-head stored values [seq_len × head_dim].

Source

pub fn o1_begin(&mut self, m: usize, w: usize, sink: usize, rect: O1Rect)

Arm query collection for a fresh prompt pass (a cleared cache).

Source

pub fn o1_push_q(&mut self, q_all: &[f32])

Record one position’s rotated queries ([num_heads × head_dim]) during the exact prompt pass. No-op unless collecting — the hook sits inside qwen_attention so every prefill flavor (sequential, batched) feeds the same trace.

Source

pub fn o1_sealed(&self) -> bool

Source

pub fn o1_seal(&mut self, num_heads: usize) -> bool

Freeze the prompt into per-KV-group Nyström states and drop this layer’s full KV. Returns false while a short collecting layer is below its deferred boundary; malformed prerequisites abort the layer instead of silently resuming exact KV growth. The seal needs f32 KV rows (CMF_KV=q8 stores int8), every group densely stored, a full q trace, and a GQA fan-out that actually divides.

Source

pub fn o1_views(&self) -> Option<Vec<O1DeviceView<'_>>>

One decode step on a sealed layer: per KV group, insert the group’s fresh (k, v) ONCE and read every Q head’s attention output. Returns [num_heads × head_dim]. Head h belongs to group h/hpk, so a group’s Q heads are contiguous in q_all/out — same math as the shared KV row the exact path appends once. Device views of the sealed o1 groups, or None when o1 is not sealed on this layer (or any group is in the degenerate exact-only mode the GPU path does not carry).

Source

pub fn o1_step( &mut self, q_all: &[f32], k_new: &[f32], v_new: &[f32], num_heads: usize, ) -> Vec<f32>

Source

pub fn o1_memory_bytes(&self) -> usize

Bytes held by the O(1) override (query trace while collecting, per-KV-group states once sealed).

Source

pub fn append(&mut self, k_new: &[f32], v_new: &[f32], alive: &[bool])

Append K/V for one position. k_new/v_new are [num_kv_heads × head_dim]; heads with alive[h] == false are skipped (their slices stay empty).

Source

pub fn attend(&self, q: &[f32], kv_head: usize) -> (Vec<f32>, Vec<f32>)

Per-head attention over its own storage: the f32 branch is bit-for-bit equal to attention_head() over slices; the q8 branch computes score = s_k·⟨q⊙col_k, k_q⟩ and the weighted sum of V in i8 with f32 accumulation. Returns (output [head_dim], probs [stored]).

Source

pub fn attend_group( &self, q_group: &[f32], kv_head: usize, out: &mut [f32], imp_acc: &mut [f32], scale: f32, first: usize, softcap: f32, sinks: &[f32], )

Grouped GQA attention: all Q-heads of one KV group in a single pass over the stored K rows and a single pass over the V rows (per-head attend re-read the shared group storage heads_per_kv times — roadmap §3 P1). Per-head score order, softmax and V accumulation are IDENTICAL to attend, so each head’s output is bit-for-bit the same.

q_group: [n_heads_in_group × head_dim] (global head order); out: same shape; imp_acc[0..stored] accumulates the probabilities of every head (attention importance), matching the caller’s former loop. scale is the score scale (1/√hd unless the arch overrides); first is the earliest visible position — sliding-window layers pass stored − window so older rows get zero probability. sinks holds one learned sink logit per head of q_group (the caller slices self.sinks to the group’s heads) or is empty for an ordinary softmax.

Source

pub fn attend_group_upto( &self, q_group: &[f32], kv_head: usize, out: &mut [f32], imp_acc: &mut [f32], scale: f32, first: usize, softcap: f32, upto: usize, sinks: &[f32], )

attend_group over the first upto stored rows only — what the same call saw when the cache held exactly upto rows. A prefill chunk appends all its rows first and then attends every position in parallel; position i passes upto = s0 + i + 1, which makes its result bit-identical to the sequential append-then-attend.

Only the visible rows [first, stored) are scored, so a sliding-window decode step costs O(window), not O(context). The rows before first get probability exactly 0 — what the former −inf-filled score row produced (exp(−inf) = 0 adds nothing to the max, the sum, V or the importance), so the result is bit-identical to scoring the whole row.

Learned sinks (sinks non-empty, one per head): the sink logit joins the softmax max and adds exp(sink − max) to the denominator; it has no value row. Identical to appending a value-less column, which is how gpt-oss and MiMo-V2 define it.

Source

pub fn truncate_last(&mut self, n_drop: usize)

Roll back the last n_drop positions (speculative-decode reject).

Source

pub fn accumulate_imp(&mut self, probs: &[f32])

Accumulate attention mass per stored position (summed over heads).

Source

pub fn head_keys(&self, kv_head: usize) -> &[f32]

Contiguous keys of one head: [stored_len × head_dim].

Source

pub fn head_values(&self, kv_head: usize) -> &[f32]

Source

pub fn head_len(&self, kv_head: usize) -> usize

Number of positions actually stored for a head (0 for dead heads).

Source

pub fn clear(&mut self)

Clear cache (e.g. on new conversation or task switch).

Source

pub fn export_wire(&self, f16: bool) -> Result<Vec<u8>, String>

Serialize this layer’s state for the wire (versioned, v2): a fixed header {magic "CMFS", version, operator identity hash64, layer, kind, f16 flag, position} followed by one record whose shape the kind fixes — per-position K/V (+ importance) for a full layer, the recurrent vector for a linear layer, the insert counter

  • ring K/V for a bounded anchor. f16 halves the K/V payloads and is the caller’s explicit choice, exactly like the hidden-state wire; recurrent vectors stay f32 whatever the wire dtype (they are the ONLY state a linear layer has — rounding them rounds the whole history).

REFUSES rather than travelling half-complete. A cache carrying frozen columns, a Nyström overlay or q8 storage holds state this format does not describe, and shipping the rest would land a plausible-looking cache that answers differently — the failure mode this whole format exists to avoid.

Source

pub fn import_wire(&mut self, buf: &[u8]) -> Result<(), String>

Install a peer’s state over this layer. The geometry must match the model both sides hold — it is checked, not assumed. Accepts the versioned wire (magic “CMFS”) and, for full/linear layers, the old unversioned per-position body.

Source

pub fn memory_bytes(&self) -> usize

Trait Implementations§

Source§

impl Clone for LayerKvCache

Source§

fn clone(&self) -> Self

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Debug for LayerKvCache

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self> ⓘ

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self> ⓘ

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, !>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self> ⓘ
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self> ⓘ

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more