pub struct LayerKvCache {
pub mode: KvMode,
pub seq_len: usize,
pub num_kv_heads: usize,
pub head_dim: usize,
pub linear_state: Vec<f32>,
pub linear_scratch: Vec<f32>,
pub o1: Option<O1State>,
pub sinks: Option<Vec<f32>>,
pub bounded: Option<BoundedState>,
pub wire_kind: WireKind,
pub wire_layer: u32,
pub wire_identity: u64,
/* private fields */
}Expand description
KV cache for a single layer, head-major.
Fields§
§mode: KvMode§seq_len: usizePositions appended so far (grows once per token, dead heads included).
num_kv_heads: usize§head_dim: usize§linear_state: Vec<f32>Linear-core recurrent state S (vmf_phase), f64; empty on full layers.
linear_scratch: Vec<f32>Tentative lane-2 state during speculative verify.
o1: Option<O1State>O(1) Nyström override (None = plain cache attention).
sinks: Option<Vec<f32>>Learned per-Q-head attention-sink logits of this layer (gpt-oss /
MiMo-V2 self_attn.sinks, one f32 per Q head). The sink is an
extra softmax column with no value: it joins the max and the
denominator of every head’s softmax and so lets a head attend to
“nothing”. These are WEIGHTS, not sequence state — clear() and
the wire import keep them. None = an ordinary softmax.
bounded: Option<BoundedState>Natively bounded anchor (swa_sink_v1): a fixed-size ring
installed from the header at load, never per prompt. A layer that
carries it stores NOTHING per position (k/v stay empty).
wire_kind: WireKindWhich state record this layer exchanges on the wire (v2).
wire_layer: u32This layer’s index in the stack (the wire header names it).
wire_identity: u64hash64 of the model’s operator identity
(ModelArch::linear_core_identity JSON); 0 = no operator record.
Implementations§
Source§impl LayerKvCache
impl LayerKvCache
pub fn new(num_kv_heads: usize, head_dim: usize) -> Self
Sourcepub fn install_bounded(&mut self, window: usize)
pub fn install_bounded(&mut self, window: usize)
Give this layer its fixed-size ring ([kvh][window][hd] K and V).
Called once at load from the header; the record is zeroed on
clear() and never reallocated.
Sourcepub fn bounded_step(
&mut self,
q: &[f32],
k: &[f32],
v: &[f32],
w: &BoundedWeights,
rope: &BoundedRope,
scale: f32,
num_heads: usize,
out: &mut [f32],
)
pub fn bounded_step( &mut self, q: &[f32], k: &[f32], v: &[f32], w: &BoundedWeights, rope: &BoundedRope, scale: f32, num_heads: usize, out: &mut [f32], )
One position of the bounded operator: insert the raw k, v
([kvh][hd]) into slot t mod W, then attend every Q head of
q ([nh][hd], raw) over sinks ∪ window into out ([nh][hd]).
Nothing is appended per position.
Sourcepub fn bounded_state_bytes(&self) -> usize
pub fn bounded_state_bytes(&self) -> usize
Bytes of the bounded ring (0 on other layers).
Sourcepub fn bounded_snapshot(&self) -> Option<BoundedSnapshot>
pub fn bounded_snapshot(&self) -> Option<BoundedSnapshot>
Bit-for-bit copy of the ring for speculation (None on other layers).
Sourcepub fn bounded_restore(&mut self, s: &BoundedSnapshot)
pub fn bounded_restore(&mut self, s: &BoundedSnapshot)
Restore a ring snapshot taken on this layer; seq_len follows the
restored insert counter.
Sourcepub fn set_linear_wire_allowed(&mut self, allowed: bool)
pub fn set_linear_wire_allowed(&mut self, allowed: bool)
Bind this layer to the legacy cache-wire policy. Delta layers must refuse untagged state exchange rather than risk a plausible additive interpretation on the peer.
Sourcepub fn discard_linear_scratch(&mut self)
pub fn discard_linear_scratch(&mut self)
Discard tentative recurrent state after a speculative rejection or any other path that abandons the lane-2 result.
Sourcepub fn k_heads(&self) -> &[Vec<f32>]
pub fn k_heads(&self) -> &[Vec<f32>]
Per-KV-head stored keys [seq_len × head_dim] (GPU token graph sync).
Sourcepub fn o1_begin(&mut self, m: usize, w: usize, sink: usize, rect: O1Rect)
pub fn o1_begin(&mut self, m: usize, w: usize, sink: usize, rect: O1Rect)
Arm query collection for a fresh prompt pass (a cleared cache).
Sourcepub fn o1_push_q(&mut self, q_all: &[f32])
pub fn o1_push_q(&mut self, q_all: &[f32])
Record one position’s rotated queries ([num_heads × head_dim])
during the exact prompt pass. No-op unless collecting — the hook
sits inside qwen_attention so every prefill flavor (sequential,
batched) feeds the same trace.
pub fn o1_sealed(&self) -> bool
Sourcepub fn o1_seal(&mut self, num_heads: usize) -> bool
pub fn o1_seal(&mut self, num_heads: usize) -> bool
Freeze the prompt into per-KV-group Nyström states and drop this
layer’s full KV. Returns false while a short collecting layer is
below its deferred boundary; malformed prerequisites abort the
layer instead of silently resuming exact KV growth. The seal needs
f32 KV rows (CMF_KV=q8 stores int8), every group densely stored, a
full q trace, and a GQA fan-out that actually divides.
Sourcepub fn o1_views(&self) -> Option<Vec<O1DeviceView<'_>>>
pub fn o1_views(&self) -> Option<Vec<O1DeviceView<'_>>>
One decode step on a sealed layer: per KV group, insert the
group’s fresh (k, v) ONCE and read every Q head’s attention
output. Returns [num_heads × head_dim]. Head h belongs to group
h/hpk, so a group’s Q heads are contiguous in q_all/out —
same math as the shared KV row the exact path appends once.
Device views of the sealed o1 groups, or None when o1 is not
sealed on this layer (or any group is in the degenerate
exact-only mode the GPU path does not carry).
pub fn o1_step( &mut self, q_all: &[f32], k_new: &[f32], v_new: &[f32], num_heads: usize, ) -> Vec<f32>
Sourcepub fn o1_memory_bytes(&self) -> usize
pub fn o1_memory_bytes(&self) -> usize
Bytes held by the O(1) override (query trace while collecting, per-KV-group states once sealed).
Sourcepub fn append(&mut self, k_new: &[f32], v_new: &[f32], alive: &[bool])
pub fn append(&mut self, k_new: &[f32], v_new: &[f32], alive: &[bool])
Append K/V for one position. k_new/v_new are
[num_kv_heads × head_dim]; heads with alive[h] == false are
skipped (their slices stay empty).
Sourcepub fn attend(&self, q: &[f32], kv_head: usize) -> (Vec<f32>, Vec<f32>)
pub fn attend(&self, q: &[f32], kv_head: usize) -> (Vec<f32>, Vec<f32>)
Per-head attention over its own storage: the f32 branch is bit-for-bit equal to attention_head() over slices; the q8 branch computes score = s_k·⟨q⊙col_k, k_q⟩ and the weighted sum of V in i8 with f32 accumulation. Returns (output [head_dim], probs [stored]).
Sourcepub fn attend_group(
&self,
q_group: &[f32],
kv_head: usize,
out: &mut [f32],
imp_acc: &mut [f32],
scale: f32,
first: usize,
softcap: f32,
sinks: &[f32],
)
pub fn attend_group( &self, q_group: &[f32], kv_head: usize, out: &mut [f32], imp_acc: &mut [f32], scale: f32, first: usize, softcap: f32, sinks: &[f32], )
Grouped GQA attention: all Q-heads of one KV group in a single
pass over the stored K rows and a single pass over the V rows
(per-head attend re-read the shared group storage
heads_per_kv times — roadmap §3 P1). Per-head score order,
softmax and V accumulation are IDENTICAL to attend, so each
head’s output is bit-for-bit the same.
q_group: [n_heads_in_group × head_dim] (global head order);
out: same shape; imp_acc[0..stored] accumulates the probabilities
of every head (attention importance), matching the caller’s former loop.
scale is the score scale (1/√hd unless the arch overrides);
first is the earliest visible position — sliding-window layers
pass stored − window so older rows get zero probability.
sinks holds one learned sink logit per head of q_group (the
caller slices self.sinks to the group’s heads) or is empty for
an ordinary softmax.
Sourcepub fn attend_group_upto(
&self,
q_group: &[f32],
kv_head: usize,
out: &mut [f32],
imp_acc: &mut [f32],
scale: f32,
first: usize,
softcap: f32,
upto: usize,
sinks: &[f32],
)
pub fn attend_group_upto( &self, q_group: &[f32], kv_head: usize, out: &mut [f32], imp_acc: &mut [f32], scale: f32, first: usize, softcap: f32, upto: usize, sinks: &[f32], )
attend_group over the first upto stored rows only — what the
same call saw when the cache held exactly upto rows. A prefill
chunk appends all its rows first and then attends every position
in parallel; position i passes upto = s0 + i + 1, which makes
its result bit-identical to the sequential append-then-attend.
Only the visible rows [first, stored) are scored, so a
sliding-window decode step costs O(window), not O(context). The
rows before first get probability exactly 0 — what the former
−inf-filled score row produced (exp(−inf) = 0 adds nothing to the
max, the sum, V or the importance), so the result is bit-identical
to scoring the whole row.
Learned sinks (sinks non-empty, one per head): the sink logit
joins the softmax max and adds exp(sink − max) to the denominator;
it has no value row. Identical to appending a value-less column,
which is how gpt-oss and MiMo-V2 define it.
Sourcepub fn truncate_last(&mut self, n_drop: usize)
pub fn truncate_last(&mut self, n_drop: usize)
Roll back the last n_drop positions (speculative-decode reject).
Sourcepub fn accumulate_imp(&mut self, probs: &[f32])
pub fn accumulate_imp(&mut self, probs: &[f32])
Accumulate attention mass per stored position (summed over heads).
Sourcepub fn head_keys(&self, kv_head: usize) -> &[f32]
pub fn head_keys(&self, kv_head: usize) -> &[f32]
Contiguous keys of one head: [stored_len × head_dim].
pub fn head_values(&self, kv_head: usize) -> &[f32]
Sourcepub fn head_len(&self, kv_head: usize) -> usize
pub fn head_len(&self, kv_head: usize) -> usize
Number of positions actually stored for a head (0 for dead heads).
Sourcepub fn export_wire(&self, f16: bool) -> Result<Vec<u8>, String>
pub fn export_wire(&self, f16: bool) -> Result<Vec<u8>, String>
Serialize this layer’s state for the wire (versioned, v2): a
fixed header {magic "CMFS", version, operator identity hash64, layer, kind, f16 flag, position} followed by one record whose
shape the kind fixes — per-position K/V (+ importance) for a full
layer, the recurrent vector for a linear layer, the insert counter
- ring K/V for a bounded anchor.
f16halves the K/V payloads and is the caller’s explicit choice, exactly like the hidden-state wire; recurrent vectors stay f32 whatever the wire dtype (they are the ONLY state a linear layer has — rounding them rounds the whole history).
REFUSES rather than travelling half-complete. A cache carrying frozen columns, a Nyström overlay or q8 storage holds state this format does not describe, and shipping the rest would land a plausible-looking cache that answers differently — the failure mode this whole format exists to avoid.
Sourcepub fn import_wire(&mut self, buf: &[u8]) -> Result<(), String>
pub fn import_wire(&mut self, buf: &[u8]) -> Result<(), String>
Install a peer’s state over this layer. The geometry must match the model both sides hold — it is checked, not assumed. Accepts the versioned wire (magic “CMFS”) and, for full/linear layers, the old unversioned per-position body.