Skip to main content

Module bounded

Module bounded 

Source
Expand description

Natively bounded softmax anchor swa_sink_v1 (Embryo-O1).

Contract: docs/EMBRYO_BOUNDED_ANCHOR.md. The operator is a TRAINED bounded attention, not a masked full attention: per token t

window keys:  j ∈ (t − W, t]                (W keys INCLUDING the token)
window score: s_j = ( R(t − j) q̂_t ) · k̂_j / √hd
sink score:   s_s = q̂_t · k̂ˢ_s / √hd,   s ∈ [0, S)      (NoPE)
one softmax over {s_s} ∪ {s_j};   out_t = Σ_s p_s v̂ˢ_s + Σ_j p_j v_j

The ring stores RAW (unrotated) keys; the query is rotated by the distance Δ = t − j ∈ [0, W) through a [W][rd/2] cos/sin table. With absolute RoPE q_rot(t) = R(t) q̂, k_rot(j) = R(j) k̂ one has q_rot(t)·k_rot(j) = q̂ᵀ R(t)ᵀ R(j) k̂ = (R(t−j) q̂)·k̂ (every 2-D block of R is a plane rotation, so R(t)ᵀ R(j) = R(j − t)), which is why the trainer may keep rotating by absolute positions and only mask, while the served operator never sees an absolute position at all.

State per layer: ring_k, ring_v [kvh][W][hd] + the insert counter — a record of fixed size derived from the header, identical in prefill, decode and across turns. Nothing here grows with the context.

Structs§

BoundedAttnCfg
Per-token configuration of a bounded layer (geometry + the shared rotation table). No position: the operator has none.
BoundedRope
Relative-rotation table: cos/sin[Δ][i] = cos/sin(Δ · inv_freq[i]) for Δ ∈ [0, W), built once per model from the layer’s inv_freq (same convention as attention::rope_rotate_scaled, angles computed in f64 so the table carries only the final f32 rounding).
BoundedSnapshot
Bit-for-bit copy of everything the operator mutates (speculation).
BoundedState
Per-layer bounded state: the ring of the last W raw keys/values per KV head plus the insert counter. len = min(seen, W), next write slot head = seen mod W, and the key of position j lives in slot j mod W — so no absolute position is ever stored.
BoundedWeights
Weights of one bounded-anchor layer.

Constants§

UNDO_DEPTH
Rows of rollback history kept beside the ring so a speculative reject (truncate_last) can restore the slots the rejected tokens overwrote. Bounded by construction; not part of the wire state.

Functions§

bounded_attention
One position: q̂ k̂ v = W x, insert, attend, W_o. The cache’s bounded record must exist (installed from the header at load).
bounded_attention_batch
A prefill chunk of b positions: the projections run as chunk GEMMs (each weight row streams once per chunk), the operator runs per position over ring + chunk — the scores of a position are against at most S + W keys whatever the chunk or the context, and the per-position arithmetic is the same code as bounded_attention.