pub struct RecurLayer {
pub conv_state: CudaSlice<f32>,
pub ssm_state: CudaSlice<f32>,
pub ssm_state_alt: CudaSlice<f32>,
}Expand description
Per-linear-attn-layer fixed recurrent state. conv_state and ssm_state are BOTH kept RESIDENT on GPU — the conv ring assemble + roll runs on-device (conv_assemble_and_roll), so there is no per-step dtoh/htod for either.
Fields§
§conv_state: CudaSlice<f32>§ssm_state: CudaSlice<f32>§ssm_state_alt: CudaSlice<f32>PERSISTENT second SSM-state buffer for the gdn-scan double buffer (DECODE DETERMINISM FIX).
gdn_scan needs DISTINCT in/out state buffers. The old eager path allocated a fresh
state_scratch via e.uninit every step and swapped its pointer into ssm_state; that
per-step alloc/free churned the stream-ordered async pool, and the freed prior ssm_state
block was recycled by the next step’s scratch while a kernel referencing the swapped-in state
was still in flight — a use-after-reuse that produced RUN-TO-RUN nondeterministic decode
(two identical prompt primes diverged). We instead PING-PONG between two STABLE resident
buffers (no per-step alloc/free, no pool churn): step writes into the spare, then swaps the
two owned buffers in place. Stable pointers, identical math. Sized like ssm_state.