pub enum StepTpGateShards<'a> {
F32(&'a [CudaSlice<f32>]),
Bf16(&'a [CudaSlice<u8>]),
}Expand description
Persistent workspace of the v2 rank-local decode-attention driver.
Buffers live in their producing rank’s CUDA context, are never freed, and events are
re-recorded per call — the pp.rs BoundarySlot discipline — so the per-token path has no
cuMemAlloc, no cross-stream free, and no host round-trip. Every buffer is fully overwritten
before its consumers run in the same call; nothing carries state between tokens.
Per-rank attn_gate row shards for the fused QKV+gate kernel, in the weight class the
fused kernels read (F32 mirror or raw checkpoint bf16).
Variants§
Auto Trait Implementations§
impl<'a> Freeze for StepTpGateShards<'a>
impl<'a> RefUnwindSafe for StepTpGateShards<'a>
impl<'a> Send for StepTpGateShards<'a>
impl<'a> Sync for StepTpGateShards<'a>
impl<'a> Unpin for StepTpGateShards<'a>
impl<'a> UnsafeUnpin for StepTpGateShards<'a>
impl<'a> UnwindSafe for StepTpGateShards<'a>
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
Mutably borrows from an owned value. Read more