Skip to main content

PagedKVCache

Struct PagedKVCache 

Source
pub struct PagedKVCache<B: Backend> { /* private fields */ }

Implementations§

Source§

impl<B: Backend> PagedKVCache<B>

Source

pub fn new(num_layers: usize, config: CacheConfig) -> Self

Creates an empty paged cache for num_layers layers (all global). Arena tensors are allocated lazily on first use of each layer.

Source

pub fn new_with_windows( num_layers: usize, config: CacheConfig, windows: Vec<Option<usize>>, ) -> Self

Creates a paged cache with a per-layer sliding-window assignment (windows[i] = Some(w) ⇒ layer i stores at most w-1 past tokens). w >= 2 — a window of 1 would leave decode steps with no past context at all.

Source

pub fn num_free_pages(&self) -> usize

Number of free pages in the arena.

Source

pub fn page_stats_inner(&self) -> PageStats

Page-table snapshot (see KVCache::page_stats).

Trait Implementations§

Source§

impl<B: Backend> KVCache<B> for PagedKVCache<B>

Source§

fn popn(&mut self, n: usize) -> usize

Rolls back the last n cached tokens, freeing trailing pages that become fully unused. K/V content of popped positions is left in the arena but is never read (writes always cover seq_len.. densely).

Sliding layers can only roll back while nothing has been evicted from their window: once eviction starts, the tokens a rollback would re-expose are gone, so the whole cache refuses (0) and the caller rebuilds from scratch — the same “prefix caching disables under sliding windows” rule HF applies. All-or-nothing: state is only mutated when the full rollback is possible.

Source§

fn attention_opts( &mut self, layer: usize, q: Tensor<B, 4>, k: Tensor<B, 4>, v: Tensor<B, 4>, pos: usize, scale: f64, window: Option<usize>, ) -> Tensor<B, 4>

KVCache::attention with an optional sliding-window span (Gemma local layers): when Some(w), query at absolute position p attends only keys in (p - w, p] — older keys stay cached but are masked out. None = full causal attention (Llama-family behavior).
Source§

fn seq_len(&self) -> usize

Total cached sequence length.
Source§

fn reset(&mut self)

Drops all cached state (session reset).
Source§

fn pages_used(&self) -> Option<usize>

Pages currently allocated to the sequence (paged cache only).
Source§

fn page_stats(&self) -> Option<PageStats>

Page-table snapshot for observability (paged cache only). Cheap: reads counters, never touches device memory.
Source§

fn attention( &mut self, layer: usize, q: Tensor<B, 4>, k: Tensor<B, 4>, v: Tensor<B, 4>, pos: usize, scale: f64, ) -> Tensor<B, 4>

Appends seq new positions of K/V for layer and computes attention of q against the full cached window (past + new). Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<T> Downcast<T> for T

Source§

fn downcast(&self) -> &T

Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self>

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self>

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> IntoComptime for T

Source§

fn comptime(self) -> Self

Source§

impl<T> Read<Exclusive, BecauseExclusive> for T
where T: ?Sized,

Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = Infallible

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, <T as TryFrom<U>>::Error>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> Upcast<T> for T

Source§

fn upcast(&self) -> Option<&T>

Source§

impl<T> WasmNotSend for T
where T: Send,

Source§

impl<T> WasmNotSendSync for T

Source§

impl<T> WasmNotSync for T
where T: Sync,

Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self>
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self>

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more