pub struct ExpertLayout {
pub offset: usize,
pub len: usize,
pub qtype: i32,
pub row_bytes: usize,
}Expand description
One layer’s stacked 256-expert tensor, raw GGUF quant bytes held HOST-RESIDENT.
EDGE-1: these bytes are NEVER uploaded at load (uploading 29.75GB would OOM a 24GB GPU — this is BUG-4). Per token, only the 8 routed experts are staged H2D into a small GPU scratch.
ne = [in_f, out_f, n_expert]; the expert axis (ne[2]) is the slowest/highest-stride axis, so
expert e occupies the CONTIGUOUS byte block bytes[e*expert_stride .. (e+1)*expert_stride].
THE 3D FIX: GpuTensor::load computes row_bytes = raw.len()/ne[1], which for a stacked 3D
tensor ignores the 256-expert axis and is 256x too large (gate_exps -> 430080 instead of 1680).
load() here uses row_bytes = raw.len() / (out_f * n_expert) (= 1680 gate/up, 544 down).
Fields§
§offset: usize§len: usize§qtype: i32§row_bytes: usizeTrait Implementations§
Source§impl Clone for ExpertLayout
impl Clone for ExpertLayout
Source§fn clone(&self) -> ExpertLayout
fn clone(&self) -> ExpertLayout
Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
Performs copy-assignment from
source. Read moreimpl Copy for ExpertLayout
Source§impl Debug for ExpertLayout
impl Debug for ExpertLayout
impl Eq for ExpertLayout
Source§impl PartialEq for ExpertLayout
impl PartialEq for ExpertLayout
impl StructuralPartialEq for ExpertLayout
Auto Trait Implementations§
impl Freeze for ExpertLayout
impl RefUnwindSafe for ExpertLayout
impl Send for ExpertLayout
impl Sync for ExpertLayout
impl Unpin for ExpertLayout
impl UnsafeUnpin for ExpertLayout
impl UnwindSafe for ExpertLayout
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
Mutably borrows from an owned value. Read more