pub struct VisionTower { /* private fields */ }Implementations§
Source§impl VisionTower
impl VisionTower
Sourcepub fn load(e: &Engine, dir: &Path) -> Result<Self, Box<dyn Error>>
pub fn load(e: &Engine, dir: &Path) -> Result<Self, Box<dyn Error>>
Load the tower from a directory containing outside.safetensors.
Sourcepub fn forward(
&self,
e: &Engine,
patches: &[f32],
gh: usize,
gw: usize,
) -> Result<CudaSlice<f32>, Box<dyn Error>>
pub fn forward( &self, e: &Engine, patches: &[f32], gh: usize, gw: usize, ) -> Result<CudaSlice<f32>, Box<dyn Error>>
Forward one image’s patches -> [ghgw/4, 5120] merged embeddings (device).
patches is host [ghgw, 1536] in the preprocessor’s (c, t, ph, pw) flat order.
Sourcepub fn forward_seq(
&self,
e: &Engine,
patches: &[f32],
groups: usize,
gh: usize,
gw: usize,
) -> Result<CudaSlice<f32>, Box<dyn Error>>
pub fn forward_seq( &self, e: &Engine, patches: &[f32], groups: usize, gh: usize, gw: usize, ) -> Result<CudaSlice<f32>, Box<dyn Error>>
Forward groups temporal groups of one video (or a single image at groups=1):
host patches [groupsghgw, 1536], frame-major -> [groupsghgw/4, 5120] merged
embeddings, frame-major. HF cu_seqlens law (vision_utils.get_vision_cu_seqlens,
merge_temporal=False — the qwen2_vl/qwen3_vl/qwen3_5 convention): EACH temporal
group is its own attention segment, and pos table / rope / merger are all
frame-local too — so a video is exactly its groups run through the single-image
forward, concatenated. (Joint clip attention is the kimi_k25 convention only;
parity receipt: joint span scored mean_cos 0.92 vs the HF oracle, per-group 1.0.)