pub struct NextDit {
pub hidden: usize,
pub in_channels: usize,
pub patch: usize,
/* private fields */
}Expand description
Next-DiT: exact f32 from a diffusers directory, or CMF-quantized (mmap-resident) from a packaged file.
Fields§
§in_channels: usize§patch: usizeImplementations§
Source§impl NextDit
impl NextDit
pub fn load_dir(dir: &Path) -> Result<Self, String>
Sourcepub fn from_cmf(model: &Arc<CmfModel>) -> Result<Self, String>
pub fn from_cmf(model: &Arc<CmfModel>) -> Result<Self, String>
Load from a packaged imagegen .cmf (dit.* tensors +
dit.config_json). Quantized projections stay mmap-resident.
Sourcepub fn forward(
&self,
latent: &[f32],
h: usize,
w: usize,
cap: &[f32],
cap_n: usize,
t: f32,
) -> Vec<f32>
pub fn forward( &self, latent: &[f32], h: usize, w: usize, cap: &[f32], cap_n: usize, t: f32, ) -> Vec<f32>
One denoising forward: latent [c, h, w] (NCHW), caption
features [cap_n, cap_feat], timestep t ∈ [0,1] (the
pipeline’s 1 − σ). Returns the velocity prediction [c, h, w].
Sourcepub fn refine_caption(&self, cap: &[f32], cap_n: usize) -> Vec<f32>
pub fn refine_caption(&self, cap: &[f32], cap_n: usize) -> Vec<f32>
Caption features → hidden, through the context refiner. Depends on
NOTHING that moves during denoising — not the timestep, not the
latents — so the whole thing is a constant of the prompt. The
denoise loop hoists it out and hands the result to
forward_with_cap; it used to be recomputed on every model call,
which for 30 steps under CFG meant 60 evaluations of a value with
two distinct instances.
Sourcepub fn forward_cfg_pair(
&self,
latent: &[f32],
h: usize,
w: usize,
cap_c: &[f32],
cap_c_n: usize,
cap_u: &[f32],
cap_u_n: usize,
t: f32,
) -> (Vec<f32>, Vec<f32>)
pub fn forward_cfg_pair( &self, latent: &[f32], h: usize, w: usize, cap_c: &[f32], cap_c_n: usize, cap_u: &[f32], cap_u_n: usize, t: f32, ) -> (Vec<f32>, Vec<f32>)
The rest of the forward, from an already-refined caption. Classifier-free guidance in ONE pass: the conditional and the unconditional sequence go through the joint stack as a single batch. Every weight is read once for both — which is the whole cost on a CPU or a phone — and the image branch (patchify, x_embedder, the noise refiners) is computed once instead of twice, because it does not depend on the caption at all. Attention stays per sequence, so the two never mix and each prediction equals what the single-sequence path returns.