pub struct AnimParams {Show 17 fields
pub width: usize,
pub height: usize,
pub frames: usize,
pub steps: usize,
pub seed: u64,
pub stock_sampler: bool,
pub max_tokens: usize,
pub first_frame: Option<(Vec<f32>, usize, usize)>,
pub last_frame: Option<(Vec<f32>, usize, usize)>,
pub mid_frames: Vec<((Vec<f32>, usize, usize), usize)>,
pub lora: Option<String>,
pub lora_strength: f32,
pub upscale: Option<String>,
pub upscale_by: f32,
pub stream_chunk: usize,
pub stream_sink: usize,
pub stream_window: usize,
}Fields§
§width: usize§height: usize§frames: usizeFrames at 24 fps; snapped up to the model’s 17k+5 grid.
steps: usize§seed: u64§stock_sampler: boolIntegrate the audio on the video’s grid, as a stock sampler would. Wrong at four steps; kept for A/B.
max_tokens: usize§first_frame: Option<(Vec<f32>, usize, usize)>RGB in [0, 1] as [3, h, w] with its size — the clip’s first
frame, and/or its last.
last_frame: Option<(Vec<f32>, usize, usize)>§mid_frames: Vec<((Vec<f32>, usize, usize), usize)>Extra reference frames with the pixel index each one stands for:
video-to-video as this architecture actually takes it — every
frame is a condition row pinned to its own time coordinate, the
same machinery --first-frame/--last-frame use for two.
lora: Option<String>A LoRA adapter (.safetensors) applied at runtime, and how hard.
lora_strength: f32§upscale: Option<String>The published latent upscaler (.safetensors) and the factor to apply: the denoised latent is resized by the learned net and the VAE decodes at the larger size, so the 5 B-parameter decode → pixel resize → encode round trip never happens.
upscale_by: f32§stream_chunk: usizeChunk-causal (streaming) generation: latent frames per chunk, and
how many chunks a chunk may see — sink from the start of the
clip and a sliding window of recent ones. 0 chunks = the
bidirectional path. This is what a streaming adapter is trained
for, and it is what stops the activation cache growing with the
clip’s length.
stream_sink: usize§stream_window: usize