Expand description
Mimi neural audio tokenizer support. Mimi neural audio tokenizer support.
Mimi is the neural audio codec used by Moshi-family realtime models. This module implements backend-neutral checkpoint parameters, the split residual vector quantizer, and the non-streaming SEANet/transformer encoder and decoder used to map between PCM and Mimi codebook tokens.
Structs§
- Checkpoint
Tensor Plan - Backend-independent loading plan for one tensor in a released Mimi checkpoint.
- Config
- Mimi codec configuration.
- Conv1x1
NoBias - Bias-free 1x1 convolution over
[batch, channels, frames]tensors. - Euclidean
Codebook - Euclidean codebook backed by EMA cluster statistics.
- Mimi
- Mimi audio tokenizer.
- Residual
Vector Quantization - Residual vector quantization layers.
- Residual
Vector Quantizer - Residual vector quantizer branch.
- Split
Residual Vector Quantizer - Split residual vector quantizer used by Mimi.
- Vector
Quantization - Single vector-quantization layer.
Enums§
- Checkpoint
Tensor Layout - Backend-independent layout conversion required by a Mimi checkpoint tensor.
- Resample
Method - Mimi resampling strategy.
Functions§
- checkpoint_
tensor_ plan - Maps a released Mimi checkpoint tensor name to its stable model parameter and required layout conversion.