rlx-mimi 0.2.11

Kyutai Mimi neural audio codec (12.5 Hz, 24 kHz) for RLX
Documentation

rlx-mimi

Native Rust inference for Kyutai Mimi — a streaming neural audio codec at 12.5 Hz / 24 kHz mono.

Quick start

# Download weights (~385 MB)
just fetch-mimi

# Encode + decode roundtrip (CPU eager — HF safetensors layout)
just mimi -- --model-dir .cache/mimi --in-wav speech.wav --out-wav /tmp/mimi-roundtrip.wav

# Candle parity oracle (DEV ONLY) — cross-checks the native codec; needs Moshi sidecar weights
export RLX_MOSHI_DIR=.cache/moshiko   # contains tokenizer-e351c8d8-checkpoint125.safetensors
cargo run -p rlx-mimi --features "parity-mimi-metal,metal" --release -- \
  --bench --in-wav speech.wav --device metal

Backends

Device Engine
cpu Native rlx graph (HF kyutai/mimi weights)
metal / mlx / cuda Native rlx graph
wgpu / vulkan Native rlx graph (gpu/vulkan)

The candle/moshi codec is retained only as a dev parity oracle behind parity-mimi (cross-check the native graph): cargo build -p rlx-mimi --features "parity-mimi-metal,metal".

GPU path loads tokenizer-e351c8d8-checkpoint125.safetensors from RLX_MOSHI_DIR or the mimi dir (not model.safetensors HF layout).

Library

use rlx_mimi::{MimiCodec, MimiCodes, SAMPLE_RATE};
use rlx_runtime::Device;

let mut codec = MimiCodec::open_on_with_moshi(".cache/mimi", Some(".cache/moshiko".as_ref()), Device::Metal, Some(8))?;
let codes: MimiCodes = codec.encode_wav("speech.wav", Some(8))?;
let recon = codec.decode_codes(&codes)?;

Set RLX_MIMI_DIR to skip --model-dir. Input WAVs are resampled to 24 kHz.

Weights

Artifact HuggingFace
HF Mimi (CPU) kyutai/mimi
Candle Mimi (GPU) sidecar in kyutai/moshiko-candle-bf16

Tests

export RLX_MIMI_DIR=.cache/mimi
export RLX_MOSHI_DIR=.cache/moshiko
cargo test -p rlx-mimi --release

RLX_MIMI_GPU_SMOKE=1 cargo test -p rlx-mimi --test gpu_codec_smoke --features "parity-mimi-metal,metal"