denoize 0.3.0

Pure-Rust audio denoiser with classical DSP and optional RNNoise
Documentation

denoize

The pursuit of the world's highest-fidelity audio denoising — in pure Rust.

denoize removes background noise from WAV recordings with maximum transparency: preserving timbre, transients, dynamics, stereo imaging, and natural "air".

Implemented technology stack

Classical DSP (always available)

  • STFT/ISTFT + Perfect Reconstruction OLA + high overlap
  • IMCRA/MCRA noise estimation + SPP + spectral-flatness profiling
  • Ephraim-Malah Decision-Directed SNR
  • 8 gain estimators: OMLSA, LogMMSE, MMSE-STSA, Wiener, SpecSub, SpecSub-NL, SpecSub-Geo
  • Transient protection, cepstral smoothing, pre-emphasis
  • Advanced windows: Kaiser, Flat-top, DPSS (+ Hann/Hamming/Sine/Blackman)
  • Multiband spectral subtraction (Bark bands)
  • Perceptual weighting (Bark-scale gain shaping)
  • Musical-noise post-filter

Optional AI backends (feature-gated)

Backend Feature Description
rnnoise --features rnnoise RNNoise via nnnoiseless (pure-Rust)
deepfilter --features deepfilter DeepFilterNet v3 (tract ONNX, embedded model)
onnx --features onnx External waveform-to-waveform ONNX model (tract, Pure Rust)
mpsenet --features mpsenet MP-SENet magnitude/phase enhancement adapter (external converted model)
bsrnn --features bsrnn ESPnet BSRNN spectral enhancement adapter (external converted model)

Build everything: cargo build --release --features full

The generic ONNX backend is the deployment foundation for future neural models. It intentionally accepts only single-input/single-output waveform models; spectral models and diffusion samplers require dedicated adapters.

The prebuilt GitHub binaries include every backend. Because DeepFilterNet 0.5.6 is not available from crates.io, the crates.io package's full feature currently includes RNNoise, generic ONNX, MP-SENet, and BSRNN, but not DeepFilterNet.

Supported input formats

Format Decoder Notes
WAV hound 8–32 bit int / float
MP3 nanomp3 (Pure Rust) ID3 skip, no resampling
M4A/AAC oxideav-aac (Pure Rust) MP4 demux + AAC-LC decode

Output formats

Format Encoder Notes
WAV hound Lossless; preserves bit depth
MP3 shine-rs (Pure Rust) --mp3-bitrate (default 192 kbps)
M4A oxideav-aac + MP4 mux GitHub/source builds; --m4a-bitrate (default 192 kbps)
# MP3 / M4A input and output — no manual ffmpeg conversion
denoize noisy.mp3 clean.mp3 -p hifi
denoize noisy.m4a clean.m4a -b deepfilter
denoize noisy.wav clean.wav --mp3-bitrate 320

# User-supplied waveform model: [1, samples] or [1, 1, samples]
denoize noisy.wav clean.wav -b onnx \
  --onnx-model model.onnx --onnx-rate 16000

# Official MP-SENet checkpoint converted with scripts/export-mpsenet.py
denoize noisy.wav clean.wav -b mpsenet \
  --onnx-model mp-senet-vb.onnx --onnx-rate 16000

# ESPnet BSRNN xtiny checkpoint converted with scripts/export-bsrnn.py
denoize noisy.wav clean.wav -b bsrnn \
  --onnx-model bsrnn-xtiny.onnx --onnx-rate 48000

To prepare the pinned official MP-SENet VoiceBank model:

git clone https://github.com/yxlu-0102/MP-SENet.git
git -C MP-SENet checkout 89932cfe90d1dacb8e170e4a331d762462c21792
python3 -m pip install torch onnx onnxscript pesq joblib matplotlib
python3 scripts/export-mpsenet.py \
  --repo MP-SENet \
  --checkpoint MP-SENet/best_ckpt/g_best_vb \
  --output mp-senet-vb.onnx

To prepare the pinned ESPnet BSRNN xtiny model (CC-BY-4.0):

curl -L \
  'https://huggingface.co/wyz/vctk_bsrnn_xtiny_causal/resolve/59e1f2263b7946b1970a222d1beef9adc5a67eaa/exp_vctk/enh_train_enh_bsrnn_xtiny_raw/58epoch.pth' \
  -o 58epoch.pth
echo 'e3cb771a452e0503144af74720b476e81b57f518b789b37ba2c253c6cc22d70b  58epoch.pth' \
  | sha256sum -c -
python3 -m pip install torch onnx onnxruntime
python3 scripts/export-bsrnn.py \
  --checkpoint 58epoch.pth \
  --output bsrnn-xtiny.onnx \
  --verify

The adapter resamples to 48 kHz and reproduces the published model's variance normalization, centered 960-point Hann STFT with a 480-sample hop, whole-utterance recurrent inference, and inverse STFT. The converted model is about 2.4 MiB. On a release build on the project reference x86-64 Linux host, the fixed two-second regression fixture took 1.58 seconds (1.3x realtime) and used 44,628 KiB maximum RSS. Runtime and memory grow with utterance length.

Run the reproducible real-speech quality gate after conversion:

python3 scripts/validate-bsrnn.py \
  --denoize target/release/denoize \
  --model bsrnn-xtiny.onnx

Quick start

cargo build --release --features full

# Best classical quality
./target/release/denoize noisy.wav clean.wav -p hifi

# RNNoise AI backend
./target/release/denoize noisy.wav clean.wav -b rnnoise

# DeepFilterNet v3 AI backend
./target/release/denoize noisy.wav clean.wav -b deepfilter

# Advanced DSP options
./target/release/denoize noisy.wav clean.wav \
  --window kaiser --kaiser-beta 10 \
  --multiband --perceptual --postfilter \
  -a specsub-nl -s 0.5

Prebuilt binaries

Each GitHub Release contains prebuilt full-feature binaries for:

  • Linux x86-64
  • macOS Intel and Apple Silicon
  • Windows x86-64

Every archive has a matching .sha256 checksum file.

Install with Cargo

The crates.io package provides the CLI and library with the classical DSP and optional RNNoise backends:

cargo install denoize --features full

For the embedded DeepFilterNet backend, use a prebuilt GitHub binary or build this repository with its primary Cargo.toml.

Publishing a release

  1. Set the same version in Cargo.toml and Cargo.crates-io.toml, then update Cargo.lock.
  2. Commit and push the version change.
  3. Create and push a matching tag:
git tag -a v0.1.0 -m "denoize v0.1.0"
git push origin v0.1.0

The GitHub Release workflow validates that the tag matches Cargo.toml, runs the full test suite, builds all supported platforms, attaches archives and checksums, and publishes generated release notes. A failed build leaves the release as a draft so it cannot expose an incomplete asset set.

CLI highlights

-b, --backend <NAME>     classical|rnnoise|deepfilter
-a, --algorithm <NAME>    omlsa|logmmse|mmse|wiener|specsub|specsub-nl|specsub-geo
--window <NAME>          hann|hamming|sine|blackman|kaiser|flattop|dpss
--kaiser-beta <B>        Kaiser β (default 8.0)
--dpss-nw <NW>           DPSS bandwidth (default 3.0)
--multiband              Multiband spectral subtraction
--perceptual             Bark perceptual gain weighting
--postfilter             Musical-noise suppression post-filter
-p hifi                   Flagship preset (Kaiser + perceptual + postfilter)
--quality ultra           Maximum fidelity settings
--onnx-model <PATH>       Waveform ONNX model used by the onnx backend
--onnx-rate <HZ>          Model sample rate (default: 16000)

Library API

use denoize::{denoise_file_with_backend, Backend, DenoiserConfig, Preset};

let cfg = Preset::HiFi.config(48000);
denoise_file_with_backend("noisy.wav", "clean.wav", cfg, Backend::Classical)?;

// With DeepFilterNet (GitHub/source build with --features full)
denoise_file_with_backend("noisy.wav", "clean.wav", cfg, Backend::DeepFilter)?;

Roadmap status

Priority Technology Status
1 DeepFilterNet v3 --features deepfilter
2 RNNoise --features rnnoise
3 Kaiser/Flat-top/DPSS windows
4 Multiband / nonlinear SpecSub
5 Perceptual weighting + musical-noise PF
6 Pure-Rust external ONNX inference foundation 🟨 waveform contract implemented
7 BSRNN / MP-SENet / MossFormer2 adapters 🟨 BSRNN implemented; MP-SENet partial; MossFormer2 researching
8 SGMSE+ 🔲 Diffusion sampler + score-model port

See ROADMAP.md for the implementation audit and the acceptance criteria for marking each named model complete.

License

The Rust project is MIT licensed. See THIRD_PARTY.md for the Apache-2.0 BSRNN conversion code and CC-BY-4.0 model attribution.