# denoize
**The pursuit of the world's highest-fidelity audio denoising — in pure Rust.**
`denoize` removes background noise from WAV recordings with **maximum transparency**:
preserving timbre, transients, dynamics, stereo imaging, and natural "air".
## Implemented technology stack
### Classical DSP (always available)
- STFT/ISTFT + Perfect Reconstruction OLA + high overlap
- IMCRA/MCRA noise estimation + SPP + spectral-flatness profiling
- Ephraim-Malah Decision-Directed SNR
- **8 gain estimators**: OMLSA, LogMMSE, MMSE-STSA, Wiener, SpecSub, SpecSub-NL, SpecSub-Geo
- Transient protection, cepstral smoothing, pre-emphasis
- **Advanced windows**: Kaiser, Flat-top, DPSS (+ Hann/Hamming/Sine/Blackman)
- **Multiband spectral subtraction** (Bark bands)
- **Perceptual weighting** (Bark-scale gain shaping)
- **Musical-noise post-filter**
### Optional AI backends (feature-gated)
| `rnnoise` | `--features rnnoise` | RNNoise via nnnoiseless (pure-Rust) |
| `deepfilter` | `--features deepfilter` | DeepFilterNet v3 (tract ONNX, embedded model) |
| `onnx` | `--features onnx` | External waveform-to-waveform ONNX model (tract, Pure Rust) |
| `mpsenet` | `--features mpsenet` | MP-SENet magnitude/phase enhancement adapter (external converted model) |
| `bsrnn` | `--features bsrnn` | ESPnet BSRNN spectral enhancement adapter (external converted model) |
| `mossformer2` | `--features mossformer2` | ClearerVoice MossFormer2 48 kHz mask adapter (external converted model) |
Build everything: `cargo build --release --features full`
The generic ONNX backend is the deployment foundation for future neural
models. It intentionally accepts only single-input/single-output waveform
models; spectral models and diffusion samplers require dedicated adapters.
> The prebuilt GitHub binaries include every backend. Because DeepFilterNet
> 0.5.6 is not available from crates.io, the crates.io package's `full` feature
> currently includes RNNoise, generic ONNX, MP-SENet, BSRNN, and MossFormer2, but not
> DeepFilterNet.
## Supported input formats
| WAV | `hound` | 8–32 bit int / float |
| MP3 | `nanomp3` (Pure Rust) | ID3 skip, no resampling |
| M4A/AAC | `oxideav-aac` (Pure Rust) | MP4 demux + AAC-LC decode |
### Output formats
| WAV | `hound` | Lossless; preserves bit depth |
| MP3 | `shine-rs` (Pure Rust) | `--mp3-bitrate` (default 192 kbps) |
| M4A | `oxideav-aac` + MP4 mux | GitHub/source builds; `--m4a-bitrate` (default 192 kbps) |
```sh
# MP3 / M4A input and output — no manual ffmpeg conversion
denoize noisy.mp3 clean.mp3 -p hifi
denoize noisy.m4a clean.m4a -b deepfilter
denoize noisy.wav clean.wav --mp3-bitrate 320
# User-supplied waveform model: [1, samples] or [1, 1, samples]
denoize noisy.wav clean.wav -b onnx \
--onnx-model model.onnx --onnx-rate 16000
# Official MP-SENet checkpoint converted with scripts/export-mpsenet.py
denoize noisy.wav clean.wav -b mpsenet \
--onnx-model mp-senet-vb.onnx --onnx-rate 16000
# ESPnet BSRNN xtiny checkpoint converted with scripts/export-bsrnn.py
denoize noisy.wav clean.wav -b bsrnn \
--onnx-model bsrnn-xtiny.onnx --onnx-rate 48000
# ClearerVoice MossFormer2 48 kHz model
denoize noisy.wav clean.wav -b mossformer2 \
--onnx-model mossformer2-se-48k.onnx --onnx-rate 48000
```
To prepare the pinned official MP-SENet VoiceBank model:
```sh
git clone https://github.com/yxlu-0102/MP-SENet.git
git -C MP-SENet checkout 89932cfe90d1dacb8e170e4a331d762462c21792
python3 -m pip install torch onnx onnxscript pesq joblib matplotlib
python3 scripts/export-mpsenet.py \
--repo MP-SENet \
--checkpoint MP-SENet/best_ckpt/g_best_vb \
--output mp-senet-vb.onnx
```
The VoiceBank graph is about 9 MiB and expects 16 kHz audio. On the reference
x86-64 Linux host, a two-second mono speech fixture took 43.67 seconds after
model loading and the complete process used 410,048 KiB maximum RSS. Run the
pinned real-speech quality gate after conversion:
```sh
python3 scripts/validate-mpsenet.py \
--denoize target/release/denoize \
--model mp-senet-vb.onnx
```
To prepare the pinned ESPnet BSRNN xtiny model (CC-BY-4.0):
```sh
curl -L \
'https://huggingface.co/wyz/vctk_bsrnn_xtiny_causal/resolve/59e1f2263b7946b1970a222d1beef9adc5a67eaa/exp_vctk/enh_train_enh_bsrnn_xtiny_raw/58epoch.pth' \
-o 58epoch.pth
echo 'e3cb771a452e0503144af74720b476e81b57f518b789b37ba2c253c6cc22d70b 58epoch.pth' \
| sha256sum -c -
python3 -m pip install torch onnx onnxruntime
python3 scripts/export-bsrnn.py \
--checkpoint 58epoch.pth \
--output bsrnn-xtiny.onnx \
--verify
```
The adapter resamples to 48 kHz and reproduces the published model's
variance normalization, centered 960-point Hann STFT with a 480-sample hop,
whole-utterance recurrent inference, and inverse STFT. The converted model is
about 2.4 MiB. On a release build on the project reference x86-64 Linux host,
the fixed two-second regression fixture took 1.58 seconds (1.3x realtime) and
used 44,628 KiB maximum RSS. Runtime and memory grow with utterance length.
Run the reproducible real-speech quality gate after conversion:
```sh
python3 scripts/validate-bsrnn.py \
--denoize target/release/denoize \
--model bsrnn-xtiny.onnx
```
To prepare the pinned Apache-2.0 MossFormer2 SE 48 kHz model:
```sh
git clone https://github.com/modelscope/ClearerVoice-Studio.git
git -C ClearerVoice-Studio checkout 6b3774dc79c46ae8bed2a4fa5f706f0ac8c75c61
curl -L \
'https://huggingface.co/alibabasglab/MossFormer2_SE_48K/resolve/eff8c97925c8bec812af707814b3e5d777fd4503/last_best_checkpoint.pt' \
-o last_best_checkpoint.pt
echo '03692b9f773bbd6bb43b9c5a41f96b1e28affd66e13796b7bec66ad3d8b227c6 last_best_checkpoint.pt' \
| sha256sum -c -
python3 -m pip install torch onnx onnxruntime numpy einops rotary-embedding-torch
python3 scripts/export-mossformer2.py \
--repo ClearerVoice-Studio \
--checkpoint last_best_checkpoint.pt \
--output mossformer2-se-48k.onnx \
--verify
```
The adapter uses 48 kHz audio, 40 ms Kaldi fbank frames at an 8 ms shift,
first- and second-order deltas, a non-centred 1,920-point symmetric-Hamming
STFT, and the official four-second/three-second-stride edge-discard
reconstruction. The converted graph is about 217 MiB. On the reference
x86-64 Linux host, a four-second mono fixture took 7.74 seconds and used
483,400 KiB maximum RSS in a release build. Model weights are not bundled.
Run the pinned real-speech quality gate after conversion:
```sh
python3 scripts/validate-mossformer2.py \
--denoize target/release/denoize \
--model mossformer2-se-48k.onnx
```
## Quick start
```sh
cargo build --release --features full
# Best classical quality
./target/release/denoize noisy.wav clean.wav -p hifi
# RNNoise AI backend
./target/release/denoize noisy.wav clean.wav -b rnnoise
# DeepFilterNet v3 AI backend
./target/release/denoize noisy.wav clean.wav -b deepfilter
# Advanced DSP options
./target/release/denoize noisy.wav clean.wav \
--window kaiser --kaiser-beta 10 \
--multiband --perceptual --postfilter \
-a specsub-nl -s 0.5
```
## Prebuilt binaries
Each [GitHub Release](https://github.com/penguin425/denoize/releases) contains
prebuilt `full`-feature binaries for:
- Linux x86-64
- macOS Intel and Apple Silicon
- Windows x86-64
Every archive has a matching `.sha256` checksum file.
## Install with Cargo
The crates.io package provides the CLI and library with the classical DSP and
optional RNNoise backends:
```sh
cargo install denoize --features full
```
For the embedded DeepFilterNet backend, use a prebuilt GitHub binary or build
this repository with its primary `Cargo.toml`.
### Publishing a release
1. Set the same version in `Cargo.toml` and `Cargo.crates-io.toml`, then update
`Cargo.lock`.
2. Commit and push the version change.
3. Create and push a matching tag:
```sh
git tag -a v0.1.0 -m "denoize v0.1.0"
git push origin v0.1.0
```
The `GitHub Release` workflow validates that the tag matches `Cargo.toml`, runs
the full test suite, builds all supported platforms, attaches archives and
checksums, and publishes generated release notes. A failed build leaves the
release as a draft so it cannot expose an incomplete asset set.
## CLI highlights
```
--window <NAME> hann|hamming|sine|blackman|kaiser|flattop|dpss
--kaiser-beta <B> Kaiser β (default 8.0)
--dpss-nw <NW> DPSS bandwidth (default 3.0)
--multiband Multiband spectral subtraction
--perceptual Bark perceptual gain weighting
--postfilter Musical-noise suppression post-filter
-p hifi Flagship preset (Kaiser + perceptual + postfilter)
--quality ultra Maximum fidelity settings
--onnx-model <PATH> Waveform ONNX model used by the onnx backend
--onnx-rate <HZ> Model sample rate (default: 16000)
```
## Library API
```rust
use denoize::{denoise_file_with_backend, Backend, DenoiserConfig, Preset};
let cfg = Preset::HiFi.config(48000);
denoise_file_with_backend("noisy.wav", "clean.wav", cfg, Backend::Classical)?;
// With DeepFilterNet (GitHub/source build with --features full)
denoise_file_with_backend("noisy.wav", "clean.wav", cfg, Backend::DeepFilter)?;
```
## Roadmap status
| 1 | DeepFilterNet v3 | ✅ `--features deepfilter` |
| 2 | RNNoise | ✅ `--features rnnoise` |
| 3 | Kaiser/Flat-top/DPSS windows | ✅ |
| 4 | Multiband / nonlinear SpecSub | ✅ |
| 5 | Perceptual weighting + musical-noise PF | ✅ |
| 6 | Pure-Rust external ONNX inference foundation | 🟨 waveform contract implemented |
| 7 | BSRNN / MP-SENet / MossFormer2 adapters | 🟨 BSRNN implemented; MP-SENet partial; MossFormer2 researching |
| 8 | SGMSE+ | 🔲 Diffusion sampler + score-model port |
See [ROADMAP.md](ROADMAP.md) for the implementation audit and the acceptance
criteria for marking each named model complete.
## License
The Rust project is MIT licensed. See [THIRD_PARTY.md](THIRD_PARTY.md) for the
Apache-2.0 BSRNN conversion code and CC-BY-4.0 model attribution.