polyvoice
WAV in, speaker turns out.
A speaker diarization crate. Powerset neural segmentation, WeSpeaker
ResNet34 embeddings, VBx clustering with automatic speaker count. One
Pipeline call from 16 kHz mono to timestamped turns. The default build
pulls no ONNX Runtime: hand-written INT8 kernels, ~8.4 MB production
model pair, MIT, ungated. ONNX Runtime is a feature (cli-ort), not a
requirement. Python, C FFI and a CLI ship from the same crate.
Examples
Library (kernels, models auto-download):
use ModelRegistry;
use ;
use ;
CLI:
SPEAKER meeting 1 0.000 12.784 <NA> <NA> SPEAKER_00 <NA> <NA>
SPEAKER meeting 1 13.005 2.530 <NA> <NA> SPEAKER_01 <NA> <NA>
SPEAKER meeting 1 15.688 10.323 <NA> <NA> SPEAKER_02 <NA> <NA>
A 1-hour meeting diarizes in about a minute on a laptop. Python:
pip install polyvoice (python/README.md). C FFI:
docs/FFI.md.
Surfaces
| Surface | Engine | Links ort |
|---|---|---|
CLI, --features cli |
INT8 kernels | no |
Rust library, pipeline-native,vbx |
INT8 kernels | no |
C FFI, --features ffi |
INT8 kernels | no |
BYO embedder, --no-default-features |
yours | no |
Python wheel, pip install polyvoice |
ONNX Runtime | yes |
CLI / library, cli-ort / pipeline-full |
ONNX Runtime | yes |
Compared to pyannote
Like-for-like, strict collar 0, VoxConverse-test (232 files). Full matrix (incl. diart, whisperx, speakrs): compare.
| polyvoice | pyannote 3.1 | |
|---|---|---|
| Job | diarization crate | research diarization |
| Runtime | Rust, CPU-only | PyTorch, GPU recommended |
| Weights | MIT, ungated | HF token required |
| Default deps | none | PyTorch stack |
| DER₀ | 15.3 % | 11.3 % |
| Speed | ~141× realtime (Ryzen AI 9 HX 370) | GPU-bound |
The trade is explicit: ~4 DER points for a CPU-only, MIT, ungated deploy with no Python. Not the accuracy leader — the deployability leader.
Speed
Kernels (product default) vs same-host ONNX Runtime, EP=cpu, INT8. Linux x86_64: Ryzen AI 9 HX 370, 2026-09-08. Darwin: Apple Silicon. DER₀ is strict collar 0. Protocol: benchmarks.
| Corpus | DER₀ | kernels Linux | ort Linux | kernels Darwin |
|---|---|---|---|---|
| VoxConverse-test (232) | 15.3 % | ~141× | ~137× | ~130× |
| AMI-test (16) | 25.5 % | ~162× | ~156× | ~109× |
| Vox-3 smoke | 7.0 % | ~103×, ~158× wall at --jobs 3 |
~129×, ~151× at --jobs 3 |
≥117× |
Peak RSS on the Vox-3 smoke: ~310 MiB kernels vs ~620 MiB ort at
jobs=1; ~470 MiB vs ~740 MiB at --jobs 3 (one shared pipeline, DER
bit-identical to jobs=1). On-disk INT8 pair: 8,414,314 bytes — a
locked scoreboard floor, as are DER and RSS (tests/native_scoreboard.json).
How it works
audio (f32 PCM)
→ powerset neural segmentation (overlap-aware)
→ WeSpeaker ResNet34 embeddings
→ VBx clustering (AHC / K-means / NME-SC alternatives, automatic speaker count)
→ overlap resegmentation → speaker turns
Install
| Platform | Get it |
|---|---|
| Linux x86_64 / ARM64, macOS, Windows | Pre-built binaries |
| Rust library (kernels, no ort) | cargo add polyvoice --features "pipeline-native,vbx" |
| Rust library (ONNX Runtime) | cargo add polyvoice --features "pipeline-full,vbx" |
| From source | cargo install polyvoice --features cli · "cli,audio-io" · cli-ort · cli-tract · ffi |
[]
= { = "0.20", = ["pipeline-native", "vbx"] }
rustc 1.94. Default features are empty: the published crate is the ort-free BYO core; models and engines are opt-in features (library mode).
benchmarks | api | architecture | library mode | ffi | python | production readiness | CHANGELOG
Batch diarization only: no ASR, no speaker identification. Beta (0.x): the public API may break between minor versions — pin an exact version in production. MIT.
Name: this project is polyvoice — speaker diarization for Rust, unrelated to ByteDance's "PolyVoice" speech-translation research.