polyvoice
Speaker diarization for Rust — who spoke when, on CPU, without Python.
Beta-quality, ONNX-powered, ~30 MB. Embeds into any Rust app, with Python, C, and CLI bindings.
Speaker_0: 0.0s - 12.3s
Speaker_1: 14.1s - 28.7s
Speaker_0: 31.2s - 45.0s
Like-for-like (collar 0, overlap-scored) VoxConverse-test DER is 15.4% (v2+VBx default) vs pyannote 3.1's 11.3% — a few DER points traded for a CPU-only, MIT, ungated engine that needs no Python — see Benchmarks.
Install
Pre-built CLI binary
Download the latest binary for your platform from the GitHub Releases page:
# Linux x86_64
# Linux ARM64
# macOS (Apple Silicon)
Docker
No pre-built image is published — build locally from the repo root (the same image CI smoke-tests):
Rust library
# Full ONNX path (Silero VAD, WeSpeaker embeddings, model download)
Library mode (no ONNX)
For BYO-embedder consumers (your own speaker embeddings + pure-Rust VAD /
clustering), disable default features so ort is never linked:
# optional pure-Rust extras: --features clusterer,vbx
You get LegacyPipeline, StreamingPipeline, EnergyVad, the Embedder trait
(implement it with Candle, tract, or any other backend), and pure clustering
math — no ONNX Runtime dylib. Runnable mock: cargo run --no-default-features --example byo_embedder. Surface inventory and the CI gate that keeps this path
ort-free: docs/library-mode.md.
Python
From source
# CLI (WAV 16 kHz mono input). Feature `cli` includes VBx (default clusterer).
# CLI + any-format audio (mp3/flac/ogg/m4a/aac, any sample rate → 16 kHz mono)
# C FFI shared library (ABI v3). Feature `ffi` includes VBx as well.
Usage
use ModelRegistry;
use ClustererKind;
use ;
use ;
# Default path: pipeline v2 + VBx (PLDA auto-downloads via the model registry;
# optional override: --vbx-plda-dir / POLYVOICE_VBX_PLDA_DIR — see docs/vbx-plda-release.md)
# Cosine AHC instead of VBx: --clusterer ahc | old path: --legacy
# With a build that includes `audio-io`:
# polyvoice diarize meeting.mp3 --output meeting.rttm
Python usage and the full API live on docs.rs.
Why polyvoice
- Maintained, pure-Rust, streaming-capable. The popular
sherpa-rsbindings are archived; polyvoice is an actively-maintained, pure-Rust diarization path (ONNX viaort, no C++ toolkit) with first-class streaming. - One library, four surfaces. Rust + Python + C FFI + CLI from a single crate.
- CPU-first, ~30 MB, MIT. No GPU, no Python runtime, no gated model access.
It is not the accuracy leader — like-for-like (collar 0, overlap-scored) VoxConverse-test DER is 15.4% (v2+VBx default) versus 11.3% for pyannote 3.1. It trades those DER points for deployability: a pure-Rust, CPU, MIT, ungated engine (pyannote's weights are gated behind an HF token) with four bindings and streaming.
How it works
audio (f32 PCM)
→ VAD / Powerset segmentation
→ WeSpeaker embeddings
→ clustering (VBx default; AHC / K-means / NME-SC alternatives, automatic speaker count)
→ speaker turns
Streaming (streaming::StreamingPipeline) and batch (crate-root Pipeline;
pipeline::LegacyPipeline on the ort-free BYO path), with a single-speaker
guard so quiet or single-voice audio does not hallucinate clusters.
Documentation
- Library mode (no ONNX) — ort-free surface for BYO embedders
- Pipeline architecture — BYO vs production ONNX paths
- Benchmarks — collar-disclosed DER numbers and provenance
- Production readiness — deployment guidance (GO / NO-GO)
- Migrating from 0.5 · Glossary
- Contributing · Changelog
License
MIT
Name: this project is polyvoice — speaker diarization for Rust, unrelated to ByteDance's "PolyVoice" speech-translation research.