audiofp 0.4.0

Pure-Rust audio fingerprinting: Wang, Panako, Haitsma–Kalker with streaming, in-memory matching, ONNX neural/watermark, no_std + alloc, Pod hash types.
docs.rs failed to build audiofp-0.4.0
Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.
Visit the last successful build: audiofp-0.3.9

audiofp

Crates.io Documentation License CI codecov Crates.io Downloads Rust Version

Audio fingerprinting library for Rust with classical landmark and band-power algorithms, in-memory matching, streaming extraction, file decoding, and AudioSeal-compatible watermark detection.

Overview

audiofp provides three complementary classical fingerprinters for music identification, each with offline and streaming variants, plus an in-memory matching layer for identification:

Method Use Case Sample Rate Frame Rate Output Size
Wang Music ID, Shazam-style matching 8 kHz 62.5 fps ~2.4 KB/s (fan-out 10)
Panako Music ID with ±5 % tempo robustness 8 kHz 62.5 fps ~2.0 KB/s (fan-out 5)
Haitsma Compact dense IDs, fastest extraction 5 kHz 78.125 fps 312 B/s
Matching In-memory ID (WangMatcher, HaitsmaMatcher, …)
Streaming Real-time hash emission (per algorithm) (per algorithm) Bit-exact offline parity
Watermark AudioSeal detection (BYO ONNX) 16 kHz (per model) Detection + 16-bit message

Perfect for:

  • Music identification ("what is this song?")
  • Audio deduplication at scale
  • Royalty / rights enforcement against re-encoded content
  • Embedding-based similarity search and cover/remix detection (BYO ONNX model via the neural feature)
  • Watermark verification on generative-AI audio

Features

  • Three Classical Algorithms - Wang (landmark pairs) + Panako (triplet hashes with tempo β) + Haitsma–Kalker (32-bit/frame band sign)
  • In-Memory Matching - WangMatcher / HaitsmaMatcher / PanakoMatcher (tempo-invariant 2-D Hough + RANSAC) / NeuralMatcher plus match_best / match_ranked and transient WangIndex / HaitsmaIndex / PanakoIndex accelerators for 1:N identification. No persistence or DB adapters.
  • Truly Incremental Streaming - Per-push CPU proportional to new samples, not total stream length. Rolling spectrogram + per-bucket finalisation + per-anchor target accumulator. Bit-exact parity with offline extract (verified by the test suite at every chunk size).
  • Bit-Exact Determinism - Same input always produces the same hashes; verified down to 1-sample-per-push streaming chunks
  • bytemuck::Pod Hash Types - Persist hashes directly to mmap'd files or ship over a C ABI without serialization
  • Audio File Decoding - MP3, FLAC, WAV, OGG-Vorbis, AAC-in-MP4, raw PCM via Symphonia
  • High-Quality Resampling - Built-in windowed-sinc Kaiser resampler with auto anti-aliasing cutoff
  • Watermark Detection - AudioSeal-compatible ONNX wrapper (Tract backend); typed model is cached per input length and rebuilt automatically when the length changes
  • Neural Embedder - Generic ONNX log-mel embedder with offline + streaming modes; build-once-runnable, zero-alloc try_push_with callback (scratch is allocated at construction, reused on every push)
  • DSP Primitives Reusable - Public dsp::stft, dsp::mel, dsp::peaks, dsp::resample, dsp::windows
  • Allocation-Free Hot Path - Streaming push reuses pre-allocated scratch after warmup
  • no_std + alloc Capable - DSP and classical fingerprinters compile without std (host-only today; bare-metal in roadmap)
  • Feature-Gated Heavy Deps - Symphonia and Tract both opt-in via Cargo features
  • Optional mimalloc - Single-flag opt-in to install mimalloc as the global allocator

Installation

[dependencies]
# WAV + MP3 decoding for the quick-start below (pick the codecs you need):
audiofp = { version = "0.4", features = ["std-wav", "std-mp3"] }

The default build is no_std + alloc with no codecs. Decoding helpers (audiofp::io) are opt-in per codec: std-wav, std-mp3, std-flac, std-ogg, std-aac, std-mp4, plus std-aiff / std-mkv / std-adpcm / std-alac for the extended formats — or all-codecs for every codec at once (the pre-0.4.0 std behavior).

Feature Flags

Feature Default Description
std-wav No WAV + raw PCM decoding via Symphonia (audiofp::io)
std-mp3 No MP3 decoding via Symphonia
std-flac No FLAC decoding via Symphonia
std-ogg No Ogg-Vorbis decoding via Symphonia
std-aac No AAC decoding via Symphonia
std-mp4 No AAC-in-MP4 / ISO-BMFF decoding via Symphonia
std-aiff / std-mkv / std-adpcm / std-alac No Extended codecs
all-codecs No Every codec at once — the pre-0.4.0 std behavior
rayon No Parallel batch fingerprinting via fingerprint_batch_parallel (implies std)
watermark No Enables audiofp::watermark via Tract ONNX runtime (implies std)
neural No Enables audiofp::neural: generic ONNX log-mel embedder via Tract (BYO model; implies std)
mimalloc No Installs mimalloc::MiMalloc as the process-wide #[global_allocator] (implies std)

Minimal build (no_std + alloc, DSP and classical only):

[dependencies]
audiofp = { version = "0.4", default-features = false }

Quick Start

Fingerprint a file

use audiofp::classical::Wang;
use audiofp::io::decode_to_mono_at;
use audiofp::{Fingerprinter, SampleRate};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Decode any supported file format and resample to Wang's 8 kHz.
    // Needs ≥ ~2 s of audio or extract returns AudioTooShort.
    let samples = decode_to_mono_at("song.mp3", 8_000)?;

    let mut wang = Wang::default();
    let fp = wang.extract(&samples, SampleRate::HZ_8000)?;

    println!("{} hashes at {:.1} fps", fp.hashes.len(), fp.frames_per_sec);
    for h in fp.hashes.iter().take(5) {
        println!("  t_anchor={.0} hash={:08x}", h.t_anchor, h.hash);
    }

    Ok(())
}

Match two fingerprints (Wang)

use audiofp::classical::Wang;
use audiofp::io::decode_to_mono_at;
use audiofp::matching::{Matcher, WangMatchConfig, WangMatcher};
use audiofp::{Fingerprinter, SampleRate};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let samples = decode_to_mono_at("clip.wav", 8_000)?;
    let query = Wang::default().extract(&samples, SampleRate::HZ_8000)?;
    let reference = query.clone(); // same recording

    let m = WangMatcher::new(WangMatchConfig::default()).match_one(&query, &reference);
    println!("is_match={} score={:.3} offset={} ms", m.is_match, m.score, m.offset.ms);
    Ok(())
}

Streaming Mode

use audiofp::classical::StreamingWang;
use audiofp::StreamingFingerprinter;

fn main() {
    let mut s = StreamingWang::default();

    // Synthetic 8 kHz mono chunks (16 ms ≈ 128 samples). Swap for mic/file chunks.
    let chunk = vec![0.0_f32; 128];
    for _ in 0..100 {
        for (timestamp, hash) in s.push(&chunk).unwrap() {
            println!("{:?} {:08x}", timestamp, hash.hash);
        }
    }

    // Drain whatever's pending at end-of-stream.
    for (timestamp, hash) in s.flush() {
        println!("{:?} {:08x}", timestamp, hash.hash);
    }

    println!("latency: {} ms", s.latency_ms());
}

Documentation

For complete API reference and usage examples, see USAGE.md.

Architecture

Fingerprint Types

Each algorithm emits a strongly-typed, bytemuck::Pod-castable result:

Wang offline                         Panako offline
┌──────────────────────────┐         ┌──────────────────────────┐
│ WangFingerprint          │         │ PanakoFingerprint        │
│   hashes: Vec<WangHash>  │         │   hashes: Vec<PanakoHash>│
│   frames_per_sec: f32    │         │   frames_per_sec: f32    │
└──────────────────────────┘         └──────────────────────────┘

WangHash (8 bytes, repr(C))          PanakoHash (16 bytes, repr(C))
├── hash: u32                        ├── hash: u32
└── t_anchor: u32                    ├── t_anchor: u32
                                     ├── t_b: u32
                                     └── t_c: u32

Haitsma offline
┌──────────────────────────┐
│ HaitsmaFingerprint       │
│   frames: Vec<u32>       │   one u32 per spectrogram frame ≥ 1
│   frames_per_sec: f32    │
└──────────────────────────┘

Performance

Offline extract (cargo bench --bench extract, 30 s of synthetic audio):

Algorithm 30 s of audio Realtime factor
Wang 79 ms 380×
Panako 81 ms 370×
Haitsma 42 ms 714×

Streaming push (cargo bench --bench streaming, 10 s of synthetic audio):

Streaming type Small chunks (256 samples) Large chunks (1 s) latency_ms()
StreamingWang 10.5 ms 10.6 ms 2 256 ms
StreamingPanako 11.6 ms 11.4 ms 2 784 ms
StreamingHaitsma 6.3 ms 6.7 ms 409 ms

Neural front-end (cargo bench --features neural --bench neural_frontend):

Path Time
log_mel_pipeline_1s_window 297 µs
strided_tensor_write 7.6 µs
l2_normalize_1024d 2.5 µs

Matching (cargo bench --bench matching, 5 s synthetic fingerprints):

Path Time Notes
WangMatcher 1:1 self-match ~111 µs Offset-histogram voting + prominence
HaitsmaMatcher 1:1 exact ~18 µs Exhaustive BER at best alignment
PanakoMatcher 1:1 ~264 µs 2-D Hough + RANSAC line-fitting
WangIndex N=100 query ~102 µs Inverted index + sliding-window peak

Latency budget (per query, default configs, Intel i5-1135G7):

Catalog size WangIndex query Throughput
100 tracks ~102 µs ~9 800 q/s
1 000 tracks ~1 ms (est.) ~1 000 q/s
10 000 tracks ~10 ms (est.) ~100 q/s

Index query scales approximately linearly with catalog size (one candidate-scoring pass per reference with hash hits). For catalogs above ~10 000 tracks, use min_votes / min_score pre-filters or shard the index.

Run benchmarks for your own host:

cargo bench --bench extract
cargo bench --bench streaming
cargo bench --bench extract -- --save-baseline main   # save for diffing later

Robustness

  • Codec-tolerant by design — Wang and Panako are spectral-peak based; Haitsma is band-power-difference based. All three survive lossy re-encoding, verified by the test suite on real music:

    Codec Wang (Jaccard) Panako (Jaccard) Haitsma (bit-sim)
    WAV/FLAC (lossless) 1.000 1.000
    MP3 128 kbps 0.40 0.45 0.93
    OGG-Vorbis 0.36 0.42 0.91
    AAC (M4A) 0.50 0.54 0.77
    AIFF (lossless) 1.000
    Cross-track (different song) 0.001

    Test audio: "Galway" and "Furious Freak" by Kevin MacLeod, 16 s each, 6 codec variants. Thresholds: Wang ≥ 0.25, Panako ≥ 0.20, Haitsma ≥ 0.75. In practice, 5–10 matching hashes suffice for confident identification.

  • Two-track discrimination verified — different songs produce <0.1% hash overlap (random collision floor), while the same song across codecs produces 25–80% overlap.

  • 606 tests including adversarial stress tests, real-audio E2E across 6 codecs, and property-based streaming/offline parity checks. See ROBUSTNESS.md for full methodology.

Comparison with Alternatives

Feature audiofp chromaprint-rust dejavu (Python)
Pure Rust Yes No (FFI to C lib) No
Wang landmarks Yes No Yes
Panako triplets (tempo-robust) Yes No No
Haitsma–Kalker Yes No No
Streaming variants Yes Limited No
Bit-exact streaming/offline parity Yes No N/A
File decoding included Yes (Symphonia) Yes (limited) Yes (FFmpeg)
Watermark detection Yes (AudioSeal) No No
no_std + alloc capable Yes (host) No N/A
bytemuck::Pod hash types Yes No N/A
Built-in resampler Yes No No
In-memory matcher (Wang/Haitsma) Yes No Yes (Dejavu)

Examples

The examples/ directory contains complete working programs that can be run with cargo run --example <name>:

  • enroll_file — fingerprint a single audio file and print the unique Wang landmark count (--features default / std).
  • match_two_files — print the number of Wang hash collisions between two files (the canonical "is this the same recording?" check).
  • compare_algorithms — run Wang, Panako, and Haitsma–Kalker over the same file and report per-algorithm timing and hash counts.
  • stream_buffer — feed Wang's streaming fingerprinter from an io::Read chunk-by-chunk.
  • dsp_starter — STFT → mel → peaks pipeline on synthetic audio (no file, no optional features).
  • neural_embed — load a BYO ONNX embedder and print embedding dim (--features neural).
  • watermark_detect — load an AudioSeal-compatible ONNX model and print confidence (--features watermark).
cargo run --example dsp_starter
cargo run --example neural_embed --features neural -- path/to/model.onnx
cargo run --example watermark_detect --features watermark,std-wav -- path/to/audioseal.onnx [audio.wav]

The doctests across the public API and USAGE.md cover the full surface for users wiring audiofp into their own binary.

Security

See SECURITY.md for the threat model (audio / PCM / ONNX / hash outputs) and how to report vulnerabilities privately. Fingerprints are perceptual, not cryptographic MACs — use DecodeLimits with decode_to_mono_limited for untrusted uploads.

Contributing

See CONTRIBUTING.md for guidelines. Quick start:

git clone https://github.com/themankindproject/audiofp && cd audiofp
cargo test --all-features
cargo clippy --all-targets --all-features -- -D warnings
cargo fmt --all -- --check

CI runs fmt, clippy, and test on ubuntu/macOS/Windows on every push and PR.

License

MIT License — see LICENSE for details.

References

  • Avery Wang, An Industrial-Strength Audio Search Algorithm (ISMIR 2003) — Wang landmarks
  • Joren Six & Marc Leman, Panako: A Scalable Acoustic Fingerprinting System (ISMIR 2014); 2021 update — triplet β hash
  • Jaap Haitsma & Ton Kalker, A Highly Robust Audio Fingerprinting System (ISMIR 2002) — band-power sign bits
  • San Roman, R., Fernandez, P., Elsahar, H., Défossez, A., Furon, T. & Tran, T. Proactive Detection of Voice Cloning with Localized Watermarking. arXiv:2401.17264, 2024 (AudioSeal) — watermark model. https://arxiv.org/abs/2401.17264