pub struct AsrOptions {
pub language: Option<String>,
pub word_timestamps: bool,
pub diarize: bool,
pub persist_speakers: bool,
pub max_speakers: Option<usize>,
pub diarize_threshold: f32,
pub translate: bool,
pub vad: bool,
pub vad_threshold: f32,
pub vad_chunk_secs: f32,
pub stream_offset_secs: f64,
}Fields§
§language: Option<String>Force a language instead of auto-detecting.
word_timestamps: boolWord-level timestamps (WhisperX-style forced alignment).
diarize: boolSpeaker diarization (WhisperX-style).
persist_speakers: boolKeep speaker identities across calls, so a voice heard in one chunk keeps its label in the next.
Off by default, and the default is the batch behaviour every
diarization system has: labels are arbitrary names for clusters within
ONE call, and SPEAKER_00 in two separate calls need not be the same
person. That is fine for a file and useless for a stream.
With this set, the engine keeps a speaker registry between calls. Call
WhisperCandle::reset_speakers() when a new recording begins — a new
session is a new set of people, and carrying identities across is
worse than starting fresh.
Matching is deliberately stricter than in-call clustering: a registry merge is permanent, and two people who share a centroid stay merged for the rest of the session.
max_speakers: Option<usize>Known speaker count, when the caller has one (“this is an interview, two people”). Overrides the clustering threshold.
Measured caution: with the threshold tuned this is not the safer choice. Blind clustering scores 4.21 % DER against 5.00 % with the true count supplied, because forcing a count forces a merge, and a bad merge attributes one speaker’s words to another. Set it when the count is certain, not as insurance.
diarize_threshold: f32Cosine-distance threshold for merging speaker clusters.
Swept against DER on a 6-conversation corpus: the minimum sits at
0.85 (2.71 %), and 0.80 (4.21 %) ships instead because over-merging
fails catastrophically (44.7 % at 0.95) while over-splitting fails
gently. See ffai_mercury::asr::diarize::DEFAULT_THRESHOLD.
translate: boolTranslate to English instead of transcribing.
vad: boolSegment on speech before transcribing, so silence never reaches the model.
On by default, for measured speed — not for quality.
- Audio with trailing silence: 2.2–4.2× faster, transcript byte-identical.
- Silent input: empty transcript, with no encoder pass at all.
- A live sliding window stops spending five encoder passes to produce nothing.
Corpus WER does move with this on (test-clean 7.99 → 6.79,
test-other 16.79 → 16.43), and that is not a quality improvement —
do not cite it as one. Per-clip decomposition over 400 clips gives 38
improved and 38 worsened, a sign test of z = 0.00. VAD shifts where
speech sits inside Whisper’s fixed 30 s context by ~0.2 s, which
re-rolls the decode on about a fifth of clips, half each way; the
aggregate moved because WER is dominated by a few high-delta clips.
Full descent: docs/whys/vad-quality.md.
Set false for the unsegmented fixed-30 s-grid behaviour.
vad_threshold: f32Speech threshold, 0..1, higher being stricter. Only read when
Self::vad is set.
vad_chunk_secs: f32Pack speech regions into windows of at most this many seconds.
stream_offset_secs: f64Where this buffer starts in the wider stream, in seconds.
Only meaningful for a streaming caller that re-sends a sliding window
(a live transcriber sending the trailing N seconds every tick). It
costs nothing to leave at 0.0.
What it buys. Diarization sub-segments each speech region into
1.5 s windows and embeds each one — the dominant cost, ~172 ms apiece.
Those windows are placed relative to the region, and a region clipped
by the buffer’s leading edge is anchored to the buffer, which moves.
So consecutive ticks re-cut the same audio at shifted offsets and every
embedding is recomputed. Measured on a 10 s window at a 1 s tick: the
window grids realign only every 3 s (lcm(1.0, 0.75)), and the cache
hit rate sat at ~24 %.
Given this, windows are placed on an ABSOLUTE grid, so the same audio yields the same window bounds no matter where the buffer happens to start — which is what makes the embedding cache actually hit.
Trait Implementations§
Source§impl Clone for AsrOptions
impl Clone for AsrOptions
Source§fn clone(&self) -> AsrOptions
fn clone(&self) -> AsrOptions
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read moreSource§impl Debug for AsrOptions
impl Debug for AsrOptions
Auto Trait Implementations§
impl Freeze for AsrOptions
impl RefUnwindSafe for AsrOptions
impl Send for AsrOptions
impl Sync for AsrOptions
impl Unpin for AsrOptions
impl UnsafeUnpin for AsrOptions
impl UnwindSafe for AsrOptions
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> ErasedDestructor for Twhere
T: 'static,
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more