Skip to main content

AsrOptions

Struct AsrOptions 

Source
pub struct AsrOptions {
    pub language: Option<String>,
    pub word_timestamps: bool,
    pub diarize: bool,
    pub persist_speakers: bool,
    pub max_speakers: Option<usize>,
    pub diarize_threshold: f32,
    pub translate: bool,
    pub vad: bool,
    pub vad_threshold: f32,
    pub vad_chunk_secs: f32,
    pub stream_offset_secs: f64,
}

Fields§

§language: Option<String>

Force a language instead of auto-detecting.

§word_timestamps: bool

Word-level timestamps (WhisperX-style forced alignment).

§diarize: bool

Speaker diarization (WhisperX-style).

§persist_speakers: bool

Keep speaker identities across calls, so a voice heard in one chunk keeps its label in the next.

Off by default, and the default is the batch behaviour every diarization system has: labels are arbitrary names for clusters within ONE call, and SPEAKER_00 in two separate calls need not be the same person. That is fine for a file and useless for a stream.

With this set, the engine keeps a speaker registry between calls. Call WhisperCandle::reset_speakers() when a new recording begins — a new session is a new set of people, and carrying identities across is worse than starting fresh.

Matching is deliberately stricter than in-call clustering: a registry merge is permanent, and two people who share a centroid stay merged for the rest of the session.

§max_speakers: Option<usize>

Known speaker count, when the caller has one (“this is an interview, two people”). Overrides the clustering threshold.

Measured caution: with the threshold tuned this is not the safer choice. Blind clustering scores 4.21 % DER against 5.00 % with the true count supplied, because forcing a count forces a merge, and a bad merge attributes one speaker’s words to another. Set it when the count is certain, not as insurance.

§diarize_threshold: f32

Cosine-distance threshold for merging speaker clusters.

Swept against DER on a 6-conversation corpus: the minimum sits at 0.85 (2.71 %), and 0.80 (4.21 %) ships instead because over-merging fails catastrophically (44.7 % at 0.95) while over-splitting fails gently. See ffai_mercury::asr::diarize::DEFAULT_THRESHOLD.

§translate: bool

Translate to English instead of transcribing.

§vad: bool

Segment on speech before transcribing, so silence never reaches the model.

On by default, for measured speed — not for quality.

  • Audio with trailing silence: 2.2–4.2× faster, transcript byte-identical.
  • Silent input: empty transcript, with no encoder pass at all.
  • A live sliding window stops spending five encoder passes to produce nothing.

Corpus WER does move with this on (test-clean 7.99 → 6.79, test-other 16.79 → 16.43), and that is not a quality improvement — do not cite it as one. Per-clip decomposition over 400 clips gives 38 improved and 38 worsened, a sign test of z = 0.00. VAD shifts where speech sits inside Whisper’s fixed 30 s context by ~0.2 s, which re-rolls the decode on about a fifth of clips, half each way; the aggregate moved because WER is dominated by a few high-delta clips. Full descent: docs/whys/vad-quality.md.

Set false for the unsegmented fixed-30 s-grid behaviour.

§vad_threshold: f32

Speech threshold, 0..1, higher being stricter. Only read when Self::vad is set.

§vad_chunk_secs: f32

Pack speech regions into windows of at most this many seconds.

§stream_offset_secs: f64

Where this buffer starts in the wider stream, in seconds.

Only meaningful for a streaming caller that re-sends a sliding window (a live transcriber sending the trailing N seconds every tick). It costs nothing to leave at 0.0.

What it buys. Diarization sub-segments each speech region into 1.5 s windows and embeds each one — the dominant cost, ~172 ms apiece. Those windows are placed relative to the region, and a region clipped by the buffer’s leading edge is anchored to the buffer, which moves. So consecutive ticks re-cut the same audio at shifted offsets and every embedding is recomputed. Measured on a 10 s window at a 1 s tick: the window grids realign only every 3 s (lcm(1.0, 0.75)), and the cache hit rate sat at ~24 %.

Given this, windows are placed on an ABSOLUTE grid, so the same audio yields the same window bounds no matter where the buffer happens to start — which is what makes the embedding cache actually hit.

Trait Implementations§

Source§

impl Clone for AsrOptions

Source§

fn clone(&self) -> AsrOptions

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Debug for AsrOptions

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more
Source§

impl Default for AsrOptions

Source§

fn default() -> Self

Returns the “default value” for a type. Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> ErasedDestructor for T
where T: 'static,

Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> IntoEither for T

Source§

fn into_either(self, into_left: bool) -> Either<Self, Self>

Converts self into a Left variant of Either<Self, Self> if into_left is true. Converts self into a Right variant of Either<Self, Self> otherwise. Read more
Source§

fn into_either_with<F>(self, into_left: F) -> Either<Self, Self>
where F: FnOnce(&Self) -> bool,

Converts self into a Left variant of Either<Self, Self> if into_left(&self) returns true. Converts self into a Right variant of Either<Self, Self> otherwise. Read more
Source§

impl<T> Pointable for T

Source§

const ALIGN: usize

The alignment of pointer.
Source§

type Init = T

The type for initializers.
Source§

unsafe fn init(init: <T as Pointable>::Init) -> usize

Initializes a with the given initializer. Read more
Source§

unsafe fn deref<'a>(ptr: usize) -> &'a T

Dereferences the given pointer. Read more
Source§

unsafe fn deref_mut<'a>(ptr: usize) -> &'a mut T

Mutably dereferences the given pointer. Read more
Source§

unsafe fn drop(ptr: usize)

Drops the object pointed to by the given pointer. Read more
Source§

impl<T> Read<Exclusive, BecauseExclusive> for T
where T: ?Sized,

Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, <T as TryFrom<U>>::Error>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<V, T> VZip<V> for T
where V: MultiLane<T>,

Source§

fn vzip(self) -> V