Skip to main content

Module parallel

Module parallel 

Source
Expand description

Chunk-parallel Opus encoding: split the input into contiguous frame ranges and encode them on separate threads.

Opus carries real inter-frame state — SILK’s LTP, noise shaping, NLSF interpolation and bit reservoir, CELT’s pre-emphasis, overlap, prefilter and energy prediction, the high-pass filter, the input resampler, and the content analysis that chooses between them — so a frame range cannot be encoded from a cold encoder and dropped into the middle of a stream. Each worker therefore primes its encoder by re-encoding the audio immediately before its chunk and discarding those packets, so that by the time it reaches its own first frame its state approximates the state a continuous encoder would have had.

Priming is neither free nor exact, and what it costs and buys is the whole subject of this module.

§What priming converges, and what it does not

The signal path converges quickly. Its memory is a handful of frames — the LTP lag, the noise-shaping delay, one MDCT overlap — and a stable filter forgets its initial conditions. Around 160 ms of priming settles it.

The content analysis does not. It keeps a hundred-entry ring of 20 ms observations and averages the music/speech probability over it, and the encoder applies hysteresis on top of that when it picks between SILK, hybrid and CELT. That memory is two seconds deep, and until it fills, a worker’s mode decision is its own rather than the one a continuous encoder would have made. The consequence is not a seam at the boundary: it is the entire chunk coded in a different mode.

Measured on 120 s of synthetic speech at 16 kb/s split four ways, where a continuous encoder settles on CELT and stays there:

warmup_msworst chunk, frames in a mode the serial encoder did not useworst frame vs serial
16036 of 1500−14.82 dB
50016 of 1500−14.53 dB
10000 of 1500−4.10 dB
2000 (the default)0 of 1500−4.10 dB

Hence DEFAULT_WARMUP_MS, which is the analysis’s own history depth rather than a tuned number. A caller who pins ParallelConfig::signal_type takes the analysis out of the mode decision entirely and can prime far less; that is the cheapest way to buy a short warm-up.

§What is left after priming

Even fully primed, a chunked encode is not the serial encode. The rate controllers — CELT’s VBR reservoir and drift, SILK’s bit reservoir — are deliberately long-memory integrators, and a worker’s differs from the continuous encoder’s for some frames after its boundary. That shows as a bitrate dip of a few percent lasting tens of frames, and as a per-frame SNR difference at the boundary of a few dB against a serial encode of the same audio. Constant bitrate removes nearly all of it, there being no reservoir to be wrong about.

This part is inherent rather than a defect awaiting a fix: at a chunk boundary one packet was produced by an encoder that did not produce the packet before it, and no amount of priming changes that. It is why this is an opt-in path and not what crate::OpusEncoder does by itself.

reference/parallel/ measures all of the above, and is where the numbers here come from.

§Cost

Every worker but the first re-encodes warmup_ms of audio it will not emit. With w workers that is (w - 1) * warmup_ms of redundant encoding, and ParallelConfig::plan reports it before any of it is done. The worker count is capped so redundancy stays at or below a quarter of the useful work, which with the default warm-up means one worker per 8 s of audio.

Deterministic: fixed chunk boundaries mean identical output across runs. Uses only std::thread.

Structs§

ParallelConfig
Configuration for a parallel encode.
ParallelPlan
How a parallel encode will be divided up, from ParallelConfig::plan.

Constants§

DEFAULT_WARMUP_MS
Default priming length, in milliseconds.

Functions§

encode_parallel
Encode pcm (interleaved f32, channels-interleaved) in frame_size samples-per-channel frames across several threads, returning one Opus packet per frame in order.