Skip to main content

OpusEncoder

Struct OpusEncoder 

Source
pub struct OpusEncoder {
    pub bitrate_bps: i32,
    pub complexity: i32,
    pub rate_control: RateControl,
    pub use_inband_fec: bool,
    pub use_dtx: bool,
    pub packet_loss_perc: i32,
    pub force_bandwidth: Option<Bandwidth>,
    pub signal_type: Option<Signal>,
    pub max_bandwidth: Bandwidth,
    pub lsb_depth: i32,
    /* private fields */
}
Expand description

An Opus encoder: PCM in, Opus packets out.

One encoder handles one stream, and carries state between packets, so the same instance has to be fed the whole stream in order. The settings below are public fields rather than setters, and every one of them may be changed between packets; only the sample rate, channel count and Application are fixed at construction.

The encoder decides per packet which of the three Opus layers to use, what audio Bandwidth to code, and how many bits to spend. Those decisions are what bitrate_bps, complexity and the rest steer.

use opus_pure::{Application, OpusEncoder};

let mut encoder = OpusEncoder::new(48_000, 2, Application::Audio)?;
encoder.bitrate_bps = 96_000;

let pcm = vec![0.0f32; 960 * 2];        // 20 ms of stereo
let mut packet = vec![0u8; 4000];
let n = encoder.encode(&pcm, 960, &mut packet)?;
assert!(n > 0);

Fields§

§bitrate_bps: i32

Target bitrate in bits per second, across all channels. Default 64000.

This is a target rather than a cap: the default is variable-rate, so an individual packet is as large as its content needs and the rate is met on average. See rate_control to make it a per-packet size instead.

It is also the single strongest input to the encoder’s own decisions. Coding mode and audio bandwidth are both chosen from it, so lowering it does not simply degrade the same signal: below roughly 20 kb/s the encoder moves to SILK and narrows the bandwidth, because spending the remaining bits on a smaller spectrum sounds better than spreading them over all of it.

§complexity: i32

How much CPU the encoder may spend, from 0 to 10. Default 9.

Lower settings take shortcuts in pitch analysis and quantisation, and below 7 the content analysis is skipped entirely, which is what otherwise informs the speech/music decision. It is not purely a speed control: the encoder scales its own idea of the bitrate by (90 + complexity) / 100 when choosing a mode and bandwidth, so a lower complexity also codes a narrower band at the same rate.

§rate_control: RateControl

How much the size of each packet may vary. Default ConstrainedVbr, which is libopus’s.

Cbr pads every packet to the same size, which is what a fixed-capacity channel wants and what a file does not: it spends bits on silence that VBR would have given to the difficult passages. It also costs quality at a given rate, and the encoder accounts for that by discounting its working bitrate by a twelfth when deciding mode and bandwidth.

Note this is the opposite polarity from libopus’s OPUS_SET_VBR: the default here is named for what it does rather than for what it is not.

§use_inband_fec: bool

Code a low-bitrate copy of the previous frame into each packet, so one lost packet can be partly recovered from the next. Default false.

The redundant copy costs bits that would otherwise go to the current frame, so this is only worth enabling when loss is actually expected, and the encoder decides per packet whether to spend them, using packet_loss_perc as its estimate of how likely that is. FEC needs SILK, so it has no effect on packets the encoder codes as CELT.

A decoder recovers the copy with OpusDecoder::decode_fec, which must be called on the packet after the missing one.

§use_dtx: bool

Discontinuous transmission: after enough consecutive inactive frames, emit a 1-byte (TOC-only) packet so the decoder runs comfort-noise/PLC.

§packet_loss_perc: i32

Expected packet loss, 0 to 100 percent. Default 0.

This is what tells the encoder how defensively to code. A non-zero value makes the quantiser less reliant on inter-frame prediction, so a lost packet corrupts less of what follows, and it is the input use_inband_fec uses to decide whether a redundant copy is worth its bits. Both cost quality on the packets that do arrive, so an estimate far above the real loss rate is not a safe default.

§force_bandwidth: Option<Bandwidth>

Overrides automatic bandwidth selection when set (OPUS_SET_BANDWIDTH).

§signal_type: Option<Signal>

OPUS_SET_SIGNAL: force the voice/music bias (None = auto from analysis).

§max_bandwidth: Bandwidth

OPUS_SET_MAX_BANDWIDTH: cap the automatically-selected bandwidth.

§lsb_depth: i32

Input bit depth assumed by the analysis noise floors. The float API default is 24; set 16 for s16-sourced content (opus_demo parity).

Implementations§

Source§

impl OpusEncoder

Source

pub fn new( sampling_rate: i32, channels: usize, application: Application, ) -> Result<Self>

Create an encoder for sampling_rate Hz and channels channels.

The rate must be one of 8000, 12000, 16000, 24000 or 48000, and the channel count 1 or 2; anything else is Error::InvalidArgument. These three arguments are the only settings fixed for the encoder’s life. See Application for which one to pass, and the fields on this type for everything that can be changed afterwards.

The rate is the rate of the PCM handed to encode, not a property of the packets produced: Opus always codes internally at one of its own rates and every packet’s duration is counted at 48 kHz regardless. Passing 48000 avoids a resampling step on the way in.

For more than two channels, see OpusMSEncoder.

Examples found in repository?
examples/encode.rs (line 42)
18fn main() -> Result<(), Box<dyn std::error::Error>> {
19    let args: Vec<String> = std::env::args().collect();
20    if args.len() < 3 {
21        eprintln!("usage: {} <input.wav> <output.opus> [bitrate_bps]", args[0]);
22        std::process::exit(2);
23    }
24    let bitrate: i32 = args
25        .get(3)
26        .map(|s| s.parse())
27        .transpose()?
28        .unwrap_or(64_000);
29
30    let input = wav::read(&args[1])?;
31    if input.samples.is_empty() {
32        return Err(format!("{}: no audio samples", args[1]).into());
33    }
34    let channels = input.channels as usize;
35    let rate = input.sample_rate as i32;
36
37    // Opus codes 20 ms frames at the encoder's own rate. The container counts
38    // everything at 48 kHz regardless, and `write_packet` reads that out of the
39    // packet rather than making us convert it.
40    let frame = (rate / 50) as usize;
41
42    let mut encoder = OpusEncoder::new(rate, channels, Application::Audio)?;
43    encoder.bitrate_bps = bitrate;
44
45    let mut tags = OpusTags::new();
46    tags.push("ENCODER", concat!("opus-pure ", env!("CARGO_PKG_VERSION")))?;
47    // Built from the encoder, so the pre-skip is that encoder's real delay
48    // rather than the conventional constant.
49    let head = OpusHead::for_encoder(&encoder, input.sample_rate);
50
51    // ---- Ending the file exactly (RFC 7845 §4.2 and §4.4) ----
52    //
53    // Every Opus rate divides 48 kHz, so one encoder-rate sample is this many
54    // granule ticks and the conversions below are exact.
55    let ticks = 48_000 / rate as usize;
56    let total = input.samples.len() / channels; // sample frames of real audio
57    //
58    // 1. The encoder runs `pre_skip` samples behind its input, so the last
59    //    `pre_skip` samples of the audio are still inside it when the input
60    //    runs out. Feeding that much extra silence is what flushes them; stop
61    //    at the audio and the tail is simply lost. Then round up to a whole
62    //    frame, because Opus has no partial ones.
63    let lookahead = (head.pre_skip as usize).div_ceil(ticks);
64    let frames = (total + lookahead).div_ceil(frame);
65    //
66    // 2. That padding now decodes as real output, so the file has to say where
67    //    the audio stopped. The final granule position is the pre-skip plus the
68    //    audio and nothing else; a player trims back to it. The last packet
69    //    carries the difference between that and what the writer has already
70    //    counted, which is always between 1 and one frame's worth.
71    let final_granule = u64::from(head.pre_skip) + (total * ticks) as u64;
72
73    let file = std::fs::File::create(&args[2])?;
74    let mut writer = OggOpusWriter::with_tags(std::io::BufWriter::new(file), head, tags)?;
75
76    let mut packet = vec![0u8; MAX_PACKET_BYTES];
77    let mut payload_bytes = 0usize;
78    let per_frame = frame * channels;
79    let mut block = vec![0.0f32; per_frame];
80
81    for i in 0..frames {
82        // Whole frames only: past the end of the input the block is silence,
83        // which is the padding that flushes the encoder's delay.
84        let start = (i * per_frame).min(input.samples.len());
85        let end = (start + per_frame).min(input.samples.len());
86        block[..end - start].copy_from_slice(&input.samples[start..end]);
87        block[end - start..].fill(0.0);
88
89        let n = encoder.encode(&block, frame, &mut packet)?;
90        if i + 1 == frames {
91            let duration = final_granule - writer.granule() as u64;
92            writer.write_packet_with_duration(&packet[..n], duration as u32)?;
93        } else {
94            writer.write_packet(&packet[..n])?;
95        }
96        payload_bytes += n;
97    }
98    writer.finish()?;
99
100    let secs = total as f64 / rate as f64;
101    println!(
102        "{} -> {}\n  {channels} ch @ {rate} Hz, {frames} frames ({secs:.2} s of audio)\n  \
103         {payload_bytes} payload bytes = {:.1} kb/s (target {:.1} kb/s)",
104        args[1],
105        args[2],
106        payload_bytes as f64 * 8.0 / secs / 1000.0,
107        bitrate as f64 / 1000.0,
108    );
109    Ok(())
110}
Source

pub fn final_range(&self) -> u32

Final range-coder state of the last encoded packet (libopus OPUS_GET_FINAL_RANGE). Stored in opus_demo .bit framing so the reference decoder can verify encoder/decoder range-coder agreement per packet.

Source

pub fn sample_rate(&self) -> i32

The sample rate this encoder was created with, in Hz.

Source

pub fn channels(&self) -> usize

The channel count this encoder was created with.

Source

pub fn application(&self) -> Application

The Application this encoder was created with.

Source

pub fn lookahead(&self) -> usize

Samples per channel of algorithmic delay, at this encoder’s sample rate (libopus OPUS_GET_LOOKAHEAD).

The encoder’s output trails its input by this much, so a decoder should discard this many samples from the start to line the two up again. It is not a constant: Application::RestrictedLowDelay gives up the 4 ms the other two spend keeping SILK and CELT aligned, so it is 120 samples at 48 kHz where Audio and Voip are 312.

For an Ogg stream this is what OpusHead::pre_skip must carry, expressed at 48 kHz — which OpusHead::for_encoder does for you, and is the reason to prefer it over OpusHead::new.

Source

pub fn reset_state(&mut self) -> Result<()>

Discard everything the encoder has learned, keeping its settings.

Equivalent to building a new encoder with the same sample rate, channel count and Application, then re-applying every setting you had changed — which is what it does. Use it to encode an unrelated second stream through the same instance, so that nothing from the first one (the filter histories, the bandwidth hysteresis, the speech/music decision) carries across and colours its opening frames.

This is libopus’s OPUS_RESET_STATE. Note it re-initialises the coding state rather than merely rewinding it, so it is not free; it is cheaper and far less error-prone than trying to keep a used encoder honest by hand, but a caller counting allocations should know it makes some.

Source

pub fn encode( &mut self, input: &[f32], frame_size: usize, output: &mut [u8], ) -> Result<usize>

Encode frame_size samples per channel into one Opus packet, returning how many bytes of output it filled.

input is interleaved and must hold frame_size * channels samples. frame_size is one of the nine durations Opus defines — 2.5, 5, 10, 20, 40, 60, 80, 100 or 120 ms at this encoder’s sample rate — and anything else is an InvalidArgument. Whether the packet ends up holding one coded frame or several is the encoder’s to decide: only SILK has configurations past 20 ms, so a longer packet in any other mode is several frames sharing one TOC byte. A caller sees the difference only in the packet’s framing, never in its duration.

output.len() is the packet’s byte budget, not merely a capacity. This is libopus’s max_data_bytes, and it means a short buffer does not produce an error — the encoder codes a smaller packet to fit it, and the stream quietly comes out under the bitrate you asked for. Pass MAX_PACKET_BYTES unless you are deliberately capping the instantaneous rate, for instance to a network MTU. The one size that is refused outright is a buffer under two bytes, which cannot hold any packet at all.

Samples outside ±1 are coded rather than rejected, so a caller working in float can drive the encoder past full scale. If your source is integer PCM, prefer encode_s16, which is the same encoder told the truth about its input’s precision.

Examples found in repository?
examples/encode.rs (line 89)
18fn main() -> Result<(), Box<dyn std::error::Error>> {
19    let args: Vec<String> = std::env::args().collect();
20    if args.len() < 3 {
21        eprintln!("usage: {} <input.wav> <output.opus> [bitrate_bps]", args[0]);
22        std::process::exit(2);
23    }
24    let bitrate: i32 = args
25        .get(3)
26        .map(|s| s.parse())
27        .transpose()?
28        .unwrap_or(64_000);
29
30    let input = wav::read(&args[1])?;
31    if input.samples.is_empty() {
32        return Err(format!("{}: no audio samples", args[1]).into());
33    }
34    let channels = input.channels as usize;
35    let rate = input.sample_rate as i32;
36
37    // Opus codes 20 ms frames at the encoder's own rate. The container counts
38    // everything at 48 kHz regardless, and `write_packet` reads that out of the
39    // packet rather than making us convert it.
40    let frame = (rate / 50) as usize;
41
42    let mut encoder = OpusEncoder::new(rate, channels, Application::Audio)?;
43    encoder.bitrate_bps = bitrate;
44
45    let mut tags = OpusTags::new();
46    tags.push("ENCODER", concat!("opus-pure ", env!("CARGO_PKG_VERSION")))?;
47    // Built from the encoder, so the pre-skip is that encoder's real delay
48    // rather than the conventional constant.
49    let head = OpusHead::for_encoder(&encoder, input.sample_rate);
50
51    // ---- Ending the file exactly (RFC 7845 §4.2 and §4.4) ----
52    //
53    // Every Opus rate divides 48 kHz, so one encoder-rate sample is this many
54    // granule ticks and the conversions below are exact.
55    let ticks = 48_000 / rate as usize;
56    let total = input.samples.len() / channels; // sample frames of real audio
57    //
58    // 1. The encoder runs `pre_skip` samples behind its input, so the last
59    //    `pre_skip` samples of the audio are still inside it when the input
60    //    runs out. Feeding that much extra silence is what flushes them; stop
61    //    at the audio and the tail is simply lost. Then round up to a whole
62    //    frame, because Opus has no partial ones.
63    let lookahead = (head.pre_skip as usize).div_ceil(ticks);
64    let frames = (total + lookahead).div_ceil(frame);
65    //
66    // 2. That padding now decodes as real output, so the file has to say where
67    //    the audio stopped. The final granule position is the pre-skip plus the
68    //    audio and nothing else; a player trims back to it. The last packet
69    //    carries the difference between that and what the writer has already
70    //    counted, which is always between 1 and one frame's worth.
71    let final_granule = u64::from(head.pre_skip) + (total * ticks) as u64;
72
73    let file = std::fs::File::create(&args[2])?;
74    let mut writer = OggOpusWriter::with_tags(std::io::BufWriter::new(file), head, tags)?;
75
76    let mut packet = vec![0u8; MAX_PACKET_BYTES];
77    let mut payload_bytes = 0usize;
78    let per_frame = frame * channels;
79    let mut block = vec![0.0f32; per_frame];
80
81    for i in 0..frames {
82        // Whole frames only: past the end of the input the block is silence,
83        // which is the padding that flushes the encoder's delay.
84        let start = (i * per_frame).min(input.samples.len());
85        let end = (start + per_frame).min(input.samples.len());
86        block[..end - start].copy_from_slice(&input.samples[start..end]);
87        block[end - start..].fill(0.0);
88
89        let n = encoder.encode(&block, frame, &mut packet)?;
90        if i + 1 == frames {
91            let duration = final_granule - writer.granule() as u64;
92            writer.write_packet_with_duration(&packet[..n], duration as u32)?;
93        } else {
94            writer.write_packet(&packet[..n])?;
95        }
96        payload_bytes += n;
97    }
98    writer.finish()?;
99
100    let secs = total as f64 / rate as f64;
101    println!(
102        "{} -> {}\n  {channels} ch @ {rate} Hz, {frames} frames ({secs:.2} s of audio)\n  \
103         {payload_bytes} payload bytes = {:.1} kb/s (target {:.1} kb/s)",
104        args[1],
105        args[2],
106        payload_bytes as f64 * 8.0 / secs / 1000.0,
107        bitrate as f64 / 1000.0,
108    );
109    Ok(())
110}
Source

pub fn encode_s16( &mut self, input: &[i16], frame_size: usize, output: &mut [u8], ) -> Result<usize>

Encode frame_size samples per channel of 16-bit PCM into one packet, returning how many bytes of output it filled.

The same encoder as encode in every respect but one: it knows the input came from 16 bits. That matters because the encoder treats anything below the source’s own noise floor as digital silence and codes it as such, and the floor sits 2^-depth from full scale. Told 24 bits when the input has 16, it holds detail no 16-bit source could carry and spends bits on the dither in the bottom bits; told 16, it drops out where the source does. libopus draws the same distinction between opus_encode and opus_encode_float, and this matches it, including honouring a lower lsb_depth if the caller has set one.

Conversion is sample / 32768, which is exact — the scale is a power of two — so this differs from converting by hand and calling encode only in the depth, never in the samples.

let mut encoder = OpusEncoder::new(48_000, 2, Application::Audio)?;
let pcm = vec![0i16; 960 * 2];               // 20 ms of stereo at 48 kHz
let mut packet = vec![0u8; 4000];
let n = encoder.encode_s16(&pcm, 960, &mut packet)?;

Trait Implementations§

Source§

impl Debug for OpusEncoder

Shows the encoder’s configuration and omits its coding state.

A derived Debug here would print every filter history and analysis buffer the encoder carries, which is tens of kilobytes of numbers that mean nothing without the codec beside them. What a caller wants from dbg! is the settings, so that is what this prints; the .. stands for the rest.

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, !>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.