dioxus-audio
Audio recording, playback, analysis, and UI components for Dioxus
Features
- Recording: microphone permissions, capture lifecycle, ordered Recording Chunks, elapsed time, peaks, and live analysis.
- Playback: Audio Data and ordered URL-addressable Playback Source alternatives, eager or on-play loading, playback lifecycle, stop/reset, whole-source repeat, network and readiness observations, buffered and seekable ranges, seeking, skipping, playback rate, and opt-in pre-gain Analysis with effective graph gain.
- Audio input devices: enumeration, selection, permission requests, and device-change handling.
- Analysis: bounded reactive snapshots, interpretable waveform and spectrum data, RMS levels, peak reduction, and PCM range trimming.
- Decoding: complete Audio Data to immutable channel-preserving planar samples with an explicit Rust-copy memory ceiling.
- Dioxus components: player and recorder controls, scrubber, input selector, microphone status, waveform views, spectrum, and level meter.
- Scoped styles: authored CSS with a
dioxus-audionamespace and stable package variables, without a Tailwind or daisyUI build dependency.
Quick Start
[]
= { = "0.x" }
Styles
The crate ships authored, namespace-scoped CSS. Load it once near the application root:
use *;
use AudioStyles;
Stable package custom properties inherit, so applications can set them on an ancestor for app-wide styling or on a local wrapper for per-instance styling. When values are omitted, components can follow an installed daisyUI theme and otherwise use standalone defaults; daisyUI is optional.
See the Style Customization Guide for complete recipes and the public token reference.
WaveformPreview normalizes each non-empty waveform against its own loudest
peak (with a quiet-signal floor), so compact previews remain legible. It is a
shape preview, not an absolute loudness comparison between recordings.
Recording
use *;
use ;
use use_audio_input_devices;
use ;
Call recorder.completed() to react to a finished RecordedAudio, then call
recorder.clear_completed() only after consuming it successfully. Keeping the
value until persistence succeeds lets an application retry without losing the
captured bytes. Consumers that maintain their own retry queue can use
recorder.take_completed() to move the recording out without cloning its
audio buffer.
For custom layouts, pass the same Recorder to RecorderStartButton,
RecorderCancelButton, RecorderPauseResumeButton, RecorderStopButton, and
RecorderClearButton. Each native button exposes command validity through its
disabled state and accepts application-specific labels. Mount
RecorderStatusAnnouncer when the application needs polite, coarse lifecycle
announcements; it does not announce elapsed time or Analysis updates.
Microphone capture requires a secure browser context. HTTPS, localhost, and
127.0.0.1 are normally accepted by browsers.
Input selection is snapshotted by recorder.start(). Let users choose a device
before starting capture; changing devices.selected() does not switch an
active recording.
Browser applications that already own a media stream can wrap it as an opaque
RecordingSource and start without another device acquisition:
use ;
let source = from_media_stream;
recorder
.start_with_source
.expect;
let disposable_source = from_media_stream
.with_shutdown;
The application retains the raw stream if it needs to use or stop it later;
Recorder does not expose that browser resource through its Controller. An
accepted supplied source must contain exactly one live audio track. Video tracks
are ignored, and a live audio track remains valid when browser-muted or
application-disabled. Recorder creates an audio-only recording view. The default
PreserveTracks agreement never calls stop() on the supplied track.
StopAudioTracks explicitly authorizes Recorder to stop the accepted audio
track exactly once during terminal cleanup, including completion, discard,
failure, source end, and unmount. Do not grant that authority unless every
consumer of the shared track may be stopped.
recorder.source_availability() reports Live or Interrupted independently
from whether Recording is paused while a source is active. Browser mute and
unmute events update availability without pausing elapsed time. Recorder does not poll
or change the track's application-controlled enabled state. If the track ends
while Recording or paused, Recorder finalizes valid partial Recorded Audio;
RecordingOutcome::Completed identifies SourceEnded separately from
Requested and UnexpectedEnd completion.
Capture constraints, microphone permission requests, effective acquired-source
settings, and Audio Input Device identity apply only to recorder.start().
Supplied-source startup enters the same source-neutral Preparing state, while
requested_constraints(), settings(), and completed input identity remain
unknown. Encoder selection, Recording Chunks, elapsed time, pause and resume,
Analysis, and completed Recorded Audio behave the same for either source.
RecorderOptions::constraints accepts portable Ideal or Exact requests for
channel count, sample rate, echo cancellation, noise suppression, and latency.
The Recorder snapshots those constraints when it accepts start(), so changing
options during a Recording only affects a future Recording.
Use recorder.requested_constraints() for that snapshot,
recorder.constraint_capabilities() for the constraint fields the browser
recognizes, and recorder.settings() for the effective settings reported by
the acquired Recording Source. Effective fields are optional because browsers
do not report all settings consistently. An exact-constraint failure has
AudioErrorKind::Overconstrained; overconstrained_constraint() identifies the
rejected browser constraint when supplied by the browser.
recorder.media_type() exposes the selected encoder format once the Recording
starts. is_recorder_mime_type_supported() can be used before starting to probe
a candidate format, but a positive probe only means the browser recognizes the
type. It does not guarantee source acquisition, Recorder construction, or a
successful Recording.
Recording Chunk delivery is opt-in and supplements the final RecordedAudio:
use Duration;
use RecordingChunk;
use RecordingChunkDelivery;
let mut pending_chunks = use_signal;
let mut options = default;
options.chunk_delivery = Some;
let recorder = use_audio_recorder;
The accepted start() snapshots the delivery configuration and selected encoder
preferences. Cadence is approximate: browsers choose actual fragment boundaries
and may produce no data at a boundary. Every delivered chunk owns non-empty
encoded bytes and carries the effective media type, a Recorder-local
RecordingId, and a zero-based contiguous sequence. Chunks are delivered
serially in browser event order and are not guaranteed to be independently
playable.
While an opted-in Recording is active or paused,
recorder.request_chunk_boundary() asks the browser for another best-effort
boundary. The request does not promise exact timing, a non-empty chunk, or a
fragment that can be played without earlier chunks.
Callback return hands the chunk to the application; the library does not add an
upload, persistence, retry, acknowledgement, or backpressure queue. Recorder
still retains every browser fragment needed for final RecordedAudio, so chunk
delivery does not reduce completion memory. On success, all final chunks are
delivered before recorder.completed() becomes populated. Discard and unmount
suppress output that has not already been handed off.
If incremental blob conversion fails, recorder.chunk_delivery_failure()
identifies the Recording and failed sequence. That failure ends chunk delivery
for the Recording, but capture continues and final RecordedAudio assembly is
still attempted from the independently retained browser fragments. A Recorder,
encoder, or final-assembly failure instead ends the Recording, suppresses future
chunks and completed output, and is reported through recorder.outcome().
RecordedAudio::recording_id correlates completion with its chunks.
recorder.outcome() exposes the same ID for completed, discarded, and failed
Recordings; configuration rejected before start() is accepted has no Recording
outcome.
Live Analysis
Use use_live_analysis with an optional AudioAnalyser from any supported
source. Recorder supplies one through recorder.analyser() while a Recording is
active, and graph-backed Playback supplies the same source-neutral handle through
player.analyser().
use ;
let analysis = use_live_analysis;
if let Some = analysis
The default cadence is 50ms. with_cadence clamps requested values to
16ms..=1s, preventing unbounded polling and rendering rates. Each hook call
has an independent schedule. Its snapshot becomes None when the Analyser is
removed or replaced, stale work cannot publish into the replacement, polling
stops on unmount, and analyser reads are suspended while the document is
hidden.
An AudioAnalyser is a weak owner handle: retaining it does not keep its audio
graph alive. is_available(), try_read(), and try_level() report
unavailability when its source is detached, its owner degrades, or its owner is
cleaned up. Convenience read() and level() retain their empty/zero fallback;
use the try_ methods whenever valid silence must be distinguished from an
unavailable source. Reactive Analysis clears its snapshot while a stable handle
is unavailable and resumes when graph-backed Playback attaches another eligible
source to that same Analyser.
Time-domain values are byte-quantized amplitudes normalized to -1.0..=1.0.
Frequency-domain values are byte-quantized magnitudes normalized to
0.0..=1.0; AnalysisMetadata supplies the effective graph sample rate, FFT
size, bin count and frequency mapping, decibel range, and smoothing constant
needed to interpret them. snapshot.level() is normalized RMS amplitude over
that snapshot's FFT-sized time-domain window. It is not peak amplitude,
perceived loudness, sound pressure level, or Playback audibility.
LiveWaveform, SpectrumVisualizer, and LevelMeter use the same bounded
scheduling behavior while collecting only the values each presentation needs.
Their changing Analysis values are not live-region announcements.
Decoded Audio
Use decode_audio_data to consume complete AudioData and copy the browser's
decoded channels into one immutable flat-planar Rust allocation:
use AudioData;
use ;
async
The reported sample rate is the browser decode context's effective rate, not a
claim about the encoded source's original rate. The media type is retained as
part of AudioData but browser decodeAudioData does not consume it; media-type
support probes therefore cannot prove that a particular payload will decode.
Unsupported codecs, malformed or truncated input, and decoder refusal all map
to the portable DecodeErrorKind::DecodeRejected outcome.
DecodeOptions defaults to a 128 MiB ceiling for the Rust-owned planar f32
copy and allows an explicit with_max_decoded_bytes override. The size uses
checked channel, frame, and sample-width arithmetic, and a resource-limit error
reports both required and configured bytes. This gate runs only after the
browser has decoded the complete file, so it cannot prevent the browser's first
PCM allocation. Successful materialization may briefly retain roughly two PCM
representations, excluding encoded data and opaque decoder internals.
Each operation owns an internal AudioContext and requests context cleanup when
it settles or its future is dropped. Dropping the future suppresses result
publication, but does not promise to abort decoding work already started by the
browser. Decoded Audio is not a streaming decoder, mutable sample buffer,
resampler, transformed output, or Playback source.
Playback
use *;
use AudioPlayer;
use PlaybackSource;
The player creates and revokes browser object URLs for Audio Data as its source changes. Playback rate changes do not reload the source or reset its position.
For URL-addressable media, construct one alternative or a non-empty ordered set. Each alternative is validated and can carry an optional media-type hint. Relative URLs are accepted; validation does not claim that a resource exists or that the browser can decode it.
use ;
let alternative = new?
.with_media_type?;
let source = url
.with_loading_policy;
# Ok::
Use PlaybackSource::url_alternatives when an application can offer multiple
alternatives:
use ;
let source = url_alternatives?;
# Ok::
Playback skips only media-type hints the browser reports as definitely
unsupported. Untyped, maybe, and probably alternatives receive real load
attempts in order. Metadata remains tentative; canplay selects and exposes the
playable alternative. Initial failures continue to the next alternative, while a
failure after selection is terminal and never switches media implicitly.
If no alternative becomes playable, PlaybackSnapshot::alternative_failures
reports each attempted or skipped alternative in order with an unsupported,
network, decode, graph-ineligible, or unknown failure kind.
Eager begins browser acquisition when the source becomes current. OnPlay
keeps the source Dormant without an attached media resource until play() is
requested. Pausing while an on-play source is still loading clears play intent
without cancelling acquisition. URL ownership remains with the application:
replacement, unload, and owner cleanup detach the media resource but never
revoke an application-supplied URL, including an application-owned blob: URL.
For custom controls, use_audio_player exposes a PlaybackSnapshot through
controller.snapshot(). Source lifecycle, transport, readiness, and recoverable
play failure are independent facets. network separately reports inactive,
unknown, loading, idle, or stalled activity, so playing transport may coexist
with waiting readiness and stalled network activity. URL selection and coarse
terminal source failure are separately available through selected_alternative
and source_failure. A play request remains PlayPending until the browser
confirms Playing, and an interaction-required rejection leaves the current
source usable for retry. That recoverable failure clears on retry, confirmed
play, source replacement, unload, or terminal source failure.
PlaybackSnapshot::buffered and seekable are immutable, sorted source-time
range snapshots for the current source attempt. Overlapping or touching ranges
are merged, but later snapshots may be empty or smaller. Treat both collections
as UI guidance: buffered time is not byte-transfer progress, and a seekable
observation does not guarantee that the same seek will remain available later.
Replacement, fallback, unload, and terminal failure clear the observations, and
late events from an older source attempt cannot republish them.
Calling controller.stop() atomically pauses Playback, resets position to zero,
invalidates an outstanding play request, and returns a loaded source to
Ready/Idle. Whole-source repeat is available through repeat(),
set_repeat(), and toggle_repeat(); it remains enabled or disabled when the
source is replaced or unloaded and applies to the next loaded source. It is
separate from Bounded Playback over a Waveform Selection.
Use play_bounded_once() or play_bounded_loop() to play one controlled
WaveformSelection once or repeatedly:
use WaveformSelection;
let selection = new;
controller.play_bounded_once?;
// Or repeatedly enforce the range independently from whole-source repeat.
controller.play_bounded_loop?;
# Ok::
The command requires a playable source, a positive authoritative duration
reported by the browser, and a non-collapsed selection wholly inside that
duration. The initial_duration hook argument is presentation fallback and does
not make a duration authoritative. A valid run pauses, seeks to the selection
start, waits for the correlated seeked outcome, requests play, and waits for
browser confirmation before PlaybackSnapshot::bounded becomes Active.
Ordinary pause() preserves the bounded range and play() resumes it. A
committed InteractiveWaveform selection edit automatically retargets an
enforcing run. retarget_bounded_playback() provides the same operation for
other controlled selection UIs: it keeps the current media position when that
position remains in the replacement range, or pause-seeks to the new start and
waits for the correlated seek before optionally resuming.
An ordinary seek() or skip() cancels boundary enforcement without changing
the application-owned Waveform Selection. cancel_bounded_playback() explicitly
cancels the bound, while a new bounded run, another retarget, source replacement,
unload, stop, and owner cleanup invalidate superseded seek, play, pause, and
deadline outcomes. At the one-shot end, Playback pauses, observable position is
clamped to the exact selection end, and the bounded phase becomes Completed.
At a loop end, Playback enters Wrapping, pauses, seeks to the range start,
waits for that seek, and requests play again. Whole-source repeat() remains a
separate preference and native whole-source looping stays disabled while a bound
is enforcing. Seek timeout or play activation rejection becomes a bounded
Failed phase without failing the Playback Source, so ordinary Playback remains
available.
Media observations and a playback-rate-adjusted deadline both enforce the boundary. Rate changes re-arm that deadline, and visibility restoration samples actual media time. If a hidden or throttled loop is observed late, Playback restarts once from the range start; it does not synthesize missed iterations or preserve modulo phase.
Pause-seek-play ordering, lifecycle publication, completion clamping, and stale outcome rejection are Controller guarantees. Audible boundary timing is browser scheduling behavior: Bounded Playback is not sample-accurate, loop wraps are not gapless, and there is no maximum-overshoot promise.
Mute is observable through muted() and can be changed with set_muted() or
toggle_muted() without pausing Playback, seeking, or changing the retained
audibility level. set_audibility_level() accepts a finite normalized value in
the inclusive range 0.0..=1.0; invalid values are rejected without changing
public state. Mute and level preferences survive source replacement and unload
and apply to the next source.
Always inspect audibility_capability() before describing the level as
effective output gain. Current direct Playback reports
PlaybackAudibilityCapability::BestEffortMediaElement: the browser media
element receives the value, but some platforms, notably iOS, may not apply it
to perceived loudness. EffectiveGraphGain is reserved for graph-backed
Playback; direct control does not claim that guarantee. Mute and all transport
commands remain independent from level capability.
For pre-gain Analysis and effective gain, create an immutable graph-backed owner:
use Duration;
use *;
use ;
let source = use_signal;
let player = use_audio_player_with_options;
let analyser = player.analyser;
The opt-in belongs to the Playback owner and cannot be switched reactively. The graph is created lazily when the first eligible Playback Source is attached, then its context, pre-gain Analyser, and gain node persist across replacement among eligible Audio Data and URL-backed Playback Sources, as well as unload. The graph state is orthogonal to source and transport state: it reports awaiting source, preparing, suspended, running, interaction required, or terminal unavailability. Direct Playback remains the default.
URL-addressable alternatives are direct-only unless they explicitly request anonymous CORS:
use ;
let alternative = new?
.with_media_type?
.with_cross_origin;
assert!;
let source = url;
# Ok::
Playback applies the declared crossOrigin policy before graph attachment and
before assigning src. The server must authorize the anonymous cross-origin
request. Rejection is an initial source-attempt failure and may fall back to the
next eligible alternative. Alternatives without a cross-origin policy and those
using UseCredentials remain available to ordinary direct Playback, but a
graph-backed owner skips them as GraphIneligible without speculative graph
attachment.
A graph-backed play request invokes context resume and media play in the same
activation turn and reports Playing only after both succeed. Rejection or a
later context suspension pauses the media and reports an interaction-required
failure that can be retried. Analysis observes the Playback Source before mute
and level gain, so visualizations remain meaningful at zero output gain. A
terminal setup failure invalidates Analysis, permanently changes that owner's
level capability to BestEffortMediaElement, and continues with direct
transport. A later failure of an already selected URL-addressable alternative is
terminal and does not silently choose another alternative.
The same Controller can drive independently arranged PlaybackSeekSlider,
PlaybackSkipButton, PlaybackStopButton, PlaybackPlayPauseButton,
PlaybackRateButton, PlaybackMuteButton, PlaybackAudibilitySlider, and
PlaybackRepeatButton components. Labels, signed skip amounts, and the rate
cycle are configurable. Mute and repeat use stable labels and native pressed
state. The seek and audibility sliders expose meaningful value text, which can
be replaced with localized value_text. PlaybackStatusAnnouncer is an optional
polite live region for coarse lifecycle, waiting, stalled, and recovery changes.
During Bounded Playback it restricts messages to discrete start, completion,
cancellation, useful wrapping, and failure transitions. It never announces
position, range snapshots, selection drafts, or audibility changes. Bounded
messages can be localized independently with
BoundedPlaybackAnnouncementLabels.
Waveform Data
WaveformData is an immutable, cheap-to-clone snapshot for duration-aware
Waveforms. A snapshot has one amplitude mode, one channel count, and one or more
resolutions ordered from finest to coarsest. Construction consumes flat
channel-major buffers and rejects invalid spans, coverage, channel alignment,
and amplitudes rather than repairing them.
use Duration;
use *;
use Waveform;
use WaveformData;
WaveformData::select accepts a half-open source-time range and a per-channel
bucket budget. It returns the finest fitting stored resolution, or the coarsest
resolution when none fit, as borrowed channel slices without copying buckets.
Clone and equality use shared snapshot identity, so independently reconstructed
data intentionally counts as changed.
For long sources, create a WaveformViewportController with
use_waveform_viewport and pass it to NavigableWaveform. The Controller
exposes the total duration and positive visible source-time range, plus clamped
pan, show_range, reset, and explicit-anchor zoom commands. It also
exposes is_following, follow_playback, and resume_follow. Following keeps
Playback inside a 15–75% safe zone and moves it toward 25% only after it leaves
that zone. Manual pan, zoom, reset, show-range, and overview movement disable
following until resume_follow is called. The composed component provides
native pan, zoom, reset, resume-follow, and optional fixed-span overview controls
with visible source-time text.
use NavigableWaveform;
use use_waveform_viewport;
let viewport = use_waveform_viewport;
rsx!
NavigableWaveform uses the fallback budget for server rendering and the first
client render. Dioxus resize observation then derives a budget from measured CSS
width with pixel density capped at 2x and a maximum of 4,096 buckets per
channel. It selects only the intersecting borrowed range and renders one compact,
amplitude-mode-aware SVG path per channel. A sparse resolution ladder can still
make the coarsest stored level exceed the requested budget; applications should
include a suitably coarse long-form level when bounded geometry matters.
When an AudioPlayerController is supplied, the component presents its visible
playhead separately from SVG geometry and follows only playable, duration-bound
positions. Position-only updates therefore do not rebuild channel paths. Zoom
buttons anchor at a visible followed playhead and otherwise use the Viewport
center. Viewport movement is immediate rather than animated, so the same stable
controls and visible range text remain the complete reduced-motion keyboard path.
InteractiveWaveform composes Waveform Data, an AudioPlayerController, and
one controlled WaveformSelection. Playback position, selection start, and
selection end remain independently named native sliders. Arrow keys use the
configurable fine source-time step, Page Up and Page Down use the coarse step,
and Home and End honor each control's valid bounds. Selection handles may meet
but do not cross or exchange identity.
Pointer movement is shown as an internal draft. A handle drag commits one selection update on release, while a track click seeks Playback immediately. Committed handle edits atomically retarget an enforcing Bounded Playback run; draft movement does not issue transport commands. A collapsed or otherwise unplayable committed selection cancels the old enforcement rather than leaving it attached to selection state that is no longer visible. Waveform duration determines rendering and selection bounds; a seek is capped by the Controller's authoritative positive Playback duration. Continuously changing position and drafts are not live-region announcements, so native slider feedback remains the continuous assistive signal.
Use from_magnitudes for normalized values in 0.0..=1.0 and
from_signed_envelopes for ordered minimum/maximum pairs in -1.0..=1.0.
from_peaks creates one evenly spaced mono Magnitude resolution; this conversion
necessarily loses Peaks cadence, channel structure, and sign information.
Platform Support
Browser recording, playback, and device hooks require the
wasm32-unknown-unknown browser target with the relevant Web APIs. On other
targets, pure analysis and visual components remain available while controllers
report AudioErrorKind::UnsupportedPlatform.
Permission prompts, available MIME types, device labels, and background-tab
capture behavior remain browser and operating-system policies.
For hydration-safe fullstack rendering, unsupported server targets use the same neutral first-render states as the browser and transition to unsupported inside the first client effect. Commands still return unsupported errors.
Domain Language
The public audio terms used by this project are defined in
CONTEXT.md. The glossary describes the domain contract without
documenting private module or backend layout.
License
Licensed under either of
- Apache License, Version 2.0 (LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0)
- MIT license (LICENSE-MIT or http://opensource.org/licenses/MIT)
at your option.
Contribution
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.