oxideav-mp4
Pure-Rust MP4 / ISO Base Media File Format container — demuxer
(probe + sample-table expansion + seek) and muxer (moov-at-end by
default, optional faststart rewrite). Three brand presets share one
implementation: mp4, mov (QuickTime), and ismv (Smooth Streaming
ftyp, non-fragmented layout). Zero C dependencies.
Part of the oxideav framework but usable standalone.
Installation
[]
= "0.1"
= "0.1"
= "0.1"
= "0.0"
Quick use
Demux an MP4 and feed packets into a codec
use CodecRegistry;
use ContainerRegistry;
let mut codecs = new;
let mut containers = new;
register;
// ... register whichever codecs you care about (aac, flac, h264, mjpeg, ...)
let input: =
Boxnew;
let mut dmx = containers.open?;
// Sample entries are resolved to concrete codec ids. For `mp4a`/`mp4v`
// tracks the esds `objectTypeIndication` is honoured, so MP3-in-mp4
// comes out as "mp3", MPEG-1 video as "mpeg1video", AAC as "aac", etc.
let stream = &dmx.streams;
let mut dec = codecs.make_decoder?;
loop
# Ok::
Mux packets into an MP4
use WriteSeek;
let f = create?;
let ws: = Boxnew;
let mut mux = open?;
mux.write_header?;
for pkt in packets
mux.write_trailer?;
# Ok::
Faststart (moov-at-front) layout
use ;
let opts = Mp4MuxerOptions ;
let mut mux = open_with_options?;
In faststart mode the muxer buffers mdat in memory and writes
[ftyp][moov][mdat] at write_trailer time, patching chunk offsets so
the file is streamable from the first byte.
Scope
Demuxer
Sample-entry FourCCs resolve to these codec ids:
| FourCC | Codec id |
|---|---|
mp4a |
aac (default); esds OTI refines to mp3 (0x69/0x6B), ac3 (0xA5), eac3 (0xA6), dts (0xA9) |
mp4v |
mpeg4video (default); esds OTI refines to mpeg1video, mpeg2video, h264, h265, mjpeg |
alac |
alac |
fLaC / flac |
flac |
Opus / opus |
opus |
ac-3 / AC-3 |
ac3 (Dolby Digital, ETSI TS 102 366 Annex F) |
ec-3 / EC-3 |
eac3 (Dolby Digital Plus, ETSI TS 102 366 Annex G) |
dtsc / dtsh / dtsl / dtse |
dts (DTS Coherent Acoustics / DTS-HD HR / DTS-HD MA / DTS Express, ETSI TS 102 114) |
ulaw / alaw |
pcm_mulaw / pcm_alaw (G.711) |
avc1 / avc3 |
h264 |
hvc1 / hev1 |
h265 |
vp08 |
vp8 |
vp09 |
vp9 |
av01 |
av1 |
jpeg / mjpa / mjpb |
mjpeg |
s263 / h263 |
h263 |
lpcm / sowt / twos |
pcm_s16le (endianness of twos is not re-swapped) |
tx3g |
mov_text (3GPP TS 26.245 timed text — "movtext") |
text |
text (QuickTime plain text) |
wvtt |
webvtt (W3C WebVTT-in-ISOBMFF) |
stpp |
ttml (XML / TTML subtitle) |
sbtt / stxt |
sbtt / stxt (BMFF §12.5–6 text) |
c608 / c708 |
eia_608 / eia_708 (closed captions) |
encv / enca / enct / encs |
original FourCC recovered from sinf/frma; params.options["protection_scheme"] carries the schm.scheme_type (e.g. cenc, cbcs) |
| any other | mp4:<fourcc> — callers can register their own decoder |
-
Codec-specific config records (
avcC,hvcC,av1C,vpcC,dfLa,dOps,dac3,dec3, esds DSI) are forwarded asextradata. -
Sample-table expansion:
stts,stsc,stsz/stz2,stco/co64,stss.next_packetserves samples in file-offset order. -
Fragmented MP4 (DASH / HLS / Smooth Streaming / CMAF):
mvex/trexper-track defaults plus zero or more trailingmoof+mdatpairs (mfhd,traf,tfhd,tfdt,trun) are stitched onto the initial sample table.default-base-is-moof, per-sample size/duration/flags/composition-time-offset overrides, andtfhd.base_data_offsetare all honoured.styp,sidx, andmfrasegment-index boxes are skipped (segment-precision seek hint consumption is a follow-up). -
Seek:
seek_to(stream, pts)lands on the nearest sync-sample ≤ pts (or the first keyframe of the stream if none qualify). -
Metadata: 3GPP
udtaboxes (titl/auth/…) and iTunes-stylemeta/ilstare surfaced viaDemuxer::metadata(). -
Extended language tag (ISO/IEC 14496-12 §8.4.6,
elng): a track'smdia/elngExtendedLanguageBox carries a NULL-terminated BCP 47 (RFC 4646) tag richer thanmdhd's packed 3-char ISO 639-2 code (region / script / variant subtags). When present it is surfaced onparams.options["language"](e.g.en-US,zh-Hant-HK,es-419) and, per §8.4.6.1, overrides themdhdlanguage. Absentelng, the option is omitted (callers fall back tomdhd). -
Handler-type recognition (ISO/IEC 14496-12 §8.4.3):
soun→ Audio,vide→ Video,subt/sbtl/text→ Subtitle,meta→ Data. Subtitle sample entries (tx3g,text,wvtt,stpp,sbtt,stxt,c608,c708) come out withMediaType::Subtitleand the per-codec id from the table above; their post-preamble payload (BMFF strings / tx3g header / vttC) is preserved verbatim inparams.extradatafor downstream renderers. -
Protected sample-entry unwrap (ISO/IEC 14496-12 §8.12): when the outer FourCC is
encv/enca/enct/encs, the demuxer walks the innersinfto recover the original codec FourCC fromfrmaand the protection scheme fromschm. The stream surfaces the un-transformed codec id so downstream decoders can be set up normally;params.options["protection_scheme"]carries the four-char scheme type (e.g.cenc,cbcs) so callers know packet payloads are still ciphertext. -
CENC metadata parsing (ISO/IEC 23001-7:2016): the three boxes that carry encryption framing on top of the §8.12 envelope are parsed structurally and surfaced to callers (no decryption — the AES key/decrypt op is left to a downstream layer with key material).
tenc(§8.2 TrackEncryptionBox) — discovered insidesinf/schi. v0 capturesdefault_isProtected,default_Per_Sample_IV_Size, anddefault_KID; v1 adds thedefault_crypt_byte_block/default_skip_byte_blockpattern pair (forcens/cbcsschemes) and thedefault_constant_IVused whenisProtected==1 && IV_size==0. Surfaced onparams.optionsascenc_default_kid(lowercase hex),cenc_default_is_protected,cenc_default_iv_size,cenc_tenc_version, and (v1 only)cenc_default_crypt_byte_block/cenc_default_skip_byte_block/cenc_default_constant_iv.pssh(§8.1 ProtectionSystemSpecificHeaderBox) — collected at moov level. Each entry captures the 16-byte SystemID UUID, optional v1 KID list, and the DRM-system-specific opaqueDatablob. Surfaced viaDemuxer::metadata()aspssh_<n>keys with value"<system_id_hex> <kid_count> <data_len>"; structured records are reachable through the publiccenc::PsshBoxtype for callers that downcast.senc(§7.2 SampleEncryptionBox) — collected from everytrafwhose matching track carried atencdefault (so the per-sample IV width is recoverable per §7.2.3). Capturesflags, the per-sampleInitializationVector, and (whenUseSubSampleEncryptionis set) the{BytesOfClearData, BytesOfProtectedData}subsample map. Surfaced viaDemuxer::metadata()assenc_<n>keys with value"track=<idx> seq=<mfhd_seq> samples=<n> flags=0x<hex>"; structured records are reachable through the publiccenc::SencBoxtype.
The standalone parsers (
cenc::parse_tenc/cenc::parse_pssh/cenc::parse_senc) are public for callers that already have the box body in hand from another path (e.g. a non-MP4 carrier of CENC framing). -
Track references (ISO/IEC 14496-12 §8.3.3,
tref): each typedTrackReferenceTypeBoxinsidetrak/trefis parsed and the resulting(reference_type → track_IDs)pairs are surfaced onparams.optionsastref_<type>keys whose value is a space-separated list of referencedtrack_IDs (e.g.tref_chap = "3",tref_subt = "10 11",tref_cdsc = "2"). Useful for wiring subtitle→video (subt), chapter (chap), content description (cdsc), font (font), hint (hint), depth / parallax auxiliary video (vdep/vplx), and hint dependency (hind) relationships.track_ID = 0entries (spec-prohibited) are dropped. -
Track Kind (ISO/IEC 14496-12 §8.10.4,
kind): eachKindBoxinside a track-leveludtais parsed as a(schemeURI, value)pair (both NULL-terminated C strings; an absent value is allowed, meaning the URI alone identifies the kind). Multiplekindboxes per track are supported — the spec explicitly allows several schemes co-labelling the same track (e.g. one DASH roleurn:mpeg:dash:role:2011 mainplus one iTunes-scheme tag). Each entry is surfaced onparams.optionsaskind_<n>(0-based encounter index); the value is the URI alone when no name follows, or"URI value"(space-separated, mirroring thetref_<type>convention) when both are present. -
Composition-to-decode (ISO/IEC 14496-12 §8.6.1.4,
cslg): a track'sstbl/cslgCompositionToDecodeBox documents the composition↔decode timeline relationship implied by a signed (v1)ctts— the DTS shift that guaranteesCTS ≥ DTS, the least / greatest composition offsets, and the composition start / end times. Both v0 (32-bit) and v1 (64-bit) layouts are read (widened toi64). The five fields are surfaced onparams.optionsascslg_composition_to_dts_shift,cslg_least_decode_to_display_delta,cslg_greatest_decode_to_display_delta,cslg_composition_start_time, andcslg_composition_end_time(decimal strings in the media timescale;composition_end_time = 0means "unknown" per §8.6.1.4.3). Absentcslg, none of the keys are emitted. -
Shadow sync samples (ISO/IEC 14496-12 §8.6.3,
stsh): a track'sstbl/stshShadowSyncSampleBox is an optional seek hint — a table of(shadowed_sample_number, sync_sample_number)pairs naming a sync sample (key frame) that can be decoded in place of a non-sync sample when seeking to or before it. Each pair is surfaced onparams.optionsasstsh_<n>(0-based encounter index) with value"shadowed sync"(both 1-based sample numbers, space-separated). The table is purely a seek optimisation — it is ignored in normal forward play and a track decodes correctly without it. Absentstsh, none of the keys are emitted. -
Sample dependency hints (ISO/IEC 14496-12 §8.6.4,
sdtp): a track'sstbl/sdtpSampleDependencyTypeBox is a per-sample table of four 2-bit fields —is_leading,sample_depends_on,sample_is_depended_on,sample_has_redundancy— packed one byte per sample (thesample_countis implicit fromstsz/stz2). The table feeds trick-mode playback (drop disposable samples on fast-forward) and refines random-access roll-forward (a sample markedsample_depends_on = 2is an I-picture without needing thestssto mark it). The raw per-sample 2-bit values are decoded and stored on the track; the demuxer surfaces a small summary onparams.optionsas five keys —sdtp_count,sdtp_leading_count(samples withis_leading ∈ {1, 3}),sdtp_independent_count(samples withsample_depends_on = 2),sdtp_disposable_count(samples withsample_is_depended_on = 2), andsdtp_redundant_count(samples withsample_has_redundancy = 1). Absentsdtp, none of the keys are emitted (the demuxer falls back tostssfor keyframe detection, as before). -
Sample groups (ISO/IEC 14496-12 §8.9,
sbgp+sgpd): a track'sstbl/sbgp(SampleToGroupBox §8.9.2) run-length map andstbl/sgpd(SampleGroupDescriptionBox §8.9.3) per-group entries are parsed. Several of each are accumulated — one pair pergrouping_type(roll,rap,sync,alst,prol, …). Eachsbgpis surfaced onparams.optionsassbgp_<n>(0-based encounter index): the grouping type, an optionalparam=<P>(v1grouping_type_parameter), then space-separatedcount:indexrun-length pairs (group_description_index0 = "no group of this type"; an index ≥ 0x10001 is a movie-fragment-local reference per §8.9.4, kept verbatim — the demuxer does not resolve fragment-local groups). Eachsgpdis surfaced assgpd_<n>: the grouping type, an optionaldefault=<D>(v2default_sample_description_index), then the per-group entry payloads as lowercase hex. Entry sizing honours §8.9.3.2 (v1 fixeddefault_length, v1 per-entrydescription_length, or the v0 deprecated no-length-signalling case captured as one combined blob). The entry payloads are grouping-type-specific and not interpreted by the container — they are surfaced verbatim for a layer that knows thegrouping_typesemantics. Absent both boxes, none of the keys are emitted. -
Sub-sample information (ISO/IEC 14496-12 §8.7.7,
subs): a track'sstbl/subsSubSampleInformationBox is an optional sparse table describing how selected samples decompose into smaller, semantically-meaningful sub-samples (e.g. NAL units / parameter sets for H.264 per ISO/IEC 14496-15, or arbitrary segment boundaries for codecs that define their own sub-sample contract). Each entry carries a sample-number delta from the previous entry, then a list of(subsample_size, subsample_priority, discardable, codec_specific_parameters)rows. Version 0 storessubsample_sizeas 16-bit; version 1 widens it to 32-bit (both layouts are read and normalised tou32).flagsdistinguishes co-residentsubsboxes with different per-codec semantics (§8.7.7.1). The container preserves the carried codec's interpretation ofsubsample_priority/discardable/codec_specific_parametersverbatim — those small ints are opaque at this layer. Eachsubsencountered on a track is surfaced onparams.optionsassubs_<n>(0-based encounter index); the value starts with"v<version> flags=<f>"and is followed by one space-separateddelta=<d>[:size,priority,discardable,csp[;...]]block per entry (decimal for everything exceptcsp, which is lowercase 8-digit hex). The trailing colon and per-sub-sample list are omitted when an entry hassubsample_count = 0. Absentsubs, no keys are emitted. -
Producer reference time (ISO/IEC 14496-12 §8.16.5,
prft): a top-level FullBox carrying a UTC wall-clock instant in NTP 64-bit format (RFC 5905 — high 32 bits = seconds since 1900-01-01 UTC, low 32 bits = fractional seconds) correlated with a media time on one reference track's media clock. Used by low-latency DASH / CMAF live streams so a consumer can match production wall-clock against media presentation time (and bound buffer occupancy without out-of-band timing signals). Eachprftencountered during the top-level walk is surfaced onDemuxer::metadata()asprft_<n>(0-based file order); the value is three space-separated decimal integers"reference_track_ID ntp_timestamp media_time". Both v0 (32-bitmedia_time) and v1 (64-bitmedia_time) layouts are read; v1media_timeis widened tou64so callers see one type regardless. Absentprft, no keys are emitted. The structured record is also reachable via the publicoxideav_mp4::demux::parse_prft_box(&[u8])entry point for tooling that wants the typedPrftRecord(reference_track_id,ntp_timestamp,media_time,version) directly.
Muxer
Only codecs with an mp4 sample-entry packaging are accepted. Codec
knowledge is confined to sample_entries::sample_entry_for; the rest
of the muxer appends opaque packet bytes.
Supported encode codec ids (produced sample entry FourCC in parentheses):
pcm_s16le→sowtflac→fLaCwithdfLaconfig (requires STREAMINFO extradata)aac→mp4awithesds(requires AudioSpecificConfig extradata)h264→avc1withavcC(requires AVCConfigurationRecord extradata)mjpeg→jpegmov_text→tx3g(3GPP TS 26.245 timed text) —texthandler +nmhdwebvtt→wvtt(BMFF §12.6.3.2 XMLSubtitleSampleEntry sibling) —subthandler +sthdttml→stpp(BMFF §12.6.3.2 XMLSubtitleSampleEntry) —subt+sthdsbtt→sbtt(BMFF §12.6.3.2 TextSubtitleSampleEntry) —subt+sthdstxt→stxt(BMFF §12.5.3.2 SimpleTextSampleEntry) —subt+sthd
For the subtitle codecs the muxer accepts the demuxer's surfaced
extradata verbatim (the post-preamble sample-entry payload: tx3g's
18-byte header, vttC config, stpp namespace strings, sbtt/stxt MIME
strings), so a demux → mux round-trip preserves the inner config.
Other codec ids fail with Error::Unsupported at open, never at
write_packet time.
Edit lists (edts/elst, ISO/IEC 14496-12 §8.6.5–6) are emitted
per-track when the first packet has a positive presentation timestamp:
a leading empty edit (media_time = -1) of the start delay (in the
movie timescale) followed by a media_time = 0 segment for the track
duration, so a player offsets the track start instead of beginning at
presentation time 0. Version 0 (32-bit) by default, auto-promoting to
version 1 (64-bit) for over-32-bit durations. Tracks starting at PTS 0
get no edts. Controlled by Mp4MuxerOptions::write_edit_list
(default true).
Chunk offsets auto-promote from stco (32-bit) to co64 (64-bit) when
any offset exceeds 4 GiB. The mdat box header stays 32-bit — files
whose mdat payload exceeds 4 GiB fail at write_trailer.
Fragmented / DASH / CMAF segment writing
The dash, cmaf, and ismv registry entries select the fragmented
muxer (oxideav_mp4::frag::open_fragmented_typed). It emits an init
segment (ftyp + moov with mvex/trex per track) followed by one
styp? + sidx? + moof + mdat segment per fragment cadence boundary, and
a trailing mfra/tfra/mfro random-access index. The per-segment
styp (ISO/IEC 14496-12 §8.16.2 Segment Type Box) is controlled by
FragmentedOptions::styp.
For caller-driven per-segment control, the
FragmentedMuxer::write_fragmented_segment_with_styp(major_brand, compat_brands) inherent method marks the next emitted segment's
styp to use the given DASH/CMAF (major, compat) pair, overriding the
preset for one segment (then consumed). The stateless byte builder is
also exposed via the public oxideav_mp4::styp module —
build_styp(major, compat) / write_styp(writer, major, compat) —
mirroring the read-side parse_styp in oxideav-mov so a producer
round-tripping a parsed Styp can emit the same byte sequence.
Sample-group muxing
Sample groups (sbgp / sgpd, ISO/IEC 14496-12 §8.9.2 / §8.9.3) are
emitted per track via Mp4MuxerOptions::track_sample_groups. Each
entry's sbgp and sgpd Vecs are placed at the end of the target
track's stbl body after the chunk-offset table; sgpd is written
before sbgp so the description table the per-sample index references
is declared first (§8.5.1 ordering). The
oxideav_mp4::sample_groups::{SampleToGroup, SampleGroupDescription, build_sbgp, build_sgpd} API also stands alone for callers that want
to assemble the raw boxes themselves.
The version pick for sgpd is automatic per §8.9.3.2: a Some(_)
default_sample_description_index with shared-length entries → v2
(no per-entry length); shared-length entries alone → v1 with fixed
default_length; mixed-length entries → v1 with per-entry
description_length. The deprecated version-0 "no length signalling"
form is not emitted. The grouping-type-specific entry payload itself
is opaque to the container — callers supply already-serialised
Vec<u8> per entry.
sbgp chooses v0 (no grouping_type_parameter) or v1 (Some(_)) per
§8.9.2; a group_description_index ≥ 0x10001 (movie-fragment-local
per §8.9.4) is written verbatim, the muxer does not resolve it.
Seek strategy
seek_to(stream, pts) tries three strategies in order, picking the
first that applies:
tfrafast-path (ISO/IEC 14496-12 §8.8.11). If the file has a trailingmfrawhosetfraindexes the requested track, the demuxer binary-searches thetfratime table for the largesttime ≤ pts, translates the result to amoof_offset, and snaps to the first keyframe at-or-after that offset. O(log N) ontfra+ a one-fragment-bounded sample-list scan.sidxfast-path (ISO/IEC 14496-12 §8.16.3). If notfracovers the track but the file carries one or moresidxboxes whosereference_idmatches the track'strack_ID, the demuxer walks every matchingsidx, expands its references into virtual(EPT, byte_offset)anchors, and picks the latest anchor whose decode-time start is at-or-beforepts(translated from the track's media timescale into the sidx timescale per §8.16.3). Both on-the-wire shapes are handled: a singlesidxindexing every subsegment (DASH on-demand profile) and onesidxper subsegment (DASH live profile / what our own muxer emits). Hierarchical (nested) sidx references are walked for byte-range accounting only — they don't carry a media-time anchor we can land on. Timescale conversion usesu128arithmetic so the multiply doesn't overflow for long-duration tracks even when the track's media timescale and the sidx's timescale differ (per the spec-permitted but DASH-IF-deprecated case).- Linear scan fallback. Walks the sample table picking the
last keyframe at-or-before
pts. This is the unconditional safety net — when neither index applies, or when the indexed offset doesn't resolve cleanly (corrupt index, mdat layout the file lied about),seek_tostill returns a correct cursor.
Not (yet) supported
- Fragmented-MP4 muxing — the demuxer reads
moof+mdatsegments, but the muxer only emits a single moov-at-end (or faststart) shape. - CENC decryption proper — the demuxer parses the CENC framing
(
tencdefaults,psshper-DRM headers, per-fragmentsencper-sample IVs + subsample maps; see "CENC metadata parsing" in the demuxer feature list above) and surfaces the metadata, but it does not run the AES-128 CTR / CBC decryption step. That belongs to a downstream layer with key material from the namedpssh.SystemID.saiz/saio(Sample Auxiliary Information Sizes / Offsets, ISO/IEC 14496-12 §8.7.8–9) wiring as an alternative IV-carriage path is also still a follow-up; for nowsencis the only IV source consumed. - Multiple sample descriptions per track (only the first entry of
stsdis used;tfhd.sample_description_indexoverrides are ignored). - mdat payloads larger than 4 GiB (the 32-bit box header is not
promoted to
largesize).
Container registry
let mut reg = new;
register;
Registers:
- Demuxer
"mp4"(also serving.mp4,.m4a,.m4v,.3gp,.mov,.ismv). - Muxers
"mp4","mov","ismv". - A content probe that recognises
ftyp/wide+ftyp/moov.
Fuzzing
A cargo-fuzz target exercises the BMFF box-tree walker on
arbitrary bytes:
The target opens, drains up to 256 packets, and re-seeks; it asserts
nothing panics, aborts, or OOMs. Seed corpus + regression artefacts
live at fuzz/corpus/demux/. The fuzz crate has its own [workspace]
and a committed Cargo.lock for reproducibility.
Pinned regressions worth calling out:
- Extended-size u64 overflow — a
size=1 largesize=u64::MAXextended box anchored at a non-zero file offset used to overflow every downstreambody_start + payload_sizearithmetic site (the §8.16.3sidxend-anchor computation is the most exposed example).read_box_headernowchecked_addsstart + total_sizeand rejects the header before any caller computes a derived end byte. Replayed bytests/largesize_overflow.rsand two boundary unit tests insrc/boxes.rs. Companion to oxideav-mov's round 187 fix on the QTFF atom walker.
License
MIT — see LICENSE.