rusty_h264 0.8.0

A ground-up, pure-Rust H.264 encoder and decoder. Decoder bit-exact vs Cisco openh264 (Baseline + B-slices + most of High profile); encoder bit-exact vs ffmpeg. forbid(unsafe) codec core; optional vendored openh264 SIMD asm (on by default — `--no-default-features` for pure safe Rust). BSD-2.
Documentation

rusty_h264

crates.io docs.rs CI License: BSD-2-Clause Remade With Rust By Mata Network

The public facade — this is the crate you depend on. A ground-up, pure-Rust H.264 encoder and decoder with a #![forbid(unsafe_code)] codec core, no C in the dependency tree, and a BSD-2 license you can embed anywhere. The decoder is validated bit-exact against Cisco's h264dec over openh264's conformance corpus; the encoder is bit-exact under ffmpeg across QP 0–51.

Part of Remade With Rust by Mata Network — and the H.264 engine inside remade_ffmpeg_rs, our memory-safe FFmpeg alternative.


Install

cargo add rusty_h264

[dependencies]

# SIMD acceleration on by default (needs `nasm` at build time; kernels are vendored):

rusty_h264 = "0.7"



# …or pure, portable, 100%-safe Rust — no nasm, no FFI, no unsafe anywhere:

rusty_h264 = { version = "0.7", default-features = false }

Quick start

use rusty_h264::{Encoder, EncoderConfig, Decoder, YuvFrame};

let mut enc = Encoder::new(EncoderConfig::new(640, 480)).unwrap();
let frame = YuvFrame::black(640, 480);
let bitstream = enc.encode(&frame);     // one Annex-B access unit

let mut dec = Decoder::new();
let decoded = dec.decode(&bitstream).unwrap().unwrap();
assert_eq!(decoded, frame);             // a flat frame has no residual → exact

The codec is lossy in general (that round-trip is exact only because the frame is flat); quality is governed by QP or the bitrate target. A moving sequence with P-frames and rate control:

use rusty_h264::{Encoder, EncoderConfig};

let mut cfg = EncoderConfig::new(640, 480);
cfg.gop_size  = 30;          // an IDR every 30 frames, P-frames between
cfg.bitrate   = 1_000_000;   // 1 Mbps average; 0 = constant-QP (cfg.qp)
cfg.framerate = 30.0;
let mut enc = Encoder::new(cfg).unwrap();
for frame in &frames { let au = enc.encode(frame); /**/ }

Decoding a whole stream in display order is one call — it splits access units, assembles multi-slice pictures and reorders by POC:

use rusty_h264::Decoder;

let frames = Decoder::new().decode_stream(&stream).unwrap();

For streaming use, the lower-level Decoder::decode returns one picture per access unit in decode order (pair it with Decoder::last_poc to reorder).

What this crate re-exports

Item From Role
Encoder, EncoderConfig, EncodeError rusty_h264-encoder the encode pipeline
Preset, LookaheadMode rusty_h264-encoder speed/quality trade-offs
Decoder, DecodeError rusty_h264-decoder the decode pipeline
YuvFrame, Profile, ChromaFormat rusty_h264-common shared types (I420 planes)
NalUnit, NalUnitType rusty_h264-common Annex-B / NAL layer
VERSION this crate the crate version string

You never need to name the sub-crates directly — that's the point of the facade.

Capabilities

Decoder — validated bit-exact vs Cisco h264dec over openh264's corpus:

  • Constrained Baseline + B-slices (temporal & spatial direct, implicit & explicit weighted prediction, L0/L1/Bi partitions, B_Skip/B_Direct).
  • Most of High profile (CAVLC): 8×8 integer transform and 8×8 intra prediction, sequence/picture scaling matrices, transform_size_8x8_flag, second chroma QP offset.
  • CABAC entropy decode (Main profile): I slices (I_4x4, I_16x16), P slices (P_Skip, all partition types + sub-types, mvd, MC, residual) and B slices (B_Skip, B_Direct_16x16, L0/L1/Bi, B_8x8, spatial + temporal direct) — brought up symbol-by-symbol against an instrumented openh264 oracle and gated pixel-exact vs ffmpeg.
  • Full intra (I_16x16/I_4x4/I_8x8/I_PCM), quarter-pel MC, in-loop deblocking (8×8-aware), multi-reference DPB with POC reordering and MMCO.
  • Fuzzed to never panic or hang on malformed input.

Encoder — every frame decodes bit-exactly under ffmpeg, QP 0–51:

  • Intra (I_16x16/I_4x4, λ-based RD mode decision), inter P-frames (P_Skip/16×16/16×8/8×16), quarter-pel MC, rate-aware ME, multi-ref DPB.
  • CABAC entropy coding (Main profile, default-on — measured −8.8…−9.0% BD-rate for 1.10–1.22× the time; RUSTY_H264_LEGACY_CAVLC=1 restores the Constrained Baseline + CAVLC bitstream byte-for-byte).
  • Adaptive quantization (default-on): per-macroblock QP finer on flat regions, coarser on busy ones — a perceptual/SSIM win that self-limits on pathological content so it never regresses.
  • Per-GOP I-frame QP cascade, in-loop deblocking, average-bitrate rate control (complexity model + leaky bucket).
  • Opt-in tools: B-frames (bframes, incl. a content-adaptive enable), the 8×8 transform (I_8x8 + inter, High profile), mb-tree temporal AQ with a lookahead, RD P_Skip. P_8x8 sub-partition motion and the adaptive wide motion search are default-on for the Quality preset.
  • Three presets — Fast (SAD, integer-pel), Balanced (adds sub-pel refinement: −42…−50% BD-rate over Fast for ~2.3–3.1× the time), Quality (full RD trial-encode, sub-partitions, full I_4x4 search).

Features

Feature Default Effect
asm Vendored openh264 BSD-2 SIMD kernels (x86-64) for MC, deblocking, transforms, SATD/SAD. Needs nasm on PATH.
(none) --no-default-features → 100% safe, portable Rust. No nasm, no FFI, no unsafe. Runs on any Rust target.

The asm kernels are x86-64 only; on other architectures (e.g. arm64 macOS) the accel crate compiles to an empty lib and the pure-Rust scalar path is selected automatically, so a default-features build works everywhere.

The codec core is #![forbid(unsafe_code)] either way. All unsafe lives in the single, optional rusty_h264-accel crate. The same acceleration boundary accepts your own custom kernels or hand-written ASM — the safe core never changes when you push for speed.

Performance

Single core, bit-exact, on the maintainer's machine:

Decode, 1800 frames of real 720p content encoded by x264 (what an encoder puts in the stream dominates decode cost), vs ffmpeg's native h264 software decoder:

x264 tool tier rusty_h264 ffmpeg native h264 gap
baseline / CAVLC (--preset veryfast) 213 Mpx/s 412 Mpx/s 1.98×
main / CABAC (--preset medium) 146 Mpx/s 294 Mpx/s 2.16×
high (--preset slower) 125 Mpx/s 255 Mpx/s 2.06×
encode workload rusty_h264 reference
Encode INTER, CIF (vs openh264) 71 Mpx/s 115 · 1.6×
Encode ALL-INTRA, CIF (vs openh264) 24 Mpx/s 88 · 3.6×

Measured 2026-08-05 after a structural-fusion campaign (same harness, same streams as the previous 2.34×/2.70×/2.49× figures — the change is decoder speed, not method): per-frame allocation pooling, stage-boundary fusion in the residual/MC paths, row-interleaved deblocking, a fused-register CABAC engine, and a parse/reconstruct loop-fission seam — all safe Rust, all byte-identical, each landed behind a paired win-rate gate (see docs/WHYS-decoder-perf.md).

These decode figures were measured with -C target-cpu=x86-64-v3 (this workspace's .cargo/config.toml). That setting is deliberately not shipped to consumers of the published crates — a library should not impose an ISA floor on its dependents — so a default cargo add rusty_h264 build compiles for baseline x86-64 and will be somewhat slower than the table above. To reproduce these numbers, build with RUSTFLAGS="-C target-cpu=x86-64-v3" (needs AVX2: Intel Haswell 2013+ / AMD Zen 2015+).

Method: pinned to one core, CPU time (not wall), arms ABBA-alternated, 9 pairs, 9/9 paired with z = 3.00 on every tier; frame counts compared between arms and every stream verified byte-identical to ffmpeg before timing. Earlier releases quoted "145 Mpx/s · 0.25×" from a differential harness that has since been refuted and replaced — see docs/WHYS-decoder-perf.md.

Decode is benched against ffmpeg's native h264 software decoder — a deliberately tougher bar than openh264's own h264dec. Full methodology, RD sweeps vs x264 and the reproducible harness: bench/ and docs/benchmarks.md.

Where this sits

Crate Role
rusty_h264 ← you are here — the public, safe facade API
rusty_h264-common bitstream I/O, Exp-Golomb, NAL/Annex-B, transforms, MC, deblock
rusty_h264-encoder the encode pipeline
rusty_h264-decoder the decode pipeline
rusty_h264-accel optional openh264 SIMD asm — the one unsafe crate

The workspace mirrors Cisco openh264's codec/ tree (common/encoder/ decoder/api/console).

Using it from remade_ffmpeg_rs

Depend on this facade and adapt to the rff-codec Encoder/Decoder traits — YuvFrame (I420 planes) ↔ VideoFrame. Note rusty_h264 speaks Annex-B (start codes), so an AVCC↔Annex-B shim is needed for MP4 inputs. Keep default-features = false in CI if you don't want a nasm build dependency there.

The Remade With Rust ecosystem

Remade With Rust is an initiative by Mata Network to rebuild essential C and C++ tools in Rust — for the memory safety, the predictable performance, and the freedom of a permissive license. Each project is a reimplementation, not a fork: same wire protocols and file formats, new code you can actually depend on. No copyleft. No surprises.

Project What it is
🎬 remade_ffmpeg_rs Our FFmpeg alternative. Drop-in ffmpeg and ffprobe binaries — demux → decode → filter → encode → mux, rebuilt as composable Rust crates with zero GPL/LGPL. Apache-2.0. rusty_h264 is its H.264 codec.
🧠 FFAI Our sister project: media for AI. "The AI media toolkit, remade with rust." Embedded ASR + TTS (Mercury), OCR (Carmenta) and vision-language captioning (Argus) behind an ffmpeg-style, swap-by-name architecture — no Python, no CUDA. MIT OR Apache-2.0.
🌐 Mata Network The home page. "Stop sacrificing your privacy for convenience." Sovereign, self-hostable privacy infrastructure — wallet & identity, password manager, contact manager, and a browser extension that stops information leaking as you browse. Remade With Rust is its open-source arm.

→ All projects: github.com/Remade-With-Rust

License

BSD-2-Clause — see LICENSE. No GPL/LGPL anywhere in the dependency tree, and no C/C++ either (CI-enforceable via cargo-deny). Embed it in closed-source software freely.