moq-video
Native video capture, encoding, decoding, and publishing for Media over QUIC.
The video counterpart to moq-audio.
Everything is native per-platform code with no ffmpeg dependency: capture, color
conversion, and the codec backends are all in-tree or thin wrappers over system
frameworks / vendored static libs. The public API is codec-agnostic, so no
signature, type, or error variant names a backend or a capture implementation;
swapping or bumping a backend crate is not a breaking change.
Capture
The opt-in capture feature exposes the device APIs and their per-platform
backends. Enable it with cargo add moq-video --features capture; the default
codec-only build accepts frames supplied by the caller without pulling Linux
V4L2 and libclang build dependencies.
Per-platform, picked at compile time:
- macOS: AVFoundation (camera) and ScreenCaptureKit (display, window, or
application), yielding zero-copy
CVPixelBuffersurfaces straight to VideoToolbox. - Linux: native V4L2 (camera; YUYV resampled, MJPEG via
zune-jpeg) and xdg-desktop-portal + PipeWire on Wayland (display; behind thepipewirefeature), with native X11 monitor/window selection and capture as the X11 fallback. The Wayland picker dialog chooses the screen, and the portal's restore token is reused so demand-driven reopens don't re-prompt. - Windows: native Media Foundation (camera;
IMFSourceReader) and DXGI Desktop Duplication (display), plus GDI single-window capture. Both convert BGRA to CPU I420 and use the ids returned by the enumerators.
capture::cameras() lists AVFoundation, V4L2, or Media Foundation cameras with
identifiers accepted by capture::Source::Camera. capture::displays() does
the same for macOS, Windows, and X11 displays. capture::windows() lists macOS,
Windows, and X11 windows. Wayland display selection stays in the desktop portal
picker, which does not expose a stable display identifier.
Embedded applications can consume raw capture without creating a MoQ broadcast:
let mut config = default;
config.source = Display;
let mut capture = open.await?;
while let Some = capture.read.await?
The stream retains only the newest unconsumed frame, so a slow encoder adds
drops rather than latency. read ends with None when the source stopped for
a benign reason, such as a window resize, so reopen to follow it. Permission
denial and a source disappearing are terminal, reported as
Error::PermissionDenied and Error::SourceUnavailable.
Encode
The codec is chosen via encode::Codec. Backends are tried in order (hardware
first, then software) and the first that opens wins; encode::Kind narrows the
choice (Auto / Hardware / Software / a named backend).
| Codec | Software | macOS | Windows | Linux | Android |
|---|---|---|---|---|---|
| H.264 | openh264 (vendored, static) | VideoToolbox | Media Foundation | NVENC (feature nvidia), VAAPI (feature vaapi) |
MediaCodec (feature mediacodec, API 26+) |
| H.265 | none | VideoToolbox | Media Foundation | NVENC (feature nvidia) |
MediaCodec (feature mediacodec, API 26+) |
Every backend emits Annex-B with in-band parameter sets (SPS/PPS, plus VPS for
H.265), so the matching moq_mux::codec importer handles framing and catalog
registration directly. There is no software H.265 encoder (it's hardware-only).
encode::Encoder::encode takes a raw Frame (a timestamp plus a Surface
holding the pixels) and returns encode::Encodeds: one whole access unit each,
carrying the timestamp of the picture it was encoded from. That matters for a
backend that buffers, which hands back an earlier frame's access unit while a
later one goes in, and for the tail finish() drains. Bring your own pixels with
Surface::rgba(...), or feed a frame straight from capture or decode.
Keyframes are automatic, at the Config::gop interval, so an application never
has to think about them. Encoder::keyframe() asks for one at the next frame when
something outside the encoder needs a decodable starting point there: opening a
new group, or resuming after an idle gap. The request is held until a frame
arrives, so it is safe to call before you have one.
Two public entry points:
encode::publish_capture(...)captures a webcam, encodes it, and publishes on demand: the track and catalog are advertised up front, but the camera opens only while a subscriber is watching and is released when the last one leaves.encode::Producerpublishes frames you encoded yourself (publish(&[Encoded])), handling the catalog and framing. Each is published at its own timestamp.
The NVENC, VAAPI, and V4L2 M2M backends are Linux-only. nvidia is on by
default: it dlopens the driver at runtime and needs nothing at build time.
vaapi and v4l2 are opt-in because their bindgen needs libclang on the build
host (plus the kernel headers for v4l2). None of them link a vendor library,
so a binary carrying them still links on a GPU-less builder and still starts on
a machine without the hardware, falling back to software.
Decode
decode::Consumer (the mirror of moq_audio::decode::Consumer) subscribes to an
H.264, H.265, or AV1 track and returns raw Frames. A hardware-decoded frame stays
on the GPU: feeding it back to a compatible hardware encode::Encoder on the
same device keeps it there (the transcode path), while into_i420() downloads
it. An encoder that can't take that surface (openh264, or a different device)
downloads it through I420 for you. Every frame carries a Surface, a
#[non_exhaustive] enum naming where the pixels live (PixelBuffer on macOS,
Texture on Windows, Cuda on Linux, HardwareBuffer on Android, or CPU
I420). Match it to take a zero-copy path for a representation you recognize, and fall back to
Surface::into_i420(), which always works. On macOS Surface::into_pixel_buffer()
is the mirror: free for a hardware-decoded frame, an upload for a CPU one.
Surface::to_rgba() and Surface::to_bgra() are the portable exits for CPU
image and UI toolkits, returning owned, tightly packed pixels with the surface's
color metadata applied. Both orders are there because toolkits disagree and the
conversion is a full pass over the frame: producing the order the caller wants
costs nothing extra, while producing the other one and swapping two channels
afterwards costs a second pass. They borrow the surface, so a frame held behind
an Arc shared with something else converts without being unwrapped first.
Backends are tried hardware-first, like encode:
| Codec | Software | macOS | Windows | Linux | Android |
|---|---|---|---|---|---|
| H.264 | openh264 (vendored, static) | VideoToolbox | Media Foundation (DXVA) | NVDEC (feature nvidia), VAAPI (feature vaapi) |
MediaCodec (feature mediacodec, API 26+) |
| H.265 | none | VideoToolbox | Media Foundation (DXVA) | NVDEC (feature nvidia) |
MediaCodec (feature mediacodec, API 26+) |
| AV1 | none | none | none | NVDEC (feature nvidia) |
MediaCodec (feature mediacodec, when the device provides it) |
On macOS VideoToolbox decodes H.264 and H.265 on hardware, pulling the parameter
sets (SPS/PPS, plus VPS for H.265) out of each keyframe to build the format
description. On Windows the Microsoft decoder MFT runs synchronously with a
Direct3D11 device bound to it, so the decode happens on the GPU through DXVA
(NVDEC / Intel / AMD). H.264 falls back to openh264 on a GPU-less host; H.265 has
no software decoder, so it needs the GPU path (on Windows, an HEVC decoder MFT:
the inbox HEVC Video Extensions or a vendor one). On Linux, NVDEC decodes H.264,
H.265, and 8-bit 4:2:0 AV1 to CUDA NV12 frames; AV1 is decode-only and is useful
for AV1 source to H.264/H.265 transcode rungs. VAAPI decodes H.264 to CPU I420 by
default; set decode::Config::gpu_frames to receive DMA-BUF surfaces that the
renderer can import without a download. A non-H.264/H.265/AV1 rendition yields
Error::UnsupportedCodec.