glass2glass (g2g)
A hardware-first, sans-IO, asynchronous multimedia graph framework in pure Rust.
One pipeline, five targets. g2g has a pure-Rust no_std core where alloc
itself is optional, so the same typed graph runs unchanged across the whole
hardware spectrum: MCU · RTOS · CPU · GPU · WASM. On the low end that is a
bare-metal Cortex-M with no heap at all; on the high end a GPU-resident
server pipeline. You write the graph once; the deployment shell picks the
executor and links the hardware.
The name reflects the metric the project optimizes for: glass-to-glass latency, the time between physical photon capture and hardware presentation.
See DESIGN.md for the architecture specification and
DEVTOOLS.md for the developer tooling (cargo xtask, the pipeline
visualizer, the caps explainer, benchmarks).
Quick start
A complete program (this exact code compiles and runs):
# Cargo.toml
[]
= { = "1", = ["macros", "rt-multi-thread"] }
= { = "https://github.com/Glass2GlassHQ/glass2glass" }
= { = "https://github.com/Glass2GlassHQ/glass2glass", = ["std"] }
// src/main.rs
use ;
use WallClock;
use default_registry;
async
With the wayland-sink feature enabled autovideosink opens a window. Headless
it falls back to fakesink, so the program runs anywhere. Or skip the project
entirely and run a pipeline from the repo:
Portability: one pipeline, five targets
The core (g2g-core) is pure Rust, no_std (with alloc an optional feature),
and sans-IO: the graph, the element traits, Caps negotiation, and the runner
are identical on every target. Only the deployment shell (which executor, which
hardware elements) changes.
| Target | What runs | How |
|---|---|---|
| MCU | a heap-free static pipeline on bare-metal Cortex-M | alloc is optional: the no-alloc build links no allocator at all, is proven panic-free, and has a budgeted ~KB-scale footprint. g2g-mcu peripheral elements (SPI display, camera / PCM capture, G.711 / ADPCM codecs, RTP egress + ingress, jitter buffer) over embedded-hal seams, plus interrupt/DMA capture and a bounded fault-recovery supervisor (retry / degrade / reset / watchdog). See the embedded section. |
| RTOS | the same static pipeline under a real RTOS task | one graph runs bit-exact under bare-metal, Embassy, FreeRTOS, and Zephyr; embassy-sync stack channels (embassy / embassy-link features) |
| CPU | the full media + protocol stack | Tokio, multi-thread on servers or current-thread on edge; the whole element library |
| GPU | zero-copy hardware pipelines | frames stay in Vulkan / CUDA / wgpu / DMABUF domains: Vulkan Video decode → wgpu::Texture, NVDEC / NVENC, CUDA ↔ wgpu bridge, no PCIe round-trip; embeddable in an app's own wgpu device (GpuContext::from_wgpu, packaged for Bevy as the bevy-g2g crate) |
| WASM | the same graph in the browser | wasm32, single-threaded (no cross-origin isolation), graphs can run in a dedicated Worker (presenting to an OffscreenCanvas): WebCodecs H.264 / H.265 decode, WebGPU present, in-browser or server-offloaded ML |
Same AsyncElement, same Caps, same runner on all five.
PORTABILITY.md runs one detection-overlay pipeline
whose processing stages come from a single shared overlay_stages() definition,
reused verbatim by the native (CPU) runner and the browser (WASM) build, and
gives reproducible evidence for each target (Cortex-M footprint, Embassy smoke,
CPU render, GPU-resident wgpu, in-browser canvas). See also
The four pillars.
OS-, GPU-, and device-coupled elements (camera, display, NVDEC/VA-API/Vulkan
Video, VideoToolbox, MediaCodec, ML device EPs) are experimental: they
compile, and some have host tests, but their runtime is not a CI promise.
g2g-inspect prints Stability experimental on those factories. See
STABILITY.md.
Also: QNX (safety-certified RTOS)
Beyond the five, the portable core is one spike away from QNX, the POSIX
microkernel that is the reference platform for ISO 26262 / IEC 62304
automotive/medical. QNX runs on application processors (aarch64 / x86-64), so it
is the std-capable path, not the MCU one. This is a Tier-0 portability spike
(compile-checked, not yet run): verified locally with no QNX SDP
(cargo +nightly ... -Zbuild-std), g2g-core (the no-alloc subset and the
full alloc + dynamic runtime layer: caps solver, autoplug, dynamic Graph),
g2g-mcu (the whole peripheral catalog), and the g2g-plugins no_std baseline
all compile for aarch64-unknown-nto-qnx800 and x86_64-pc-nto-qnx800 with zero
code changes. It stays clean because every OS/HW element is gated by a specific
target_os ("linux" / "windows" / "macos" / "android"), never
cfg(unix), so the Linux HW paths (VAAPI, DRM/KMS, dma-buf, v4l2, ALSA/PipeWire)
are excluded on nto rather than pulled in. The Tier 1 (std transports over
the free SDP; tokio-on-QNX is the one open dependency question) and Tier 2 (QNX
Screen display sink + vendor VPU via the C-seam) roadmap is in
PORTABILITY.md.
Migrating an existing pipeline?
Many gst-launch-1.0 lines run unchanged through g2g-launch. Paste one in:
Element names mostly match (with aliases: avdec_h264→ffmpegdec,
qtmux→mp4mux, autovideosink→waylandsink/kmssink, autovideosrc→v4l2src/libcamerasrc, ...). Inline caps
filters, tee name=t fan-out, muxer fan-in, decodebin/uridecodebin/playbin
and the encode-side encodebin/transcodebin
(encodebin profile="video/x-matroska:video/x-vp8,width=1280,height=720,bitrate=2000000:audio/x-opus",
which expands into those encoders plus matroskamux, splicing the scaler a
pinned geometry needs) all parse. When a line doesn't port, you get a hint, not a bare error:
$ g2g-launch videotestsrc ! theoraenc ! fakesink
parse error: unknown element: theoraenc
hint: `theoraenc` has no g2g element: no Theora encoder; use `vpxenc` (VP8/VP9) or `av1enc`
g2g-launch -v ...prints each link's negotiated caps + memory domain (thegst-launch -vanalog);--dotdumps a Graphviz graph.g2g-inspectisgst-inspect-1.0: list elements, dump one's properties/pads, or map a GStreamer name withg2g-inspect --gst x264enc. Scan an app's source with--gst-scan app.c.g2g-discoverisgst-discoverer-1.0:g2g-discover clip.mkvprints the container, each elementary stream with its codec and shape (video geometry, audio channels and sample rate), the duration, and the container's metadata. It sniffs the type, runs a headless probe graph to the first frame, and reads the demuxer's stream collection off the pipeline bus, so nothing is decoded.--jsonfor tooling. Local files only (a path or afile://URI); any other scheme is refused by name rather than fetched.g2g-device-monitorisgst-device-monitor-1.0: list cameras, audio devices, PipeWire nodes, and (a g2g extension)Compute/GPUdevices, with probed caps and the launch fragment that opens each (v4l2src device=/dev/video0). Filter by class (g2g-device-monitor Video/Source),--jsonfor tooling,--followfor live hotplug (native PipeWire and WASAPI events, poll-and-diff elsewhere). Backends: V4L2 / ALSA / PipeWire / GPU on Linux, Media Foundation + WASAPI on Windows, AVFoundation + Core Audio on macOS. Every device's id is what its element's selection property takes, so a saved launch line reopens the same hardware after a replug (v4l2srctakes the id asdevice-id=, since a/dev/videoNpath is not stable). A V4L2 camera's listing also carries the controls it reports and the range each accepts, under the namesv4l2src extra-controls=matches them by.- Migrate incrementally in either direction:
g2g-bridgeembeds a g2g sub-graph inside a GStreamer pipeline;gstwraphosts an un-ported GStreamer element inside a g2g graph.
Full guide, including the equivalence cookbook and application/element porting: PORTING.md.
Scripting: config files and Rhai
A gst-launch string is the one-liner. For version-controllable, generated, or
computed pipelines there are three more surfaces, all built on the same registry
and negotiation as parse_launch (so any element / caps / policy works):
-
Declarative graphs (JSON / YAML),
--graph.nodes+edges, with a{ id, caps }capsfilter shorthand and a top-levelpipeline:escape hatch. Roles follow link degree (auto source / sink / muxer / auto-tee), and property values are typed by the target element exactly as in a launch string.# pipe.yaml nodes: - - # a capsfilter shorthand - edges: - - -
Rhai builder scripts,
--script. Where a document is a fixed graph, a script computes one (loops, parameters, conditionals) via a small builder API (add/caps/set/link/link_leaky) that emits the same graph model. Rhai is pure Rust, so this works on the same wasm / embedded targets the core reaches (--features script-rhai).// Fan N cameras into one funnel, sized at runtime. add("funnel", "mix"); for i in 0..num_cams { let id = "cam" + i; add("rtspsrc", id); set(id, "location", cams[i]); link(id, "mix"); } add("autovideosink", "screen"); link("mix", "screen"); -
scriptelement: per-frame logic in Rhai. A raw-video transform whoseprocess(frame)runs on every frame, the pure-Rust cousin ofpyelement. Theframeis a zero-copy handle to the live buffer: index it in place and read its geometry.g2g-launch videotestsrc ! scriptelement script="fn process(f){ f.invert(); }" ! autovideosinkfn process(frame) { // frame.width / .height / .format / .pts / .sequence / .len frame[3] = 128; // per-pixel edit in place (convenient; interpreted) frame.invert(); // whole-frame native ops: fill(v) / invert() / apply_lut(lut) }Performance model (the NumPy rule): the script is the control plane, native code is the data plane. A per-pixel Rhai loop over an HD frame is milliseconds per pixel-thousand (interpreted — inherent to any embedded scripting language), so use it for logic, metadata, and small regions. For whole-frame transforms call a native bulk op (
invert()~1 ms/frame vs a per-pixel loop's seconds), and build the general per-value transform (brightness, gamma, threshold, ...) as a 256-entryapply_lut(lut). For heavy per-pixel math, write a compiled element instead. -
scriptrouter: script-decided routing to N outputs. A 1-to-N demux whoseroute(frame)picks which output port each buffer goes to: an index (negative drops it), or an array like[0, 1]to multicast one buffer to several ports at once (a shared copy per port). Put anappsink channel=...on each branch and you have per-buffer routing into your own code / pipelines — the buffers go where the script says, and each channel ispull()ed live from Python/C/Rust just like a GStreamerappsink. Media-agnostic (audio, video, byte streams);routereadsframe.pts/.sequence/.keyframe/.lenand can peek bytes (frame[i]).# Split an audio stream to two consumers by parity; pull each from your app. g2g-launch whepsrc uri=... ! opusdec ! audioconvert ! \ scriptrouter name=r script="fn route(f){ f.sequence % 2 }" \ r.0 ! appsink channel=even r.1 ! appsink channel=odd, = , # pull() each, feed anywhereA runnable end-to-end demo (routes to two pull channels drained live, with real
pull()timing):cargo run -p g2g-plugins --features script-rhai --example scriptrouter_appsink_egress
Embedded: heap-free pipelines on a bare-metal MCU
The MCU end of the spectrum is not a stripped-down build, it is the same graph
with a hard guarantee: alloc is an optional feature, and the default build
links no allocator at all. That makes it a fit for safety-critical, no-heap
targets (MISRA, certification processes).
- Static element model. A heap-free pipeline is a compile-time-static graph
of concrete typed elements (
g2g_core::staticelem:StaticSource/StaticTransform/StaticSinkwithasync fnin trait, const-arity runners, aChaincombinator), so every stage's future is unboxed, nodyn, noBox, no allocation. Buffers are lent zero-copy from a const-genericStaticLendRingsized at compile time. The full dynamic runner carries the same steady-state contract on the host:run_graphpushes and processes without a single per-frame heap allocation (counting-allocator proven over 100k frames). - The guarantees are machine-checked in CI, not asserted. The linked archive
carries zero allocator symbols and zero panic symbols (
tools/noalloc-check.sh); a gc-sectioned ELF is measured for ROM / static RAM / worst-case stack and budget-enforced (tools/footprint-report.sh); the pipeline then executes on an emulated Cortex-M (tools/qemu-check.sh) and a per-frame timing / jitter report runs under deterministic QEMU-icount(tools/timing-report.sh). App code on this surface needs zerounsafe. - One graph, four executors. The same static pipeline runs bit-exact under a bare poll loop, Embassy, FreeRTOS (C-ABI staticlib), and Zephyr (a drop-in Zephyr module the app lists in its west manifest).
- Integrates with your existing C, both directions. A C/RTOS app can link
the pipeline as a static library and call in, or your existing C drivers can
be the peripheral:
g2g-mcu::cffi'sCFrameGrabber/CPacketSenderwrap C capture/send function pointers, andstep_source_sinkhands control back to your superloop after each frame. Zero Rust on the driver side; proven from a real C caller (examples/g2g-cffi), still heap-free and data-panic-free. g2g-mcuperipheral elements. Heap-free elements written againstembedded-haltrait seams rather than chip registers, so the driver logic is host-tested against the datasheet with mock peripherals and a board port is just the vendor HAL's trait impls:SpiDisplaySink(ST7789 / ILI9341, whole frame or banded streaming for panels too large to ring-buffer),GrabberSrc(DCMI/CSI camera),PcmSink(I2S/SAI), the fixed-point G.711 and IMA ADPCM codecs (bit-exact vs ffmpeg), the hardware JPEG-decode and H.264-encode seams (HwJpegDec/HwH264Enc, the peripheral reached over anembedded-halor a C-function-pointer driver),YuyvToI420(heap-free camera-4:2:2 → encoder-4:2:0 convert), andRtpSink.- Interrupt/DMA-driven capture. Real capture runs in interrupt context, so
SpscFrameRingis a lock-free, heap-free single-producer/single-consumer FIFO a DMA-completion ISR fills while the pipeline drains it (SpscCaptureSrc, sleeping onwfibetween frames), with bounded back-pressure (a full ring drops and counts, never stalls the interrupt). Proven on emulated Cortex-M: a SysTick interrupt feeds acapture → G.711 → checksumpipeline lossless and in order, bit-exact against synchronous delivery. Atomic load/store only, so it works on cores without CAS (thumbv6m). - Runtime fault recovery. A supervisor (
g2g_core::supervise) turns a returned peripheral fault into a bounded, deterministic action instead of aborting: aFaultPolicychooses retry / skip (degraded mode) / reset / escalate, aRecoverseam re-initializes the faulting stage (re-arm the camera, re-open the socket), and aWatchdogis petted only on real forward progress, so a wedged or escalated pipeline stops petting and a hardware watchdog resets the chip. ASupervisorReportaccounts every fault for the safety case. Proven on emulated Cortex-M: the pipeline recovers a latched mid-stream capture fault (all frames still delivered) and escalates a dead peripheral within its bounded ladder without hanging. - Receive direction (RTP ingress + jitter buffer). The inverse of the
capture→egress reference pipeline:
RtpSrcreceives and parses RTP (a wire-tolerant, bounds-checked header parser shared with the std depayloader), a heap-freeJitterBuffer<N, BYTES>absorbs arrival jitter and reorders by sequence number, handling reorder / duplicate / late / loss explicitly and countably, andG711Decdecodes. Proven on emulated Cortex-M: a reordered RTP wire is reconstructed to the ordered PCM stream, verified by an order-sensitive hash against an independent in-order decode. - I2C sensors + UART transport. Beyond the media pipeline:
Sht3xSrcis a real SHT3x temperature/humidity driver overembedded-halI2C (datasheet single-shot command, CRC-8 validation with the datasheet's0xBEEF→0x92vector, fixed-point conversion), andUartSink/UartSrcare a byte-stream egress / ingress over local serial seams. An I2C-sensor→UART telemetry pipeline is proven on emulated Cortex-M. - A checkable safety case.
docs/safety/carries a requirements traceability matrix (15 requirements, each linked to the proof script / test / CI job that verifies it) and a safety manual (conditions of use, assumptions, theunsafeinventory).tools/traceability-check.shfails in CI if any cited evidence goes missing, so the matrix can't drift from the code;tools/qualification-kit.shruns the whole proof set into one report. A down-payment on a functional-safety case, not a certificate (pre-1.0, emulated not silicon). - ARM and RISC-V. The static element model is ISA-agnostic pure Rust, so
the no-alloc core and
g2g-mcubuild unchanged forriscv32imafc(ESP32-P4 class). The heap-free + panic-free symbol proofs and the footprint report run on boththumbv7emand RISC-V — the portability claim is machine-checked, not asserted. - The reference deterministic-audio graph.
capture -> convert -> resample -> mix -> encode -> RTPcomposed as one static heap-free pipeline, fully fixed-point, so its RTP wire bytes are bit-exact across every target: pinned by a host test (DSP validated against an independent float reference) and re-verified on-target on all four executors.
g2g-mcugen, the host graph compiler. Develop and test on Linux, ship a
bounded static build to the MCU: a declarative graph document compiles to the
monomorphized static pipeline, with every ring sized from the graph's frame
geometry and a total ring-memory budget reported. It is a general MCU graph
compiler, not audio-only, spanning an audio catalog and a video / display one:
# camera -> SPI panel, one static pipeline. `g2g-mcugen display.yaml -o graph.rs`
name: display
frame_ns: 33333333 # ~30 fps
frames: 64
nodes:
-
-
edges:
-
A mis-wired graph (an encoder fed the wrong sample width, a mixer whose inputs
disagree, a display fed the wrong pixel format) is rejected with a diagnostic
before a line of Rust is emitted, and the generated pipeline reproduces the
hand-written reference's wire output byte-for-byte (checked in CI for both
catalogs, tools/mcugen-check.sh).
The four pillars
- Async execution. Every element is a cooperative
Future. The framework is runtime-agnostic (Tokio on servers, Embassy on RTOS,wasm-bindgen-futuresin the browser). - Hardware-first, zero-copy. Buffers live in DMABUF / Vulkan /
CUDA / D3D11 / WebGPU memory domains. Negotiation settles a zero-copy
path per link where one exists and auto-plugs a converter where none
does, so every remaining copy is explicit (
g2g-launch -vshows each link's domain). no_std,alloc-optional, sans-IO core. The same pipeline shape runs on a bare-metal Cortex-M with no heap, an RTOS (Embassy / FreeRTOS / Zephyr), a multi-threaded server, a GPU-resident pipeline, orwasm32(see Portability).- First-class ML. Tensor allocation, reshaping, and pipeline batching are part of graph orchestration.
Workspace
| Crate | Role | Profile |
|---|---|---|
g2g-core |
Traits, Frame/PipelinePacket, caps algebra, clock, runner, static element model. |
no_std, alloc optional |
g2g-mcu |
Heap-free MCU peripheral elements (SPI display incl. banded streaming, camera / PCM capture, I2C sensor, UART, G.711 / ADPCM codecs, hardware JPEG-decode / H.264-encode seams, RTP egress + ingress, jitter buffer, fault-recovery watchdog) over embedded-hal / C-callback seams. |
no_std, no alloc |
g2g-mcugen |
Host graph compiler: a declarative MCU graph (YAML/JSON, audio or video/display) → a monomorphized static pipeline (heap-free Rust). | std (host tool) |
g2g-plugin |
SDK for dynamically loadable plugins: same-toolchain (declare_plugin! + ABI tag) and the frozen C ABI v2, so cross-toolchain Rust and plain-C plugins load too. |
no_std + alloc |
g2g-plugins |
Sources/sinks/transforms (RTSP, RTP in/out, HTTP/HLS/DASH/RTMP ingest, V4L2 / PipeWire / MF capture, ffmpeg, VAAPI, MF, VideoToolbox (macOS), MediaCodec (Android), Wayland, KMS, WASAPI, ALSA / PulseAudio / PipeWire audio, compositor, Embassy, web), container mux/demux (MP4, MPEG-TS, MPEG-PS, Matroska/WebM, FLV, Ogg), codec parsers + encoders (AV1, VP8/9, MJPEG), the tag system, and the gst-launch text DSL. |
mixed |
g2g-ml |
ORT, Burn, WgpuPreprocess, TensorPostprocess, multi-stream tensor batcher. | std |
g2g-bridge |
GStreamer C-FFI bridge. | std |
g2g-python |
Hosts gst-python-ml elements in-process (embedded CPython via pyo3). | std |
g2g-capi |
C ABI (cdylib/staticlib + g2g.h): launch pipelines + bus + appsrc/appsink from any language. |
std |
g2g-pyapi |
Python (pyo3) bindings: drive pipelines + bus + appsrc/appsink. | std |
Build
Stable Rust, resolver = "2". MSRV 1.92, except the embedded-facing crates
(g2g-core, g2g-mcu, g2g-mcugen, g2g-plugin), which build on 1.86 so a
vendor-pinned toolchain can consume the portable core (see STABILITY.md).
OS-coupled elements live behind cargo features:
| Element | Feature | Platform / system dep |
|---|---|---|
RtspSrc |
rtsp |
retina |
H264Parse |
(default) | — |
FfmpegH264Dec (sw / NvdecCuvid / NvdecCuda / Vaapi) |
ffmpeg |
Linux + libavcodec |
VaapiH264Dec |
vaapi |
Linux + libva + GBM |
MfDecode / MfEncode / MfAacEncode / MfAacDecode |
mf-decode, mf-encode, mf-aac |
Windows + Media Foundation |
VtDecode / VtEncode (H.264 / H.265, validated on the CI Mac; zero-copy CVPixelBuffer output via cv-output) |
vtdecode, vtencode |
macOS + VideoToolbox |
MediaCodecDec (H.264 / H.265, on-device validated; zero-copy GPU output via with_gpu_output) |
mediacodec, mediacodec-wgpu |
Android + NDK MediaCodec (+ wgpu / Vulkan for GPU output) |
WaylandSink |
wayland-sink |
Linux + Wayland |
KmsSink |
kms-sink |
Linux + libdrm; needs DRM master / tty |
D3D11Sink |
d3d11-sink |
Windows |
MetalVideoSink (zero-copy from CVPixelBuffer, validated on the CI Mac) |
metal-sink |
macOS + Metal |
WgpuPresentSink (wgpusink: owns its Wayland window, presents GPU-resident frames with no upload) |
wgpu-present |
Linux + Wayland + wgpu |
NvDec (native NVDEC H.264/H.265/AV1 → CUDA NV12 or 10-bit P010, NVCUVID) |
nvdec |
Linux + NVIDIA driver (libnvcuvid) |
NvEnc (native NVENC CUDA NV12/P010 → H.264/H.265, incl. HEVC Main 10) |
nvenc |
Linux + NVIDIA driver (libnvidia-encode) |
CudaDownload (CUDA → System), CudaUpload (System → CUDA) |
cuda |
Linux + NVIDIA driver (libcuda) |
CudaGlSink (CUDA-GL present), CudaKmsSink (CUDA-GL on KMS) |
cuda-gl, cuda-kms |
Linux + NVIDIA + EGL + GL (+ libdrm for KMS) |
CudaToWgpu / WgpuToCuda (CUDA ↔ wgpu zero-copy bridge) |
cuda-wgpu |
Linux + NVIDIA + Vulkan |
UdpSink + RTP packetizer, or raw datagrams (multiudpsink clients=) |
udp-egress |
— |
UdpSrc (RTP ingest + jitter buffer + RTCP/NACK, or raw MPEG-TS datagrams) |
udp-ingress |
— |
SrtpEnc / SrtpDec (RFC 3711 / RFC 7714 SRTP and SRTCP, per-SSRC receive contexts) |
srtp |
— |
DtlsSrtpEnc / DtlsSrtpDec (DTLS-SRTP handshake over the media socket keys SRTP) |
dtls-srtp |
— |
TcpServerSrc / TcpClientSrc / TcpServerSink / TcpClientSink (plain TCP byte streams) |
tcp |
— |
ShmSink / ShmSrc (GStreamer's shm protocol: shared-memory frames + unix control socket) |
shm |
Linux |
RtmpSrc (RTMP publisher ingest) |
rtmp |
— |
WebRtcSink (WHIP egress, H.264 + Opus) / WebRtcWhepSrc (WHEP ingest, H.264), via str0m: ICE/DTLS/SRTP, trickle ICE + ICE restart, NACK/RTX |
webrtc |
str0m (rust-crypto) + reqwest |
WebRtcDataSrc / WebRtcDataSink (P2P data channels on SCTP) |
webrtc |
str0m |
MoqtSink (IETF MoQ Transport draft-16/18 publisher: fMP4 → groups / objects over WebTransport, on subgroup streams or datagrams) |
moqt |
web-transport-quinn |
MoqtSrc (IETF MoQ Transport draft-16/18 subscriber: catalog read, stream + datagram reassembly → fMP4) |
moqt |
web-transport-quinn |
LiveKitSink (publish into a LiveKit room: JWT + protobuf signalling) |
webrtc-livekit |
+ tokio-tungstenite |
HttpSrc (HTTP(S) byte-stream source) |
http-src |
reqwest |
HlsSrc (HLS: TS + fMP4/CMAF, live, AES-128 / SAMPLE-AES) |
hls |
reqwest + aes |
DashSrc (DASH: SegmentTemplate / SegmentTimeline, live, CMAF chunked low latency) |
dash |
reqwest + roxmltree |
V4l2Src |
v4l2 |
Linux + V4L2 (/dev/videoN) |
WasapiSink / WasapiSrc |
wasapi-sink, wasapi-src |
Windows |
AlsaSink |
alsa-sink |
Linux + libasound |
PulseSink |
pulse-sink |
Linux + libpulse |
PipeWireSink / PipeWireSrc (audio) |
pipewire |
Linux + libpipewire |
PipeWireVideoSrc (video capture, io-mode=mmap or dmabuf) |
pipewire |
Linux + libpipewire |
PipeWireVideoSrc portal=true (screen capture via xdg-desktop-portal) |
portal |
Linux + libpipewire + a desktop portal |
MfVideoSrc (camera) |
mf-video-src |
Windows + Media Foundation |
Av1Enc (pure-Rust rav1e) |
av1-encode |
— |
VpxEnc (VP8 / VP9 via libvpx) |
vpx |
libvpx |
MjpegDec / MjpegEnc (pure Rust) |
mjpeg, mjpeg-encode |
— |
PngDec / PngEnc (still images, pure Rust) |
png |
— |
WebPDec (lossy + lossless stills, pure Rust) |
webp |
— |
AnalyticsOverlay (CPU) / VelloAnalyticsOverlay (GPU) (detection boxes, segmentation masks, ROIs) / WgpuSink |
analytics, vello-overlay, wgpu-sink |
wgpu (GPU variants) |
VelloTextOverlay (subtitle cues drawn on the GPU, WgpuTexture out) |
vello-text-overlay |
wgpu |
OrtInference (+ CUDA / DirectML EPs) |
ort, cuda, directml (in g2g-ml) |
onnxruntime |
BurnInference (linear layer, or an ONNX topology imported by burn-onnx codegen) |
burn (in g2g-ml) |
wgpu (Vulkan / Metal / DX12) |
WgpuPreprocess (NV12 / YUYV system bytes, a dma-buf, or a GPU texture in, NCHW tensor out) |
wgpu, dmabuf-wgpu, mediacodec-wgpu (in g2g-ml) |
wgpu (Vulkan for the dma-buf import) |
| Embassy / RTOS pool + clock | embassy, embassy-link |
— |
| Browser elements | web, web-codecs |
wasm32-unknown-unknown |
The container parsers and muxers (mp4src / mp4mux, tsdemux / mpegtsmux,
matroskademux / matroskamux, flvdemux / flvmux, oggdemux / oggmux,
avidemux / avimux, fmp4demux, mpegpsdemux, multipartdemux / multipartmux,
y4mdec / y4menc, aiffparse / aiffmux, auparse / avmux_au), the bitstream parsers (h264parse, h265parse, aacparse,
mpegaudioparse + id3demux / apedemux, ac3parse, opusparse, vp8parse, vp9parse, av1parse,
jpegparse, pngparse) and the headerless framers (rawvideoparse /
rawaudioparse, a .yuv / .pcm dump cut into buffers from declared
properties),
the G.711 / IMA ADPCM codecs (mulawenc / mulawdec, alawenc / alawdec,
adpcmenc / adpcmdec), the software video/audio
transforms (videoscale / videorate / imagefreeze / videocrop / videoflip /
videobalance / videobox / alpha / gamma / deinterlace / timeoverlay /
aspectratiocrop / gaussianblur / videomedian / smooth / coloreffects /
chromahold / zebrastripe / videodiff / solarize / chromium / dilate / dodge /
exclusion / burn,
audioconvert / audioresample / audiorate / audiomixer / interleave /
deinterleave / scaletempo / volume / audiopanorama /
audioamplify / audioecho / audiodynamic / audiowsinclimit / audiocheblimit /
audiochannelmix / audiomixmatrix / stereo / audiofirfilter / audioiirfilter /
removesilence / audiobuffersplit / speed /
level / cutter / equalizer-3bands / spectrum), the KLV telemetry codec (klvdecode, MISB ST 0601 / STANAG 4609),
the bitmap-subtitle decoders (vobsubdec, alias dvdsubdec, dvbsubdec and
pgsdec) with the subpictureoverlay that blends their cues onto video,
and the EBU teletext subtitle decoder (teletextdec),
the subtitle reader and writers (subparse, srtenc, webvttenc) and the
closed-caption elements (ccextract / ccinsert and the ccconverter that
moves captions between the cc_data, CDP, S334-1A and raw CEA-608 layouts),
the flow-control and debug elements (concat / input-selector /
output-selector / valve / fakesrc / fdsrc / fdsink / watchdog /
capssetter / taginject / rndbuffersize / errorignore / breakmydata /
chopmydata / checksumsink / fakevideosink / fakeaudiosink / progressreport), the compositor, the tag system, and the
gst-launch text DSL (parse_launch / gst-inspect) are all in the pure
no_std + alloc default build. The std build adds clockoverlay, fpsdisplaysink, the
multifilesink / multifilesrc image-sequence pair (imagesequencesrc when it
stamps a framerate), splitfilesrc (the parts of a cut recording read as one
byte stream), dataurisrc (a data: URI's payload), vobsubsrc (a DVD
subtitle .idx / .sub sidecar pair), splitmuxsink
(segmented recording, muxer=mp4|matroska|mpegts), and hlssink (HLS
packaging: segment files plus an .m3u8 playlist, fed by tsmux or mp4mux).
Sample pipelines
The graph API is run_source_transform_sink / run_linear_chain /
run_source_fanout / run_muxer_sink over typed elements. Examples are
condensed; full versions live in the integration tests under
g2g-plugins/tests/.
RTSP → ffmpeg decode → Wayland window
let src = new;
let dec = new.with_output_format;
let sink = new;
run_source_transform_sink.await?;
Features: rtsp ffmpeg wayland-sink.
RTSP → NVDEC (CUDA device memory) → CUDA-GL display
Zero-copy after decode: NV12 stays in CUDA device memory until the GL fragment shader samples it.
let src = new;
let dec = with_backend; // MemoryDomain::Cuda
let sink = new; // EGL on Wayland, NV12 shader
run_source_transform_sink.await?;
Features: rtsp ffmpeg cuda cuda-gl. Linux + NVIDIA only. See
DESIGN.md §4.11.5.
Native NVDEC → NVENC transcode, GPU-resident, with domain auto-plug
NvDec / NvEnc drive NVCUVID / NVENC directly (no libavcodec). The decode
stays in MemoryDomain::Cuda straight into the encoder. Memory-domain
negotiation settles a shared domain when one exists; where it can't (a CPU-side
NV12 source feeding the CUDA-only NvEnc), auto_plug_cuda_converters splices a
CudaUpload automatically — no hand-wiring.
let mut g: = new;
let src = g.add_source; // System NV12
let enc = g.add_transform; // CUDA NV12 → H.264
let snk = g.add_sink;
g.link.unwrap;
g.link.unwrap;
let g = auto_plug_cuda_converters; // splices CudaUpload: src → [CudaUpload] → enc → snk
run_graph.await?;
Features: nvenc (nvdec for the decoder). Linux + NVIDIA only. NvDec
itself is multi-domain: driven by downstream demand it keeps frames on the GPU
(zero-copy) or downloads to System. It decodes H.264 / H.265 / AV1, emits P010
for a 10-bit stream, reconfigures in place on a mid-stream resolution change, and
takes max-display-delay to trade latency for decode/display pipelining and
cuda-device-id to pick the GPU on a multi-card host. NvEnc
encodes P010 as HEVC Main 10 and takes gop-size / repeat-sequence-header for
periodic IDRs carrying their own SPS/PPS.
RTSP → decode → KMS (tty / no compositor)
let src = new;
let dec = new.with_output_format;
let sink = new.with_device;
run_source_transform_sink.await?;
Features: rtsp ffmpeg kms-sink. Run from a tty after stopping the
display manager (KMS sink needs DRM master).
RTSP → decode → ML preprocess → ORT inference → postprocess
let src = new;
let mut dec = new.with_output_format;
let mut preproc = new; // NV12 -> f32 NCHW on GPU
let mut inference = from_memory_with_cuda?;
let mut post = topk_classification;
run_linear_chain.await?;
Features: rtsp ffmpeg (plugins) + wgpu cuda (g2g-ml). The CUDA execution
provider falls back to CPU if no CUDA runtime is present.
Android: MediaCodec decode → GPU → ML preprocess (zero-copy)
// Decode on the NDK MediaCodec and keep the frame on the GPU as an RGBA wgpu
// texture (no CPU NV12 pack); WgpuPreprocess samples it straight into a tensor.
let dec = h264.with_gpu_output; // MemoryDomain::WgpuTexture (RGBA)
let preproc = new; // samples the texture -> f32 NCHW
// dec -> preproc -> OrtInference / BurnInference, all on the GPU
Features: mediacodec-wgpu (plugins) + mediacodec-wgpu (g2g-ml). Android only,
validated on a Pixel 10a. The decoded AHardwareBuffer is imported into Vulkan
and converted to RGBA through an immutable VkSamplerYcbcrConversion compute
pass (the conversion wgpu's bind-group API cannot express), then handed
downstream as a wgpu::Texture — the frame never touches the CPU.
File → H.264 parse → fMP4 record
let graph = parse_launch?;
run_graph.await?;
MPEG-TS file → demux → H.264 parse → decode → Wayland
The container demuxers (tsdemux, matroskademux, flvdemux, oggdemux,
fmp4demux, mpegpsdemux) accept a Caps::ByteStream and split out elementary
streams. mpegpsdemux reads .mpg / .vob program streams, including their DVD
subpicture tracks. Every playbin video branch carries a deinterlace mode=auto (yadif): the decoder declares interlaced streams in its output caps
(interlace-mode=interleaved) and the filter weaves only those, so interlaced
MPEG-2 or H.264 plays clean from any container while progressive content passes
through untouched.
let src = new;
let mut demux = new.with_stream; // PAT/PMT/PES -> Annex-B
let mut parse = new;
let mut dec = new.with_output_format;
let sink = new;
run_linear_chain.await?;
Features: ffmpeg wayland-sink.
STANAG 4609 (drone / ISR): KLV telemetry alongside the video
tsdemux stream=klv splits the MISB metadata stream out of the same multiplex
(private PES with the KLVA registration, or metadata-in-PES 0x15), and
klvdecode parses each ST 0601 UAS Datalink Local Set into a timed
key=value text line, ready for textoverlay or an app sink. The tag table
covers the telemetry core plus identity strings, target geometry, and the
nested ST 0102 security local set; the parser is validated against the
published MISMMS reference packet with klvdata as the oracle. The mux
direction takes Caps::Klv packets (built with UasDatalink::encode) on a
mpegtsmux input, and rtpklv carries KLV over RTP (RFC 6597) for
low-latency links. ffmpeg-validated bit-exact both ways, and validated against
a real UAS capture (the public "Day Flight" sample from samples.ffmpeg.org:
point G2G_STANAG_SAMPLE at it to run the local klv_stanag_sample test).
The rest of the ISR stack sits on the same codec: vmti for ST 0903
moving-target reports with their nested mask / ontology / tracker / chip sets
(including vmti_from_analytics, which turns an in-pipeline detector's output
into VTargets), ST 1204 MIIS identifiers (standard text form included),
misptimeinsert / misptimeextract for ST 0604 timestamps in H.264 / H.265
SEI, st2022fec for SMPTE 2022-1 loss recovery on a contribution link, and
cotsink to put a drone track on a TAK / ATAK network as Cursor-on-Target
events, optionally with the ST 0805.1 sensor point of interest alongside. SRT
carries it encrypted (passphrase= on srtsink / srtsrc).
let src = new;
let demux = new.with_stream;
let dec = new; // -> "ts=.. lat=.. lon=.. alt=.. heading=.." lines
Adaptive streaming: HLS / DASH → decode → display
let src = new; // or DashSrc::new(mpd_url)
let mut demux = new.with_stream;
let mut parse = new;
let mut dec = new.with_output_format;
let sink = new;
run_linear_chain.await?;
Features: hls ffmpeg wayland-sink (dash for the DASH front end). HlsSrc
follows live playlist reloads and decrypts AES-128 / SAMPLE-AES segments;
DashSrc handles SegmentTemplate / SegmentTimeline and dynamic (live) MPDs,
and with low-latency=true consumes a CMAF segment chunk by chunk as the packager
writes it. Both prebuffer ahead by duration (prebuffer-ms), posting Buffering
bus levels while they fill, like HttpSrc's byte window (prebuffer-bytes).
gst-launch text pipeline
parse_launch builds a runnable Graph from a GStreamer-style string against
the default_registry, including caps filters, tee branching, and muxer
fan-in. Registry::inspect(name) is the gst-inspect analog.
let graph = parse_launch?;
run_graph.await?;
Registered launch elements include videotestsrc / audiotestsrc, the SW
transforms, the demuxers (tsdemux, matroskademux, flvdemux, oggdemux, mpegpsdemux)
and muxers (mpegtsmux, matroskamux, flvmux, oggmux, funnel, audiomixer),
filesrc / filesink, and fakesink. Feature-gated capture / decode / display
elements register their launch factories when their feature is enabled, and the
autovideosink / autoaudiosink aliases resolve to whichever sink is built
(falling back to fakesink, so a tutorial line runs headless).
The ML elements live in a separate crate, so an app opts them in after building
the registry: g2g_ml::register(&mut reg) (the launch feature) adds
ortinfer, wgpupreprocess, detectionpostprocess, and ortsegment (instance
segmentation), making
... ! ortinfer model=yolov8n.onnx ! detectionpostprocess conf-threshold=0.3 ! ... parse.
Camera → encode → RTP egress over UDP
let src = new; // RGBA test pattern, unbounded
let enc = new.with_hardware; // Windows; on Linux use NvEnc / ffmpeg
let sink = new
.with_rtp; // payload type, SSRC
run_source_transform_sink.await?;
Features: udp-egress (plus the platform encoder feature). UdpSink honors
receive-side NACK by retransmitting from a bounded send history
(with_retransmit).
RTP ingress over UDP → ffmpeg decode → Wayland
The receive-side inverse, with a jitter buffer (reorder / bounded-latency loss handling) and RTCP feedback (periodic receiver reports, NACK on gaps) built in.
let src = new
.with_jitter // 50 ms hold, 64-packet depth
.with_rtcp; // 1 s reports, NACK enabled
let dec = new.with_output_format;
let sink = new;
run_source_transform_sink.await?;
Features: udp-ingress ffmpeg wayland-sink.
Picture-in-picture: webcam over a test pattern (compositor)
let bg = new.with_pattern;
let cam = new.with_size; // -> VideoConvert(RGBA) -> VideoScale
let comp = new;
// bg -> comp.input(0); cam -> rgba -> scale -> comp.input(1); comp -> sink (see tests).
Or as a launch line (M876), placement via the flattened pad properties:
videotestsrc ! c. v4l2src device=/dev/video0 ! videoconvert ! videoscale ! c. \
compositor name=c width=1280 height=720 sink1-xpos=940 sink1-ypos=460 sink1-zorder=1 ! waylandsink
Features: v4l2 wayland-sink. Full graph in
g2g-plugins/tests/pip_smoke.rs.
WgpuCompositor is the bit-exact GPU sibling (composites WgpuTexture frames
in place, zero-copy); with_timed_output() (timed-output=true) holds the output
rate over a stalled input whenever the pipeline clock can sleep on a deadline.
Camera → MoQ Transport → a browser
libcamerasrc width=640 height=480 framerate=30 ! videoconvert ! x264enc ! mp4mux \
! moqtsink location=https://127.0.0.1:4443/ namespace=live
tools/moqt-demo/ runs that end to end: node watch-live.mjs
starts a local moq-relay-ietf, publishes the camera into it and opens a browser
that subscribes and plays. The page's MoQT client is the third-party
MOQtail draft-16 implementation, so the
browser decodes our bytes with nothing shared from the Rust side.
node headless/run-moqt-play.mjs is the same path in headless Chromium with
assertions on the decoded frames. Features: libcamera moqt ffmpeg.
Running smoke tests
Most integration tests are marked #[ignore] because they need a live RTSP
feed and/or a display. The pattern is the same across recipes:
A standing RTSP feed
A typical loopback setup uses mediamtx as the relay and ffmpeg as the
publisher. In one terminal:
In a second terminal, push a synthetic H.264 feed into it:
A public RTSP feed (Wowza demo, IP camera on the LAN, etc.) also works — the smoke tests don't care where the stream comes from.
Software decode + Wayland
G2G_RTSP_TEST_URL=rtsp://localhost:8554/pattern \
A window titled "glass2glass" appears showing the feed.
NVIDIA NVDEC (system memory) + Wayland
G2G_DECODER=nvdec \
G2G_RTSP_TEST_URL=rtsp://localhost:8554/pattern \
G2G_TARGET_FRAMES=300 \
G2G_TARGET_FRAMES >= 300 is needed to amortize cuvid startup (libnvcuvid
load, CUDA context, surface pool alloc) for meaningful p50 / p95 latency
numbers. Compare against G2G_DECODER=software on the same feed.
NVIDIA NVDEC → CUDA → CUDA-GL zero-copy display
G2G_RTSP_TEST_URL=rtsp://localhost:8554/pattern \
KMS scanout (no compositor)
Drop to a tty, stop the display manager, then:
G2G_RTSP_TEST_URL=rtsp://localhost:8554/pattern \
ML inference
# ORT with the CUDA execution provider (silently falls back to CPU):
# Pure-Rust Burn over wgpu (any Vulkan/Metal/DX12 adapter):
# An ONNX topology imported into that element by build-time codegen. Standalone
# (workspace-excluded): keeps burn's codegen tree out of the workspace lockfile.
&&
UDP egress (loopback, no network)
Binds a UDP receiver on localhost, drives the H.264 RTP packetizer, and asserts the datagrams parse back byte-exactly, with sequence numbers, marker bit, and FU-A reassembly all correct.
UDP ingress + resilience (loopback, no network)
End-to-end over localhost: depayload round-trip, the jitter buffer reordering out-of-order packets, and NACK-driven recovery (a lossy relay drops chosen sequences; the receiver NACKs, the sender retransmits, every access unit is recovered in order).
Android on-device testing
The Android elements (mediacodec decode/encode, mediacodec-wgpu zero-copy
decode->GPU, aaudio audio, camera2 capture) are cross-compiled in CI but
validated on a real device. Each ships an on-device probe + a smoke script that
builds just that test binary, adb pushes it to /data/local/tmp, runs it, and
checks the libtest summary.
Prerequisites:
adbonPATHwith a phone attached and USB debugging authorised (adb deviceslists it asdevice).cargo-ndk(cargo install cargo-ndk).- The rustup target:
rustup target add aarch64-linux-android. - The Android NDK, with
ANDROID_NDK_HOMEpointing at it (cargo-ndk uses that to find the toolchain), e.g.export ANDROID_NDK_HOME=$HOME/android-ndk-r27c.
Run a probe:
Each takes an optional ABI argument (default arm64-v8a; also x86_64,
armeabi-v7a). To drive an element by hand, build the test the same way and push
the binary yourself:
(--platform: 24 for AImageReader, 26 for AHardwareBuffer / AAudio.)
The APK harness (examples/g2g-android-present, M742) is the one probe that is
a real app, not a pushed binary: a NativeActivity whose window WgpuSink
presents to (gradle-free: aapt2 + zipalign + apksigner from the SDK
build-tools, plus keytool for the one-time debug keystore; point
ANDROID_SDK_ROOT at an SDK with build-tools/ and platforms/). The device
must be unlocked while it runs. Its manifest declares RECORD_AUDIO / CAMERA
(granted by the script via pm grant) so permission-gated capture can run
in-app later.
Permission caveats. A bare /data/local/tmp binary has no app manifest, so
the permission-gated capture paths can't run there: mic capture needs
RECORD_AUDIO and camera capture needs CAMERA. Those probes assert the
parts they can check headlessly (device open, caps, FFI linkage, encode/render)
and report the denial rather than failing; full capture and a true on-screen
SurfaceView present need an APK harness. If adb reports "insufficient
permissions", run adb kill-server && adb start-server and re-accept the prompt
on the phone.
System dependencies
The cargo features pull pure Rust crates; OS-level dependencies must be present on the host.
| Distro | Decoder (ffmpeg) |
Wayland sink | KMS sink | VAAPI |
|---|---|---|---|---|
| Fedora | ffmpeg-devel (RPM Fusion) or ffmpeg-free-devel |
wayland-devel |
libdrm-devel |
libva-devel |
| Debian / Ubuntu | libavcodec-dev libavformat-dev libavutil-dev libswscale-dev |
libwayland-dev |
libdrm-dev |
libva-dev |
| Arch | ffmpeg |
wayland |
libdrm |
libva |
For the CUDA path: install the NVIDIA driver and CUDA runtime your
distribution ships, and ensure libnvcuvid.so and libcuda.so are on the
linker path. The ffmpeg build must include cuvid support if you intend
to use Backend::NvdecCuvid / Backend::NvdecCuda.
mediamtx for the loopback RTSP server is available as a single binary
from https://github.com/bluenviron/mediamtx/releases; some distros also
package it.
Layout
g2g-core/ # traits, runner, solver, frame, caps, clock, static element model
g2g-mcu/ # heap-free MCU peripheral elements over embedded-hal seams
g2g-mcugen/ # host graph compiler: a declarative MCU graph -> a static pipeline (heap-free Rust)
g2g-plugin/ # dynamic-plugin SDK (declare_plugin! + ABI tag)
g2g-plugins/ # all source/sink/transform elements
g2g-ml/ # ORT, Burn, WgpuPreprocess, batcher
g2g-bridge/ # GStreamer C-FFI bridge (libgstglass2glass.so)
g2g-python/ # gst-python-ml element host (embedded CPython)
g2g-capi/ # C ABI (cdylib/staticlib + include/g2g.h)
g2g-pyapi/ # Python (pyo3) bindings
xtask/ # dev-command crate (cargo xtask ci | test --here | size | wasm | bench | ffi-probe)
g2g-bench/ # criterion benchmarks (excluded from the workspace)
DESIGN.md # architecture specification
DEVTOOLS.md # developer tooling reference
docs/ # GitHub Pages site
License
The whole repository is MPL-2.0: every crate, including the examples, tools, and test fixtures.
See LICENSE.