media-pp 0.2.0

A small, GStreamer-flavored media pipeline library built on FFmpeg. Capture, composite and encode without leaving the GPU, on D3D11 and CUDA.
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
# media-pp


`media-pp` is a small, GStreamer-flavored media pipeline library for Rust,
built on [`ffmpeg-next`]. It provides synchronous pipeline stages by default
and explicit thread boundaries through bounded queues.

The library crate lives in `lib/`. Each directory below `examples/` is an
independent example crate, so platform-specific dependencies do not leak into
the core library.

## Quick start


FFmpeg development libraries must be installed and discoverable by
`ffmpeg-sys-next`.

Add the crate to your project:

```toml
[dependencies]
media-pp = "0.2"
```

0.2.0 renames two elements and takes a pair of binding methods away. See
[`CHANGELOG.md`] for what to write instead.

That is the only dependency you need. `ffmpeg-next` is part of this crate's
API — `MediaBuffer` carries its frames and packets, and an encoder's
`parameters()`/`time_base` are its types — so it is re-exported as
`media_pp::ffmpeg`:

```rust
use media_pp::ffmpeg;

let time_base = ffmpeg::Rational::new(1, 30);
```

Depending on `ffmpeg-next` separately works only while that dependency
resolves to the same version this crate uses. When it does not, the compiler
sees two unrelated crates and every type above stops matching, without naming
the version as the cause.

This minimal pipeline generates video for one second and counts the frames:

```rust
use std::{sync::atomic::Ordering, time::Duration};
use media_pp::{
    elements::{FrameCounter, TestVideoOptions, TestVideoSource},
    pipeline::Pipeline,
};

fn main() -> media_pp::Result<()> {
    media_pp::init()?;
    let source = TestVideoSource::new("source", TestVideoOptions::default());
    let (counter, frames) = FrameCounter::new("counter");
    let pipeline = Pipeline::new("demo", source, |source, ctx| {
        let branch = ctx.branch().to(Box::new(counter))?;
        ctx.attach(source, 0, branch)?;
        Ok(())
    })?;
    pipeline.run()?;
    std::thread::sleep(Duration::from_secs(1));
    pipeline.stop();
    println!("frames: {}", frames.load(Ordering::Relaxed));
    Ok(())
}
```

To work with this repository directly:

```sh
cargo test -p media-pp
cargo run -p decode -- path/to/video.mp4
```

Stress and leak scenarios live in `lib/tests/soak.rs`. Each runs for tens of
seconds, so they are `#[ignore]`d and stay out of the command above:

```sh
cargo test -p media-pp --features d3d11,d3d12,cuda --test soak -- --ignored --nocapture
```

On Linux, `pipewire-screen-capture` takes the place of `d3d11`. Its two capture
scenarios also need `MEDIA_PP_SOAK_RESTORE_TOKEN`, since xdg-desktop-portal
would otherwise show its picker and block; any run of `screen_record_software`
prints a token to reuse.

File-based examples require a media path. No media files are checked into the
repository and examples do not use a default path.

## How pipelines work


A pipeline connects a source to filters and a terminal sink:

```text
FileDemuxer → SwDecoder → Queue → Pacer → FrameCounter
```

The core types are deliberately small:

- `MediaBuffer` carries packets, video, audio, and EOS.
- `Sink::consume` is a synchronous call and may return an error.
- `Sink::ready_consume` propagates downstream readiness back toward sources
  and queue workers, so pause/preroll backpressure does not consume or drop the
  next buffered item.
- `SrcPad` connects one source output to one downstream sink.
- `Queue` introduces a bounded worker-thread boundary. Downstream errors are
  reported through the pipeline `Bus`, and the worker continues.
- An element takes what it *is* — its name, the stream it handles, its own
  options — and the pipeline gives it the rest. `Element::attach_context`
  hands over the clock, the playback clock and the bus when the element is
  wired, the same way the pipeline's identity is stamped onto its log. This
  is not tidiness: a `Pacer` handed a clock from somewhere else goes on
  pacing through a paused pipeline, and an audio renderer registered on a
  foreign `PlaybackClock` claims a master slot nothing reads. Both fail
  silently, and neither is now expressible.
- `Pipeline` owns source threads, control flow, the shared clock, bus, and
  topology graph.
- `PipelineBridge` carries buffers from one `Pipeline` into another, so a
  source that dies takes only its own pipeline with it. `AudioMixer` and the
  video compositors already join pipelines that meet at one of them; a bridge
  is the case where nothing should be mixed or composited on the way — it
  carries anything, and passes timestamps through untouched, which is why
  what follows one is a muxer or a `TimestampOrigin` rather than a `Pacer`.
  One input at a time: connecting again replaces it, which is how a
  reconnection works.
- `Tee` provides fan-out; `AudioMixer` and the video compositors provide
  fan-in. `TeeHandle` adds and removes branches while the pipeline runs:
  `attach` joins one, `finish_branch` ends one cleanly — an ordered EOS so
  codecs flush and muxers finalize, then detach — and `detach` abandons one
  outright, for a branch that is failing rather than finishing. Recording
  while a preview keeps running is `finish_branch`; it returns without
  waiting for the drain, and the terminal's `BusEvent::Eos` says when the
  output is actually complete.
- Every video compositor and every screen capture emits at a rate that can be
  changed while it runs. The compositors take it on their existing handle
  (`set_frame_rate`, and `frame_rate` to read back what was actually kept); a
  capture has no handle of its own, so it hands out a `FrameRateHandle` from
  `frame_rate()` before it is moved into a `Pipeline`. Their `time_base` is
  the reciprocal of that rate and their output `pts` a tick counter in those
  units, so a change re-means every timestamp after it while the ones already
  downstream were stamped under the old rate. It is therefore safe exactly
  while nothing downstream reads timestamps — a preview, a frame counter — and
  a branch attached after the change is consistent, because it takes its
  `time_base` when it is built. Changing it during a recording is the caller's
  to refuse.
- `AudioMixer` sums into a format its handle can change while it runs —
  `set_mix_format`, and `mix_format` to read back what was kept. Every input
  rebuilds its own resampler when it next pushes, because each remembers what
  its own was built for, so none has to be found and invalidated from
  outside and an input registered after the change is correct without being
  told. Its `time_base` is `1/sample_rate` and its `pts` a running sample
  count in those units, so the same caveat as the compositors' rate applies —
  safe while nothing downstream reads timestamps, and the caller's to refuse
  during a recording. The mixer itself keeps running either way, which is the
  point: a format change is not a reason to restart the one element every
  audio source is registered with.
- `FileDemuxer` plays its file once unless a `FileDemuxerHandle`
  (`looping_handle()`, taken before the demuxer is moved into a `Pipeline`)
  says otherwise. Looping rewinds at the end of the file instead of ending
  the stream, and carries the timestamps: each lap starts where the last one
  reached, so a `Pacer` still paces the second lap rather than dumping it as
  fast as it can be read, and a muxer still sees its timestamps advance. The
  flag is read only at the end of the file, so switching it off part way
  through plays that lap out and ends at the file's own end. A looping source
  otherwise never ends on its own, and the two existing ways still apply
  mid-lap: `Pipeline::finish` for ordered EOS now, `Pipeline::stop` to
  abandon. The consequence of carrying the timeline is that a looping
  source's timestamps are no longer positions in the file. `seek` still
  speaks in the file's own, and `lap_offset` on the same handle is what turns
  one of those timestamps back into a position in it — a progress bar's
  number.

Every `SourceElement` explicitly classifies whether it is live and whether it
can reposition its own input timeline through `is_live()` and
`is_seekable()`. These are independent source capabilities: seekability does
not imply that every attached downstream branch can accept a pipeline seek.
Pipeline seeks first run a synchronous `CheckSeek` cascade, then execute
`Pause -> Flush -> Seek -> Preroll` under one operation lock. A live or
non-seekable source and a recording muxer branch reject the check before any
mutation. Once every terminal in the starting topology snapshot reports its
first new-timeline sample, the pipeline restores the caller-requested state:
paused stays paused, while playing resumes.
For decoded playback branches, decoders use the seek target carried by
`PrerollContext` to retain the frame covering the requested instant while
discarding earlier warm-up output. At EOS, the last pre-target frame becomes
the preview fallback. `Pacer` and `VideoSynchronizer` only bypass their paused
clocks during preroll. Packet-only branches still preroll on their first
post-seek packet. `Pipeline::seek(target, SeekMode::Keyframe)` skips the
decoded target gate and previews the first sample at the demuxer's keyframe
landing point; `SeekMode::Accurate` decodes forward to the target.
`Preroll` is a distinct control phase: it releases workers parked by `Pause`
and carries a shared `PrerollContext` that can wait for every expected
terminal to report its first sample (or EOS) without blocking the synchronous
control cascade itself.

Buffers use shared ownership, so fan-out clones references rather than media
payloads. PTS, duration, packet time bases, video color information, and EOS
are preserved through stages that do not intentionally create a new timeline.

### Link contracts


Building or attaching a branch rejects a connection that could never carry
data — feeding encoded packets to an encoder that takes decoded frames, or a
D3D11 texture to a CPU filter. The check runs before the pipeline starts and
returns `GraphError::IncompatibleLink`:

```text
decoder produces VideoFrame (System), which rec cannot accept
(it takes VideoPacket|AudioPacket)
```

It compares only what an element already knows when it is constructed. A
`PortContract` is either `Packets` — which `MediaKind`s of encoded media
(`VideoPacket`, `AudioPacket`) — or `Frames`, which decoded kinds
(`VideoFrame`, `AudioFrame`) plus the `MemoryDomain`s they may live in
(`System`, `Cuda`, `D3d11`, `D3d12`). Encoded media is always host memory, so
a packet contract has nowhere to put a domain and nowhere to forget one.
Pixel format, resolution, stride, color space, and the identity of a specific
GPU device are not part of it and stay validated against the real buffer when
it arrives.

Both halves of the kind separate buffers the `MediaBuffer` variant cannot.
The medium splits encoded data, because a container's audio and video pads
emit the same `Packet` — so wiring the audio stream into a video decoder is
caught rather than failing inside libavcodec on the first packet. The memory
domain splits decoded data, because a frame in system memory and one holding
a D3D11 texture are both `Video` — so a `SwDecoder` wired straight into a
`D3d11Scaler` with no `D3d11Upload` between them is caught too. The domain
names the backend rather than just marking a frame as "on a GPU", so a D3D11
texture handed to a CUDA filter is caught the same way.

This is not caps negotiation. Nothing selects a codec, inserts a converter,
renegotiates mid-stream, or reallocates a pool. Declaring a contract is
opt-in, and these elements declare one:

- Packet path: `FileDemuxer`, `SwDecoder`, `SwEncoder`, `SwAudioEncoder`,
  `FileMuxer`, `SegmentedFileMuxer`, `HlsMuxer`, `RtspSink`, `PacketCounter`.
- Video: `SwScaler`, `SwChromaKey`, `SwVideoCompositor`, `OrtDetector`, and
  every backend's upload, download, scaler, converter, chroma key, decoder,
  encoder, renderer, and compositor (`D3d11*`, `D3d12*`, `Cuda*`).
- Audio: `AudioResampler`, `AudioVolume`, `AudioMixer`, `WasapiRenderer`,
  `PipeWireAudioRenderer`.
- Sources: `FileDemuxer`, `RtspSource`, `TestVideoSource`, `TestAudioSource`,
  the capture sources, and inbound WebRTC tracks.
- Either decoded medium: `FrameCounter`.
- Passthrough: `Queue`, `Tee`, `Pacer`, `VideoSynchronizer`, `ChangeGate`,
  `TimestampOrigin`, `PipelineBridge`. `AppSink` accepts anything.

`AppSource` stays undeclared, since only the application knows what it will
push. Anything else undeclared defaults to "unknown", which always links and
leaves the runtime check in charge. A passthrough element carries its
upstream contract forward, so a mismatch is still caught across a thread
boundary and still names the element that actually produces the data.

An element that genuinely handles any backend says so — `VideoSynchronizer`
paces a system frame and a device texture alike, because it never reads the
pixels, so it declares `MemoryDomainSet::ALL`. That is a claim, not a blank:
claiming a narrower domain than an element needs would refuse a pipeline that
works, which is worse than the runtime error the contract was meant to
pre-empt.

Use `Pipeline::finish` to stop a live source with ordered EOS and drain queued
buffers, codecs, and muxers; `Pipeline::stop` abandons buffered work immediately.

## Element inventory


| Kind | Elements |
|---|---|
| Sources | `FileDemuxer`, `AppSource`, `RtspSource`, `TestVideoSource`, `TestAudioSource`, `DxgiCaptureSource`, `WgcCaptureSource`, `MfCaptureSource`, `V4l2CaptureSource`, `PipeWireScreenCaptureSource`, `PipeWireAudioCaptureSource`, `WasapiCaptureSource`, `AudioMixer`, `SwVideoCompositor`, `CudaVideoCompositor`, `D3d11VideoCompositor`, `WebRtcTrackSource` |
| Filters | `SwDecoder`, `CudaDecoder`, `D3d11Decoder`, `D3d12Decoder`, `SwEncoder`, `CudaEncoder`, `D3d11VideoEncoder`, `SwAudioEncoder`, `AudioResampler`, `AudioVolume`, `SwScaler`, `SwChromaKey`, `D3d11ChromaKey`, `Pacer`, `VideoSynchronizer`, `CudaScaler`, `D3d11Scaler`, `D3d12Scaler`, `CudaUpload`, `CudaDownload`, `CudaConverter`, `D3d11Upload`, `D3d11Download`, `D3d12Upload`, `D3d12Download`, `Tee`, `ChangeGate`, `TimestampOrigin` |
| Sinks | `FrameCounter`, `PacketCounter`, `AppSink`, `FileMuxer`, `SegmentedFileMuxer`, `HlsMuxer`, `RtspSink`, `CudaRenderer`, `D3d11Renderer`, `D3d12Renderer`, `PipeWireAudioRenderer`, `WasapiRenderer`, `OrtDetector`, `WebRtcTrackSink` |

Backend-specific elements require their corresponding Cargo feature and are
available only on that backend's platform. See each type's Rust documentation
for buffer requirements, ownership, error behavior, and runtime-control
semantics — for example, why `DxgiCaptureSource` and
`PipeWireScreenCaptureSource` are separate types rather than one struct with a
platform switch is explained on `PipeWireScreenCaptureSource` itself.

On Windows, `DxgiCaptureSource` captures a monitor or desktop region, while
`WgcCaptureSource` captures one application window selected by its `HWND`.
Enable `wgc-capture`, then either build downstream D3D11 elements from the
device returned by `WgcCaptureSource::open`, or inject an existing shared
device through `open_with_device`. `DxgiCaptureSource` exposes the same choice.
Every element here that accepts an `ID3D11Device` — both capture sources, the
D3D11 decoder, scaler, chroma key, download, NVENC encoder, video compositor,
and renderer — rejects `D3D11_CREATE_DEVICE_SINGLETHREADED` and enables the
shared immediate context's runtime multithread protection before issuing any
command, because a `Queue` puts the elements on either side of it on different
threads and that context is not free-threaded by default.
The WGC source intentionally does not show `GraphicsCapturePicker`; selecting
a window in application UI and resolving its `HWND` remain the application's
responsibility. See `screen_preview_gpu` for both shared-device paths.

## Examples


The examples are grouped by purpose:

- `examples/core`: decoding, queues, fan-out, dynamic tees, app sources/sinks,
  audio, muxing, HLS, and CPU compositing.
- `examples/cuda`: headless CUDA recording and GPU text compositing. CUDA is a
  vendor backend rather than a platform one, so these build and run on both
  Windows and Linux; the `examples/render` crates of the same shape are their
  D3D11 counterparts.
- `examples/render`: D3D11/D3D12 playback, desktop/window capture, synchronization,
  GPU scaling/compositing, chroma keying, NVENC hardware encoding, and
  recording. The CUDA halves of the display and screen-capture examples stay
  here because their renderer (Vulkan external memory over an fd) and capture
  source (PipeWire) are genuinely Linux-only. Start with the
  [render example index]examples/render/README.md when choosing among the
  screen preview and recording variants.
- `examples/rtsp`: publishing, seeking, and receiving RTSP streams.
- `examples/vision`: scaling and ONNX object detection.
- `examples/webrtc`: data and encoded A/V loopback pipelines, a two-way video
  call that presents both incoming tracks on Windows and Linux, and an
  all-platform H.264/Opus receive-record example that muxes both WebRTC tracks
  into MP4.

Useful starting points:

```sh
cargo run -p probe -- path/to/video.mp4
cargo run -p fanout -- path/to/video.mp4
cargo run -p app_sink -- path/to/video.mp4
cargo run -p scale -- path/to/video.mp4
cargo run -p d3d11_scale_render -- path/to/video.mp4
```

Backend-specific examples enable their required library features in their own
`Cargo.toml` files, per target where an example covers more than one
platform. Each such example's module docs explain how the backends differ;
run an example without arguments to see its usage line.

## Feature flags


The library has no default features.

| Feature | Adds | Platform |
|---|---|---|
| `cuda` | NVDEC decode, NVENC encode, scaling, compositing, upload/download, and rendering, all on CUDA-resident frames | Linux, Windows |
| `d3d11` | D3D11 decode, scaling, upload/download, rendering, GPU compositing, and NVENC encoding | Windows |
| `d3d12` | D3D12VA decode, scaling, upload/download, and rendering interfaces | Windows |
| `dxgi-capture` | Desktop capture; also enables `d3d11` | Windows |
| `wgc-capture` | Individual-window capture through Windows Graphics Capture; also enables `d3d11` | Windows |
| `mf-capture` | Camera capture through Media Foundation | Windows |
| `pipewire-audio-capture` | System-audio and microphone capture through PipeWire | Linux |
| `pipewire-audio-renderer` | Audio playback through PipeWire | Linux |
| `pipewire-screen-capture` | Desktop capture through xdg-desktop-portal and PipeWire | Linux |
| `v4l2-capture` | Camera capture through Video4Linux2 | Linux |
| `wasapi-capture` | System-audio and microphone capture | Windows |
| `wasapi-renderer` | Shared-mode audio playback | Windows |
| `ort` | ONNX Runtime object detection | All supported targets |
| `webrtc` | `str0m`-based WebRTC peer and track elements | All supported targets |

Each attached WebRTC source and sink exposes the codec families retained by
SDP negotiation. `WebRtcTrackSource::codec()` separately reports the codec
actually observed after RTP starts arriving. A `WebRtcTrackSink` is told what
feeds it through `set_source_parameters`, taking the `parameters()` of an
encoder, a demuxer's stream, or another track: that one value settles the
outbound payload type, validated against the negotiated list, and the codec
headers H.264 keeps outside its bitstream, which the sink then puts in front
of every keyframe because RTP has no container to carry them. A caller
pushing packets it assembled itself, with no parameters to hand, declares the
codec alone through `set_codec` and must carry its own parameter sets in-band;
`consume` refuses a keyframe that has neither. `set_source_parameters` also
accepts a demuxer's length-prefixed H.264, rewriting each packet as Annex-B.
A receiver that must configure its graph
from the sender's actual payload can call `WebRtcTrackSource::wait_stream_info`
with an explicit timeout; received packets stay buffered while the downstream
graph is built. H.264 waits until actual SPS/PPS have arrived. The returned
`WebRtcStreamInfo` derives the RTP time base and purpose-independent FFmpeg
codec parameters. H.264 parameters include received SPS/PPS and dimensions,
and Opus parameters include its negotiated channel layout and `OpusHead`;
decoder and muxer compatibility is decided by the consuming element.

For example, build all Windows API documentation locally. Nightly rustdoc is
what labels each item with the feature that enables it:

```powershell
$env:RUSTDOCFLAGS = "--cfg docsrs"
cargo +nightly doc -p media-pp --open --features d3d11,d3d12,dxgi-capture,wgc-capture,mf-capture,wasapi-capture,wasapi-renderer,webrtc
```

[docs.rs] builds this crate for Linux, so it documents only the
backend-independent API and omits Windows-only types. The complete API,
including D3D11, D3D12, DXGI, and WASAPI, is available in the
[Windows API documentation] published on GitHub Pages.

## Logging


Library diagnostics use a private, opt-in logger and never install a global
`log` logger or `tracing` subscriber:

```rust
let _log_guard = media_pp::log::init(
    "media-pp",
    "./logs",
    media_pp::log::Level::Info,
    7,
)?;
```

Keep the returned guard alive until logging is no longer needed. Pipeline
starts and dynamic `Tee` changes include a stable-ID topology diagram; detailed
EOS and control propagation is available at `Trace` level. Ordinary media
buffers are not logged one record per buffer.

## Requirements and platform notes


- Install FFmpeg 8.0 or newer development headers and libraries in a location
  discoverable by `ffmpeg-sys-next`. The build script reads the version
  `ffmpeg-sys-next` detected and fails with an explicit message on anything
  older, rather than letting the mismatch surface as a link or runtime error.
- Rust 1.88 or newer is required.
- D3D11VA/D3D12VA require compatible FFmpeg builds, Windows drivers, and GPU
  hardware. Check available accelerators with `ffmpeg -hwaccels`.
- D3D11 elements in one pipeline must share the same `ID3D11Device` and
  immediate context.
- `D3d11Decoder` uses a fixed-size FFmpeg surface pool; its downstream-frame
  budget must cover the deepest buffering. The decoder reserves its accurate-
  seek candidate surface internally.
- `PipeWireScreenCaptureSource` needs `libpipewire-0.3` development files, a
  running PipeWire session, and an `xdg-desktop-portal` backend implementing
  `org.freedesktop.portal.ScreenCast`. See its own Rust documentation for the
  interactive portal dialog, restore tokens, window-vs-monitor stall behavior,
  and closed-window detection this implies.
- `PipeWireScreenCaptureSource::open_gpu` (needs `cuda` as well) captures into
  CUDA surfaces instead of CPU frames, so `screen_record_nvenc` records with
  no upload element. It negotiates DMA-BUF only and fails rather than falling
  back, and it `dlopen`s the driver's `libEGL.so.1`/`libGLESv2.so.2` at run
  time — no development packages are needed to build it.
- `PipeWireAudioCaptureSource`/`PipeWireAudioRenderer` need PipeWire 0.3.50 or
  newer development files and a running session, but no portal.
- CUDA surfaces carry either NV12 or BGRA (`CudaFrameFormat`). Recording needs
  no conversion between them: NVENC ingests BGRA as directly as NV12,
  converting in hardware, so a capture recorded through `CudaEncoder` stays
  BGRA end to end. `CudaVideoCompositor` and `CudaRenderer` work in NV12
  instead, and `CudaConverter` is what a BGRA capture goes through to reach
  them — with a kernel of this crate's own, since `scale_cuda` resizes but has
  no RGB-to-YUV kernel and `CudaScaler` therefore does not convert.
- A `CudaDevice` opens the device's primary CUDA context, so create one per
  process before starting pipelines rather than per pipeline: creating or
  dropping one while another thread is decoding or encoding can crash inside
  the NVIDIA driver.
- `CudaVideoCompositor` composites NV12 CUDA surfaces with `scale_cuda`, 2D
  device-to-device copies, and one small blend kernel, so every `VideoFit`
  and `opacity` works as it does on the other backends — `Cover` needs
  cropping that no CUDA filter offers, and translucency needs arithmetic no
  copy can do. The kernel ships as PTX text that the driver JIT-compiles at
  startup, so no CUDA toolkit is involved. Layer placement and size are
  aligned to even pixels, since NV12 chroma is subsampled. It also draws text
  layers (`CudaTextLayerHandle`), sharing the glyph rasterizer with the D3D11
  compositor and blending the coverage with the same kernel.
- The `cuda` feature links the NVIDIA driver library directly (`libcuda.so`
  on Linux, `nvcuda.dll` on Windows) for those copies and for the blend
  kernel. No CUDA toolkit is needed — the driver ships both the library and
  the PTX compiler.
- `D3d11VideoEncoder` reaches whichever encode hardware the machine has, and
  which of its codecs open depends on that. The `*Nvenc` variants need an
  NVIDIA GPU and an FFmpeg build with NVENC; the `*MediaFoundation` ones need
  a driver that registers a hardware H.264/HEVC transform, which Intel, AMD
  and NVIDIA all do — so those are what give an Intel or AMD machine hardware
  encoding rather than a fall back to the CPU. Either fails to open with a
  typed error, not a panic, so a caller can probe the list. The other `d3d11`
  elements are vendor-neutral.
- RTSP publishing requires an external server that accepts publishing, such as
  MediaMTX.
- Tests needing media build their own: `cargo test` synthesizes a fixture from
  the crate's own synthetic sources, so nothing has to be installed or
  downloaded and every machine tests the same file. The soak scenarios in
  `lib/tests/soak.rs` are the exception and still read `MEDIA_PP_TEST_VIDEO`,
  where a real recording is the point.
- Windows-backed examples compile as unsupported stubs on other targets.

## License


Licensed under either the [Apache License, Version 2.0](LICENSE-APACHE) or the
[MIT License](LICENSE-MIT), at your option.

`media-pp` does not bundle FFmpeg. Users are responsible for complying with
the license of their FFmpeg build and optional codecs.

[`CHANGELOG.md`]: https://github.com/hash1018/media-pp/blob/main/CHANGELOG.md
[`ffmpeg-next`]: https://github.com/zmwangx/rust-ffmpeg
[docs.rs]: https://docs.rs/media-pp
[Windows API documentation]: https://hash1018.github.io/media-pp/media_pp/