frust_engine/renderer.rs
1//! [`EngineRenderer`]: the public seam a host drives one surface's 2D frames
2//! through.
3//!
4//! # The encode contract
5//!
6//! [`EngineRenderer::encode`] **records into a `wgpu::CommandEncoder` the
7//! caller owns and never submits it.** That is the whole point of the seam: a
8//! host compositing 2D over its own 3D content records its passes before and
9//! after the engine's into one encoder and submits once, and a submit hidden
10//! inside the engine would split that into two command buffers with a pipeline
11//! flush between them. Two obligations follow, and both are contract rather
12//! than preference:
13//!
14//! - every pass the engine begins is ended before `encode` returns, so the
15//! caller's next `begin_render_pass` on the same encoder is legal; and
16//! - the engine issues no `queue.submit` for scene work. It does issue
17//! `queue.write_texture`/`write_buffer` uploads, which are ordered ahead of
18//! the command buffers submitted after them and so land before the passes
19//! that read them.
20//!
21//! One thing the engine *does* submit, and it is worth naming precisely because
22//! the rule above is otherwise absolute: growing the image atlas array submits a
23//! command buffer of its own, holding one texture copy and nothing else, on the
24//! rare frame that grows it. That is **maintenance, not scene work** — it never
25//! touches the caller's encoder, records no pass and no draw, and exists because
26//! the queued writes it has to precede would otherwise be flushed ahead of it
27//! (see [`crate::gpu::atlas`]'s *Why growth submits a command buffer of its
28//! own*). The caller's encoder is still never submitted by the engine, and the
29//! frame's own passes still reach the queue only when the caller submits it.
30//!
31//! # Sharing the depth attachment
32//!
33//! A caller recording its own depth-writing passes into that encoder hands the
34//! same attachment in as `EngineTarget::depth`, and the two renderers then
35//! occlude each other correctly in either order. Three rules make that work,
36//! and [`crate::gpu::depth`] is where they are stated in full: the shared
37//! comparison and which end of the range is near
38//! ([`DEPTH_COMPARE`](crate::gpu::depth::DEPTH_COMPARE) over a buffer whose far
39//! plane is [`DEPTH_CLEAR`](crate::gpu::depth::DEPTH_CLEAR)); the depth
40//! attachment's extent matching the colour target's; and who owns the clear —
41//! whichever pass runs first in the encoder, which the caller states through
42//! [`EngineRenderer::set_depth_pre_cleared`].
43//!
44//! What that contract does *not* extend to is colour. The clear pass below
45//! clears the frame's colour target unconditionally, so content painted into
46//! that target before `encode` keeps its depth and loses its pixels: a host
47//! compositing over its own content records that content after the frame, or
48//! into a target of its own the frame composites.
49//!
50//! # The frame's passes
51//!
52//! A frame records into that encoder, in this order.
53//!
54//! 1. **Clear.** Clears the colour target to the frame's base colour, and the
55//! depth attachment to the far plane unless the caller stated it is already
56//! populated (see [`crate::gpu::depth`]). It draws nothing; separating it
57//! from the strip passes is what lets a frame with no draws at all still
58//! resolve to a clean surface.
59//! 2. **Opaque strips**, depth-tested and depth-writing, unblended. Once per
60//! frame, ahead of every round: only the fully-covered interior spans of
61//! opaque draws *targeting the surface* reach it — an anti-aliased edge is by
62//! definition not opaque, and a page carries no depth attachment for the
63//! split to be sound against. The depth it establishes is what lets the alpha
64//! passes reject fragments an opaque draw in front of them already covered.
65//!
66//! Once, not once per surface round, and that is a correctness requirement:
67//! the surface can take several rounds (a [cut](crate::schedule::cut_at)
68//! round is how a wide sibling fan is served), and re-recording this pass
69//! ahead of each would re-draw opaque coverage at equal stored depth over
70//! composites the round before it had already blended. Running it ahead of
71//! the layer rounds rather than after them changes nothing they do: a layer
72//! round writes a pooled page and reads neither the surface nor the depth
73//! attachment.
74//! 3. **One pass per [round](crate::schedule).** A page round renders one
75//! isolated layer into a pooled intermediate
76//! [page](crate::schedule::pages) it clears to transparent, at the page's
77//! own origin — so every instance is shifted by the page's tile-aligned
78//! bounds, and clipped to them, which is what makes a
79//! [banded](crate::schedule::pages::page_bands) layer's column pages tile
80//! their layer instead of each holding a clamped copy of it (see
81//! [`PageWindow`]) — and through the page's own viewport uniform; a round
82//! continuing a page an earlier round of the same layer opened loads it
83//! instead. A surface round draws **alpha strips**, premultiplied-blended,
84//! in painter order: depth-tested but not depth-writing when a depth
85//! attachment is in play, a plain painter's-algorithm pass when it is not.
86//! A round's ops run in the order [`Schedule::build`] listed them, so a
87//! finished child page composites into its parent exactly where the
88//! recording entered it.
89//!
90//! A **filter round** is the one round that draws no strip at all: it runs
91//! one pass of a [filter](crate::filters)'s sequence, one instanced quad
92//! through [`EnginePipeline::Filter`], reading the layer's other pooled page
93//! through the engine's only sampler and clearing the page it writes (see
94//! [`FilterResources`]). It is an ordinary round of this walk in every other
95//! respect — recorded into the caller's own encoder, in the order the
96//! scheduler listed it, with its pages handed back the moment its pass ends.
97//! It is *not* an own-encoder exception; the atlas replay remains the only
98//! one of those.
99//! 4. **The hole punch**, destination-out, when the frame recorded a
100//! `ClearRect` — see [`crate::compile::clear`] for the whole contract this
101//! pass implements. It is issued at the punch's own painter-order position,
102//! not at the end of the frame: a surface round is *cut* where the punch
103//! was recorded, the punch pass goes into that cut, and the round's
104//! remaining ops resume in a pass of their own after it. Everything drawn
105//! over the slot is therefore recorded after the erase and survives it,
106//! whether or not it wrote depth. A punch past every op of the frame — the
107//! ordinary case, a `ClearRect` recorded last — cuts nothing and lands after
108//! the last round exactly as it always did.
109//!
110//! With depth unavailable — no attachment, or `FRUST_ENGINE_NO_DEPTH` set —
111//! pass 2 disappears and every instance travels through the surface rounds,
112//! blended in painter order. That is a correctness requirement rather than a
113//! fallback detail: routing the opaque spans into a separate, earlier pass is
114//! only sound because the depth buffer re-establishes their ordering against the
115//! blended ones.
116//!
117//! # Compositing a layer
118//!
119//! A finished page reaches its parent as ONE instanced quad through the same
120//! strip program every draw goes through, flagged as a whole rectangle and
121//! naming the layer colour source: the fragment stage then reads the page
122//! bound as `layer_input_texture` at the quad's own texel and scales it by the
123//! opacity packed into the instance's low byte. The page is bound through a
124//! bind group of its own rather than the frame's shared one, because group 0
125//! carries both the pass's viewport uniform and that layer input, and a page
126//! round's viewport is its own.
127//!
128//! A composite carries the deepest painter's-order index of everything inside
129//! the layer it composites, nested layers included. That is what keeps a
130//! translucent layer correctly ordered against the root round's own draws
131//! without giving the scheduler a depth model: every draw recorded before the
132//! layer sits behind that index and every draw recorded after it sits in front.
133//!
134//! # Paint resolution
135//!
136//! A solid colour travels inside the strip instance itself. Anything else the
137//! compiler encoded — a gradient, an image, a blurred rounded rectangle — is
138//! resolved once per frame before a single instance is built, in three steps
139//! that have to happen in this order:
140//!
141//! 1. residency is settled for the whole frame: the frame's LUT requests are
142//! serviced through the [`GradientCache`] so every gradient's colour ramp
143//! has an offset into the packed LUT buffer, and every image paint's atlas
144//! rectangle is looked up (or learned, the first time it is drawn) in the
145//! renderer's own image registry — see [`FrameResources::resolve_paints`];
146//! 2. each encoded paint is lowered into the [`GpuEncodedPaint`] record the
147//! fragment shader samples, carrying that residency; and
148//! 3. the records are serialized back to back, which fixes the texel each one
149//! starts at — the index a strip instance names its paint by.
150//!
151//! Ramp offsets are only valid within the frame that took them: the cache
152//! compacts and rewrites them in [`EngineRenderer::end_frame`], which is why
153//! residency is decided here rather than at compile time. An image's atlas
154//! rectangle, by contrast, is stable for as long as the image stays resident
155//! (the compiler's [`crate::cache::images::ImageResidency`] does not move a
156//! live image), so the renderer's own registry only ever forgets an entry
157//! when the frame that compiled it reports the entry's region evicted.
158//!
159//! A paint that still cannot be resolved — an image the atlas has no room
160//! for, a gradient whose ramp could not be baked, an external texture (nothing
161//! binds one yet) — leaves its draw skipped rather than stamped in a wrong
162//! colour: the same "a frame draws less, never wrong" rule the compiler
163//! follows for the commands it does not lower.
164//!
165//! An image the atlas holds a *minified* copy of is the one paint whose lowered
166//! record needs a correction here. The compiler composed its natural-to-device
167//! transform against the source's declared extent, before residency was
168//! consulted and so before the fit was known; the shader samples the resident
169//! rectangle. [`ResidentImage::minify_scale`] is the ratio between the two, and
170//! folding it into the lowered record's transform is what keeps a downsampled
171//! image landing on the destination rectangle the display list asked for.
172
173use core::ops::Range;
174use std::collections::{HashMap, HashSet};
175use std::sync::{Arc, Once};
176// Durations and the compile's own phase record are read by the `perf-trace`
177// encode window and by the tests that pin its arithmetic, and by nothing else.
178#[cfg(any(test, feature = "perf-trace"))]
179use std::time::Duration;
180
181use frust_gpu::{PipelineCache, PooledTexture, SceneTextureId, ShaderLibrary, TierCaps};
182use frust_scene::{Scene, SceneBuilder};
183use glifo::{AtlasCommand, AtlasCommandRecorder, AtlasPaint};
184use kurbo::Affine;
185use peniko::{Brush, Color};
186use vello_common::encode::{EncodedImage, EncodedPaint};
187use vello_common::fearless_simd::Level;
188use vello_common::paint::{ImageId, ImageSource, Paint};
189use vello_common::strip::Strip;
190
191use crate::cache::images::{ATLAS_PADDING, AtlasRegion};
192use crate::cache::{
193 AtlasBudget, BYTES_PER_TEXEL, CachedRamp, GradientCache, GradientTextureLayout, ResidentImage,
194};
195#[cfg(any(test, feature = "perf-trace"))]
196use crate::compile::CompileSpans;
197use crate::compile::paint::resolve_lut_request;
198use crate::compile::{ClearPunch, CompiledFrame, PhaseClock, SceneCompiler};
199use crate::config;
200use crate::diag::{EngineSpan, FrameTimestamps};
201use crate::error::EngineError;
202use crate::filters::blur::{FilterInstanceData, GpuFilterData, GpuGaussianBlur};
203use crate::filters::drop_shadow::GpuDropShadow;
204use crate::filters::{FilterStep, ServedFilter, served_filter};
205use crate::gpu::atlas::{
206 AtlasPageBuffers, AtlasRenderReport, AtlasRenderer, lower_encoded_image, push_solid_strips,
207};
208use crate::gpu::bindings::{ExternalRuns, ExternalTextures, lower_encoded_external};
209use crate::gpu::depth::DepthAttachment;
210use crate::gpu::paint_texture::lower_encoded_paint;
211use crate::gpu::pipelines::{EnginePipeline, EngineShaders, atlas_strip_desc, warm_up_descs};
212use crate::gpu::strips::{PaintType, pack_paint_descriptor};
213use crate::gpu::targets::{
214 IntermediateTargets, IntermediateTexture, filter_data_texture_descriptor,
215 filter_data_texture_height, filter_sampler,
216};
217use crate::gpu::{self, AtlasArray, GpuConfig, GpuEncodedPaint, GpuStrip, StripDraw};
218use crate::schedule::pages::{PageConfig, PageSize};
219use crate::schedule::{
220 Composite, MAX_LIVE_PAGES, PageParity, PageTarget, Round, RoundOp, Schedule,
221};
222use crate::{EngineTarget, OutputAlpha};
223use vello_common::geometry::SizeU16;
224use vello_common::record::RecordedLayerKind;
225
226/// The packed paint descriptor of an inline premultiplied solid colour.
227///
228/// A solid paint indexes no encoded-paint record, so its descriptor carries
229/// only the colour source and paint type — both of which are zero, which is
230/// why this is named rather than written as a bare `0` at its call site: the
231/// zero is a coincidence of the layout, not an absence of information.
232const SOLID_PAINT: u32 = pack_paint_descriptor(PaintType::Solid, 0);
233
234/// The smallest resource-texture dimension the engine can address.
235///
236/// A resource texture's row stride is its width times its texel size, and
237/// `wgpu::COPY_BYTES_PER_ROW_ALIGNMENT` is 256; the narrowest of the three
238/// resources is the gradient LUT at 4 bytes per texel, so 64 texels is the
239/// point below which an upload's `bytes_per_row` stops being legal.
240const MIN_RESOURCE_TEXTURE_DIM: u32 = 64;
241
242/// The smallest strip instance buffer the engine allocates, in instances.
243///
244/// A first frame with a handful of strips should not force a second allocation
245/// on the second frame.
246const MIN_INSTANCE_CAPACITY: u64 = 4096;
247
248/// The colour source a composite instance names its finished page by:
249/// `COLOR_SOURCE_LAYER` in bits 29-30 of the packed paint descriptor.
250///
251/// Deliberately not built through
252/// [`pack_paint_descriptor`](crate::gpu::strips::pack_paint_descriptor): that
253/// helper packs a paint type and a paint-record index into the low bits, and a
254/// composite spends the same bits on a constant opacity instead. Two readings
255/// of one word, so each is written where its own reading is obvious.
256const LAYER_PAINT_SOURCE: u32 = 1 << 29;
257
258/// The premultiplied source colour a hole-punch instance erases with.
259///
260/// Opaque white. Destination-out weights the erase by the source's own *alpha*
261/// and multiplies its colour by zero, so the colour channels never reach the
262/// target and only full alpha matters — it is what makes a fully covered pixel
263/// read exactly `(0, 0, 0, 0)`.
264const PUNCH_SOURCE: u32 = u32::MAX;
265
266/// The label every intermediate page is acquired from the pool under.
267const PAGE_LABEL: &str = "frust-engine layer page";
268
269/// The label every filter round's pass is recorded under.
270const FILTER_LABEL: &str = "frust-engine filter pass";
271
272/// Vertices one filter pass's quad is built from, the same four-vertex
273/// triangle strip every engine program expands an instance into.
274const FILTER_QUAD_VERTICES: u32 = 4;
275
276/// The smallest filter instance buffer the engine allocates, in instances.
277///
278/// A σ-32 blur is ten passes, so this is roughly "one deep blur costs no second
279/// allocation"; the buffer is 32 bytes an instance and grows from here.
280const MIN_FILTER_INSTANCE_CAPACITY: u64 = 16;
281
282static INDEXED_PAINT_WARNING: Once = Once::new();
283
284/// Raised the first time an atlas region is declined, so a renderer whose
285/// budget and array have gone out of agreement says so once rather than every
286/// frame.
287static ATLAS_REFUSAL_WARNING: Once = Once::new();
288
289/// The CPU phases one [`EngineRenderer::encode_traced`] call splits into.
290///
291/// The GPU half of a frame is [`EngineSpan`]'s, measured by the timestamp ring
292/// at pass boundaries; this is the CPU half, measured by lapping a
293/// [`PhaseClock`] between the encode's own steps. The two answer different
294/// questions and neither substitutes for the other — a text-heavy scene can
295/// cost milliseconds here while its passes cost a fraction of one there.
296///
297/// The phases partition the call in the order they run, so [`Self::total`] is
298/// the whole encode. [`counts`](Self::counts) trails them and is not one of
299/// them.
300///
301/// - [`compile`](Self::compile) — [`SceneCompiler::compile`] end to end,
302/// itself split six ways by [`CompileSpans`].
303/// - [`schedule`](Self::schedule) — building the frame's pass plan and
304/// settling its depth attachment.
305/// - [`paints`](Self::paints) — resolving every draw's paint to a texel of the
306/// encoded-paint texture.
307/// - [`instances`](Self::instances) — building the per-round GPU instance
308/// arrays from the recording.
309/// - [`resize`](Self::resize) — the fallible capacity checks, the image
310/// atlas's own growth, the glyph-rectangle clears an earlier frame's
311/// eviction owes, and the three resource-texture resizes.
312/// - [`upload`](Self::upload) — the frame's queue writes: alpha coverage,
313/// encoded paints, gradient ramps, image texels and the instance arrays.
314/// - [`replay`](Self::replay) — the render-to-atlas pass that draws newly
315/// cached glyphs into their array layers (zero on a steady frame, which
316/// caches none).
317/// - [`pipelines`](Self::pipelines) — taking (or building) this frame's
318/// pipelines and bind groups, and preparing the filter blocks.
319/// - [`record`](Self::record) — recording the frame's passes into the caller's
320/// encoder.
321///
322/// Compiled only under `perf-trace`, with everything that reads it: the whole
323/// `frust-perf enc` route is a `#[cfg]` island rather than a runtime branch the
324/// optimizer is trusted to fold, so a release-lean binary carries none of its
325/// field names, format strings or window (see [`EncodeTrace`]).
326#[cfg(feature = "perf-trace")]
327#[derive(Debug, Clone, Copy, Default)]
328struct EncodeSpans {
329 /// Scene compilation, split further by its own six phases.
330 compile: CompileSpans,
331 /// Whole-call cost of [`SceneCompiler::compile`], call overhead included.
332 ///
333 /// Deliberately measured from outside rather than summed from
334 /// [`Self::compile`]: the difference between the two is the part of the
335 /// call the six inner phases do not cover, and a breakdown that could only
336 /// report its own sum could never show that gap.
337 compile_total: Duration,
338 /// Pass planning and depth settlement.
339 schedule: Duration,
340 /// Paint resolution.
341 paints: Duration,
342 /// Instance-array building.
343 instances: Duration,
344 /// Capacity checks, atlas growth, glyph-rectangle clears, resource resizes.
345 resize: Duration,
346 /// The frame's queue writes.
347 upload: Duration,
348 /// The render-to-atlas pass for newly cached glyphs.
349 replay: Duration,
350 /// Pipeline/bind-group acquisition and filter preparation.
351 pipelines: Duration,
352 /// Recording the frame's passes into the caller's encoder.
353 record: Duration,
354 /// This frame's own draw, strip, alpha and glyph counts.
355 ///
356 /// Last because it is not a phase: it is what the phases above were spent
357 /// on. Carried here because a phase's cost is only readable against the
358 /// work it did — "the walk cost two milliseconds" says nothing on its own,
359 /// while "two milliseconds over three thousand strips and three hundred
360 /// glyphs" says where a lever would have to bite.
361 counts: EncodeCounts,
362}
363
364/// The columns one `frust-perf enc` line reports, in row order.
365///
366/// The first seven are [`CompileSpans`]' own — six phases plus the `glyphs`
367/// subset of the walk — and `compile` after them is the whole compile they sit
368/// inside (so the six sum to at most it, never past it). The next eight are
369/// [`EncodeSpans`]' remaining phases, and `total` closes the line. Pinned as
370/// one list because the line's field order is what a capture is graded by —
371/// the same contract the `frust-perf img` line keeps.
372#[cfg(feature = "perf-trace")]
373const ENCODE_TRACE_COLUMNS: [&str; 17] = [
374 "validate",
375 "prepare",
376 "classify",
377 "admit",
378 "walk",
379 "glyphs",
380 "finish",
381 "compile",
382 "schedule",
383 "paints",
384 "instances",
385 "resize",
386 "upload",
387 "replay",
388 "pipelines",
389 "record",
390 "total",
391];
392
393/// The unitless per-frame counts the line reports after its phases, in row
394/// order — what the phases above were spent on.
395#[cfg(feature = "perf-trace")]
396const ENCODE_TRACE_COUNTS: [&str; 5] = ["draws", "strips", "alphas", "glyph_draws", "atlas_glyphs"];
397
398/// Width of one window row: every phase column followed by every count.
399#[cfg(feature = "perf-trace")]
400const ENCODE_TRACE_ROW: usize = ENCODE_TRACE_COLUMNS.len() + ENCODE_TRACE_COUNTS.len();
401
402/// What one frame drew, carried beside its phases.
403///
404/// Counts rather than durations, and so reported without a unit suffix.
405/// `perf-trace`-only, with [`EncodeSpans`], which is all that carries it.
406#[cfg(feature = "perf-trace")]
407#[derive(Debug, Clone, Copy, Default)]
408struct EncodeCounts {
409 /// Recorded draws.
410 draws: u32,
411 /// Strips generated.
412 strips: u32,
413 /// Bytes of alpha coverage those strips index — the frame's largest
414 /// single queue write.
415 alphas: u32,
416 /// Draws that painted one glyph outline.
417 glyph_draws: u32,
418 /// How many of those sampled the glyph atlas rather than rasterizing.
419 atlas_glyphs: u32,
420}
421
422#[cfg(feature = "perf-trace")]
423impl EncodeCounts {
424 /// What `frame` drew.
425 fn of(frame: &CompiledFrame) -> Self {
426 fn count(len: usize) -> u32 {
427 u32::try_from(len).unwrap_or(u32::MAX)
428 }
429 Self {
430 draws: count(frame.draws().len()),
431 strips: count(frame.strip_buf().len()),
432 alphas: count(frame.alphas().len()),
433 glyph_draws: frame.glyph_draws,
434 atlas_glyphs: frame.atlas_glyph_draws,
435 }
436 }
437}
438
439/// Frames one `frust-perf enc` line summarises — one line a second at 60 Hz.
440///
441/// A window rather than a line per frame: the encode is the very thing being
442/// measured, so formatting and logging inside it once per frame would charge
443/// the measurement to its own subject. One line per sixty frames keeps that
444/// charge under a microsecond a frame while still resolving a scenario's
445/// phases (a thirty-second run reports thirty times), and it matches the
446/// cadence `frust-shell-common`'s own rate-limited frame summary already
447/// emits at.
448#[cfg(feature = "perf-trace")]
449const ENCODE_TRACE_WINDOW: usize = 60;
450
451#[cfg(feature = "perf-trace")]
452impl EncodeSpans {
453 /// The whole encode.
454 ///
455 /// Saturating throughout: a diagnostic sum must not take the frame with it
456 /// on overflow (E17).
457 fn total(&self) -> Duration {
458 // `compile.glyphs` is deliberately absent: it is a subset of the walk
459 // inside `compile_total`, and adding it would count text twice.
460 [
461 self.schedule,
462 self.paints,
463 self.instances,
464 self.resize,
465 self.upload,
466 self.replay,
467 self.pipelines,
468 self.record,
469 ]
470 .iter()
471 .fold(self.compile_total, |acc, span| acc.saturating_add(*span))
472 }
473
474 /// This frame's row, in [`ENCODE_TRACE_COLUMNS`] order, in nanoseconds.
475 ///
476 /// Nanoseconds in a `u32` rather than a `Duration` per column: a window of
477 /// sixty rows is then three kilobytes of plain integers to sort, and no
478 /// single phase of one frame reaches the four-second ceiling that would
479 /// saturate one.
480 fn row(&self) -> [u32; ENCODE_TRACE_ROW] {
481 fn ns(span: Duration) -> u32 {
482 u32::try_from(span.as_nanos()).unwrap_or(u32::MAX)
483 }
484 [
485 ns(self.compile.validate),
486 ns(self.compile.prepare),
487 ns(self.compile.classify),
488 ns(self.compile.admit),
489 ns(self.compile.walk),
490 ns(self.compile.glyphs),
491 ns(self.compile.finish),
492 ns(self.compile_total),
493 ns(self.schedule),
494 ns(self.paints),
495 ns(self.instances),
496 ns(self.resize),
497 ns(self.upload),
498 ns(self.replay),
499 ns(self.pipelines),
500 ns(self.record),
501 ns(self.total()),
502 self.counts.draws,
503 self.counts.strips,
504 self.counts.alphas,
505 self.counts.glyph_draws,
506 self.counts.atlas_glyphs,
507 ]
508 }
509}
510
511/// The rolling window of per-frame [`EncodeSpans`] one `frust-perf enc` line
512/// is computed over.
513///
514/// Absent altogether from a build without `perf-trace`, rather than inert in
515/// one. The type, its window, its column names and the `frust-perf enc` literal
516/// are all inside a `#[cfg]` island — including the renderer's own field — so
517/// "no frame pays a push, a sort or a format" is what the compiler emitted and
518/// not what the optimizer was expected to prove about a constant branch. That
519/// distinction is the one the release-lean gate measures: it greps a shipping
520/// binary for `frust-perf` and expects to find nothing.
521#[cfg(feature = "perf-trace")]
522#[derive(Debug, Default)]
523struct EncodeTrace {
524 /// One row per frame in the window, in [`ENCODE_TRACE_COLUMNS`] order.
525 window: Vec<[u32; ENCODE_TRACE_ROW]>,
526 /// The column being ranked, kept across windows so a steady stream of
527 /// them allocates once.
528 ranked: Vec<u32>,
529 /// Frames recorded since this renderer was built — the line's own `n`,
530 /// which is the engine's count and deliberately not the host's raw-line
531 /// `n` (a different emitter counting different frames).
532 frames: u64,
533}
534
535#[cfg(feature = "perf-trace")]
536impl EncodeTrace {
537 /// Adds one frame's phases to the window, emitting the window's line when
538 /// it fills.
539 ///
540 /// Called at the very end of the encode, so the formatting a full window
541 /// costs falls outside every phase the line reports — the numbers describe
542 /// the encode, not the encode plus its own accounting.
543 fn record(&mut self, spans: &EncodeSpans) {
544 // Saturating like every other counter on this path: a renderer that
545 // outlived `u64::MAX` frames would report a wrapped `n`, and a wrong
546 // number in a capture is worse than a stuck one.
547 self.frames = self.frames.saturating_add(1);
548 self.window.push(spans.row());
549 if self.window.len() < ENCODE_TRACE_WINDOW {
550 return;
551 }
552 let line = self.line();
553 self.window.clear();
554 log::info!("{line}");
555 }
556
557 /// The `frust-perf enc` line the current window reports.
558 ///
559 /// Each column is its own median over the window and the trailing field is
560 /// the whole encode's 95th percentile, so a column answers "what does this
561 /// phase usually cost" while the tail answers "how bad is a bad frame".
562 /// Medians are taken per column and so do not sum to the `total_us`
563 /// median exactly — they are close on a steady scene, and the gap is
564 /// itself the signal that the phases are not moving together.
565 fn line(&mut self) -> String {
566 // Sized for the longest line the loops below write: a
567 // `<name>_us=<7 digits>.<1>` field per phase, a `<name>=<10 digits>`
568 // field per count, the two `n`/`w` fields and the p95 tail, with room
569 // to spare rather than a byte-exact fit.
570 let mut line = String::with_capacity(640);
571 line.push_str("frust-perf enc n=");
572 line.push_str(&self.frames.to_string());
573 line.push_str(" w=");
574 line.push_str(&self.window.len().to_string());
575 for (column, name) in ENCODE_TRACE_COLUMNS.iter().enumerate() {
576 let p50 = self.percentile(column, 50);
577 line.push(' ');
578 line.push_str(name);
579 line.push_str("_us=");
580 line.push_str(&format_us(p50));
581 }
582 let tail = self.percentile(ENCODE_TRACE_COLUMNS.len() - 1, 95);
583 line.push_str(" total_p95_us=");
584 line.push_str(&format_us(tail));
585 for (offset, name) in ENCODE_TRACE_COUNTS.iter().enumerate() {
586 let p50 = self.percentile(ENCODE_TRACE_COLUMNS.len() + offset, 50);
587 line.push(' ');
588 line.push_str(name);
589 line.push('=');
590 line.push_str(&p50.to_string());
591 }
592 line
593 }
594
595 /// The `pct`-th percentile of `column` over the window, in nanoseconds.
596 ///
597 /// Nearest-rank on the sorted column, which needs no interpolation and so
598 /// reports a value the window really contains — the same choice
599 /// `frust-shell-common`'s frame summary makes, and the one that keeps a
600 /// sixty-sample window honest.
601 fn percentile(&mut self, column: usize, pct: usize) -> u32 {
602 self.ranked.clear();
603 self.ranked
604 .extend(self.window.iter().filter_map(|row| row.get(column)));
605 self.ranked.sort_unstable();
606 let len = self.ranked.len();
607 if len == 0 {
608 return 0;
609 }
610 let rank = (len * pct).div_ceil(100).clamp(1, len);
611 self.ranked.get(rank - 1).copied().unwrap_or(0)
612 }
613}
614
615/// Nanoseconds as microseconds with one decimal — the unit every
616/// `frust-perf enc` field is reported in.
617///
618/// Fixed-point rather than a float format: a phase can be tens of nanoseconds
619/// (an empty punch pass) or milliseconds (a ten-thousand-row table's walk), and
620/// one decimal microsecond reads the same either way without a float's
621/// locale-dependent formatting reaching a capture the harness greps.
622#[cfg(feature = "perf-trace")]
623fn format_us(nanos: u32) -> String {
624 let tenths = u64::from(nanos).div_ceil(100);
625 format!("{}.{}", tenths / 10, tenths % 10)
626}
627
628/// One surface's 2D render engine: a scene in, recorded passes out.
629///
630/// Create one per surface and keep it across frames — the retained scene
631/// compiler, gradient cache, pipeline cache, resource textures and intermediate
632/// pool are the reason a steady-state frame allocates nothing.
633#[derive(Debug)]
634pub struct EngineRenderer {
635 caps: TierCaps,
636 format: wgpu::TextureFormat,
637 shaders: EngineShaders,
638 pipelines: PipelineCache,
639 compiler: SceneCompiler,
640 gradients: GradientCache,
641 depth: DepthAttachment,
642 targets: IntermediateTargets,
643 /// The bounds an intermediate page is sized between — the policy half of
644 /// page sizing, kept beside the pool the extents are requested from.
645 pages: PageConfig,
646 /// The caller-owned textures a `Command::SceneTexture` resolves against
647 /// (see [`crate::gpu::bindings`]). Its extent half lives on the compiler,
648 /// written by the same two calls that write this.
649 textures: ExternalTextures,
650 resources: FrameResources,
651 scratch: Scratch,
652 /// The render-to-atlas pass, created by the first frame that caches a
653 /// glyph.
654 ///
655 /// Lazy because it owns a coverage texture, an instance buffer and five
656 /// stand-in bindings of its own, and a renderer that never draws text —
657 /// or one running with `FRUST_ENGINE_NO_ATLAS` — should pay for none of
658 /// them.
659 atlas_glyphs: Option<AtlasRenderer>,
660 /// The compiler the atlas replay lowers a page's recorded commands
661 /// through, sized to the atlas page rather than to the surface.
662 ///
663 /// A second compiler rather than this renderer's own: the frame's compiler
664 /// is mid-frame (its glyph entry map is exactly what the replay is
665 /// draining) and its viewport is the surface's, while a page's commands are
666 /// in page space. Created on the first replay and kept, so a steady stream
667 /// of first-seen glyphs allocates a strip generator once.
668 atlas_lowering: Option<SceneCompiler>,
669 /// What the last frame's replay serviced, for
670 /// [`Self::atlas_render_report`].
671 atlas_report: AtlasRenderReport,
672 /// The filter-data texture, sampler and instance buffer a filter round is
673 /// executed with.
674 ///
675 /// Lazy for the same reason [`Self::atlas_glyphs`] is: a renderer that
676 /// never blurs should own neither a sampler nor a resource texture it will
677 /// not read. Created by the first frame that schedules a filter round.
678 filters: Option<FilterResources>,
679 /// The rolling CPU-phase window `frust-perf enc` lines are emitted from.
680 ///
681 /// Absent entirely from a build without `perf-trace`, along with the type
682 /// itself — see [`EncodeTrace`].
683 #[cfg(feature = "perf-trace")]
684 encode_trace: EncodeTrace,
685}
686
687impl EngineRenderer {
688 /// Builds a renderer for `format` targets on `caps`' adapter, warming
689 /// every engine pipeline in the background.
690 ///
691 /// `pipeline_cache` is the host's persisted driver cache when it has one —
692 /// `None` on every backend but Vulkan.
693 ///
694 /// # Errors
695 ///
696 /// [`EngineError::AtlasError`] when the adapter's `resource_texture_dim`
697 /// is not a power of two of at least [`MIN_RESOURCE_TEXTURE_DIM`]. The
698 /// strip shader reconstructs a resource texture's width as `1 << bits` and
699 /// an upload's row stride must satisfy wgpu's copy alignment, so neither
700 /// the shader's addressing nor the uploads would be valid otherwise.
701 pub fn new(
702 device: &wgpu::Device,
703 caps: &TierCaps,
704 format: wgpu::TextureFormat,
705 pipeline_cache: Option<&wgpu::PipelineCache>,
706 ) -> Result<Self, EngineError> {
707 let dim = caps.resource_texture_dim;
708 if !dim.is_power_of_two() || dim < MIN_RESOURCE_TEXTURE_DIM {
709 return Err(EngineError::AtlasError);
710 }
711
712 let mut library = ShaderLibrary::new();
713 let shaders = EngineShaders::register(&mut library, device);
714 let mut pipelines = PipelineCache::new(Arc::new(library), pipeline_cache.cloned());
715 pipelines.warm_up(device, &warm_up_descs(&shaders, format));
716
717 let level = Level::try_detect().unwrap_or(Level::baseline());
718 let gradients = GradientCache::for_texture(GradientTextureLayout::square(dim), level);
719
720 Ok(Self {
721 caps: caps.clone(),
722 format,
723 shaders,
724 pipelines,
725 // The viewport is re-asserted on every compile, so the extent is
726 // only an initial allocation hint. `caps` is not: it is what fixes
727 // the image atlas budget for this renderer's whole life (mobile or
728 // desktop tier, the adapter's own ceilings, and any
729 // `FRUST_ENGINE_ATLAS_SIZE` override), and the adapter is known
730 // exactly here — a compiler built without it would silently keep
731 // the mobile budget on every device.
732 compiler: SceneCompiler::for_caps(1, 1, caps),
733 gradients,
734 depth: DepthAttachment::new(),
735 targets: IntermediateTargets::new(caps),
736 pages: PageConfig::default(),
737 textures: ExternalTextures::new(),
738 resources: FrameResources::new(device, dim),
739 scratch: Scratch::default(),
740 atlas_glyphs: None,
741 atlas_lowering: None,
742 atlas_report: AtlasRenderReport::default(),
743 filters: None,
744 #[cfg(feature = "perf-trace")]
745 encode_trace: EncodeTrace::default(),
746 })
747 }
748
749 /// The target format this renderer warmed its pipelines for.
750 #[must_use]
751 pub fn format(&self) -> wgpu::TextureFormat {
752 self.format
753 }
754
755 /// The adapter capabilities every sizing decision is made against.
756 #[must_use]
757 pub fn caps(&self) -> &TierCaps {
758 &self.caps
759 }
760
761 /// The pool the engine's off-screen intermediates come from.
762 #[must_use]
763 pub fn targets(&self) -> &IntermediateTargets {
764 &self.targets
765 }
766
767 /// The bounds this renderer sizes intermediate layer pages between.
768 #[must_use]
769 pub fn page_config(&self) -> PageConfig {
770 self.pages
771 }
772
773 /// The atlas geometry image residency allocates within.
774 ///
775 /// Derived from the adapter in [`Self::new`], so this is the tier's budget
776 /// narrowed to what the adapter can create — not a constant.
777 #[must_use]
778 pub fn atlas_budget(&self) -> AtlasBudget {
779 self.compiler.images().budget()
780 }
781
782 /// Re-budget image residency, dropping every image currently resident and
783 /// the atlas array holding them.
784 ///
785 /// An atlas rectangle only means anything against the geometry it was
786 /// allocated in, so a new budget invalidates every one already handed out —
787 /// which is why the array, the registry of where each image lives and the
788 /// bind groups naming that array all go in the same step, and each image
789 /// re-uploads on the next frame that draws it. A start-up or adapter-change
790 /// operation, never a per-frame one.
791 pub fn set_atlas_budget(&mut self, budget: AtlasBudget) {
792 self.compiler.set_atlas_budget(budget);
793 self.resources.reset_atlas();
794 }
795
796 /// Replace image residency wholesale, on the same invalidation terms as
797 /// [`Self::set_atlas_budget`].
798 ///
799 /// The programmatic counterpart to `FRUST_ENGINE_NO_ATLAS`: an
800 /// [`ImageResidency::disabled`] residency takes both atlas classes out of
801 /// the frame — images are skipped and every glyph is drawn as outline
802 /// strips — without a process-global environment variable, which is what a
803 /// caller comparing the two paths on one device needs.
804 pub fn set_image_residency(&mut self, images: crate::cache::images::ImageResidency) {
805 self.compiler.set_image_residency(images);
806 self.resources.reset_atlas();
807 }
808
809 /// How many atlas regions this renderer has declined to write or clear.
810 ///
811 /// Zero on every sound frame: residency allocates inside the budget the
812 /// array is created at, and the array is grown to the depth the frame
813 /// reports before its regions are written, so a refusal means those two
814 /// went out of agreement. The count exists so that disagreement is
815 /// measurable rather than silent — a refused write is a region the frame
816 /// believed it had filled.
817 #[must_use]
818 pub fn refused_atlas_regions(&self) -> u64 {
819 self.resources.refused_regions
820 }
821
822 /// Finishes pipeline warm-up on the calling thread, returning only once
823 /// every engine pipeline exists.
824 ///
825 /// Warm-up is started in the background by [`Self::new`] and normally
826 /// needs no attention. Two callers want it forced: a host that must not
827 /// let the *first* frame pay for a compile, and anything about to drop the
828 /// `wgpu::Device` shortly after building a renderer — the warm-up worker
829 /// holds its own handle on that device, and tearing it down while the
830 /// worker is mid-compile is a driver-level hazard rather than a clean
831 /// cancellation.
832 ///
833 /// Each variant is requested through the cache, which builds a queued one
834 /// inline and waits for one the worker has already started, so nothing is
835 /// compiled twice and nothing is left for the worker to claim afterwards.
836 pub fn finish_warm_up(&mut self, device: &wgpu::Device) {
837 for pipeline in EnginePipeline::ALL {
838 let _ = self
839 .pipelines
840 .get_or_create(device, &pipeline.desc(&self.shaders, self.format));
841 }
842 }
843
844 /// How many render pipelines this renderer has compiled so far.
845 ///
846 /// Counts *distinct* pipelines, which is why it is a diagnostic rather
847 /// than something to wait on: two entries of [`EnginePipeline::ALL`] that
848 /// differ only in their colour format describe the same pipeline whenever
849 /// the frame's target format happens to equal
850 /// [`crate::gpu::pipelines::INTERMEDIATE_FORMAT`], and the cache compiles
851 /// that one variant once. Use [`Self::finish_warm_up`] to wait.
852 #[must_use]
853 pub fn compiled_pipelines(&self) -> u64 {
854 self.pipelines.compiled_variants()
855 }
856
857 /// States whether a caller-supplied depth attachment already holds the
858 /// depth this frame should test against — see
859 /// [`DepthAttachment::set_pre_cleared`].
860 pub fn set_depth_pre_cleared(&mut self, pre_cleared: bool) {
861 self.depth.set_pre_cleared(pre_cleared);
862 }
863
864 /// Whether a caller-supplied depth attachment is treated as already
865 /// populated.
866 ///
867 /// The statement is sticky and set once, so a host driving several
868 /// surfaces (or re-establishing one after a device loss) can read back
869 /// what this renderer is on rather than tracking it a second time.
870 #[must_use]
871 pub fn depth_pre_cleared(&self) -> bool {
872 self.depth.is_pre_cleared()
873 }
874
875 /// Registers a caller-owned texture so a
876 /// [`Command::SceneTexture`](frust_scene::Command::SceneTexture) naming
877 /// `id` draws it, returning whatever was registered under `id` before.
878 ///
879 /// `id` is the [`SceneTextureId`] the texture minted for itself
880 /// (`frust_gpu::Texture::as_scene_texture`), and `size` is its extent in
881 /// texels — the rectangle a display list's destination is mapped onto.
882 /// `view` must be a non-array 2D view of a float-sampleable texture
883 /// carrying `wgpu::TextureUsages::TEXTURE_BINDING`; `wgpu` rejects
884 /// anything else when the frame's bind group is built.
885 ///
886 /// Both halves of the registration land here: the view a pass samples and
887 /// the extent the compiler composes a paint transform against. Registering
888 /// is idempotent — re-registering the same id replaces the view and drops
889 /// the bind groups naming the old one.
890 ///
891 /// An extent past `u16::MAX` on either axis, or a zero one, registers
892 /// nothing and answers `None`: the record the shader reads packs the
893 /// source region into `u16` halves, so there is no honest rectangle to
894 /// name. Scenes drawing that id go on drawing nothing.
895 pub fn bind_texture(
896 &mut self,
897 id: SceneTextureId,
898 size: (u32, u32),
899 view: wgpu::TextureView,
900 ) -> Option<wgpu::TextureView> {
901 self.resources.forget_external(id.get());
902 if !self.compiler.bind_external_texture(id.get(), size) {
903 return self.textures.unbind(id);
904 }
905 self.textures.bind(id, view)
906 }
907
908 /// Removes the texture registered under `id`, returning its view.
909 pub fn unbind_texture(&mut self, id: SceneTextureId) -> Option<wgpu::TextureView> {
910 self.compiler.unbind_external_texture(id.get());
911 self.resources.forget_external(id.get());
912 self.textures.unbind(id)
913 }
914
915 /// The view registered under `id`, if any.
916 #[must_use]
917 pub fn bound_texture(&self, id: SceneTextureId) -> Option<&wgpu::TextureView> {
918 self.textures.get(id)
919 }
920
921 /// How many external textures are currently bound.
922 #[must_use]
923 pub fn bound_texture_count(&self) -> usize {
924 self.textures.len()
925 }
926
927 /// Releases everything sized against the old surface extent and
928 /// re-establishes what the next frame needs at the new one.
929 ///
930 /// Pipelines, shader modules, the gradient cache and the resource textures
931 /// are all extent-independent and deliberately survive: a resize must not
932 /// cost a pipeline rebuild or a ramp re-bake. What goes is the intermediate
933 /// pool's parked entries — every one keyed on an extent nothing will ask
934 /// for again — and the engine-owned depth attachment, which has to match
935 /// its colour attachment exactly. Reallocating the depth buffer here rather
936 /// than on the next frame keeps it off the frame path.
937 pub fn resize(&mut self, device: &wgpu::Device, width: u32, height: u32) {
938 self.targets.drop_parked();
939 self.depth.resize(device, width, height);
940 }
941
942 /// Closes the frame out: ages the intermediate pool by one frame and
943 /// evicts the gradient cache down to its capacity.
944 ///
945 /// Call once per frame, after the frame's commands have been submitted and
946 /// including frames that drew nothing — those are the frames a parked
947 /// intermediate ages on. Gradient eviction compacts the packed LUT buffer
948 /// and rewrites the offsets of the survivors, which is why it belongs at
949 /// the frame boundary rather than mid-frame, where it would invalidate
950 /// offsets the frame's own encoded paints already carry.
951 ///
952 /// `_queue` is part of the signature because the end-of-frame maintenance
953 /// this method owns grows queue writes as the engine does (an atlas region
954 /// cleared after the frame that consumed it, in the reference renderer);
955 /// it has none of them yet.
956 pub fn end_frame(&mut self, _queue: &wgpu::Queue) {
957 self.targets.end_frame();
958 self.gradients.maintain();
959 }
960
961 /// Compiles `scene` and records the frame's passes into `encoder`.
962 ///
963 /// `root` is applied ahead of every command's own transform and
964 /// `base_color` is what the target is cleared to before anything is drawn.
965 /// Neither the encoder nor the queue is submitted — see the module header.
966 ///
967 /// # Errors
968 ///
969 /// [`EngineError::TargetTooLarge`] for a target outside the `u16` device
970 /// grid the strip pipeline addresses, [`EngineError::InvalidTransform`] for
971 /// a non-finite transform, [`EngineError::InvalidGeometry`] for non-finite
972 /// command geometry (rect extents, radii, path points, stroke or dash
973 /// values), [`EngineError::SchedulerEscalation`] for a layer shape the
974 /// engine's scheduler does not serve,
975 /// [`EngineError::IntermediateTextureTooLarge`] for a layer no page can be
976 /// sized to, [`EngineError::AlphaCapacity`] when a frame's coverage
977 /// outgrows the alpha texture, and [`EngineError::PaintCapacity`] when its
978 /// encoded paints or colour ramps outgrow theirs. Every one of them is
979 /// returned before anything is recorded, uploaded, allocated or submitted,
980 /// so a refused frame leaves `encoder` exactly as it was found and the
981 /// renderer's own resources — the atlas array included — exactly as they
982 /// were. That is what lets the caller skip the frame cleanly (nothing is
983 /// presented and the previously presented content persists) rather than
984 /// present it half-drawn, and what keeps a refused frame's image uploads
985 /// alive for the next frame that is not refused.
986 #[expect(
987 clippy::too_many_arguments,
988 reason = "the seam frust-render drives: device, queue, encoder, scene, \
989 target, base colour and root transform are each supplied by a \
990 different owner, so bundling them would only move the \
991 assembly to every call site"
992 )]
993 pub fn encode(
994 &mut self,
995 device: &wgpu::Device,
996 queue: &wgpu::Queue,
997 encoder: &mut wgpu::CommandEncoder,
998 scene: &Scene,
999 target: EngineTarget<'_>,
1000 base_color: Color,
1001 root: Affine,
1002 ) -> Result<(), EngineError> {
1003 self.encode_traced(
1004 device,
1005 queue,
1006 encoder,
1007 scene,
1008 target,
1009 base_color,
1010 root,
1011 FrameTimestamps::inert(),
1012 )
1013 }
1014
1015 /// [`Self::encode`], with each pass's GPU time stamped into `timestamps`.
1016 ///
1017 /// The one difference is the sink: every pass this records asks
1018 /// `timestamps` for its own `timestamp_writes` and takes `None` for an
1019 /// answer, so a frame encoded with [`FrameTimestamps::inert`] — which is
1020 /// exactly what [`Self::encode`] passes — records byte-identical work.
1021 /// Which pass is charged to which span is [`EngineSpan`]'s own
1022 /// documentation; the host owns the ring behind the sink and reads the
1023 /// frame's spans back out of it some frames later (see
1024 /// [`frust_gpu::diag::TimestampRing`]).
1025 ///
1026 /// # Errors
1027 ///
1028 /// Exactly [`Self::encode`]'s, on exactly its terms — a refused frame has
1029 /// recorded no pass, so it has taken no timestamp either and the host
1030 /// abandons the ring's slot rather than mapping it.
1031 #[expect(
1032 clippy::too_many_arguments,
1033 reason = "[`Self::encode`]'s argument list plus the timestamp sink, \
1034 each still supplied by a different owner"
1035 )]
1036 pub fn encode_traced(
1037 &mut self,
1038 device: &wgpu::Device,
1039 queue: &wgpu::Queue,
1040 encoder: &mut wgpu::CommandEncoder,
1041 scene: &Scene,
1042 target: EngineTarget<'_>,
1043 base_color: Color,
1044 root: Affine,
1045 timestamps: FrameTimestamps<'_>,
1046 ) -> Result<(), EngineError> {
1047 // The CPU half of the frame's accounting, lapped phase by phase
1048 // through to the end of the call (see [`EncodeSpans`]). Free without
1049 // `perf-trace` — no clock is read at all.
1050 let mut clock = PhaseClock::start();
1051 let size = grid_size(target.width, target.height)?;
1052 let mut frame = self.compiler.compile(scene, root, size)?;
1053 let compile_total = clock.lap();
1054
1055 // The frame's pass plan, settled before anything is allocated or
1056 // recorded: a layer shape this scheduler does not serve, or one larger
1057 // than a page can be sized to, refuses the whole frame here so the
1058 // caller can skip it cleanly rather than present it half-drawn.
1059 let rounds = Schedule::build(&frame.recorder, &self.caps, &self.pages)?;
1060 let ceiling = self.targets.max_texture_size();
1061 for round in &rounds {
1062 if let Some(page) = round.page()
1063 && (page.size.width > ceiling || page.size.height > ceiling)
1064 {
1065 return Err(EngineError::IntermediateTextureTooLarge);
1066 }
1067 }
1068
1069 // Depth is available only when there is an attachment to use and the
1070 // kill switch is off. Settling that before a single instance is built
1071 // is what keeps the opaque/alpha split and the pass shape in agreement.
1072 let depth_enabled = !config::depth_disabled();
1073 if depth_enabled && target.depth.is_none() {
1074 self.depth.ensure(device, target.width, target.height);
1075 }
1076 // Cloned out of the renderer rather than borrowed from it: an engine-
1077 // owned attachment lives in `self.depth`, and recording the frame needs
1078 // `self` mutably (the page pool, the pipeline cache). A `wgpu`
1079 // texture view is a reference-counted handle, so the clone is a
1080 // refcount bump once per frame rather than an allocation.
1081 let depth_view = depth_enabled
1082 .then(|| target.depth.or_else(|| self.depth.owned_view()))
1083 .flatten()
1084 .cloned();
1085 let depth_view = depth_view.as_ref();
1086
1087 // Destination-out erases colour as well as alpha, so a target whose
1088 // alpha is disregarded would take a black rectangle where the display
1089 // list says nothing changes. A frame cleared to an opaque base colour
1090 // is exactly that target — every pixel of it presents opaquely — and is
1091 // the only such statement `encode` is handed, so it is what the skip
1092 // `compile::clear`'s contract calls for is decided on.
1093 let punches = !frame.clears.is_empty() && !is_opaque(base_color);
1094 let schedule_span = clock.lap();
1095
1096 // Paints are resolved before instances are built: an instance names
1097 // its paint by the texel its record starts at, which only exists once
1098 // the frame's ramps are resident and its records are laid out.
1099 let atlas_budget = self.compiler.images().budget();
1100 self.resources
1101 .resolve_paints(&frame, &mut self.gradients, atlas_budget, &self.textures);
1102 let paints_span = clock.lap();
1103 self.scratch.build(
1104 &frame,
1105 &rounds,
1106 depth_view.is_some(),
1107 punches,
1108 &self.resources.paint_slots,
1109 );
1110 let instances_span = clock.lap();
1111
1112 // Everything that can fail does so here, ahead of the first
1113 // `begin_render_pass` AND ahead of the first thing this frame changes
1114 // about the renderer's frame-visible state: apart from the engine's
1115 // own depth attachment (re-sized above, an internal resource no pass
1116 // has read yet), a frame refused below has allocated nothing, replaced
1117 // no texture and submitted nothing, which is what lets the caller
1118 // skip it cleanly.
1119 let dim = self.caps.resource_texture_dim;
1120 let alphas_grown = gpu::grow_alpha_texture_height(
1121 self.resources.alphas.height,
1122 frame.alphas().len(),
1123 dim,
1124 )?;
1125 let paints_grown = gpu::paint_texture::grow_encoded_paints_texture_height(
1126 self.resources.paints.height,
1127 self.resources.paint_texels(),
1128 dim,
1129 )?;
1130 let gradients_grown = self.resources.grown_gradient_height(&self.gradients)?;
1131
1132 // Past the last fallible step. The atlas is created or grown first, so
1133 // the growth copy's own submit precedes the frame's atlas writes below
1134 // (see `gpu::atlas`) and the array is deep enough for every region they
1135 // name.
1136 self.resources
1137 .ensure_atlas(device, queue, atlas_budget, frame.atlas_layers);
1138
1139 // The glyph atlas, in the order [`crate::gpu::atlas`] documents: the
1140 // rectangles last frame's eviction freed are zeroed first, ahead of
1141 // every write this frame issues, so a rectangle handed straight back
1142 // out cannot be erased after its new occupant landed in it.
1143 //
1144 // Acknowledged only once the writes were really issued — the same
1145 // re-offer contract the image plan keeps below. A frame refused before
1146 // this point leaves every rectangle pending, so the next frame that
1147 // gets here still zeroes it.
1148 if self.clear_glyph_rects(device, queue, &frame) {
1149 self.compiler.acknowledge_glyph_clears();
1150 }
1151
1152 self.resources.resize_alphas(device, alphas_grown);
1153 self.resources.resize_paints(device, paints_grown);
1154 self.resources.resize_gradients(device, gradients_grown);
1155 let resize_span = clock.lap();
1156 let atlas_serviced =
1157 self.resources
1158 .upload(queue, &mut frame, &mut self.gradients, size, dim);
1159 self.resources
1160 .upload_instances(device, queue, &self.scratch);
1161
1162 // Residency is committed exactly here: the frame passed every fallible
1163 // step and its evictions and uploads have reached the array, so the
1164 // compiler may stop re-offering them. A frame that returned early above
1165 // never gets here, and its plan is re-offered on the next frame that
1166 // does (see `cache::images`).
1167 if atlas_serviced {
1168 self.compiler.acknowledge_image_plan();
1169 }
1170 let upload_span = clock.lap();
1171
1172 // The glyph pixels themselves, last of the atlas work and strictly
1173 // before the scene pass: every page `glifo` dirtied this frame is
1174 // lowered to strips and drawn into its own array layer, on an encoder
1175 // this call owns and submits (the sanctioned exception to the encode
1176 // contract — see this module's header and `gpu::atlas`). The queue
1177 // writes issued above are flushed ahead of that submit, so the pass
1178 // composites onto a layer whose clears and image uploads have landed.
1179 //
1180 // Driven by the atlas's own pending work, never by this frame's
1181 // surviving draws: `glifo` dirties a page when it *inserts* an entry,
1182 // so a run culled away behind a clip records fills while drawing
1183 // nothing, and a draw-gated replay would leave those commands recorded
1184 // until some later frame happened to run one — by which time eviction
1185 // may have re-let the rectangles they name.
1186 //
1187 // Acknowledged only when the pass really ran, the same way the clears
1188 // above are: acknowledging is what lifts the eviction deferral, so an
1189 // acknowledgement for a replay that returned early would let `glifo`
1190 // free and re-let the very rectangles those commands still name.
1191 if self.compiler.glyph_replay_pending()
1192 && self.replay_glyph_pages(device, queue, timestamps)
1193 {
1194 self.compiler.acknowledge_glyph_replay();
1195 }
1196
1197 let replay_span = clock.lap();
1198
1199 let format = target.format;
1200 let pipelines = self.frame_pipelines(device, format, depth_view.is_some());
1201
1202 // The filter rounds' own resources, created by the first frame that
1203 // schedules one. Only the parameter blocks are uploaded here: a pass's
1204 // instance names the extent the *pool* quantized its destination page
1205 // up to, which only the round that acquires it knows, so the instances
1206 // are written round by round in `record_frame`.
1207 if let Some(pipeline) = pipelines.filter.as_ref() {
1208 let scratch = &self.scratch;
1209 self.filters
1210 .get_or_insert_with(|| FilterResources::new(device))
1211 .prepare(
1212 device,
1213 queue,
1214 pipeline,
1215 &scratch.filter_blocks,
1216 scratch.filter_passes,
1217 );
1218 }
1219
1220 let pipelines_span = clock.lap();
1221
1222 self.record_frame(
1223 device, queue, encoder, &target, depth_view, base_color, &pipelines, timestamps,
1224 );
1225
1226 // The laps above still run without `perf-trace` — `PhaseClock` reads
1227 // no clock there, so each one is `Duration::ZERO` — but nothing records
1228 // them, and the trace they would have fed is not compiled at all.
1229 #[cfg(not(feature = "perf-trace"))]
1230 let _ = (
1231 compile_total,
1232 schedule_span,
1233 paints_span,
1234 instances_span,
1235 resize_span,
1236 upload_span,
1237 replay_span,
1238 pipelines_span,
1239 clock.lap(),
1240 );
1241
1242 // Last, so the window's own formatting is charged to no phase it
1243 // reports. A frame refused above records nothing: its phases are a
1244 // partial encode and would drag every percentile toward a frame that
1245 // was never presented.
1246 #[cfg(feature = "perf-trace")]
1247 self.encode_trace.record(&EncodeSpans {
1248 compile: frame.compile_spans,
1249 compile_total,
1250 schedule: schedule_span,
1251 paints: paints_span,
1252 instances: instances_span,
1253 resize: resize_span,
1254 upload: upload_span,
1255 replay: replay_span,
1256 pipelines: pipelines_span,
1257 record: clock.lap(),
1258 counts: EncodeCounts::of(&frame),
1259 });
1260
1261 Ok(())
1262 }
1263
1264 /// Creates the render-to-atlas pass on first use.
1265 fn ensure_atlas_renderer(&mut self, device: &wgpu::Device) {
1266 if self.atlas_glyphs.is_none() {
1267 self.atlas_glyphs = Some(AtlasRenderer::new(device, &self.caps));
1268 }
1269 }
1270
1271 /// Zero every atlas rectangle an earlier frame's glyph eviction freed,
1272 /// answering whether they were serviced.
1273 ///
1274 /// Queue writes, issued before this frame's image uploads and before the
1275 /// replay pass's submit — the first of the three orderings
1276 /// [`crate::gpu::atlas`] states. A rectangle the array will not take is
1277 /// counted rather than dropped silently, on the same terms an image region
1278 /// it refuses is.
1279 ///
1280 /// `true` means every rectangle was *offered* to the array — including one
1281 /// it refused, which no later frame could place either — so the caller may
1282 /// stop re-offering them. `false` means there was no array to write to at
1283 /// all, which is the one case where trying again later can succeed. An
1284 /// empty list is serviced trivially.
1285 fn clear_glyph_rects(
1286 &mut self,
1287 device: &wgpu::Device,
1288 queue: &wgpu::Queue,
1289 frame: &CompiledFrame,
1290 ) -> bool {
1291 if frame.glyph_clears.is_empty() {
1292 return true;
1293 }
1294 if self.resources.atlas.is_none() {
1295 return false;
1296 }
1297 self.ensure_atlas_renderer(device);
1298
1299 let refused = {
1300 let Self {
1301 atlas_glyphs,
1302 resources,
1303 ..
1304 } = self;
1305 let (Some(glyphs), Some(atlas)) = (atlas_glyphs.as_ref(), resources.atlas.as_ref())
1306 else {
1307 return false;
1308 };
1309 frame
1310 .glyph_clears
1311 .iter()
1312 .filter(|rect| !glyphs.clear_rect(queue, atlas, **rect))
1313 .count()
1314 };
1315
1316 if refused > 0 {
1317 self.resources.note_refused_regions(refused as u64);
1318 }
1319 true
1320 }
1321
1322 /// Draw every atlas page `glifo` dirtied this frame into its own array
1323 /// layer.
1324 ///
1325 /// The pixels of a newly cached glyph, and the last atlas work before the
1326 /// scene pass. Each page's recorded commands are lowered to strips by
1327 /// [`lower_atlas_page`] and drawn through the pipeline
1328 /// [`atlas_strip_desc`] describes; a page the lowering declines is left
1329 /// undrawn and counted, so a glyph whose shape this tier cannot express
1330 /// goes *missing* rather than landing half-painted.
1331 ///
1332 /// Answers whether the pass really ran, on the same terms
1333 /// [`clear_glyph_rects`](Self::clear_glyph_rects) does and for the same
1334 /// reason: `false` means there was no atlas array to draw into at all, so
1335 /// the recorded commands are still recorded and the caller must go on
1336 /// offering them. A page the lowering *declined* is not a `false` — it was
1337 /// offered to the array and counted refused, and no later frame could lower
1338 /// it either.
1339 ///
1340 /// `timestamps` is passed straight through to
1341 /// [`gpu::atlas::AtlasRenderer::render_pending`], which charges each dirty
1342 /// page's own pass to [`EngineSpan::Prepass`] — this is the frame's own
1343 /// Prepass recording site.
1344 #[must_use]
1345 fn replay_glyph_pages(
1346 &mut self,
1347 device: &wgpu::Device,
1348 queue: &wgpu::Queue,
1349 timestamps: FrameTimestamps<'_>,
1350 ) -> bool {
1351 if self.resources.atlas.is_none() {
1352 return false;
1353 }
1354 self.ensure_atlas_renderer(device);
1355 let pipeline = self
1356 .pipelines
1357 .get_or_create(device, &atlas_strip_desc(&self.shaders))
1358 .clone();
1359
1360 let report = {
1361 let Self {
1362 atlas_glyphs,
1363 atlas_lowering,
1364 resources,
1365 compiler,
1366 ..
1367 } = self;
1368 let (Some(glyphs), Some(atlas)) = (atlas_glyphs.as_mut(), resources.atlas.as_ref())
1369 else {
1370 return false;
1371 };
1372
1373 let (width, height) = atlas.size();
1374 let page = (
1375 u16::try_from(width).unwrap_or(u16::MAX),
1376 u16::try_from(height).unwrap_or(u16::MAX),
1377 );
1378 let lowering = atlas_lowering.get_or_insert_with(|| SceneCompiler::new(page.0, page.1));
1379
1380 glyphs.render_pending(
1381 device,
1382 queue,
1383 &pipeline,
1384 atlas,
1385 compiler.glyph_atlas_mut(),
1386 timestamps,
1387 |recorder, buffers| lower_atlas_page(recorder, buffers, lowering, page),
1388 )
1389 };
1390
1391 if report.refused > 0 {
1392 self.resources
1393 .note_refused_regions(u64::from(report.refused));
1394 }
1395 self.atlas_report = report;
1396 true
1397 }
1398
1399 /// What the last frame's render-to-atlas pass serviced.
1400 ///
1401 /// Zero across the board on a steady-state frame: text that hit the cache
1402 /// on every glyph frees no rectangle, queues no pixmap and dirties no page.
1403 #[must_use]
1404 pub fn atlas_render_report(&self) -> AtlasRenderReport {
1405 self.atlas_report
1406 }
1407
1408 /// Builds (or takes from the cache) every pipeline this frame's passes
1409 /// need, and the bind groups each of them will be bound through.
1410 ///
1411 /// All of it happens before the first `begin_render_pass`: a pipeline
1412 /// compiled mid-recording would be the very stall the warm-up exists to
1413 /// avoid, and a bind group is only valid against the pipeline that derived
1414 /// its layout.
1415 fn frame_pipelines(
1416 &mut self,
1417 device: &wgpu::Device,
1418 format: wgpu::TextureFormat,
1419 depth: bool,
1420 ) -> FramePipelines {
1421 let alpha_variant = if depth {
1422 EnginePipeline::StripDepthAlpha
1423 } else {
1424 EnginePipeline::StripAlpha
1425 };
1426 let punch_variant = if depth {
1427 EnginePipeline::StripDepthDestOut
1428 } else {
1429 EnginePipeline::StripDestOut
1430 };
1431
1432 let mut frame = FramePipelines {
1433 alpha: (
1434 alpha_variant,
1435 self.pipelines
1436 .get_or_create(device, &alpha_variant.desc(&self.shaders, format))
1437 .clone(),
1438 ),
1439 opaque: None,
1440 page: None,
1441 punch: None,
1442 filter: None,
1443 };
1444 if depth && !self.scratch.opaque.is_empty() {
1445 frame.opaque = Some(
1446 self.pipelines
1447 .get_or_create(
1448 device,
1449 &EnginePipeline::StripOpaque.desc(&self.shaders, format),
1450 )
1451 .clone(),
1452 );
1453 }
1454 if self.scratch.page_rounds() > 0 {
1455 frame.page = Some(
1456 self.pipelines
1457 .get_or_create(
1458 device,
1459 &EnginePipeline::StripIntermediate.desc(&self.shaders, format),
1460 )
1461 .clone(),
1462 );
1463 }
1464 if self.scratch.punches() {
1465 frame.punch = Some((
1466 punch_variant,
1467 self.pipelines
1468 .get_or_create(device, &punch_variant.desc(&self.shaders, format))
1469 .clone(),
1470 ));
1471 }
1472 if self.scratch.filter_passes > 0 {
1473 // Takes the frame's format like every other variant and ignores
1474 // it: a filter pass only ever writes a pooled page, so its own
1475 // description is pinned to `INTERMEDIATE_FORMAT`. One filter
1476 // pipeline therefore serves a renderer for its whole life,
1477 // whatever its surface is reconfigured to.
1478 frame.filter = Some(
1479 self.pipelines
1480 .get_or_create(device, &EnginePipeline::Filter.desc(&self.shaders, format))
1481 .clone(),
1482 );
1483 }
1484
1485 self.resources
1486 .ensure_bind_groups(device, frame.alpha.0, &frame.alpha.1, format);
1487 self.resources.ensure_external_groups(
1488 device,
1489 frame.alpha.0,
1490 &frame.alpha.1,
1491 &self.textures,
1492 );
1493 if let Some(pipeline) = frame.opaque.as_ref() {
1494 self.resources.ensure_bind_groups(
1495 device,
1496 EnginePipeline::StripOpaque,
1497 pipeline,
1498 format,
1499 );
1500 }
1501 if let Some(pipeline) = frame.page.as_ref() {
1502 self.resources.ensure_bind_groups(
1503 device,
1504 EnginePipeline::StripIntermediate,
1505 pipeline,
1506 format,
1507 );
1508 // A layer's own round draws through this variant, so an external
1509 // texture inside an isolated layer needs its group here too.
1510 self.resources.ensure_external_groups(
1511 device,
1512 EnginePipeline::StripIntermediate,
1513 pipeline,
1514 &self.textures,
1515 );
1516 }
1517 if let Some((variant, pipeline)) = frame.punch.as_ref() {
1518 self.resources
1519 .ensure_bind_groups(device, *variant, pipeline, format);
1520 }
1521 self.resources
1522 .ensure_page_configs(device, self.scratch.page_rounds());
1523
1524 frame
1525 }
1526
1527 /// Records the clear pass, every round's own pass, and the hole punch.
1528 ///
1529 /// Every pass opened here is ended before the method returns, which is the
1530 /// half of the encode contract a caller cannot check for itself.
1531 ///
1532 /// Each pass names the [`EngineSpan`] it is charged to: the frame's own
1533 /// surface passes are [`EngineSpan::Main`], a layer page round and a
1534 /// filter pass are [`EngineSpan::Composite`]. Several passes per span is
1535 /// the ordinary case and they sum.
1536 #[expect(
1537 clippy::too_many_arguments,
1538 reason = "one frame's full recording state, each piece owned by a \
1539 different part of the renderer; bundling them would move the \
1540 same assembly one call up"
1541 )]
1542 fn record_frame(
1543 &mut self,
1544 device: &wgpu::Device,
1545 queue: &wgpu::Queue,
1546 encoder: &mut wgpu::CommandEncoder,
1547 target: &EngineTarget<'_>,
1548 depth_view: Option<&wgpu::TextureView>,
1549 base_color: Color,
1550 pipelines: &FramePipelines,
1551 timestamps: FrameTimestamps<'_>,
1552 ) {
1553 let depth_load = self.depth.load_op(target.depth.is_some());
1554
1555 // Drawing nothing is the point: this pass exists so a frame with no
1556 // instances at all still resolves to a clean surface. It is timed
1557 // alongside the frame's other surface passes even though a backend
1558 // that samples its counters at the vertex/fragment stage boundaries
1559 // may write nothing for it — an untimed pass is a measurement gap, not
1560 // a wrong measurement (see `frust_gpu::diag`).
1561 drop(encoder.begin_render_pass(&wgpu::RenderPassDescriptor {
1562 label: Some("frust-engine clear"),
1563 color_attachments: &[Some(color_attachment(
1564 target.view,
1565 wgpu::LoadOp::Clear(clear_color(base_color, target.output)),
1566 ))],
1567 depth_stencil_attachment: depth_view.map(|view| depth_attachment(view, depth_load)),
1568 timestamp_writes: timestamps.writes(EngineSpan::Main),
1569 occlusion_query_set: None,
1570 multiview_mask: None,
1571 }));
1572
1573 // Every field below is reached through `self.<field>` rather than
1574 // through a method: the page pool is borrowed mutably for the whole
1575 // walk while the instance buffer, the bind groups and the plan are
1576 // borrowed immutably, and only disjoint field borrows let those
1577 // coexist.
1578 let Some(instances) = self.resources.instances.as_ref() else {
1579 return;
1580 };
1581 let dim = self.caps.resource_texture_dim;
1582 let opaque_count = self.scratch.opaque.len() as u32;
1583 // The alpha region starts where the opaque one ends, so every segment's
1584 // own index is relative to that.
1585 let base = opaque_count;
1586
1587 // The opaque pass, once for the whole frame and ahead of every round:
1588 // the depth it writes is what every blended instance the frame draws
1589 // onto its own target, composites included, is then tested against.
1590 //
1591 // Hoisted out of the round walk rather than recorded ahead of each
1592 // surface round, because a frame can take several of those (the
1593 // scheduler cuts one short to hand a page group back) and a second
1594 // recording of this pass would re-draw opaque coverage at equal stored
1595 // depth over composites the round before it had already blended. Ahead
1596 // of the layer rounds costs them nothing: a layer round writes a pooled
1597 // page and reads neither this target nor the depth attachment.
1598 if opaque_count > 0
1599 && let Some(opaque) = pipelines.opaque.as_ref()
1600 && let Some(opaque_groups) =
1601 self.resources.bind_groups.get(&EnginePipeline::StripOpaque)
1602 {
1603 record_pass(
1604 encoder,
1605 &PassPlan {
1606 label: "frust-engine opaque strips",
1607 view: target.view,
1608 load: wgpu::LoadOp::Load,
1609 depth: depth_view,
1610 pipeline: opaque,
1611 groups: opaque_groups,
1612 resources: &opaque_groups.resources,
1613 composites: &[],
1614 // An external paint is never claimed opaque, so the
1615 // depth-writing pass never holds a run.
1616 external_keys: &[],
1617 instances,
1618 base: 0,
1619 segments: &[Segment::Strips(0, opaque_count)],
1620 timestamps: timestamps.writes(EngineSpan::Main),
1621 },
1622 );
1623 }
1624
1625 // The destination-out pass's own recording state, resolved once for the
1626 // frame: a punch can now land at any of the walk's cuts, and every one
1627 // of them erases the same target through the same pipeline.
1628 let punch_pass = pipelines.punch.as_ref().and_then(|(variant, pipeline)| {
1629 Some(PunchPass {
1630 view: target.view,
1631 depth: depth_view,
1632 pipeline,
1633 groups: self.resources.bind_groups.get(variant)?,
1634 instances,
1635 base,
1636 })
1637 });
1638
1639 // The frame's live pages — the two ping-pong groups and the one spill
1640 // page beside them — each holding the finished page a later round
1641 // composites (see [`crate::schedule`]).
1642 let mut live: [Option<PooledTexture>; MAX_LIVE_PAGES] = [const { None }; MAX_LIVE_PAGES];
1643 let mut page_slot = 0_usize;
1644
1645 for plan in &self.scratch.rounds {
1646 let own = match plan.page {
1647 None => None,
1648 // A round continuing a page an earlier round of the same layer
1649 // opened takes that very texture back out of its group: a fresh
1650 // one from the pool would hold the previous holder's pixels
1651 // instead of the half already drawn.
1652 Some(page) if page.continued => match live[page.parity.index()].take() {
1653 Some(pooled) => Some((page.parity, pooled)),
1654 // Unreachable: the round that opened the page put it in
1655 // this group, and no round between the two releases it.
1656 None => continue,
1657 },
1658 Some(page) => match self.targets.acquire(
1659 device,
1660 page.size.width,
1661 page.size.height,
1662 PAGE_LABEL,
1663 ) {
1664 IntermediateTexture::Texture(pooled) => Some((page.parity, pooled)),
1665 // Unreachable: every page extent was checked against this
1666 // pool's own ceiling before the first pass was recorded.
1667 // Skipping the round draws less rather than taking a device
1668 // error mid-frame.
1669 IntermediateTexture::TooLarge { .. } => continue,
1670 },
1671 };
1672
1673 // A filter round is only a filter pass: no strip instance, no
1674 // composite, no viewport uniform of its own — the pass maps NDC
1675 // against the destination extent its own instance carries, which is
1676 // why that instance is written here, where the pool's quantized
1677 // extent is finally known.
1678 if let Some(filter) = plan.filter {
1679 if let Some((_, pooled)) = own.as_ref()
1680 && let Some(pipeline) = pipelines.filter.as_ref()
1681 && let Some(filters) = self.filters.as_ref()
1682 {
1683 let (width, height) = pooled.size();
1684 filters.write_instance(
1685 queue,
1686 filter.instance,
1687 &FilterInstanceData::new(
1688 &filter.step,
1689 filter.data_offset,
1690 // Both pages hold the layer at their own origin, so
1691 // neither region is offset within its page.
1692 (0, 0),
1693 (0, 0),
1694 SizeU16::from_wh(
1695 u16::try_from(width).unwrap_or(u16::MAX),
1696 u16::try_from(height).unwrap_or(u16::MAX),
1697 ),
1698 filter.original,
1699 ),
1700 );
1701 // A group with no live page samples the transparent
1702 // placeholder, which filters nothing — the same "draw less,
1703 // never wrong" answer an unresolvable paint gets.
1704 // Unreachable: the round that wrote this pass's source is
1705 // the one before it, and nothing between the two releases
1706 // that group.
1707 let source = live[filter.source.index()].as_ref().map_or(
1708 &self.resources.placeholders.layer_input,
1709 PooledTexture::view,
1710 );
1711 filters.record_pass_timed(
1712 device,
1713 encoder,
1714 &FilterPassPlan {
1715 label: FILTER_LABEL,
1716 pipeline,
1717 dest: pooled.view(),
1718 source,
1719 instance: filter.instance,
1720 },
1721 timestamps.writes(EngineSpan::Composite),
1722 );
1723 }
1724
1725 settle_pages(&mut self.targets, &mut live, plan.released, own);
1726 continue;
1727 }
1728
1729 // The round's viewport uniform. A page's is written here rather
1730 // than with the frame's other uploads because only the pool knows
1731 // the extent it quantized the request up to, and NDC is computed
1732 // against the attachment's real extent. Distinct buffers, so the
1733 // write ordering against the frame's own config never matters.
1734 let config = match &own {
1735 None => &self.resources.config,
1736 Some((_, pooled)) => {
1737 let Some(config) = self.resources.page_configs.get(page_slot) else {
1738 continue;
1739 };
1740 let (width, height) = pooled.size();
1741 queue.write_buffer(
1742 config,
1743 0,
1744 bytemuck::bytes_of(&GpuConfig::new(width, height, dim, dim)),
1745 );
1746 page_slot = page_slot.saturating_add(1);
1747 config
1748 }
1749 };
1750
1751 let variant = match &own {
1752 None => pipelines.alpha.0,
1753 Some(_) => EnginePipeline::StripIntermediate,
1754 };
1755 let pipeline = match (&own, pipelines.page.as_ref()) {
1756 (None, _) => &pipelines.alpha.1,
1757 (Some(_), Some(page)) => page,
1758 (Some(_), None) => continue,
1759 };
1760 let Some(groups) = self.resources.bind_groups.get(&variant) else {
1761 continue;
1762 };
1763
1764 // A page round binds its own viewport uniform; the root round's is
1765 // already the one the shared set carries.
1766 let page_resources = own.as_ref().map(|_| {
1767 resources_bind_group(
1768 device,
1769 pipeline,
1770 &self.resources.alphas.view,
1771 config,
1772 &self.resources.placeholders.layer_input,
1773 )
1774 });
1775 let resources = page_resources.as_ref().unwrap_or(&groups.resources);
1776
1777 let segments = self
1778 .scratch
1779 .segments
1780 .get(plan.segments.clone())
1781 .unwrap_or(&[]);
1782 let mut composites: Vec<wgpu::BindGroup> = Vec::new();
1783 for segment in segments {
1784 if let Segment::Composite(_, parity) = *segment {
1785 // A group with no live page samples the transparent
1786 // placeholder, which composites nothing — the same "draw
1787 // less, never wrong" answer an unresolvable paint gets.
1788 let view = live[parity.index()].as_ref().map_or(
1789 &self.resources.placeholders.layer_input,
1790 PooledTexture::view,
1791 );
1792 composites.push(resources_bind_group(
1793 device,
1794 pipeline,
1795 &self.resources.alphas.view,
1796 config,
1797 view,
1798 ));
1799 }
1800 }
1801
1802 // A page round draws into an off-screen layer page, the frame's own
1803 // rounds into the surface — the split the two spans name.
1804 let span = if own.is_some() {
1805 EngineSpan::Composite
1806 } else {
1807 EngineSpan::Main
1808 };
1809
1810 let (view, label, load, depth) = match &own {
1811 Some((_, pooled)) => (
1812 pooled.view(),
1813 PAGE_LABEL,
1814 // A pooled texture holds whatever its last holder left
1815 // there, so a layer's first round into a page clears it —
1816 // and a round continuing that same page loads it, because
1817 // clearing again would wipe the half already drawn.
1818 if plan.page.is_some_and(|page| page.continued) {
1819 wgpu::LoadOp::Load
1820 } else {
1821 wgpu::LoadOp::Clear(wgpu::Color::TRANSPARENT)
1822 },
1823 None,
1824 ),
1825 None => (
1826 target.view,
1827 "frust-engine alpha strips",
1828 wgpu::LoadOp::Load,
1829 depth_view,
1830 ),
1831 };
1832
1833 // A surface plan with nothing of its own to draw records no pass:
1834 // loading and storing the target unchanged is exactly nothing. A
1835 // punch cut falling at the very start of a round — a `ClearRect`
1836 // recorded before anything the round draws — is how one arises. A
1837 // *page* plan with no segments still records its pass, because that
1838 // is what clears the pooled page.
1839 if !segments.is_empty() || own.is_some() {
1840 record_pass(
1841 encoder,
1842 &PassPlan {
1843 label,
1844 view,
1845 load,
1846 depth,
1847 pipeline,
1848 groups,
1849 resources,
1850 composites: &composites,
1851 external_keys: self.resources.external_runs.keys(),
1852 instances,
1853 base,
1854 segments,
1855 timestamps: timestamps.writes(span),
1856 },
1857 );
1858 }
1859
1860 settle_pages(&mut self.targets, &mut live, plan.released, own);
1861
1862 // The cut this plan ends at, erased once the ops before it have
1863 // been recorded and before the plan after it draws a thing. Inert
1864 // on every plan the walk did not cut, which is every plan of a
1865 // frame that punches nothing.
1866 if let Some(punch) = punch_pass.as_ref() {
1867 punch.record(encoder, plan.punch, timestamps.writes(EngineSpan::Main));
1868 }
1869 }
1870
1871 // The punches past every op of the frame, in the position the pass held
1872 // unconditionally before painter order was restored.
1873 if let Some(punch) = punch_pass.as_ref() {
1874 punch.record(
1875 encoder,
1876 self.scratch.punch,
1877 timestamps.writes(EngineSpan::Main),
1878 );
1879 }
1880
1881 for page in live.into_iter().flatten() {
1882 self.targets.release(page);
1883 }
1884 }
1885}
1886
1887/// The pipelines one frame's passes are recorded with, resolved once before the
1888/// first `begin_render_pass`.
1889///
1890/// Only `alpha` is unconditional: the rest exist exactly when the frame has
1891/// work for them, so a plain frame compiles and binds nothing it will not draw.
1892#[derive(Debug)]
1893struct FramePipelines {
1894 /// The blended pass over the frame's own target, with its variant — which
1895 /// of the two it is depends on whether depth is in play.
1896 alpha: (EnginePipeline, wgpu::RenderPipeline),
1897 /// The depth-writing pass, when depth is available and the frame has
1898 /// fully covered spans to route into it.
1899 opaque: Option<wgpu::RenderPipeline>,
1900 /// The pass a layer page is rendered through, when the frame has one.
1901 page: Option<wgpu::RenderPipeline>,
1902 /// The destination-out pass, with its variant, when the frame punches.
1903 punch: Option<(EnginePipeline, wgpu::RenderPipeline)>,
1904 /// The pass one filter round runs through, when the frame filters a layer.
1905 filter: Option<wgpu::RenderPipeline>,
1906}
1907
1908/// One pass's full recording state, assembled before the pass is begun.
1909///
1910/// A struct rather than an argument list because the composite groups have to
1911/// be built (and so borrowed) before `begin_render_pass` takes the encoder, and
1912/// naming them together is what makes that ordering obvious at the call site.
1913struct PassPlan<'a> {
1914 label: &'a str,
1915 view: &'a wgpu::TextureView,
1916 load: wgpu::LoadOp<wgpu::Color>,
1917 depth: Option<&'a wgpu::TextureView>,
1918 pipeline: &'a wgpu::RenderPipeline,
1919 /// The pipeline variant's own groups 1-3.
1920 groups: &'a StripBindGroups,
1921 /// Group 0 for the pass's ordinary strip segments.
1922 resources: &'a wgpu::BindGroup,
1923 /// Group 0 per composite segment, in the order the segments name them.
1924 composites: &'a [wgpu::BindGroup],
1925 /// The externally bound texture each [`Segment::External`] slot names,
1926 /// indexed by slot — the frame's own
1927 /// [`ExternalRuns::keys`](crate::gpu::bindings::ExternalRuns::keys).
1928 /// Empty for a pass that draws no external texture, which is every pass of
1929 /// every frame that records no `SceneTexture`.
1930 external_keys: &'a [u64],
1931 instances: &'a wgpu::Buffer,
1932 /// The instance index every segment's own index is relative to.
1933 base: u32,
1934 segments: &'a [Segment],
1935 /// The query pair this pass's GPU time is stamped into, `None` for an
1936 /// untimed pass — which is every pass of every frame encoded without a
1937 /// timestamp sink (see [`crate::diag`]).
1938 timestamps: Option<wgpu::RenderPassTimestampWrites<'a>>,
1939}
1940
1941/// Records one pass: load the colour target, then draw each segment in order.
1942///
1943/// Two groups are re-bound per segment. Group 0, because a composite reads its
1944/// page through it while an ordinary strip reads the placeholder; and group 1,
1945/// because a run of instances sampling an externally bound texture needs that
1946/// texture bound where the atlas placeholder otherwise sits. Groups 2-3 are the
1947/// variant's own and never change within a pass. A pass with one segment —
1948/// every frame that records no layer and no external texture — sets them all
1949/// exactly once.
1950///
1951/// A segment naming a slot with no bind group behind it draws nothing rather
1952/// than drawing with whatever group 1 last held: the texture was unbound
1953/// between the frame's paint resolution and its recording, and a wrongly
1954/// sampled run is worse than a missing one.
1955fn record_pass(encoder: &mut wgpu::CommandEncoder, plan: &PassPlan<'_>) {
1956 let mut pass = encoder.begin_render_pass(&wgpu::RenderPassDescriptor {
1957 label: Some(plan.label),
1958 color_attachments: &[Some(color_attachment(plan.view, plan.load))],
1959 depth_stencil_attachment: plan
1960 .depth
1961 .map(|view| depth_attachment(view, wgpu::LoadOp::Load)),
1962 timestamp_writes: plan.timestamps.clone(),
1963 occlusion_query_set: None,
1964 multiview_mask: None,
1965 });
1966 pass.set_pipeline(plan.pipeline);
1967 pass.set_vertex_buffer(0, plan.instances.slice(..));
1968
1969 let mut composite = 0_usize;
1970 for segment in plan.segments {
1971 let (group, images, range) = match *segment {
1972 Segment::Strips(first, count) => {
1973 if count == 0 {
1974 continue;
1975 }
1976 (
1977 plan.resources,
1978 None,
1979 GpuStrip::instance_range(plan.base.saturating_add(first), count),
1980 )
1981 }
1982 Segment::External(first, count, slot) => {
1983 if count == 0 {
1984 continue;
1985 }
1986 let Some(images) = plan
1987 .external_keys
1988 .get(slot as usize)
1989 .and_then(|key| plan.groups.externals.get(key))
1990 else {
1991 continue;
1992 };
1993 (
1994 plan.resources,
1995 Some(images),
1996 GpuStrip::instance_range(plan.base.saturating_add(first), count),
1997 )
1998 }
1999 Segment::Composite(first, _) => {
2000 let Some(group) = plan.composites.get(composite) else {
2001 continue;
2002 };
2003 composite = composite.saturating_add(1);
2004 (
2005 group,
2006 None,
2007 GpuStrip::instance_range(plan.base.saturating_add(first), 1),
2008 )
2009 }
2010 };
2011 plan.groups.bind_with(&mut pass, group, images);
2012 pass.draw(GpuStrip::vertex_range(), range);
2013 }
2014}
2015
2016/// Everything the frame's hole-punch passes are recorded with, resolved once
2017/// before the round walk begins.
2018///
2019/// A frame can now record several: the punches are issued at the painter-order
2020/// positions they were hoisted from, so a surface round carrying two of them is
2021/// cut twice and each cut erases through this same pipeline and these same
2022/// groups (see [`crate::compile::clear`]). Resolving them once is what keeps
2023/// that from becoming a per-cut lookup, and holding the borrows in one value is
2024/// what keeps them out of the page pool's way — the walk holds that mutably
2025/// throughout.
2026struct PunchPass<'a> {
2027 view: &'a wgpu::TextureView,
2028 depth: Option<&'a wgpu::TextureView>,
2029 pipeline: &'a wgpu::RenderPipeline,
2030 /// The destination-out variant's own groups 1-3.
2031 groups: &'a StripBindGroups,
2032 instances: &'a wgpu::Buffer,
2033 /// The instance index the punch's own `(first, count)` is relative to — the
2034 /// alpha region's start, the same one every round's segments use.
2035 base: u32,
2036}
2037
2038impl PunchPass<'_> {
2039 /// Records `punch`'s instances as one destination-out pass, or nothing at
2040 /// all when the cut issued none — which is every cut of every frame that
2041 /// records no `ClearRect`.
2042 fn record(
2043 &self,
2044 encoder: &mut wgpu::CommandEncoder,
2045 punch: (u32, u32),
2046 timestamps: Option<wgpu::RenderPassTimestampWrites<'_>>,
2047 ) {
2048 let (first, count) = punch;
2049 if count == 0 {
2050 return;
2051 }
2052 record_pass(
2053 encoder,
2054 &PassPlan {
2055 label: "frust-engine hole punch",
2056 view: self.view,
2057 load: wgpu::LoadOp::Load,
2058 depth: self.depth,
2059 pipeline: self.pipeline,
2060 groups: self.groups,
2061 resources: &self.groups.resources,
2062 composites: &[],
2063 // A punch erases with a solid source; it samples nothing.
2064 external_keys: &[],
2065 instances: self.instances,
2066 base: self.base,
2067 segments: &[Segment::Strips(first, count)],
2068 timestamps,
2069 },
2070 );
2071 }
2072}
2073
2074/// Hands back the page groups a finished round consumed, and parks the page it
2075/// wrote in its own group.
2076///
2077/// The same bookkeeping after every round, strip and filter alike: a page is
2078/// free the moment the pass that sampled it ends, which is what bounds a chain
2079/// of any depth — and a filter layer's own pair of pages — to the two ping-pong
2080/// groups, and every shape this scheduler serves to
2081/// [`MAX_LIVE_PAGES`](crate::schedule::MAX_LIVE_PAGES) live intermediates.
2082///
2083/// A round that is not continuing a page of its own takes a *fresh* texture out
2084/// of the pool rather than the group's current occupant, and the occupant it
2085/// displaces goes back here. That is what keeps a filter pass from ever holding
2086/// one texture as both its attachment and its source: a filter round never
2087/// continues a page — it clears — so what it writes is always a different
2088/// texture from the one the pass before it wrote and this pass reads.
2089fn settle_pages(
2090 targets: &mut IntermediateTargets,
2091 live: &mut [Option<PooledTexture>; MAX_LIVE_PAGES],
2092 released: [bool; MAX_LIVE_PAGES],
2093 own: Option<(PageParity, PooledTexture)>,
2094) {
2095 for (index, slot) in live.iter_mut().enumerate() {
2096 if released.get(index).copied().unwrap_or(false)
2097 && let Some(page) = slot.take()
2098 {
2099 targets.release(page);
2100 }
2101 }
2102 if let Some((parity, pooled)) = own
2103 && let Some(previous) = live[parity.index()].replace(pooled)
2104 {
2105 targets.release(previous);
2106 }
2107}
2108
2109/// The GPU resources a frame's [filter](crate::filters) rounds are executed
2110/// with, beyond the two pooled pages they ping-pong between.
2111///
2112/// Three of them, and each is the engine's only one of its kind: the
2113/// filter-data texture holding every filter in the frame's 48-byte parameter
2114/// block, the bilinear sampler
2115/// ([`filter_sampler`](crate::gpu::targets::filter_sampler)) the blur kernels
2116/// read their source page through, and the instance buffer one quad per pass is
2117/// drawn from.
2118///
2119/// Public because this *is* executing a filter pass — the renderer holds one
2120/// and drives it over the pages the scheduler named, and `tests/filters.rs`
2121/// drives the same type over pages of its own on real hardware. That second
2122/// caller is not a convenience: `frust_scene` carries no filter command yet
2123/// (the scene seam is a later plan), so a filter layer cannot reach
2124/// [`EngineRenderer::encode`] through a `Scene` at all, and driving this type
2125/// directly is the only way the ported WGSL is exercised on a device.
2126#[derive(Debug)]
2127pub struct FilterResources {
2128 /// Every filter in the frame's parameter block, back to back.
2129 data: ResourceTexture,
2130 /// Group 0, naming [`Self::data`]'s view. Rebuilt whenever that texture is,
2131 /// and valid for the renderer's whole life otherwise: a filter pipeline's
2132 /// description does not depend on the frame's target format, so there is
2133 /// only ever one layout to have derived it from.
2134 data_group: Option<wgpu::BindGroup>,
2135 sampler: wgpu::Sampler,
2136 instances: Option<wgpu::Buffer>,
2137 instance_capacity: u64,
2138 /// Reusable staging for the parameter-block upload, padded to the
2139 /// texture's own footprint.
2140 staging: Vec<u8>,
2141}
2142
2143impl FilterResources {
2144 /// A renderer's filter resources, with nothing uploaded yet.
2145 #[must_use]
2146 pub fn new(device: &wgpu::Device) -> Self {
2147 Self {
2148 data: ResourceTexture::new(
2149 device,
2150 &filter_data_texture_descriptor(gpu::MIN_RESOURCE_TEXTURE_HEIGHT),
2151 ),
2152 data_group: None,
2153 sampler: filter_sampler(device),
2154 instances: None,
2155 instance_capacity: 0,
2156 staging: Vec::new(),
2157 }
2158 }
2159
2160 /// Uploads this frame's parameter `blocks` and reserves room for `passes`
2161 /// pass instances, growing either resource if the frame outgrew it.
2162 ///
2163 /// `pipeline` is the one [`EnginePipeline::Filter`] describes; it is needed
2164 /// because wgpu derives a pipeline's bind-group layouts from its shader
2165 /// module, so the group naming the filter-data texture can only be built
2166 /// against the pipeline that will bind it.
2167 ///
2168 /// A block count no filter-data texture could hold (see
2169 /// [`filter_data_texture_height`]) leaves the group unbuilt, which leaves
2170 /// every filter pass of the frame issuing no draw — the page is still
2171 /// cleared, so the layer composites as transparent rather than as whatever
2172 /// its page last held. Unreachable in practice, and "draw less, never
2173 /// wrong" when it is not.
2174 pub fn prepare(
2175 &mut self,
2176 device: &wgpu::Device,
2177 queue: &wgpu::Queue,
2178 pipeline: &wgpu::RenderPipeline,
2179 blocks: &[GpuFilterData],
2180 passes: u32,
2181 ) {
2182 let Some(height) = filter_data_texture_height(blocks.len()) else {
2183 self.data_group = None;
2184 return;
2185 };
2186 if height > self.data.height {
2187 self.data = ResourceTexture::new(device, &filter_data_texture_descriptor(height));
2188 self.data_group = None;
2189 }
2190 if self.data_group.is_none() {
2191 self.data_group = Some(device.create_bind_group(&wgpu::BindGroupDescriptor {
2192 label: Some("frust-engine filter data"),
2193 layout: &pipeline.get_bind_group_layout(0),
2194 entries: &[wgpu::BindGroupEntry {
2195 binding: 0,
2196 resource: wgpu::BindingResource::TextureView(&self.data.view),
2197 }],
2198 }));
2199 }
2200
2201 // A queue write covers the whole texture extent, so the blocks are
2202 // padded out to its footprint; the trailing texels are never addressed,
2203 // because a pass names its own block by a texel offset the host handed
2204 // it.
2205 let footprint = gpu::resource_texture_bytes(self.data.width, self.data.height);
2206 self.staging.clear();
2207 self.staging
2208 .resize(usize::try_from(footprint).unwrap_or(usize::MAX), 0);
2209 let bytes: &[u8] = bytemuck::cast_slice(blocks);
2210 if let Some(head) = self.staging.get_mut(..bytes.len()) {
2211 head.copy_from_slice(bytes);
2212 }
2213 queue.write_texture(
2214 self.data.copy_target(),
2215 &self.staging,
2216 resource_layout(
2217 gpu::resource_bytes_per_row(self.data.width),
2218 self.data.height,
2219 ),
2220 self.data.extent(),
2221 );
2222
2223 self.reserve_instances(device, passes);
2224 }
2225
2226 /// Grows the instance buffer if this frame's pass count outgrew it.
2227 fn reserve_instances(&mut self, device: &wgpu::Device, passes: u32) {
2228 let stride = size_of::<FilterInstanceData>() as u64;
2229 let required = u64::from(passes).saturating_mul(stride).max(stride);
2230 if self.instances.is_some() && self.instance_capacity >= required {
2231 return;
2232 }
2233 let capacity = required
2234 .checked_next_power_of_two()
2235 .unwrap_or(required)
2236 .max(MIN_FILTER_INSTANCE_CAPACITY.saturating_mul(stride));
2237 self.instance_capacity = capacity;
2238 self.instances = Some(device.create_buffer(&wgpu::BufferDescriptor {
2239 label: Some("frust-engine filter instances"),
2240 size: capacity,
2241 usage: wgpu::BufferUsages::VERTEX | wgpu::BufferUsages::COPY_DST,
2242 mapped_at_creation: false,
2243 }));
2244 }
2245
2246 /// Writes one pass's instance into slot `index` of the instance buffer.
2247 ///
2248 /// Separate from [`Self::prepare`] because a pass's `dest_texture_size` is
2249 /// the extent the texture pool *quantized* its destination page up to, and
2250 /// NDC is computed against the attachment's real extent — only the round
2251 /// that acquires the page knows it. A queue write issued while the frame is
2252 /// being recorded still lands ahead of the command buffers it is submitted
2253 /// with, which is the same ordering a page's viewport uniform already
2254 /// relies on.
2255 ///
2256 /// A slot past the reserved capacity is dropped rather than written, which
2257 /// leaves that pass drawing whatever the slot last held; unreachable, since
2258 /// `prepare` reserved one slot per pass of this very frame.
2259 pub fn write_instance(&self, queue: &wgpu::Queue, index: u32, instance: &FilterInstanceData) {
2260 let Some(buffer) = self.instances.as_ref() else {
2261 return;
2262 };
2263 let stride = size_of::<FilterInstanceData>() as u64;
2264 let offset = u64::from(index).saturating_mul(stride);
2265 if offset.saturating_add(stride) > self.instance_capacity {
2266 return;
2267 }
2268 queue.write_buffer(buffer, offset, bytemuck::bytes_of(instance));
2269 }
2270
2271 /// Records one filter pass into `encoder`: clear the destination page, then
2272 /// draw the one instanced quad that filters `plan`'s source into it.
2273 ///
2274 /// The clear is unconditional and the draw is not. A filter pass writes only
2275 /// the region its step names — a decimated one a quarter of the texels the
2276 /// pass before it did — and the kernels sample past that region without
2277 /// bounds checks, so whatever surrounds it has to be transparent rather than
2278 /// a previous holder's pixels. That has to hold even on the path where the
2279 /// pass itself cannot be issued, or the layer's composite would sample the
2280 /// page's previous tenant instead of nothing.
2281 ///
2282 /// The pass is opened and closed here, on the caller's own encoder: a filter
2283 /// round is not an exception to [`EngineRenderer::encode`]'s contract.
2284 pub fn record_pass(
2285 &self,
2286 device: &wgpu::Device,
2287 encoder: &mut wgpu::CommandEncoder,
2288 plan: &FilterPassPlan<'_>,
2289 ) {
2290 self.record_pass_timed(device, encoder, plan, None);
2291 }
2292
2293 /// [`Self::record_pass`], stamping the pass's GPU time into `timestamps`.
2294 ///
2295 /// A separate method rather than a field on [`FilterPassPlan`]: the plan is
2296 /// public and built by struct literal outside this crate, so a new required
2297 /// field would break every one of those call sites to serve a diagnostic
2298 /// they do not use. `None` records exactly what [`Self::record_pass`] does.
2299 pub fn record_pass_timed(
2300 &self,
2301 device: &wgpu::Device,
2302 encoder: &mut wgpu::CommandEncoder,
2303 plan: &FilterPassPlan<'_>,
2304 timestamps: Option<wgpu::RenderPassTimestampWrites<'_>>,
2305 ) {
2306 let mut pass = encoder.begin_render_pass(&wgpu::RenderPassDescriptor {
2307 label: Some(plan.label),
2308 color_attachments: &[Some(color_attachment(
2309 plan.dest,
2310 wgpu::LoadOp::Clear(wgpu::Color::TRANSPARENT),
2311 ))],
2312 // A pooled page carries no depth attachment, which is also why the
2313 // filter pipeline declares no depth state.
2314 depth_stencil_attachment: None,
2315 timestamp_writes: timestamps,
2316 occlusion_query_set: None,
2317 multiview_mask: None,
2318 });
2319
2320 let (Some(data_group), Some(instances)) =
2321 (self.data_group.as_ref(), self.instances.as_ref())
2322 else {
2323 return;
2324 };
2325
2326 let source = device.create_bind_group(&wgpu::BindGroupDescriptor {
2327 label: Some("frust-engine filter source"),
2328 layout: &plan.pipeline.get_bind_group_layout(1),
2329 entries: &[
2330 wgpu::BindGroupEntry {
2331 binding: 0,
2332 resource: wgpu::BindingResource::TextureView(plan.source),
2333 },
2334 wgpu::BindGroupEntry {
2335 binding: 1,
2336 resource: wgpu::BindingResource::Sampler(&self.sampler),
2337 },
2338 ],
2339 });
2340
2341 pass.set_pipeline(plan.pipeline);
2342 pass.set_vertex_buffer(0, instances.slice(..));
2343 pass.set_bind_group(0, data_group, &[]);
2344 pass.set_bind_group(1, &source, &[]);
2345 pass.draw(
2346 0..FILTER_QUAD_VERTICES,
2347 plan.instance..plan.instance.saturating_add(1),
2348 );
2349 }
2350}
2351
2352/// One filter pass's full recording state.
2353///
2354/// The two pages are named as views rather than as parities because
2355/// [`FilterResources`] holds no pool: which texture each group is holding is the
2356/// caller's bookkeeping, whether that caller is the frame path or a test.
2357#[derive(Debug)]
2358pub struct FilterPassPlan<'a> {
2359 /// Label for captures and validation messages.
2360 pub label: &'a str,
2361 /// The pipeline [`EnginePipeline::Filter`] describes.
2362 pub pipeline: &'a wgpu::RenderPipeline,
2363 /// The page this pass writes. Cleared to transparent before it is written.
2364 pub dest: &'a wgpu::TextureView,
2365 /// The page this pass reads — the one the pass before it wrote.
2366 pub source: &'a wgpu::TextureView,
2367 /// Which instance of the filter instance buffer this pass draws.
2368 pub instance: u32,
2369}
2370
2371/// A colour attachment over `view` with `load`, keeping what it stores.
2372fn color_attachment(
2373 view: &wgpu::TextureView,
2374 load: wgpu::LoadOp<wgpu::Color>,
2375) -> wgpu::RenderPassColorAttachment<'_> {
2376 wgpu::RenderPassColorAttachment {
2377 view,
2378 depth_slice: None,
2379 resolve_target: None,
2380 ops: wgpu::Operations {
2381 load,
2382 store: wgpu::StoreOp::Store,
2383 },
2384 }
2385}
2386
2387/// A depth attachment over `view` with `load`, keeping what it stores so the
2388/// pass after it tests against the same buffer.
2389fn depth_attachment(
2390 view: &wgpu::TextureView,
2391 load: wgpu::LoadOp<f32>,
2392) -> wgpu::RenderPassDepthStencilAttachment<'_> {
2393 wgpu::RenderPassDepthStencilAttachment {
2394 view,
2395 depth_ops: Some(wgpu::Operations {
2396 load,
2397 store: wgpu::StoreOp::Store,
2398 }),
2399 stencil_ops: None,
2400 }
2401}
2402
2403/// The frame's base colour as a clear value.
2404///
2405/// The strip pipelines blend premultiplied and the shader emits premultiplied
2406/// colour, so a [`OutputAlpha::Premultiplied`] target's clear value has to be
2407/// premultiplied too — otherwise the background sits in a different alpha
2408/// convention from everything drawn over it. `Straight` is honoured here and
2409/// only here: the pipelines themselves are fixed premultiplied, so a straight
2410/// target's *drawn* content is premultiplied regardless.
2411fn clear_color(base_color: Color, output: OutputAlpha) -> wgpu::Color {
2412 let components = match output {
2413 OutputAlpha::Premultiplied => base_color.premultiply().components,
2414 OutputAlpha::Straight => base_color.components,
2415 };
2416 wgpu::Color {
2417 r: f64::from(components[0]),
2418 g: f64::from(components[1]),
2419 b: f64::from(components[2]),
2420 a: f64::from(components[3]),
2421 }
2422}
2423
2424/// Whether a frame cleared to `base_color` presents opaquely, and so whether
2425/// its alpha channel carries anything a hole punch could reveal.
2426///
2427/// The engine's only statement about the target's own alpha handling:
2428/// [`EngineTarget`] describes how the alpha it produces is *interpreted*
2429/// (premultiplied or straight), never whether it is used at all. A base colour
2430/// at full alpha seals every pixel of the surface, which is exactly the
2431/// presentation [`crate::compile::clear`]'s contract says to skip the
2432/// destination-out pass on.
2433fn is_opaque(base_color: Color) -> bool {
2434 base_color.components[3] >= 1.0
2435}
2436
2437/// The target extent on the `u16` device grid the strip pipeline addresses.
2438fn grid_size(width: u32, height: u32) -> Result<(u16, u16), EngineError> {
2439 let width = u16::try_from(width).map_err(|_| EngineError::TargetTooLarge)?;
2440 let height = u16::try_from(height).map_err(|_| EngineError::TargetTooLarge)?;
2441 Ok((width, height))
2442}
2443
2444/// One unit of drawing inside a round's pass, in execution order.
2445///
2446/// Instance indices are relative to the alpha region of the shared instance
2447/// buffer, which is why [`PassPlan::base`] exists rather than the indices being
2448/// absolute: the opaque region is laid out first and its own segment addresses
2449/// from zero.
2450#[derive(Debug, Clone, Copy, PartialEq, Eq)]
2451enum Segment {
2452 /// Ordinary strip instances, as `(first, count)`, drawn with the frame's
2453 /// own atlas binding.
2454 Strips(u32, u32),
2455 /// Strip instances sampling an externally bound texture, as
2456 /// `(first, count, slot)` — a *run*, in the sense
2457 /// [`crate::gpu::bindings::ExternalRuns`] gives the word: the maximal span
2458 /// of consecutive instances that share one texture, drawn with that
2459 /// texture's own group 1 bound.
2460 External(u32, u32, u32),
2461 /// One composite quad at `first`, sampling the finished page in the named
2462 /// group.
2463 Composite(u32, PageParity),
2464}
2465
2466/// One scheduled round, resolved to the instances and target it draws with.
2467///
2468/// One *plan* rather than one scheduled round: a surface round carrying a hole
2469/// punch is cut into a plan per span between its punches, so the punch pass can
2470/// be recorded at the painter-order position it was hoisted from (see
2471/// [`Scratch::cut_for_punches`]).
2472#[derive(Debug, Clone)]
2473struct RoundPlan {
2474 /// The page this round renders into, or `None` for the frame's own target.
2475 page: Option<PagePlan>,
2476 /// This round's slice of [`Scratch::segments`], in execution order.
2477 segments: Range<usize>,
2478 /// Page groups this round consumed, indexed by
2479 /// [`PageParity::index`]; each returns to the pool once the pass ends.
2480 released: [bool; MAX_LIVE_PAGES],
2481 /// The filter pass this round runs, on a filter round — which draws no
2482 /// strip and composites nothing, so its `segments` range is empty.
2483 filter: Option<FilterPlan>,
2484 /// The hole punches to erase with once this plan's own pass has been
2485 /// recorded, as `(first, count)` into the alpha region — the cut this plan
2486 /// ends at.
2487 ///
2488 /// Only a *surface* plan ever carries one: a punch recorded inside an
2489 /// isolated layer was hoisted out of it to the frame root (see
2490 /// [`crate::compile::clear`]), so a page round is never cut and a filter
2491 /// round — which draws nothing of the frame's own — never is either.
2492 punch: (u32, u32),
2493}
2494
2495/// One filter pass, resolved to the instance slot it draws and the page group
2496/// it reads.
2497#[derive(Debug, Clone, Copy)]
2498struct FilterPlan {
2499 /// Which pass of the filter's sequence this is, and at what extents.
2500 step: FilterStep,
2501 /// The group holding the page this pass reads.
2502 source: PageParity,
2503 /// The texel this filter's parameter block starts at in the filter-data
2504 /// texture — where the fragment stage reads its kernel from.
2505 data_offset: u32,
2506 /// The filter layer's own extent, before any decimation; it bounds the
2507 /// transparent border a decimated pass overdraws.
2508 original: SizeU16,
2509 /// This pass's slot in the frame's filter instance buffer.
2510 instance: u32,
2511}
2512
2513/// The pooled page one round renders into.
2514#[derive(Debug, Clone, Copy)]
2515struct PagePlan {
2516 parity: PageParity,
2517 size: PageSize,
2518 /// Whether an earlier round of the same layer already rendered into this
2519 /// page, so this round takes that texture back rather than acquiring a
2520 /// fresh one, and loads it rather than clearing it.
2521 continued: bool,
2522}
2523
2524/// The frame's instance buffers and pass plan, retained so a steady-state frame
2525/// refills them rather than reallocating them.
2526///
2527/// The segments of every round live in one flat vector rather than a vector per
2528/// round, so a frame with layers costs no allocation once the first one has
2529/// grown these buffers.
2530#[derive(Debug, Default)]
2531struct Scratch {
2532 /// Fully-covered spans of the opaque draws targeting the frame's own
2533 /// surface, drawn unblended with depth write in one pass ahead of every
2534 /// round — including the surface's own, of which a frame may have several.
2535 opaque: Vec<GpuStrip>,
2536 /// Everything else, drawn premultiplied-blended in painter order and laid
2537 /// out round by round in execution order.
2538 alpha: Vec<GpuStrip>,
2539 /// Every round's segments, back to back; [`RoundPlan::segments`] slices it.
2540 segments: Vec<Segment>,
2541 /// The frame's rounds, each layer's before the round that composites it and
2542 /// the surface's last round last.
2543 rounds: Vec<RoundPlan>,
2544 /// The instances of the punches no cut reached — those recorded past every
2545 /// op of the frame — as `(first, count)` into the alpha region.
2546 ///
2547 /// Their pass is the last thing the frame records, which is where a
2548 /// `ClearRect` recorded last belongs in painter order anyway. Punches the
2549 /// walk *did* cut at are held on their own [`RoundPlan::punch`] instead.
2550 ///
2551 /// No round's segments name a punch instance, wherever it sits in the
2552 /// buffer, so dropping the punch passes leaves the frame exactly as the
2553 /// display list would read without the clear.
2554 punch: (u32, u32),
2555 /// The deepest painter's-order index inside each recorded layer, which is
2556 /// the depth its composite carries. Filled as the rounds are walked, which
2557 /// is sound because a layer's own round always precedes the round that
2558 /// composites it.
2559 layer_depth: Vec<u32>,
2560 /// One parameter block per *filter layer* of the frame, in the order the
2561 /// rounds first named them — which is the order the texel offsets in
2562 /// [`FilterPlan::data_offset`] were taken from.
2563 filter_blocks: Vec<GpuFilterData>,
2564 /// The layers `filter_blocks` holds, parallel to it, so a filter's second
2565 /// and later passes reuse the block its first one packed rather than
2566 /// repacking one per pass.
2567 filter_layers: Vec<u32>,
2568 /// How many filter passes this frame runs, and so how many instances its
2569 /// filter instance buffer has to hold.
2570 filter_passes: u32,
2571}
2572
2573impl Scratch {
2574 /// Turns a scheduled frame into instances and a per-round pass plan,
2575 /// routing each instance to the pass that can draw it.
2576 ///
2577 /// A draw's anti-aliased spans always land in the blended buffer: partial
2578 /// coverage is not opaque however opaque the paint is. Its fully-covered
2579 /// spans land in the opaque buffer only when the paint is opaque, depth is
2580 /// available to re-establish their ordering against the blended ones, *and*
2581 /// the draw targets the frame's own surface — a page has no depth
2582 /// attachment, so a page round is a plain painter's-algorithm walk.
2583 fn build(
2584 &mut self,
2585 frame: &CompiledFrame,
2586 rounds: &[Round],
2587 depth_active: bool,
2588 punches: bool,
2589 paint_slots: &[Option<ResolvedPaint>],
2590 ) {
2591 self.opaque.clear();
2592 self.alpha.clear();
2593 self.segments.clear();
2594 self.rounds.clear();
2595 self.punch = (0, 0);
2596 self.layer_depth.clear();
2597 self.layer_depth.resize(frame.recorder.layers.len(), 0);
2598 self.filter_blocks.clear();
2599 self.filter_layers.clear();
2600 self.filter_passes = 0;
2601
2602 let draws = frame.draws();
2603 let strips = frame.strip_buf();
2604 // The punches still to be issued, in the order the compiler hoisted
2605 // them — which is depth order, since each took its own index from the
2606 // frame's monotonic painter-order counter. Empty on a frame with no
2607 // clear and on a target that disregards alpha (`punches`), and then
2608 // nothing below ever cuts: a non-punching frame is planned exactly as
2609 // it always was.
2610 let mut pending: &[ClearPunch] = if punches { &frame.clears } else { &[] };
2611
2612 for round in rounds {
2613 let page = round.page();
2614 let filter = self.plan_filter(frame, round);
2615 // A page holds its layer — or one column band of it — at the page's
2616 // own origin, so every instance of the round is shifted into that
2617 // column and clipped to it (see `PageWindow`).
2618 let window = page.map_or(PageWindow::ROOT, PageWindow::of);
2619 let split_opaque = page.is_none() && depth_active;
2620 // Only the frame's own target is punched — a punch inside an
2621 // isolated layer was hoisted out of it — so only a surface round is
2622 // ever cut.
2623 let cuts = page.is_none();
2624
2625 let mut first_segment = self.segments.len();
2626 let mut deepest = 0_u32;
2627 let mut run_start = self.alpha.len() as u32;
2628 // The external texture the open run is drawn with, `None` while it
2629 // is drawn with the frame's own atlas binding. A draw that names a
2630 // different one closes the run: there is a single external binding
2631 // to set (see [`crate::gpu::bindings`]).
2632 let mut run_external: Option<u32> = None;
2633
2634 for op in &round.ops {
2635 match op {
2636 RoundOp::Draws(range) => {
2637 let batch = draws
2638 .get(range.start as usize..range.end as usize)
2639 .unwrap_or(&[]);
2640 for draw in batch {
2641 deepest = deepest.max(draw.depth);
2642 // Ahead of the paint lookup rather than after it, so
2643 // the cut follows the order the display list was
2644 // recorded in rather than the subset of it that
2645 // survived lowering.
2646 if cuts && !pending.is_empty() {
2647 self.cut_for_punches(
2648 &mut pending,
2649 strips,
2650 draw.depth,
2651 &mut first_segment,
2652 &mut run_start,
2653 &mut run_external,
2654 );
2655 }
2656 let Some(paint) = pack_paint(&draw.paint, draw.depth, paint_slots)
2657 else {
2658 continue;
2659 };
2660 let Some(run) = strips.get(draw.strip_range.clone()) else {
2661 continue;
2662 };
2663 // Before the first of this draw's instances lands,
2664 // so the run that closes holds exactly the
2665 // instances drawn with the texture it names.
2666 if paint.external != run_external {
2667 let end = self.alpha.len() as u32;
2668 self.push_run(run_start, end, run_external);
2669 run_start = end;
2670 run_external = paint.external;
2671 }
2672 let to_opaque = paint.opaque && split_opaque;
2673
2674 // A generation's last strip is its sentinel, which
2675 // is what carries the preceding strip's extent — so
2676 // every instance comes from a pair, and the
2677 // sentinel itself never becomes one.
2678 for pair in run.windows(2) {
2679 self.push_span(&pair[0], &pair[1], paint, to_opaque, window);
2680 }
2681 }
2682 }
2683 RoundOp::Composite(composite) => {
2684 let depth = self
2685 .layer_depth
2686 .get(composite.layer as usize)
2687 .copied()
2688 .unwrap_or(0);
2689 // A composite carries the deepest index inside its
2690 // layer, so a layer recorded after a punch is cut
2691 // against it exactly like a draw would be — which is
2692 // what keeps a translucent layer over the slot from
2693 // being erased by it.
2694 if cuts && !pending.is_empty() {
2695 self.cut_for_punches(
2696 &mut pending,
2697 strips,
2698 depth,
2699 &mut first_segment,
2700 &mut run_start,
2701 &mut run_external,
2702 );
2703 }
2704
2705 // Consecutive draw batches merge into one segment; a
2706 // composite is what breaks the run, because it binds a
2707 // different page as its colour source.
2708 let end = self.alpha.len() as u32;
2709 self.push_run(run_start, end, run_external);
2710
2711 deepest = deepest.max(depth);
2712 self.segments
2713 .push(Segment::Composite(end, composite.parity));
2714 self.alpha
2715 .push(composite_instance(composite, window.origin(), depth));
2716 run_start = self.alpha.len() as u32;
2717 // A composite draws through group 1's own atlas
2718 // binding, so the run after it starts un-bound again.
2719 run_external = None;
2720 }
2721 }
2722 }
2723
2724 let end = self.alpha.len() as u32;
2725 self.push_run(run_start, end, run_external);
2726
2727 // Accumulated rather than assigned: a layer whose round was cut
2728 // renders in several rounds, and the depth its composite carries is
2729 // the deepest index across all of them.
2730 if let Some(page) = page
2731 && let Some(slot) = self.layer_depth.get_mut(page.layer as usize)
2732 {
2733 *slot = (*slot).max(deepest);
2734 }
2735
2736 let mut released = [false; MAX_LIVE_PAGES];
2737 for parity in &round.released {
2738 if let Some(slot) = released.get_mut(parity.index()) {
2739 *slot = true;
2740 }
2741 }
2742
2743 self.rounds.push(RoundPlan {
2744 page: page.map(|page| PagePlan {
2745 parity: page.parity,
2746 size: page.size,
2747 continued: page.continued,
2748 }),
2749 segments: first_segment..self.segments.len(),
2750 released,
2751 filter,
2752 // The round's last plan draws to its own end; a punch past
2753 // every op of the frame is issued after the whole walk instead.
2754 punch: (0, 0),
2755 });
2756 }
2757
2758 // Whatever no op was recorded after: a `ClearRect` recorded last, which
2759 // is the ordinary shape of a platform-view slot.
2760 if !pending.is_empty() {
2761 self.punch = self.push_punches(pending, strips);
2762 }
2763 }
2764
2765 /// Closes the open span of a surface round at `depth`, issuing every punch
2766 /// recorded before it as a pass of its own.
2767 ///
2768 /// This is the whole painter-order restoration: the ops recorded before the
2769 /// punch become a plan that ends here, the punch's own instances follow
2770 /// them in the buffer, and the ops recorded after it start a plan of their
2771 /// own — so the erase lands between the two rather than after both. A punch
2772 /// pass is a pass of its own because it draws through a different pipeline
2773 /// (destination-out) than the strips around it, not because of what it
2774 /// covers.
2775 ///
2776 /// Does nothing when no pending punch is shallower than `depth`, which is
2777 /// every op of every frame that records no clear.
2778 fn cut_for_punches(
2779 &mut self,
2780 pending: &mut &[ClearPunch],
2781 strips: &[Strip],
2782 depth: u32,
2783 first_segment: &mut usize,
2784 run_start: &mut u32,
2785 run_external: &mut Option<u32>,
2786 ) {
2787 let cut = pending
2788 .iter()
2789 .take_while(|punch| punch.depth < depth)
2790 .count();
2791 if cut == 0 {
2792 return;
2793 }
2794 let (issued, rest) = pending.split_at(cut);
2795 *pending = rest;
2796
2797 // Close the run this cut interrupts, so the plan's segments name only
2798 // what was recorded before the punch.
2799 let end = self.alpha.len() as u32;
2800 self.push_run(*run_start, end, *run_external);
2801 // The punch's own instances go in next and are drawn solid, so the run
2802 // resumed after the cut starts with nothing bound.
2803 *run_external = None;
2804
2805 let punch = self.push_punches(issued, strips);
2806 self.rounds.push(RoundPlan {
2807 page: None,
2808 segments: *first_segment..self.segments.len(),
2809 // The pages a cut round consumed are handed back when its LAST plan
2810 // ends: nothing between two plans of one round acquires a page, and
2811 // a composite in an earlier plan still has to sample the page it
2812 // names.
2813 released: [false; MAX_LIVE_PAGES],
2814 filter: None,
2815 punch,
2816 });
2817 *first_segment = self.segments.len();
2818 *run_start = self.alpha.len() as u32;
2819 }
2820
2821 /// Records the open run of instances `first..end` as one segment, drawn
2822 /// with the texture in `external` bound, or nothing at all when the run is
2823 /// empty.
2824 ///
2825 /// The one place a strip segment is created, so the run-breaking rule —
2826 /// a segment holds instances sharing one group-1 binding — is stated once
2827 /// rather than at each of the four points a run can close.
2828 fn push_run(&mut self, first: u32, end: u32, external: Option<u32>) {
2829 let Some(count) = end.checked_sub(first).filter(|count| *count > 0) else {
2830 return;
2831 };
2832 self.segments.push(match external {
2833 Some(slot) => Segment::External(first, count, slot),
2834 None => Segment::Strips(first, count),
2835 });
2836 }
2837
2838 /// Turns `punches` into destination-out instances, appended to the alpha
2839 /// region and named by no round's segments, as `(first, count)`.
2840 ///
2841 /// Each punch is a strip run like any other, drawn with an opaque source so
2842 /// the blend state's `1 − src.a` reaches zero exactly where the coverage is
2843 /// full, and carrying the painter-order depth it was hoisted from — which
2844 /// still orders it against the frame's one depth-writing pass, whose opaque
2845 /// coverage is recorded once ahead of every round and so cannot be cut.
2846 fn push_punches(&mut self, punches: &[ClearPunch], strips: &[Strip]) -> (u32, u32) {
2847 let first = self.alpha.len() as u32;
2848
2849 for punch in punches {
2850 let paint = PackedPaint {
2851 payload: PaintPayload::Solid(PUNCH_SOURCE),
2852 paint: SOLID_PAINT,
2853 depth_index: punch.depth,
2854 // An erase never joins the depth-writing pass: it establishes
2855 // no colour for a later fragment to be rejected against.
2856 opaque: false,
2857 external: None,
2858 };
2859 let Some(run) = strips.get(punch.strip_range.clone()) else {
2860 continue;
2861 };
2862 for pair in run.windows(2) {
2863 // Only the frame's own target is punched, so a punch's
2864 // instances are never shifted and never clipped.
2865 self.push_span(&pair[0], &pair[1], paint, false, PageWindow::ROOT);
2866 }
2867 }
2868
2869 (first, (self.alpha.len() as u32).saturating_sub(first))
2870 }
2871
2872 /// Emits the instances the `strip`/`next` pair describes: the strip's own
2873 /// alpha-sampled span, plus the solid span filling the gap to `next` when
2874 /// the winding between them says there is one.
2875 ///
2876 /// `window` shifts each instance's *geometry* into the round's target and
2877 /// clips it to the column that target holds, while the paint is still
2878 /// sampled at the scene position the strip was rasterized at — a layer's
2879 /// contents move into its page, the gradient or image painting them does
2880 /// not. An instance the window culls entirely is not emitted at all.
2881 fn push_span(
2882 &mut self,
2883 strip: &Strip,
2884 next: &Strip,
2885 paint: PackedPaint,
2886 to_opaque: bool,
2887 window: PageWindow,
2888 ) {
2889 let values = paint.values_at(strip.x, strip.y);
2890 let mut span = GpuStrip::from_strip_pair(strip, next, values);
2891 if window.place(&mut span, paint) {
2892 self.alpha.push(span);
2893 }
2894 // A gap starts where the strip ends rather than where it begins, so a
2895 // position-sampled paint is re-evaluated at the gap's own origin — the
2896 // gap is a different piece of the scene, not a continuation of the
2897 // span's sampling. `place` re-evaluates it once more if the window's
2898 // own left edge moves the instance again.
2899 if let Some(mut gap) = GpuStrip::gap_fill(strip, next, values) {
2900 gap.payload = paint.payload_at(gap.x, gap.y);
2901 if window.place(&mut gap, paint) {
2902 if to_opaque {
2903 self.opaque.push(gap);
2904 } else {
2905 self.alpha.push(gap);
2906 }
2907 }
2908 }
2909 }
2910
2911 /// Resolves `round`'s filter pass, if it has one, packing the layer's
2912 /// parameter block on first sight and claiming the pass's instance slot.
2913 ///
2914 /// `None` for an ordinary round, and also for the two shapes that cannot
2915 /// occur: a filter round with no page (the scheduler always gives one a
2916 /// page) and a recorded kind [`served_filter`] does not recognise (the
2917 /// scheduler refuses one before it plans a round). Both answer by leaving
2918 /// the round's pass unissued — its
2919 /// page is still cleared — rather than by asserting (E17).
2920 fn plan_filter(&mut self, frame: &CompiledFrame, round: &Round) -> Option<FilterPlan> {
2921 let pass = round.filter_pass()?;
2922 let page = round.page()?;
2923 let recorded = frame.recorder.layers.get(pass.layer as usize)?;
2924 let data_offset = self.filter_block(pass.layer, &recorded.kind)?;
2925
2926 let instance = self.filter_passes;
2927 self.filter_passes = self.filter_passes.saturating_add(1);
2928 Some(FilterPlan {
2929 step: pass.step,
2930 source: pass.source,
2931 data_offset,
2932 original: SizeU16::from(page.bounds),
2933 instance,
2934 })
2935 }
2936
2937 /// The texel offset of `layer`'s parameter block, packing the block on
2938 /// first sight.
2939 ///
2940 /// A linear scan rather than a map: a frame's filter layers are counted in
2941 /// ones (a filter layer is served only directly under the surface), and one
2942 /// allocation-free vector beats a hash map that would have to be cleared
2943 /// every frame.
2944 fn filter_block(&mut self, layer: u32, kind: &RecordedLayerKind) -> Option<u32> {
2945 let index = match self.filter_layers.iter().position(|id| *id == layer) {
2946 Some(index) => index,
2947 None => {
2948 // The one dispatch every filter-recognising site in this crate
2949 // shares (`schedule::layer_role`/`filter_rounds`), so the block
2950 // packed here is for the same filter the scheduler planned the
2951 // passes of.
2952 let block = match served_filter(layer, kind).ok()? {
2953 ServedFilter::Blur(blur) => GpuFilterData::from(GpuGaussianBlur::from(&blur)),
2954 ServedFilter::DropShadow(shadow) => {
2955 GpuFilterData::from(GpuDropShadow::from(&shadow))
2956 }
2957 };
2958 self.filter_layers.push(layer);
2959 self.filter_blocks.push(block);
2960 self.filter_layers.len().saturating_sub(1)
2961 }
2962 };
2963
2964 u32::try_from(index)
2965 .ok()?
2966 .checked_mul(GpuFilterData::SIZE_TEXELS)
2967 }
2968
2969 /// Whether this frame records a hole-punch pass at all — at a cut inside a
2970 /// surface round, or after every round of the frame.
2971 ///
2972 /// Derived rather than counted alongside the instances: the two places a
2973 /// punch can land are the two places its instances are recorded from, and
2974 /// one answer read off both is one fewer field to keep in step.
2975 fn punches(&self) -> bool {
2976 self.punch.1 > 0 || self.rounds.iter().any(|plan| plan.punch.1 > 0)
2977 }
2978
2979 /// How many of this frame's rounds render *strips* into a pooled page, and
2980 /// so how many viewport uniforms of their own the frame needs.
2981 ///
2982 /// A filter round targets a page too and is deliberately not counted: it
2983 /// binds no viewport uniform at all, mapping NDC against the destination
2984 /// extent its own instance carries.
2985 fn page_rounds(&self) -> usize {
2986 self.rounds
2987 .iter()
2988 .filter(|plan| plan.page.is_some() && plan.filter.is_none())
2989 .count()
2990 }
2991
2992 /// The instance bytes, opaque buffer first, so one vertex buffer serves
2993 /// every pass and each is a contiguous instance range into it.
2994 fn instance_bytes(&self) -> (&[u8], &[u8]) {
2995 (
2996 bytemuck::cast_slice(&self.opaque),
2997 bytemuck::cast_slice(&self.alpha),
2998 )
2999 }
3000}
3001
3002/// The column of device space one round's target holds: the origin its
3003/// instances are shifted to, and the edges they are clipped against.
3004///
3005/// A page holds its layer at the page's own origin, and its composite samples
3006/// exactly `(0, 0)`-to-its-own-extent back out
3007/// ([`Composite::source`](crate::schedule::Composite::source)). For a layer
3008/// [banded](crate::schedule::pages::page_bands) into column pages that makes
3009/// the band's own rectangle two things at once: the shift, and the *clip*. The
3010/// scheduler replays the layer's whole op list into every band — a band differs
3011/// only in which page it writes and where that page lands — so this is the site
3012/// that decides which part of each replayed strip belongs to the band at hand.
3013/// Emission is where the decision lives because it is the one place that knows
3014/// both the origin and the width; the alternative, clipping the ops
3015/// scheduler-side, would have to re-rasterize geometry the scheduler only holds
3016/// as draw ranges.
3017///
3018/// Two rules, and both halves of each matter:
3019///
3020/// - a span entirely outside the column contributes **no instance** — shifting
3021/// it by a saturating subtraction instead would clamp it onto the page's own
3022/// edge, stretching a span the scene drew elsewhere across content this band
3023/// really holds;
3024/// - a span straddling an edge has its geometry **and** its first alpha column
3025/// ([`GpuStrip::col_idx_or_rect_frac`]) advanced by the same amount, because
3026/// the shader reads that column unshifted and steps one column per pixel of
3027/// the instance (`shaders/strip.wgsl`'s `col_offset`): advancing the x
3028/// without the column would sample another pixel's coverage.
3029///
3030/// The frame's own surface is [`ROOT`](Self::ROOT), a window over the whole
3031/// device grid that shifts nothing and clips nothing. A layer that fits one
3032/// page is one band covering all of it, and a layer's bounds are the
3033/// tile-aligned union of its own draws' bounds, so the clip is a no-op for
3034/// every layer that is not banded.
3035#[derive(Debug, Clone, Copy, PartialEq, Eq)]
3036struct PageWindow {
3037 /// Left edge of the column in device space — what every instance's `x` is
3038 /// shifted by, and what one left of it is culled against.
3039 x0: u16,
3040 /// Right edge of the column in device space, exclusive.
3041 x1: u16,
3042 /// Top edge of the column in device space — what every instance's `y` is
3043 /// shifted by. A band spans its layer's whole height, so the vertical axis
3044 /// is shifted and never clipped.
3045 y0: u16,
3046}
3047
3048impl PageWindow {
3049 /// The whole device grid: the window a surface round carries.
3050 const ROOT: Self = Self {
3051 x0: 0,
3052 x1: u16::MAX,
3053 y0: 0,
3054 };
3055
3056 /// The window `page` renders through — its own tile-aligned bounds, which
3057 /// are one column band's for a banded layer and the whole layer's for
3058 /// every other.
3059 fn of(page: &PageTarget) -> Self {
3060 Self {
3061 x0: page.bounds.x0,
3062 x1: page.bounds.x1,
3063 y0: page.bounds.y0,
3064 }
3065 }
3066
3067 /// The origin instances are shifted by.
3068 fn origin(self) -> (u16, u16) {
3069 (self.x0, self.y0)
3070 }
3071
3072 /// Clips `span` — a strip instance still carrying its scene coordinates —
3073 /// to this window and shifts it into the target, answering whether any of
3074 /// it survived.
3075 ///
3076 /// `paint` is the draw's own paint, needed because a left-clipped instance
3077 /// starts at a different scene position than the one it was rasterized at:
3078 /// a position-sampled paint (a gradient, an image) is re-evaluated there,
3079 /// exactly as a gap fill is at its own origin, so the paint keeps landing
3080 /// where the scene put it rather than being squeezed into the clipped span.
3081 fn place(self, span: &mut GpuStrip, paint: PackedPaint) -> bool {
3082 let end = span.x.saturating_add(span.width);
3083 if span.width == 0 || end <= self.x0 || span.x >= self.x1 {
3084 return false;
3085 }
3086
3087 let cut = self.x0.saturating_sub(span.x);
3088 if cut > 0 {
3089 span.x = self.x0;
3090 span.width = span.width.saturating_sub(cut);
3091 // The dense part of an instance starts at its left edge, so the
3092 // columns cut off the geometry are cut off the coverage too.
3093 let dense = cut.min(span.dense_width_or_rect_height);
3094 span.dense_width_or_rect_height = span.dense_width_or_rect_height.saturating_sub(dense);
3095 // A sparse instance names no column at all: the fragment stage
3096 // reads `col + dense_width` as "does this instance sample
3097 // coverage", so leaving a non-zero column on one whose dense part
3098 // is gone would send it to the alpha texture for coverage it never
3099 // wrote.
3100 span.col_idx_or_rect_frac = if span.dense_width_or_rect_height == 0 {
3101 0
3102 } else {
3103 span.col_idx_or_rect_frac.saturating_add(u32::from(dense))
3104 };
3105 span.payload = paint.payload_at(span.x, span.y);
3106 }
3107
3108 // Past the right edge is outside the region the composite samples, so
3109 // it reaches no pixel of the parent either way; clipping it keeps a
3110 // band's instances inside the band's own rectangle rather than relying
3111 // on the page's quantized extent to swallow the overhang.
3112 let over = end.saturating_sub(self.x1);
3113 if over > 0 {
3114 span.width = span.width.saturating_sub(over);
3115 span.dense_width_or_rect_height = span.dense_width_or_rect_height.min(span.width);
3116 }
3117
3118 span.x = span.x.saturating_sub(self.x0);
3119 span.y = span.y.saturating_sub(self.y0);
3120 span.width > 0
3121 }
3122}
3123
3124/// The single quad that composites a finished page onto `origin`-shifted
3125/// target.
3126///
3127/// A whole-rectangle instance with no fractional edges: a layer's bounds are
3128/// tile-aligned, so the quad covers whole pixels and the fragment stage leaves
3129/// its coverage at one. The payload is the page texel the quad's top-left
3130/// corner samples — the page holds the layer at its own origin, so that is
3131/// `Composite::source`'s origin and not the layer's device position.
3132fn composite_instance(composite: &Composite, origin: (u16, u16), depth: u32) -> GpuStrip {
3133 let bounds = composite.bounds;
3134 let source = composite.source();
3135
3136 GpuStrip::from_rect(
3137 bounds.x0.saturating_sub(origin.0),
3138 bounds.y0.saturating_sub(origin.1),
3139 bounds.width(),
3140 bounds.height(),
3141 0,
3142 StripDraw {
3143 payload: pack_u16_pair(source.x0, source.y0),
3144 paint: LAYER_PAINT_SOURCE | u32::from(pack_opacity(composite.opacity)),
3145 depth_index: depth,
3146 },
3147 )
3148}
3149
3150/// A layer's constant opacity as the eight bits a composite instance carries.
3151///
3152/// Clamped rather than refused: the scheduler already admits only opacities
3153/// strictly between zero and one, and rounding is what keeps a 0.5 layer at
3154/// exactly the 128 the reference renderer's own packing produces.
3155fn pack_opacity(opacity: f32) -> u8 {
3156 (opacity.clamp(0.0, 1.0) * 255.0).round() as u8
3157}
3158
3159/// Where a strip instance's payload comes from.
3160///
3161/// A solid paint carries its colour there; every other paint reads its colour
3162/// from a record instead, and spends the payload on the scene position the
3163/// paint is sampled at — which is why the payload is per instance rather than
3164/// per draw.
3165#[derive(Debug, Clone, Copy)]
3166enum PaintPayload {
3167 /// A premultiplied RGBA8 colour, the same for every instance of the draw.
3168 Solid(u32),
3169 /// The instance's own scene-space origin, packed as a `u16` pair.
3170 Position,
3171}
3172
3173/// One draw's paint, resolved to the values its instances repeat.
3174#[derive(Debug, Clone, Copy)]
3175struct PackedPaint {
3176 payload: PaintPayload,
3177 /// The packed paint descriptor, without [`RECT_STRIP_FLAG`](crate::gpu::strips::RECT_STRIP_FLAG).
3178 paint: u32,
3179 depth_index: u32,
3180 /// Whether every pixel this paint produces is opaque, and so whether its
3181 /// fully-covered spans may take the depth-writing pass.
3182 opaque: bool,
3183 /// The external-texture slot this paint's instances have to be drawn with
3184 /// bound, or `None` for every paint the frame's atlas binding serves.
3185 ///
3186 /// What breaks a pass's instances into runs: the strip pipelines have one
3187 /// external binding, so consecutive instances may share a segment only
3188 /// while this stays equal (see [`crate::gpu::bindings::ExternalRuns`]).
3189 external: Option<u32>,
3190}
3191
3192impl PackedPaint {
3193 /// The payload an instance whose scene origin is `(x, y)` carries.
3194 fn payload_at(self, x: u16, y: u16) -> u32 {
3195 match self.payload {
3196 PaintPayload::Solid(rgba) => rgba,
3197 PaintPayload::Position => pack_u16_pair(x, y),
3198 }
3199 }
3200
3201 /// The per-instance values for an instance whose scene origin is `(x, y)`.
3202 fn values_at(self, x: u16, y: u16) -> StripDraw {
3203 StripDraw {
3204 payload: self.payload_at(x, y),
3205 paint: self.paint,
3206 depth_index: self.depth_index,
3207 }
3208 }
3209}
3210
3211/// Two `u16`s in one word, low half first — the packing the shader's
3212/// `unpack_u16_pair` reverses.
3213fn pack_u16_pair(x: u16, y: u16) -> u32 {
3214 u32::from(x) | (u32::from(y) << 16)
3215}
3216
3217/// The shader values for `paint`, resolved against the frame's paint slots.
3218///
3219/// Answers `None` for an indexed paint the frame could not resolve — one whose
3220/// entry is missing from `paint_slots` entirely, or one the lowering refused
3221/// because the engine has nowhere to hold its pixels yet (see the module
3222/// header). The draw is then skipped rather than stamped in a wrong colour.
3223fn pack_paint(
3224 paint: &Paint,
3225 depth: u32,
3226 paint_slots: &[Option<ResolvedPaint>],
3227) -> Option<PackedPaint> {
3228 match paint {
3229 Paint::Solid(color) => Some(PackedPaint {
3230 payload: PaintPayload::Solid(color.as_premul_rgba8().to_u32()),
3231 paint: SOLID_PAINT,
3232 depth_index: depth,
3233 opaque: color.is_opaque(),
3234 external: None,
3235 }),
3236 Paint::Indexed(indexed) => {
3237 let Some(Some(resolved)) = paint_slots.get(indexed.index()) else {
3238 INDEXED_PAINT_WARNING.call_once(|| {
3239 log::warn!(
3240 "an indexed paint could not be lowered to a GPU record; those draws are \
3241 skipped (logged once)"
3242 );
3243 });
3244 return None;
3245 };
3246 Some(PackedPaint {
3247 payload: PaintPayload::Position,
3248 paint: pack_paint_descriptor(resolved.paint_type, resolved.texel_offset),
3249 depth_index: depth,
3250 opaque: resolved.opaque,
3251 external: resolved.external,
3252 })
3253 }
3254 }
3255}
3256
3257/// One encoded paint after lowering: where its record landed and how a strip
3258/// instance names it.
3259#[derive(Debug, Clone, Copy)]
3260struct ResolvedPaint {
3261 /// How the fragment shader reads the record.
3262 paint_type: PaintType,
3263 /// The texel the record starts at in the encoded-paint texture.
3264 texel_offset: u32,
3265 /// Whether every pixel the paint produces is opaque.
3266 opaque: bool,
3267 /// The external-texture slot this paint samples, for a paint that reads a
3268 /// caller-owned texture rather than the atlas (see
3269 /// [`crate::gpu::bindings`]). `None` — the ordinary case — means the
3270 /// frame's atlas binding serves it.
3271 external: Option<u32>,
3272}
3273
3274/// One resource texture and the extent it currently holds.
3275#[derive(Debug)]
3276struct ResourceTexture {
3277 texture: wgpu::Texture,
3278 view: wgpu::TextureView,
3279 width: u32,
3280 height: u32,
3281}
3282
3283impl ResourceTexture {
3284 fn new(device: &wgpu::Device, descriptor: &wgpu::TextureDescriptor<'_>) -> Self {
3285 let texture = device.create_texture(descriptor);
3286 let view = texture.create_view(&wgpu::TextureViewDescriptor::default());
3287 Self {
3288 texture,
3289 view,
3290 width: descriptor.size.width,
3291 height: descriptor.size.height,
3292 }
3293 }
3294
3295 /// The whole texture as a copy destination.
3296 fn copy_target(&self) -> wgpu::TexelCopyTextureInfo<'_> {
3297 wgpu::TexelCopyTextureInfo {
3298 texture: &self.texture,
3299 mip_level: 0,
3300 origin: wgpu::Origin3d::ZERO,
3301 aspect: wgpu::TextureAspect::All,
3302 }
3303 }
3304
3305 /// The extent one full-texture upload covers.
3306 fn extent(&self) -> wgpu::Extent3d {
3307 wgpu::Extent3d {
3308 width: self.width,
3309 height: self.height,
3310 depth_or_array_layers: 1,
3311 }
3312 }
3313}
3314
3315/// The GPU resources one renderer holds across frames.
3316#[derive(Debug)]
3317struct FrameResources {
3318 alphas: ResourceTexture,
3319 paints: ResourceTexture,
3320 gradients: ResourceTexture,
3321 /// Stand-ins for the strip shader's bindings a given pass has nothing real
3322 /// for: the layer input (a real page when a composite is being drawn, this
3323 /// otherwise), the glyph/image atlas array until the first image is
3324 /// resident, and an externally bound texture, which nothing writes yet.
3325 /// Every declared binding has to be bound for a pass to validate, whether
3326 /// or not an instance samples it.
3327 placeholders: Placeholders,
3328 config: wgpu::Buffer,
3329 /// One viewport uniform per page round of the busiest frame so far.
3330 ///
3331 /// A round's NDC mapping is against the extent of the attachment it writes,
3332 /// and a page's extent is neither the frame's nor the same from one round
3333 /// to the next, so each needs a buffer of its own — a bind group holds the
3334 /// whole buffer, not an offset into one. Grown only, like every other
3335 /// retained resource here.
3336 page_configs: Vec<wgpu::Buffer>,
3337 instances: Option<wgpu::Buffer>,
3338 instance_capacity: u64,
3339 /// The GPU records this frame's indexed paints resolve against, in
3340 /// serialization order — which is the order the texel offsets in
3341 /// `paint_slots` were taken from.
3342 paints_data: Vec<GpuEncodedPaint>,
3343 /// One slot per encoded paint of the frame, indexed by
3344 /// [`Paint::Indexed`](vello_common::paint::Paint::Indexed): `None` for one
3345 /// the engine could not lower. Kept parallel to the *encoded* paints
3346 /// rather than compacted to the lowered ones, because a draw names its
3347 /// paint by the compiler's index.
3348 paint_slots: Vec<Option<ResolvedPaint>>,
3349 /// Reusable per-paint ramp residency, filled from a frame's LUT requests
3350 /// before its paints are lowered.
3351 paint_ramps: Vec<Option<CachedRamp>>,
3352 /// Reusable staging for the encoded-paint upload, padded to the texture's
3353 /// footprint.
3354 paint_staging: Vec<u8>,
3355 /// The distinct external textures this frame's paints sample, in the order
3356 /// they were first named — the slot numbering the pass segments and the
3357 /// group-1 bind groups both address by.
3358 external_runs: ExternalRuns,
3359 /// One set of bind groups per strip pipeline variant. Not one shared set:
3360 /// every engine pipeline uses wgpu's derived layout, and a derived layout
3361 /// is exclusive to the pipeline that derived it.
3362 bind_groups: HashMap<EnginePipeline, StripBindGroups>,
3363 /// The target format the live bind groups were built against; a change
3364 /// means different pipelines, so the whole map is dropped rather than
3365 /// accumulating a set per format ever rendered to.
3366 bind_group_format: Option<wgpu::TextureFormat>,
3367 /// The real image atlas array, created lazily by the first frame that
3368 /// makes an image resident and grown from then on — `None` is exactly
3369 /// [`Placeholders::atlas_array`]'s domain, a renderer that has never
3370 /// drawn an image.
3371 atlas: Option<AtlasArray>,
3372 /// The atlas rectangle every image the renderer has ever drawn currently
3373 /// holds, keyed by the stable [`ImageId`] the compiler's own residency
3374 /// minted it. See [`FrameResources::resolve_paints`] for how this is
3375 /// kept in step with residency across frames without reaching into the
3376 /// compiler's own cache.
3377 image_registry: HashMap<ImageId, ResidentImage>,
3378 /// How many atlas regions have been declined — a write or clear the array
3379 /// refused, or one the budget says the array could not hold. Counted rather
3380 /// than dropped silently, because each one is a region the frame believed
3381 /// it had filled.
3382 refused_regions: u64,
3383}
3384
3385impl FrameResources {
3386 fn new(device: &wgpu::Device, dim: u32) -> Self {
3387 let min = gpu::MIN_RESOURCE_TEXTURE_HEIGHT;
3388 Self {
3389 alphas: ResourceTexture::new(device, &gpu::alpha_texture_descriptor(dim, min)),
3390 paints: ResourceTexture::new(
3391 device,
3392 &gpu::paint_texture::encoded_paints_texture_descriptor(dim, min),
3393 ),
3394 gradients: ResourceTexture::new(device, &gradient_texture_descriptor(dim, min)),
3395 placeholders: Placeholders::new(device),
3396 config: device.create_buffer(&config_descriptor("frust-engine config uniform")),
3397 page_configs: Vec::new(),
3398 instances: None,
3399 instance_capacity: 0,
3400 paints_data: Vec::new(),
3401 paint_slots: Vec::new(),
3402 paint_ramps: Vec::new(),
3403 paint_staging: Vec::new(),
3404 external_runs: ExternalRuns::new(),
3405 bind_groups: HashMap::new(),
3406 bind_group_format: None,
3407 atlas: None,
3408 image_registry: HashMap::new(),
3409 refused_regions: 0,
3410 }
3411 }
3412
3413 /// Drops the atlas array, everything recorded about what lives in it, and
3414 /// the bind groups naming it.
3415 ///
3416 /// What a re-budget needs: the rectangles the registry holds were allocated
3417 /// in a geometry that no longer exists, so keeping any of them would point
3418 /// a paint at a rectangle of a texture that is gone.
3419 fn reset_atlas(&mut self) {
3420 self.atlas = None;
3421 self.image_registry.clear();
3422 self.bind_groups.clear();
3423 }
3424
3425 /// Count `regions` atlas regions as declined, saying so once.
3426 ///
3427 /// Once, not per region: a budget and an array that disagree disagree about
3428 /// every region, and a per-frame line would bury the fact under itself. The
3429 /// count on [`EngineRenderer::refused_atlas_regions`] is the measure.
3430 fn note_refused_regions(&mut self, regions: u64) {
3431 self.refused_regions = self.refused_regions.saturating_add(regions);
3432 ATLAS_REFUSAL_WARNING.call_once(|| {
3433 log::warn!(
3434 "an atlas region was refused by the image atlas array; those images are skipped \
3435 and their uploads re-offered on a later frame (logged once — see \
3436 EngineRenderer::refused_atlas_regions for the count)"
3437 );
3438 });
3439 }
3440
3441 /// Services `frame`'s LUT and image residency, then lowers its encoded
3442 /// paints into the records the shader samples.
3443 ///
3444 /// Leaves `paints_data` holding the lowered records in serialization order
3445 /// and `paint_slots` naming, per *encoded* paint, the texel its record
3446 /// starts at — or `None` where the paint could not be lowered.
3447 ///
3448 /// A solid-only frame leaves both empty and touches neither the cache nor
3449 /// the paint texture, so it costs exactly what it did before paints were
3450 /// wired up. Image residency is still serviced even then, since an image
3451 /// can be evicted on a frame that draws nothing at all (see
3452 /// [`Self::update_image_registry`]).
3453 fn resolve_paints(
3454 &mut self,
3455 frame: &CompiledFrame,
3456 cache: &mut GradientCache,
3457 budget: AtlasBudget,
3458 externals: &ExternalTextures,
3459 ) {
3460 self.paints_data.clear();
3461 self.paint_slots.clear();
3462 self.external_runs.clear();
3463
3464 // Kept in step every frame, not only when this frame's own paints
3465 // need it: an image reaped by the compiler's age-based eviction while
3466 // nothing draws it must still be forgotten here, or a later draw that
3467 // reuses its freed rectangle's `ImageId` would read the stale entry.
3468 self.update_image_registry(frame, budget);
3469
3470 if frame.encoded_paints.is_empty() {
3471 return;
3472 }
3473
3474 // Ramp residency next, for the whole frame: a ramp's offset is only
3475 // meaningful once the cache has finished baking this frame's misses,
3476 // and a record built before that would name a ramp that had not been
3477 // packed yet.
3478 self.paint_ramps.clear();
3479 self.paint_ramps.resize(frame.encoded_paints.len(), None);
3480 for request in &frame.lut_requests {
3481 let ramp = resolve_lut_request(*request, &frame.encoded_paints, cache);
3482 if let Some(slot) = self.paint_ramps.get_mut(request.paint_index) {
3483 *slot = ramp;
3484 }
3485 }
3486
3487 let mut texel_offset = 0;
3488 for (index, paint) in frame.encoded_paints.iter().enumerate() {
3489 // An external texture's slot travels with its record: the record
3490 // itself only says "sample the external binding", and which
3491 // texture that binding holds is settled per run when the pass is
3492 // recorded rather than per paint.
3493 let mut external = None;
3494 let lowered = match paint {
3495 EncodedPaint::Image(image) => image_id(image)
3496 .and_then(|id| self.image_registry.get(&id))
3497 .and_then(|resident| {
3498 lower_encoded_image(image, resident)
3499 .map(|record| fit_minified(record, resident))
3500 }),
3501 // A paint naming a texture nothing is bound under lowers to
3502 // nothing, so its draws are skipped rather than sampling
3503 // whichever texture the binding happens to hold.
3504 EncodedPaint::ExternalTexture(entry) => externals
3505 .view(entry.texture_id.0)
3506 .and_then(|_| self.external_runs.slot_of(entry.texture_id.0))
3507 .map(|slot| {
3508 external = Some(slot);
3509 lower_encoded_external(entry)
3510 }),
3511 _ => {
3512 let ramp = self.paint_ramps.get(index).copied().flatten();
3513 lower_encoded_paint(paint, ramp)
3514 }
3515 };
3516 let slot = lowered.map(|record| {
3517 let resolved = ResolvedPaint {
3518 paint_type: record.paint_type(),
3519 texel_offset,
3520 opaque: !paint.may_have_transparency(),
3521 external,
3522 };
3523 texel_offset += record.texel_len();
3524 self.paints_data.push(record);
3525 resolved
3526 });
3527 self.paint_slots.push(slot);
3528 }
3529 }
3530
3531 /// Keeps [`Self::image_registry`] in step with the compiler's own image
3532 /// residency, without reaching into it: the residency's rectangles are
3533 /// not reachable from here (see [`crate::gpu::paint_texture::lower_encoded_paint`]'s
3534 /// doc for why), so this reconstructs the same information from what a
3535 /// compiled frame already reports.
3536 ///
3537 /// Two passes, and each reads the frame's plan directly rather than
3538 /// inferring anything from draw order. First, every region this frame's
3539 /// residency reaped is forgotten — `frame.image_evictions` names it by
3540 /// rectangle, and a rectangle uniquely identifies the one image that held it
3541 /// (padding is always zero in this engine, so the reported and the stored
3542 /// rectangle are the same value; see [`crate::cache::images`]'s module doc).
3543 /// Second, every entry of `frame.image_uploads` is registered under the
3544 /// [`ImageId`] it carries.
3545 ///
3546 /// The order matters and the id does. Evictions run first so a same-frame
3547 /// evict-then-reallocate that reuses a rectangle registers the new tenant
3548 /// rather than having it removed again. And the upload naming its own id is
3549 /// what makes this sound under a *re-offered* plan: an upload the previous
3550 /// frame did not service is reported again alongside no new draw of its own,
3551 /// so a walk pairing uploads positionally against this frame's encoded
3552 /// paints would hand a fresh image the stale upload's rectangle.
3553 ///
3554 /// A region the atlas array could not hold at `budget`'s geometry and this
3555 /// frame's depth is counted and left unregistered instead. Its draws are
3556 /// then skipped, which is the whole point: a registered rectangle nothing
3557 /// wrote would be sampled as whatever the texture happened to contain.
3558 fn update_image_registry(&mut self, frame: &CompiledFrame, budget: AtlasBudget) {
3559 if !frame.image_evictions.is_empty() {
3560 // A set rather than `Vec::contains`: both sides of this scan scale
3561 // with content — the registry with how many images and glyph slots
3562 // are resident, the plan with how hard the atlas is churning — and
3563 // their product is the frame path's, not a report's.
3564 let cleared: HashSet<AtlasRegion> = frame.image_evictions.iter().copied().collect();
3565 self.image_registry
3566 .retain(|_, resident| !cleared.contains(&resident.region));
3567 }
3568
3569 let mut refused = 0_u64;
3570 for upload in &frame.image_uploads {
3571 if !budget.contains(upload.region, frame.atlas_layers) {
3572 refused = refused.saturating_add(1);
3573 continue;
3574 }
3575 self.image_registry.insert(
3576 upload.id,
3577 ResidentImage {
3578 id: upload.id,
3579 region: upload.region,
3580 natural: upload.natural,
3581 padding: u32::from(ATLAS_PADDING),
3582 may_have_transparency: upload.may_have_transparency,
3583 },
3584 );
3585 }
3586 // The glyph half of the same registry, and the reason it is a second
3587 // loop rather than a branch inside the first: a glyph slot carries no
3588 // pixels and is *not* re-offered across frames, because the pixels are
3589 // produced by the replay pass rather than uploaded from here. Its
3590 // rectangle is reported by every draw that names it (see
3591 // [`crate::compile::GlyphSlot`]), so registering it here is what makes
3592 // a handle `glifo` recycled resolve against its current occupant.
3593 for slot in &frame.glyph_slots {
3594 if !budget.contains(slot.region, frame.atlas_layers) {
3595 refused = refused.saturating_add(1);
3596 continue;
3597 }
3598 self.image_registry.insert(
3599 slot.id,
3600 ResidentImage {
3601 id: slot.id,
3602 region: slot.region,
3603 // Never minified: a glyph is rasterized straight into the
3604 // rectangle it was allocated, so the natural extent and the
3605 // resident one are the same value by construction.
3606 natural: slot.region.size,
3607 padding: slot.padding,
3608 may_have_transparency: true,
3609 },
3610 );
3611 }
3612
3613 if refused > 0 {
3614 self.note_refused_regions(refused);
3615 }
3616 }
3617
3618 /// Grows or creates the image atlas array to hold `layers` layers at
3619 /// `budget`'s per-layer extent, clearing the live bind groups when it
3620 /// does — a bind group built against the old (or absent) atlas view would
3621 /// otherwise sample nothing, or a freed texture.
3622 ///
3623 /// A frame that has never made an image resident (`layers == 0`) leaves
3624 /// the atlas unset, so the strip shader's binding stays on
3625 /// [`Placeholders::atlas_array`] until the first one is.
3626 ///
3627 /// Growth submits a maintenance command buffer of its own rather than
3628 /// recording into the frame's encoder — see [`crate::gpu::atlas`]. So this
3629 /// must be called after the frame's last fallible step (a refused frame
3630 /// must submit nothing) and before its atlas writes are issued (the copy has
3631 /// to precede them, and a submit flushes whatever is already queued).
3632 fn ensure_atlas(
3633 &mut self,
3634 device: &wgpu::Device,
3635 queue: &wgpu::Queue,
3636 budget: AtlasBudget,
3637 layers: u32,
3638 ) -> bool {
3639 let grew = match &mut self.atlas {
3640 None if layers == 0 => false,
3641 None => {
3642 self.atlas = Some(AtlasArray::with_layers(
3643 device,
3644 budget.atlas_size.0,
3645 budget.atlas_size.1,
3646 layers,
3647 ));
3648 true
3649 }
3650 Some(atlas) => atlas.ensure_layers(device, queue, layers),
3651 };
3652 if grew {
3653 self.bind_groups.clear();
3654 }
3655 grew
3656 }
3657
3658 /// Flushes `frame`'s atlas evictions and uploads against the live atlas
3659 /// array, evictions first — a rectangle this frame's residency freed may
3660 /// already hold a fresh upload by the time this runs (see
3661 /// [`crate::cache::images`]'s module doc), so clearing after writing
3662 /// would erase the image that just moved in.
3663 ///
3664 /// Answers whether the whole plan reached the array — which is what tells
3665 /// the compiler it may stop re-offering it. A region the array declined is
3666 /// counted (see [`Self::note_refused_regions`]) and the plan stays pending,
3667 /// so a later frame that has grown the array writes it rather than the
3668 /// image being lost.
3669 ///
3670 /// An absent array with a plan to service is that same disagreement rather
3671 /// than a quiet no-op: both are driven by the same `frame.atlas_layers`, so
3672 /// an empty plan and an absent atlas normally agree.
3673 fn upload_atlas(&mut self, queue: &wgpu::Queue, frame: &CompiledFrame) -> bool {
3674 let refused = match self.atlas.as_ref() {
3675 None => (frame.image_evictions.len() + frame.image_uploads.len()) as u64,
3676 Some(atlas) => {
3677 let mut refused = 0_u64;
3678 for region in &frame.image_evictions {
3679 if !atlas.clear_region(queue, *region) {
3680 refused = refused.saturating_add(1);
3681 }
3682 }
3683 for upload in &frame.image_uploads {
3684 if !atlas.write_region(queue, upload.region, upload.pixels.data_as_u8_slice()) {
3685 refused = refused.saturating_add(1);
3686 }
3687 }
3688 refused
3689 }
3690 };
3691
3692 if refused == 0 {
3693 return true;
3694 }
3695 self.note_refused_regions(refused);
3696 false
3697 }
3698
3699 /// Makes sure there is one viewport uniform per page round of this frame.
3700 ///
3701 /// Grown only: the buffers are 32 bytes each and a frame that once needed
3702 /// four keeps them rather than reallocating on the next frame that does.
3703 fn ensure_page_configs(&mut self, device: &wgpu::Device, rounds: usize) {
3704 while self.page_configs.len() < rounds {
3705 self.page_configs
3706 .push(device.create_buffer(&config_descriptor("frust-engine page config uniform")));
3707 }
3708 }
3709
3710 /// The texels this frame's encoded paints occupy.
3711 fn paint_texels(&self) -> u32 {
3712 self.paints_data
3713 .iter()
3714 .map(GpuEncodedPaint::texel_len)
3715 .sum()
3716 }
3717
3718 /// The height the gradient LUT texture must grow to for every ramp the
3719 /// cache has packed, or `None` when it already fits.
3720 ///
3721 /// # Errors
3722 ///
3723 /// [`EngineError::PaintCapacity`]: a LUT set past the resource dimension
3724 /// squared has nowhere to live, and the frame path reports it rather than
3725 /// asserting.
3726 fn grown_gradient_height(&self, cache: &GradientCache) -> Result<Option<u32>, EngineError> {
3727 let width = self.gradients.width;
3728 let texels = u32::try_from(cache.luts_size() / BYTES_PER_TEXEL as usize)
3729 .map_err(|_| EngineError::PaintCapacity)?;
3730 let required = texels.div_ceil(width).max(gpu::MIN_RESOURCE_TEXTURE_HEIGHT);
3731 if required > width {
3732 return Err(EngineError::PaintCapacity);
3733 }
3734 Ok((required > self.gradients.height).then_some(required))
3735 }
3736
3737 fn resize_alphas(&mut self, device: &wgpu::Device, height: Option<u32>) {
3738 if let Some(height) = height {
3739 self.alphas = ResourceTexture::new(
3740 device,
3741 &gpu::alpha_texture_descriptor(self.alphas.width, height),
3742 );
3743 self.bind_groups.clear();
3744 }
3745 }
3746
3747 fn resize_paints(&mut self, device: &wgpu::Device, height: Option<u32>) {
3748 if let Some(height) = height {
3749 self.paints = ResourceTexture::new(
3750 device,
3751 &gpu::paint_texture::encoded_paints_texture_descriptor(self.paints.width, height),
3752 );
3753 self.bind_groups.clear();
3754 }
3755 }
3756
3757 fn resize_gradients(&mut self, device: &wgpu::Device, height: Option<u32>) {
3758 if let Some(height) = height {
3759 self.gradients = ResourceTexture::new(
3760 device,
3761 &gradient_texture_descriptor(self.gradients.width, height),
3762 );
3763 self.bind_groups.clear();
3764 }
3765 }
3766
3767 /// Uploads the frame's coverage, paint records, colour ramps, atlas
3768 /// evictions/uploads and config, answering whether the atlas plan was
3769 /// serviced in full (see [`Self::upload_atlas`]).
3770 fn upload(
3771 &mut self,
3772 queue: &wgpu::Queue,
3773 frame: &mut CompiledFrame,
3774 cache: &mut GradientCache,
3775 size: (u16, u16),
3776 dim: u32,
3777 ) -> bool {
3778 let atlas_serviced = self.upload_atlas(queue, frame);
3779
3780 let alphas = &self.alphas;
3781 gpu::with_padded_alphas(
3782 &mut frame.strips.alphas,
3783 alphas.width,
3784 alphas.height,
3785 |bytes| {
3786 queue.write_texture(
3787 alphas.copy_target(),
3788 bytes,
3789 resource_layout(gpu::resource_bytes_per_row(alphas.width), alphas.height),
3790 alphas.extent(),
3791 );
3792 },
3793 );
3794
3795 if !self.paints_data.is_empty() {
3796 let footprint = gpu::resource_texture_bytes(self.paints.width, self.paints.height);
3797 self.paint_staging
3798 .resize(usize::try_from(footprint).unwrap_or(usize::MAX), 0);
3799 // The height was checked to fit before any pass was recorded, so a
3800 // buffer too short for the records is unreachable; skipping the
3801 // upload rather than unwrapping keeps the frame path total anyway.
3802 if GpuEncodedPaint::serialize_to_buffer(&self.paints_data, &mut self.paint_staging)
3803 .is_ok()
3804 {
3805 queue.write_texture(
3806 self.paints.copy_target(),
3807 &self.paint_staging,
3808 resource_layout(
3809 gpu::resource_bytes_per_row(self.paints.width),
3810 self.paints.height,
3811 ),
3812 self.paints.extent(),
3813 );
3814 }
3815 }
3816
3817 if cache.has_changed() {
3818 let layout = GradientTextureLayout {
3819 width: self.gradients.width,
3820 height: self.gradients.height,
3821 };
3822 if let Some(upload) = cache.begin_upload(layout) {
3823 queue.write_texture(
3824 self.gradients.copy_target(),
3825 &upload,
3826 resource_layout(upload.bytes_per_row(), self.gradients.height),
3827 self.gradients.extent(),
3828 );
3829 }
3830 cache.mark_synced();
3831 }
3832
3833 let config = GpuConfig::new(u32::from(size.0), u32::from(size.1), dim, dim);
3834 queue.write_buffer(&self.config, 0, bytemuck::bytes_of(&config));
3835
3836 atlas_serviced
3837 }
3838
3839 /// Grows the instance buffer if this frame outgrew it, then uploads the
3840 /// opaque and alpha instances back to back.
3841 fn upload_instances(&mut self, device: &wgpu::Device, queue: &wgpu::Queue, scratch: &Scratch) {
3842 let (opaque, alpha) = scratch.instance_bytes();
3843 let required = (opaque.len() + alpha.len()) as u64;
3844 if required == 0 {
3845 return;
3846 }
3847
3848 let buffer = match &mut self.instances {
3849 Some(buffer) if self.instance_capacity >= required => buffer,
3850 slot => {
3851 let floor = MIN_INSTANCE_CAPACITY * size_of::<GpuStrip>() as u64;
3852 let capacity = required
3853 .checked_next_power_of_two()
3854 .unwrap_or(required)
3855 .max(floor);
3856 self.instance_capacity = capacity;
3857 slot.insert(device.create_buffer(&wgpu::BufferDescriptor {
3858 label: Some("frust-engine strip instances"),
3859 size: capacity,
3860 usage: wgpu::BufferUsages::VERTEX | wgpu::BufferUsages::COPY_DST,
3861 mapped_at_creation: false,
3862 }))
3863 }
3864 };
3865 if !opaque.is_empty() {
3866 queue.write_buffer(buffer, 0, opaque);
3867 }
3868 if !alpha.is_empty() {
3869 queue.write_buffer(buffer, opaque.len() as u64, alpha);
3870 }
3871 }
3872
3873 /// Builds `variant`'s bind groups if it has none, first dropping every
3874 /// set when the target format changed under them.
3875 fn ensure_bind_groups(
3876 &mut self,
3877 device: &wgpu::Device,
3878 variant: EnginePipeline,
3879 pipeline: &wgpu::RenderPipeline,
3880 format: wgpu::TextureFormat,
3881 ) {
3882 if self.bind_group_format != Some(format) {
3883 self.bind_groups.clear();
3884 self.bind_group_format = Some(format);
3885 }
3886 if self.bind_groups.contains_key(&variant) {
3887 return;
3888 }
3889 let atlas_view = self
3890 .atlas
3891 .as_ref()
3892 .map(AtlasArray::view)
3893 .unwrap_or(&self.placeholders.atlas_array);
3894 let groups = StripBindGroups::new(
3895 device,
3896 pipeline,
3897 &self.alphas.view,
3898 &self.config,
3899 &self.placeholders,
3900 atlas_view,
3901 &self.paints.view,
3902 &self.gradients.view,
3903 );
3904 self.bind_groups.insert(variant, groups);
3905 }
3906
3907 /// Builds `variant`'s group 1 for every external texture this frame draws
3908 /// that it has none for yet.
3909 ///
3910 /// Called after [`Self::ensure_bind_groups`] has built the variant's own
3911 /// set, and only for the variants the frame records with, so a frame that
3912 /// draws no external texture builds nothing. What is built is retained
3913 /// alongside the rest of the variant's groups and dropped with them — on a
3914 /// format change, an atlas growth or a re-budget, all of which change the
3915 /// atlas view this group also holds.
3916 ///
3917 /// A texture with no registered view is skipped rather than substituted:
3918 /// its paint did not resolve either, so nothing in the frame names its
3919 /// slot.
3920 fn ensure_external_groups(
3921 &mut self,
3922 device: &wgpu::Device,
3923 variant: EnginePipeline,
3924 pipeline: &wgpu::RenderPipeline,
3925 externals: &ExternalTextures,
3926 ) {
3927 if self.external_runs.is_empty() {
3928 return;
3929 }
3930 // Destructured rather than reached through `self`: the group map is
3931 // borrowed mutably while the atlas view and the placeholders are read.
3932 let Self {
3933 bind_groups,
3934 external_runs,
3935 atlas,
3936 placeholders,
3937 ..
3938 } = self;
3939 let atlas_view = atlas
3940 .as_ref()
3941 .map(AtlasArray::view)
3942 .unwrap_or(&placeholders.atlas_array);
3943 let Some(groups) = bind_groups.get_mut(&variant) else {
3944 return;
3945 };
3946 for key in external_runs.keys() {
3947 if groups.externals.contains_key(key) {
3948 continue;
3949 }
3950 let Some(view) = externals.view(*key) else {
3951 continue;
3952 };
3953 groups
3954 .externals
3955 .insert(*key, images_bind_group(device, pipeline, atlas_view, view));
3956 }
3957 }
3958
3959 /// Drops every bind group naming the texture registered under `key`.
3960 ///
3961 /// Called when that registration changes, because a group holds its view
3962 /// by value: keeping one past a re-bind would go on sampling the texture
3963 /// the caller replaced, and keeping one past an unbind would hold the
3964 /// caller's texture alive for as long as this renderer lives.
3965 fn forget_external(&mut self, key: u64) {
3966 for groups in self.bind_groups.values_mut() {
3967 groups.externals.remove(&key);
3968 }
3969 }
3970}
3971
3972/// Lower one atlas page's recorded commands into the strips that draw it,
3973/// answering whether the whole page could be expressed.
3974///
3975/// The caller-supplied half of the render-to-atlas seam: `gpu::atlas` owns the
3976/// pass, the orderings and the submit, while turning a command stream into
3977/// strips is compiler work and stays on this side of the edge — `compile`
3978/// already depends on `gpu::atlas`, so taking the reverse dependency would make
3979/// the two mutually recursive.
3980///
3981/// The stream is replayed as a scene in *page* space and compiled by
3982/// `lowering`, which is why that compiler is sized to the page rather than to
3983/// the surface. Only the four commands an outline glyph produces are lowered;
3984/// anything else — a clip path, a blend layer, a gradient paint, which is to
3985/// say every COLR shape — refuses the page whole rather than drawing part of
3986/// it. An indexed paint coming back out of the compile means the same thing:
3987/// the atlas pass binds no paint texture, so a record it would have to sample
3988/// cannot be drawn.
3989///
3990/// Refusing a page is a *last* line rather than the design, because refusal
3991/// cannot be made harmless here: `glifo` clears a recorder's commands whether
3992/// or not this answered `true`, and it offers no way to withdraw the entries
3993/// whose pixels those commands were going to be. Nothing that would reach this
3994/// refusal is therefore admitted to the atlas in the first place — a colour
3995/// face never takes the atlas route at all (`crate::text::atlas_policy`), so
3996/// what arrives here is the solid outline stream this lowers.
3997fn lower_atlas_page(
3998 recorder: &AtlasCommandRecorder,
3999 buffers: &mut AtlasPageBuffers,
4000 lowering: &mut SceneCompiler,
4001 page: (u16, u16),
4002) -> bool {
4003 let mut scene = Scene::new();
4004 {
4005 let mut builder = SceneBuilder::new(&mut scene);
4006 let mut transform = Affine::IDENTITY;
4007 let mut brush = Brush::Solid(Color::BLACK);
4008
4009 for command in &recorder.commands {
4010 match command {
4011 AtlasCommand::SetTransform(next) => transform = *next,
4012 AtlasCommand::SetPaint(AtlasPaint::Solid(color)) => brush = Brush::Solid(*color),
4013 AtlasCommand::FillPath(path) => {
4014 builder.push_transform(transform);
4015 builder.fill_path((**path).clone(), brush.clone());
4016 builder.pop_transform();
4017 }
4018 AtlasCommand::FillRect(rect) => {
4019 builder.push_transform(transform);
4020 builder.fill_rect(*rect, brush.clone());
4021 builder.pop_transform();
4022 }
4023 _ => return false,
4024 }
4025 }
4026 }
4027
4028 let Ok(frame) = lowering.compile(&scene, Affine::IDENTITY, page) else {
4029 return false;
4030 };
4031
4032 let strips = frame.strip_buf();
4033 for draw in frame.draws() {
4034 let Paint::Solid(color) = &draw.paint else {
4035 return false;
4036 };
4037 let Some(run) = strips.get(draw.strip_range.clone()) else {
4038 continue;
4039 };
4040 push_solid_strips(
4041 run,
4042 color.as_premul_rgba8().to_u32(),
4043 draw.depth,
4044 &mut buffers.instances,
4045 );
4046 }
4047 buffers.alphas.extend_from_slice(frame.alphas());
4048 true
4049}
4050
4051/// The [`ImageId`] an encoded image paint names, or `None` for the one
4052/// [`ImageSource`] variant no residency ever mints — the paint carrying its
4053/// pixels inline as a [`vello_common::pixmap::Pixmap`] rather than through a
4054/// handle. The compiler's own image encoding always produces the handle form
4055/// (see [`crate::compile::paint::encode_image`]), so a compiled frame never
4056/// exercises the `None` arm; it exists because the type itself admits both.
4057fn image_id(image: &EncodedImage) -> Option<ImageId> {
4058 match image.source {
4059 ImageSource::OpaqueId { id, .. } => Some(id),
4060 ImageSource::Pixmap(_) => None,
4061 }
4062}
4063
4064/// A lowered image record corrected for an atlas rectangle that holds a
4065/// *minified* copy of the source.
4066///
4067/// The record's transform maps a device position back onto the image's own
4068/// texels, and the compiler composed it against the source's declared extent —
4069/// it had to, since that is all it knows before residency is consulted. Scaling
4070/// its output by [`ResidentImage::minify_scale`] retargets it at the smaller
4071/// rectangle actually uploaded, which is the whole correction a downsampled
4072/// image needs: the record's `image_size` and `image_offset` already describe
4073/// the resident rectangle.
4074///
4075/// A record stored at full size is returned untouched, which is every image but
4076/// one larger than an atlas layer.
4077fn fit_minified(record: GpuEncodedPaint, resident: &ResidentImage) -> GpuEncodedPaint {
4078 let Some((x, y)) = resident.minify_scale() else {
4079 return record;
4080 };
4081 let GpuEncodedPaint::Image(mut image) = record else {
4082 return record;
4083 };
4084
4085 // `[a, b, c, d, tx, ty]`, mapping `(u, v)` to `(a·u + c·v + tx, b·u + d·v +
4086 // ty)`: the x row is scaled by one factor and the y row by the other.
4087 image.transform[0] *= x;
4088 image.transform[2] *= x;
4089 image.transform[4] *= x;
4090 image.transform[1] *= y;
4091 image.transform[3] *= y;
4092 image.transform[5] *= y;
4093
4094 GpuEncodedPaint::Image(image)
4095}
4096
4097/// The descriptor every `Config` uniform buffer is created with.
4098fn config_descriptor(label: &str) -> wgpu::BufferDescriptor<'_> {
4099 wgpu::BufferDescriptor {
4100 label: Some(label),
4101 size: GpuConfig::SIZE,
4102 usage: wgpu::BufferUsages::UNIFORM | wgpu::BufferUsages::COPY_DST,
4103 mapped_at_creation: false,
4104 }
4105}
4106
4107/// The texel copy layout of a full resource-texture upload.
4108fn resource_layout(bytes_per_row: u32, rows: u32) -> wgpu::TexelCopyBufferLayout {
4109 wgpu::TexelCopyBufferLayout {
4110 offset: 0,
4111 bytes_per_row: Some(bytes_per_row),
4112 rows_per_image: Some(rows),
4113 }
4114}
4115
4116/// The descriptor for a gradient LUT texture of `width` x `height` texels.
4117fn gradient_texture_descriptor(width: u32, height: u32) -> wgpu::TextureDescriptor<'static> {
4118 wgpu::TextureDescriptor {
4119 label: Some("frust-engine gradient texture"),
4120 size: wgpu::Extent3d {
4121 width,
4122 height,
4123 depth_or_array_layers: 1,
4124 },
4125 mip_level_count: 1,
4126 sample_count: 1,
4127 dimension: wgpu::TextureDimension::D2,
4128 format: GradientTextureLayout::FORMAT,
4129 usage: gpu::RESOURCE_TEXTURE_USAGES,
4130 view_formats: &[],
4131 }
4132}
4133
4134/// The 1x1 stand-ins for the strip shader's not-yet-written bindings.
4135#[derive(Debug)]
4136struct Placeholders {
4137 layer_input: wgpu::TextureView,
4138 atlas_array: wgpu::TextureView,
4139 external: wgpu::TextureView,
4140}
4141
4142impl Placeholders {
4143 fn new(device: &wgpu::Device) -> Self {
4144 Self {
4145 layer_input: placeholder_view(device, "frust-engine layer input placeholder", false),
4146 atlas_array: placeholder_view(device, "frust-engine atlas placeholder", true),
4147 external: placeholder_view(device, "frust-engine external placeholder", false),
4148 }
4149 }
4150}
4151
4152/// A 1x1 transparent `Rgba8Unorm` texture's view, as a plain 2D texture or as
4153/// a 2D array.
4154///
4155/// The array variant allocates **two** layers, not one, matching
4156/// [`gpu::atlas::atlas_texture_descriptor`]'s own floor: wgpu-hal 30.0.1's
4157/// GLES backend derives the GL target from the descriptor's layer count
4158/// alone, and a one-layer array descriptor binds as `GL_TEXTURE_2D` rather
4159/// than `GL_TEXTURE_2D_ARRAY` (see that function's doc comment for the
4160/// file:line and upstream issue refs). Only layer zero of the two is ever
4161/// sampled here.
4162fn placeholder_view(device: &wgpu::Device, label: &str, array: bool) -> wgpu::TextureView {
4163 let texture = device.create_texture(&wgpu::TextureDescriptor {
4164 label: Some(label),
4165 size: wgpu::Extent3d {
4166 width: 1,
4167 height: 1,
4168 depth_or_array_layers: if array { 2 } else { 1 },
4169 },
4170 mip_level_count: 1,
4171 sample_count: 1,
4172 dimension: wgpu::TextureDimension::D2,
4173 format: wgpu::TextureFormat::Rgba8Unorm,
4174 usage: wgpu::TextureUsages::TEXTURE_BINDING,
4175 view_formats: &[],
4176 });
4177 texture.create_view(&wgpu::TextureViewDescriptor {
4178 label: Some(label),
4179 dimension: Some(if array {
4180 wgpu::TextureViewDimension::D2Array
4181 } else {
4182 wgpu::TextureViewDimension::D2
4183 }),
4184 ..Default::default()
4185 })
4186}
4187
4188/// The four bind groups every strip pass sets.
4189///
4190/// The grouping is the shader's, not this module's: coverage plus config plus
4191/// layer input, atlas array plus external texture, encoded paints, gradient
4192/// ramps — four groups exactly, which is the downlevel ceiling with no
4193/// headroom left.
4194#[derive(Debug)]
4195struct StripBindGroups {
4196 resources: wgpu::BindGroup,
4197 images: wgpu::BindGroup,
4198 paints: wgpu::BindGroup,
4199 gradients: wgpu::BindGroup,
4200 /// Group 1 again, once per externally bound texture this variant has
4201 /// drawn: the same atlas array beside that texture's view instead of the
4202 /// placeholder. Keyed by the id a display list names the texture by.
4203 ///
4204 /// A second group rather than a fifth: the four-group ceiling has no
4205 /// headroom (see [`crate::gpu::pipelines`]), so an external texture is
4206 /// bound by re-setting the group the atlas already occupies — which is why
4207 /// a pass's instances are split into runs at all.
4208 externals: HashMap<u64, wgpu::BindGroup>,
4209}
4210
4211impl StripBindGroups {
4212 /// Takes the eight resources one by one rather than a `&FrameResources`
4213 /// so the caller can build a set while holding the map it lands in
4214 /// mutably.
4215 ///
4216 /// `atlas` is the live [`AtlasArray`] view once the renderer has one, or
4217 /// [`Placeholders::atlas_array`] until then — the caller picks, since
4218 /// only it knows which the frame's own `atlas_layers` calls for.
4219 #[expect(
4220 clippy::too_many_arguments,
4221 reason = "one bind-group build's full resource list; a struct would \
4222 only rename the same borrows the caller already holds \
4223 mutably in `FrameResources`"
4224 )]
4225 fn new(
4226 device: &wgpu::Device,
4227 pipeline: &wgpu::RenderPipeline,
4228 alphas: &wgpu::TextureView,
4229 config: &wgpu::Buffer,
4230 placeholders: &Placeholders,
4231 atlas: &wgpu::TextureView,
4232 paints: &wgpu::TextureView,
4233 gradients: &wgpu::TextureView,
4234 ) -> Self {
4235 let resources_group =
4236 resources_bind_group(device, pipeline, alphas, config, &placeholders.layer_input);
4237 let images = images_bind_group(device, pipeline, atlas, &placeholders.external);
4238 let paints = device.create_bind_group(&wgpu::BindGroupDescriptor {
4239 label: Some("frust-engine strip paints"),
4240 layout: &pipeline.get_bind_group_layout(2),
4241 entries: &[wgpu::BindGroupEntry {
4242 binding: 0,
4243 resource: wgpu::BindingResource::TextureView(paints),
4244 }],
4245 });
4246 let gradients = device.create_bind_group(&wgpu::BindGroupDescriptor {
4247 label: Some("frust-engine strip gradients"),
4248 layout: &pipeline.get_bind_group_layout(3),
4249 entries: &[wgpu::BindGroupEntry {
4250 binding: 0,
4251 resource: wgpu::BindingResource::TextureView(gradients),
4252 }],
4253 });
4254
4255 Self {
4256 resources: resources_group,
4257 images,
4258 paints,
4259 gradients,
4260 externals: HashMap::new(),
4261 }
4262 }
4263
4264 /// Sets all four groups on `pass`, taking group 0 from `resources` and
4265 /// group 1 from `images` when a segment names one, rather than from this
4266 /// set.
4267 ///
4268 /// Two of the four vary within a pass. Group 0 carries both the pass's
4269 /// viewport uniform and the layer input a composite samples, so a page
4270 /// round and every composite in it substitute their own. Group 1 carries
4271 /// the external texture a run is drawn with, so a segment that samples one
4272 /// substitutes the group holding it; every other segment takes this set's
4273 /// own, which pairs the atlas array with a placeholder nothing reads.
4274 /// Groups 2-3 are frame-wide and belong to the pipeline variant.
4275 fn bind_with(
4276 &self,
4277 pass: &mut wgpu::RenderPass<'_>,
4278 resources: &wgpu::BindGroup,
4279 images: Option<&wgpu::BindGroup>,
4280 ) {
4281 pass.set_bind_group(0, resources, &[]);
4282 pass.set_bind_group(1, images.unwrap_or(&self.images), &[]);
4283 pass.set_bind_group(2, &self.paints, &[]);
4284 pass.set_bind_group(3, &self.gradients, &[]);
4285 }
4286}
4287
4288/// Group 1 of a strip pass: the image atlas array, and the caller-owned
4289/// texture an external image paint samples.
4290///
4291/// One function for both shapes the group takes — `external` is
4292/// [`Placeholders::external`] for the frame-wide group, or a registered view
4293/// for the group a run of external instances is drawn with (see
4294/// [`crate::gpu::bindings`]). Built against one pipeline for the same reason
4295/// [`resources_bind_group`] is: every engine pipeline uses wgpu's derived
4296/// layout, which is exclusive to the pipeline that derived it.
4297fn images_bind_group(
4298 device: &wgpu::Device,
4299 pipeline: &wgpu::RenderPipeline,
4300 atlas: &wgpu::TextureView,
4301 external: &wgpu::TextureView,
4302) -> wgpu::BindGroup {
4303 device.create_bind_group(&wgpu::BindGroupDescriptor {
4304 label: Some("frust-engine strip images"),
4305 layout: &pipeline.get_bind_group_layout(1),
4306 entries: &[
4307 wgpu::BindGroupEntry {
4308 binding: 0,
4309 resource: wgpu::BindingResource::TextureView(atlas),
4310 },
4311 wgpu::BindGroupEntry {
4312 binding: 1,
4313 resource: wgpu::BindingResource::TextureView(external),
4314 },
4315 ],
4316 })
4317}
4318
4319/// Group 0 of a strip pass: the frame's coverage, the pass's own viewport
4320/// uniform, and the texture a composite reads its finished page from.
4321///
4322/// Built against one pipeline rather than shared, because every engine pipeline
4323/// uses wgpu's *derived* layout — a layout derived from a shader module is
4324/// exclusive to the pipeline that derived it, so a bind group built against one
4325/// variant's layout is rejected by another's even when the two layouts are
4326/// structurally identical.
4327fn resources_bind_group(
4328 device: &wgpu::Device,
4329 pipeline: &wgpu::RenderPipeline,
4330 alphas: &wgpu::TextureView,
4331 config: &wgpu::Buffer,
4332 layer_input: &wgpu::TextureView,
4333) -> wgpu::BindGroup {
4334 device.create_bind_group(&wgpu::BindGroupDescriptor {
4335 label: Some("frust-engine strip resources"),
4336 layout: &pipeline.get_bind_group_layout(0),
4337 entries: &[
4338 wgpu::BindGroupEntry {
4339 binding: 0,
4340 resource: wgpu::BindingResource::TextureView(alphas),
4341 },
4342 wgpu::BindGroupEntry {
4343 binding: 1,
4344 resource: config.as_entire_binding(),
4345 },
4346 wgpu::BindGroupEntry {
4347 binding: 2,
4348 resource: wgpu::BindingResource::TextureView(layer_input),
4349 },
4350 ],
4351 })
4352}
4353
4354#[cfg(test)]
4355mod tests {
4356 use super::*;
4357 use frust_gpu::DownlevelProfile;
4358 use kurbo::Rect;
4359 use peniko::color::palette::css::{BLUE, RED};
4360 use std::collections::BTreeMap;
4361 use vello_common::geometry::RectU16;
4362 use vello_common::paint::IndexedPaint;
4363
4364 /// The `frust-perf enc` line is a contract: a capture is graded by
4365 /// grepping its fields, so the field order and the names are pinned here
4366 /// rather than only by the formatter. `perf-trace`-only, with the line —
4367 /// `cargo test -p frust-engine --features perf-trace` is where it runs.
4368 #[cfg(feature = "perf-trace")]
4369 #[test]
4370 fn an_encode_trace_line_reports_every_column_in_order() {
4371 let mut trace = EncodeTrace {
4372 frames: 7,
4373 ..EncodeTrace::default()
4374 };
4375 // One row per column, each carrying a distinct value in its own
4376 // column, so a transposed or dropped field shows up as a wrong number
4377 // rather than only as a wrong name.
4378 let mut row = [0_u32; ENCODE_TRACE_ROW];
4379 for (column, slot) in row.iter_mut().enumerate() {
4380 *slot = (column as u32 + 1) * 1_000;
4381 }
4382 trace.window.push(row);
4383
4384 let line = trace.line();
4385 let phases: Vec<String> = ENCODE_TRACE_COLUMNS
4386 .iter()
4387 .enumerate()
4388 .map(|(column, name)| format!("{name}_us={}.0", column + 1))
4389 .collect();
4390 // The counts carry no unit suffix and are reported as the raw values
4391 // the row holds, not scaled to microseconds like the phases above.
4392 let counts: Vec<String> = ENCODE_TRACE_COUNTS
4393 .iter()
4394 .enumerate()
4395 .map(|(offset, name)| {
4396 format!(
4397 "{name}={}",
4398 (ENCODE_TRACE_COLUMNS.len() + offset + 1) * 1_000
4399 )
4400 })
4401 .collect();
4402 assert_eq!(
4403 line,
4404 format!(
4405 "frust-perf enc n=7 w=1 {} total_p95_us={}.0 {}",
4406 phases.join(" "),
4407 ENCODE_TRACE_COLUMNS.len(),
4408 counts.join(" "),
4409 ),
4410 );
4411 }
4412
4413 /// A window's percentile is nearest-rank over the column it names, so the
4414 /// value reported is one the window really contains.
4415 #[cfg(feature = "perf-trace")]
4416 #[test]
4417 fn an_encode_trace_percentile_is_nearest_rank_per_column() {
4418 let mut trace = EncodeTrace::default();
4419 for value in [4_000_u32, 1_000, 3_000, 2_000] {
4420 let mut row = [0_u32; ENCODE_TRACE_ROW];
4421 row[0] = value;
4422 trace.window.push(row);
4423 }
4424 assert_eq!(trace.percentile(0, 50), 2_000);
4425 assert_eq!(trace.percentile(0, 95), 4_000);
4426 // An empty window reports zero rather than reaching past its end.
4427 trace.window.clear();
4428 assert_eq!(trace.percentile(0, 50), 0);
4429 }
4430
4431 /// Nanoseconds round up to a tenth of a microsecond, so a phase that cost
4432 /// anything at all never reports as free.
4433 #[cfg(feature = "perf-trace")]
4434 #[test]
4435 fn a_span_that_cost_anything_never_formats_as_zero() {
4436 assert_eq!(format_us(0), "0.0");
4437 assert_eq!(format_us(1), "0.1");
4438 assert_eq!(format_us(4_170_000), "4170.0");
4439 }
4440
4441 /// The compile phases partition the compile they sit inside, so their sum
4442 /// is the whole call minus the overhead the renderer measures around it.
4443 #[test]
4444 fn compile_phases_sum_without_overflowing() {
4445 let spans = CompileSpans {
4446 validate: Duration::from_micros(10),
4447 prepare: Duration::from_micros(20),
4448 classify: Duration::from_micros(30),
4449 admit: Duration::from_micros(40),
4450 walk: Duration::from_micros(50),
4451 // A subset of `walk`, so the sum must not grow by it.
4452 glyphs: Duration::from_micros(45),
4453 finish: Duration::from_micros(60),
4454 };
4455 assert_eq!(spans.total(), Duration::from_micros(210));
4456 assert_eq!(
4457 CompileSpans {
4458 validate: Duration::MAX,
4459 walk: Duration::MAX,
4460 ..CompileSpans::default()
4461 }
4462 .total(),
4463 Duration::MAX,
4464 );
4465 }
4466
4467 #[test]
4468 fn a_target_past_the_u16_grid_is_refused() {
4469 assert_eq!(
4470 grid_size(1920, 1080).expect("a normal surface"),
4471 (1920, 1080)
4472 );
4473 assert_eq!(
4474 grid_size(u32::from(u16::MAX), 16).expect("the grid ceiling itself"),
4475 (u16::MAX, 16)
4476 );
4477 assert!(matches!(
4478 grid_size(u32::from(u16::MAX) + 1, 16),
4479 Err(EngineError::TargetTooLarge)
4480 ));
4481 assert!(matches!(
4482 grid_size(16, u32::from(u16::MAX) + 1),
4483 Err(EngineError::TargetTooLarge)
4484 ));
4485 }
4486
4487 #[test]
4488 fn a_premultiplied_clear_folds_alpha_into_the_colour() {
4489 let half = RED.with_alpha(0.5);
4490 let premultiplied = clear_color(half, OutputAlpha::Premultiplied);
4491 let straight = clear_color(half, OutputAlpha::Straight);
4492
4493 assert!((premultiplied.r - 0.5).abs() < 1e-6);
4494 assert!((premultiplied.a - 0.5).abs() < 1e-6);
4495 assert!((straight.r - 1.0).abs() < 1e-6);
4496 assert!((straight.a - 0.5).abs() < 1e-6);
4497
4498 // An opaque colour is identical under both conventions.
4499 assert_eq!(
4500 clear_color(BLUE, OutputAlpha::Premultiplied),
4501 clear_color(BLUE, OutputAlpha::Straight)
4502 );
4503 }
4504
4505 #[test]
4506 fn a_solid_paint_travels_premultiplied_in_the_instance_payload() {
4507 let paint = pack_paint(&Paint::from(RED), 3, &[]).expect("a solid paint always resolves");
4508 let values = paint.values_at(40, 12);
4509 assert_eq!(values.paint, SOLID_PAINT);
4510 assert_eq!(values.depth_index, 3);
4511 assert_eq!(values.payload, RED.premultiply().to_rgba8().to_u32());
4512 assert!(paint.opaque);
4513 assert_eq!(
4514 paint.values_at(0, 0).payload,
4515 values.payload,
4516 "a solid colour is the same wherever it is stamped"
4517 );
4518
4519 let paint = pack_paint(&Paint::from(RED.with_alpha(0.5)), 0, &[])
4520 .expect("a translucent solid paint still resolves");
4521 assert!(
4522 !paint.opaque,
4523 "a translucent paint never reaches the opaque pass"
4524 );
4525 }
4526
4527 #[test]
4528 fn an_unresolvable_indexed_paint_skips_its_draw() {
4529 let indexed = Paint::Indexed(IndexedPaint::new(0));
4530 assert!(
4531 pack_paint(&indexed, 0, &[]).is_none(),
4532 "an index past the frame's slots resolves to nothing"
4533 );
4534 assert!(
4535 pack_paint(&indexed, 0, &[None]).is_none(),
4536 "so does a slot the lowering refused"
4537 );
4538 }
4539
4540 #[test]
4541 fn a_resolved_indexed_paint_samples_at_each_instances_own_position() {
4542 let slots = [Some(ResolvedPaint {
4543 paint_type: PaintType::LinearGradient,
4544 texel_offset: 6,
4545 opaque: true,
4546 external: None,
4547 })];
4548 let paint = pack_paint(&Paint::Indexed(IndexedPaint::new(0)), 2, &slots)
4549 .expect("a lowered paint resolves");
4550
4551 assert_eq!(
4552 paint.paint,
4553 pack_paint_descriptor(PaintType::LinearGradient, 6)
4554 );
4555 assert_eq!(paint.depth_index, 2);
4556 // The payload is the instance's scene origin, not a colour, so two
4557 // instances of the same draw carry different payloads.
4558 assert_eq!(paint.payload_at(8, 4), 8 | (4 << 16));
4559 assert_eq!(paint.payload_at(40, 4), 40 | (4 << 16));
4560 }
4561
4562 #[test]
4563 fn a_u16_pair_packs_low_half_first() {
4564 assert_eq!(pack_u16_pair(0, 0), 0);
4565 assert_eq!(pack_u16_pair(1, 0), 1);
4566 assert_eq!(pack_u16_pair(0, 1), 1 << 16);
4567 assert_eq!(pack_u16_pair(u16::MAX, u16::MAX), u32::MAX);
4568 }
4569
4570 #[test]
4571 fn the_minimum_resource_dimension_keeps_every_upload_row_legal() {
4572 // The gradient LUT is the narrowest resource at four bytes per texel,
4573 // so it is the one that fixes the floor.
4574 assert_eq!(
4575 MIN_RESOURCE_TEXTURE_DIM * BYTES_PER_TEXEL,
4576 wgpu::COPY_BYTES_PER_ROW_ALIGNMENT
4577 );
4578 assert!(MIN_RESOURCE_TEXTURE_DIM.is_power_of_two());
4579 }
4580
4581 // -----------------------------------------------------------------
4582 // Where the hole punch lands in the pass plan
4583 //
4584 // What a frame's pass plan COSTS, and where the erase sits inside it,
4585 // is a decision over the recording's shape — no device, no pixels.
4586 // The pixel half of the same contract is
4587 // `frust-testing`'s `tests/aa_over_punch.rs`, against `vello_cpu`.
4588 // -----------------------------------------------------------------
4589
4590 /// Viewport every plan below is built against.
4591 const PLAN_VIEWPORT: (u16, u16) = (64, 48);
4592
4593 /// The slot, and a chip straddling its right edge — so a chip drawn over
4594 /// the punch is half inside it and half over the backdrop.
4595 const PLAN_SLOT: Rect = Rect::new(4.0, 4.0, 32.0, 44.0);
4596 const PLAN_CHIP: Rect = Rect::new(16.0, 12.0, 56.0, 32.0);
4597
4598 /// The plan `Scratch::build` produces for `scene`, alongside the rounds it
4599 /// was scheduled from, under a target that carries alpha.
4600 fn plan_of(scene: &Scene) -> (Scratch, Vec<Round>) {
4601 let frame = SceneCompiler::new(PLAN_VIEWPORT.0, PLAN_VIEWPORT.1)
4602 .compile(scene, Affine::IDENTITY, PLAN_VIEWPORT)
4603 .expect("an in-range scene compiles");
4604 let rounds = Schedule::build(
4605 &frame.recorder,
4606 &TierCaps::fake(DownlevelProfile::Full),
4607 &PageConfig::default(),
4608 )
4609 .expect("a scene of solid fills schedules");
4610
4611 let mut scratch = Scratch::default();
4612 // Depth on and punching on: the shape a translucent presentation
4613 // carrying a platform-view slot is planned under.
4614 scratch.build(&frame, &rounds, true, !frame.clears.is_empty(), &[]);
4615 (scratch, rounds)
4616 }
4617
4618 /// A scene recorded through the public builder.
4619 fn plan_scene(record: impl FnOnce(&mut SceneBuilder<'_>)) -> Scene {
4620 let mut scene = Scene::new();
4621 let mut builder = SceneBuilder::new(&mut scene);
4622 record(&mut builder);
4623 scene
4624 }
4625
4626 fn plan_backdrop(builder: &mut SceneBuilder<'_>) {
4627 builder.fill_rect(
4628 Rect::new(
4629 0.0,
4630 0.0,
4631 f64::from(PLAN_VIEWPORT.0),
4632 f64::from(PLAN_VIEWPORT.1),
4633 ),
4634 Brush::Solid(RED),
4635 );
4636 }
4637
4638 #[test]
4639 fn a_frame_that_punches_nothing_is_planned_exactly_as_its_rounds() {
4640 let (scratch, rounds) = plan_of(&plan_scene(|builder| {
4641 plan_backdrop(builder);
4642 builder.fill_rect(PLAN_CHIP, Brush::Solid(BLUE));
4643 }));
4644
4645 assert_eq!(
4646 scratch.rounds.len(),
4647 rounds.len(),
4648 "a frame with no clear is cut nowhere, so it costs exactly its \
4649 scheduled rounds' passes"
4650 );
4651 assert!(!scratch.punches(), "and records no punch pass at all");
4652 }
4653
4654 #[test]
4655 fn a_clear_recorded_last_keeps_its_pass_after_every_round() {
4656 let (scratch, rounds) = plan_of(&plan_scene(|builder| {
4657 plan_backdrop(builder);
4658 builder.clear_rect(PLAN_SLOT);
4659 }));
4660
4661 assert_eq!(
4662 scratch.rounds.len(),
4663 rounds.len(),
4664 "nothing is recorded after the clear, so nothing is cut"
4665 );
4666 assert!(
4667 scratch.rounds.iter().all(|plan| plan.punch.1 == 0),
4668 "no round carries the punch"
4669 );
4670 assert!(
4671 scratch.punch.1 > 0,
4672 "it is issued after the last round instead — the position an \
4673 unconditionally-trailing pass would have put it in, which is why \
4674 a frame shaped like this renders byte-identically"
4675 );
4676 }
4677
4678 #[test]
4679 fn a_draw_over_a_clear_cuts_the_surface_round_and_the_punch_goes_in_the_cut() {
4680 let (scratch, rounds) = plan_of(&plan_scene(|builder| {
4681 plan_backdrop(builder);
4682 builder.clear_rect(PLAN_SLOT);
4683 builder.fill_rect(PLAN_CHIP, Brush::Solid(BLUE));
4684 }));
4685
4686 assert_eq!(
4687 scratch.rounds.len(),
4688 rounds.len() + 1,
4689 "the surface round is cut in two — one extra pass, and one only"
4690 );
4691 assert_eq!(
4692 scratch.punch,
4693 (0, 0),
4694 "with nothing left over for a trailing pass"
4695 );
4696
4697 let cut = scratch
4698 .rounds
4699 .iter()
4700 .position(|plan| plan.punch.1 > 0)
4701 .expect("the cut plan carries the punch");
4702 assert_eq!(cut, scratch.rounds.len() - 2, "and it is the cut plan");
4703
4704 // The buffer says the same thing the plan does: the backdrop's
4705 // instances precede the punch's, and the chip's follow them.
4706 let before = scratch.rounds[cut].segments.clone();
4707 let after = scratch.rounds[cut + 1].segments.clone();
4708 let end_of = |range: Range<usize>| {
4709 scratch.segments[range]
4710 .iter()
4711 .map(|segment| match *segment {
4712 Segment::Strips(first, count) | Segment::External(first, count, _) => {
4713 first + count
4714 }
4715 Segment::Composite(first, _) => first + 1,
4716 })
4717 .max()
4718 .expect("a plan of this frame draws something")
4719 };
4720 let (punch_first, punch_count) = scratch.rounds[cut].punch;
4721 assert!(
4722 end_of(before) <= punch_first,
4723 "everything recorded before the clear is erased by the punch"
4724 );
4725 assert!(
4726 scratch.segments[after]
4727 .iter()
4728 .all(|segment| match *segment {
4729 Segment::Strips(first, _)
4730 | Segment::External(first, _, _)
4731 | Segment::Composite(first, _) => first >= punch_first + punch_count,
4732 }),
4733 "and everything recorded after it lands on top of the erase"
4734 );
4735 }
4736
4737 // -----------------------------------------------------------------
4738 // A banded layer's column pages hold their own column, and only it
4739 //
4740 // The scheduler hands every band of a layer the same op list, so the
4741 // instances a band emits are where a column split is made or lost. Every
4742 // case here is host-only: the plan is built with no device, and the pixel
4743 // half of the same contract is `tests/desktop_stress.rs`'s 5K cases.
4744 // -----------------------------------------------------------------
4745
4746 /// Viewport the band plans below are built against: wide enough that one
4747 /// layer over all of it needs several column pages under [`BAND_PAGES`],
4748 /// and one tile row tall, so a draw costs one strip row per column.
4749 const BAND_VIEWPORT: (u16, u16) = (200, 8);
4750
4751 /// A page ceiling small enough to band a test-sized layer, so no case here
4752 /// has to allocate — or even name — a 5K one.
4753 const BAND_PAGES: PageConfig = PageConfig {
4754 min_page_size: 64,
4755 max_page_size: 64,
4756 };
4757
4758 /// The plan `scene` produces under `config`, alongside the rounds it was
4759 /// scheduled into.
4760 ///
4761 /// Depth is off, so every instance of the frame lands in the one blended
4762 /// buffer in painter order and a case can read the whole plan out of it;
4763 /// the split into the depth-writing pass is a surface-round decision and a
4764 /// page round never takes it.
4765 fn band_plan_of(scene: &Scene, config: &PageConfig) -> (Scratch, Vec<Round>) {
4766 let frame = SceneCompiler::new(BAND_VIEWPORT.0, BAND_VIEWPORT.1)
4767 .compile(scene, Affine::IDENTITY, BAND_VIEWPORT)
4768 .expect("an in-range scene compiles");
4769 let rounds = Schedule::build(
4770 &frame.recorder,
4771 &TierCaps::fake(DownlevelProfile::Full),
4772 config,
4773 )
4774 .expect("a layer wider than the ceiling bands rather than refusing");
4775
4776 let mut scratch = Scratch::default();
4777 scratch.build(&frame, &rounds, false, false, &[]);
4778 assert_eq!(
4779 scratch.rounds.len(),
4780 rounds.len(),
4781 "a frame with no clear is cut nowhere, so the plans and the rounds \
4782 line up one for one"
4783 );
4784 (scratch, rounds)
4785 }
4786
4787 /// A layer at half opacity — so it cannot be inlined and has to take a page
4788 /// — holding one solid rectangle per entry of `rects`.
4789 fn band_scene(rects: &[(Rect, Color)]) -> Scene {
4790 let mut scene = Scene::new();
4791 let mut builder = SceneBuilder::new(&mut scene);
4792 builder.push_layer(
4793 Rect::new(
4794 0.0,
4795 0.0,
4796 f64::from(BAND_VIEWPORT.0),
4797 f64::from(BAND_VIEWPORT.1),
4798 ),
4799 0.5,
4800 );
4801 for (rect, color) in rects {
4802 builder.fill_rect(*rect, Brush::Solid(*color));
4803 }
4804 builder.pop_layer();
4805 scene
4806 }
4807
4808 /// The bounds of every page round, in order — one band's column each.
4809 fn band_bounds(rounds: &[Round]) -> Vec<RectU16> {
4810 rounds
4811 .iter()
4812 .filter_map(|round| round.page().map(|page| page.bounds))
4813 .collect()
4814 }
4815
4816 /// The strip instances one round plan issues, in execution order.
4817 fn plan_strips(scratch: &Scratch, plan: &RoundPlan) -> Vec<GpuStrip> {
4818 scratch.segments[plan.segments.clone()]
4819 .iter()
4820 .filter_map(|segment| match *segment {
4821 Segment::Strips(first, count) | Segment::External(first, count, _) => {
4822 Some((first as usize, count as usize))
4823 }
4824 Segment::Composite(..) => None,
4825 })
4826 .flat_map(|(first, count)| scratch.alpha[first..first + count].iter().copied())
4827 .collect()
4828 }
4829
4830 /// Every pixel column the frame's *page* rounds paint, mapped back out of
4831 /// the pages they were shifted into, and valued by what the shader reads
4832 /// there: the instance's paint payload, plus the exact alpha column that
4833 /// pixel samples (`None` for a sparse instance, which samples none).
4834 ///
4835 /// Keyed by the draw's own depth as well as the position, so a page painted
4836 /// by two different draws is not conflated — and so a *second* instance of
4837 /// one draw covering a pixel it already covered, which is precisely what a
4838 /// clamped band replay produces, is caught here rather than silently
4839 /// overwriting the first.
4840 fn painted_pixels(
4841 scratch: &Scratch,
4842 rounds: &[Round],
4843 ) -> BTreeMap<(u32, u16, u16), (u32, Option<u32>)> {
4844 let mut painted = BTreeMap::new();
4845
4846 for (plan, round) in scratch.rounds.iter().zip(rounds) {
4847 let Some(page) = round.page() else {
4848 continue;
4849 };
4850 for span in plan_strips(scratch, plan) {
4851 for offset in 0..span.width {
4852 let x = page.bounds.x0 + span.x + offset;
4853 let y = page.bounds.y0 + span.y;
4854 let column = (offset < span.dense_width_or_rect_height)
4855 .then(|| span.col_idx_or_rect_frac + u32::from(offset));
4856 assert!(
4857 painted
4858 .insert((span.depth_index, y, x), (span.payload, column))
4859 .is_none(),
4860 "one draw covers a device pixel at most once, however the \
4861 layer holding it was split"
4862 );
4863 }
4864 }
4865 }
4866
4867 painted
4868 }
4869
4870 #[test]
4871 fn a_banded_layer_paints_exactly_what_one_page_would_have() {
4872 // The whole point of the split, asserted as an equality rather than as
4873 // a rectangle property: the same scene planned onto one page and onto
4874 // column bands has to paint the same device pixels, from the same
4875 // paints, sampling the same alpha columns. A band replay that clamped
4876 // a strip left of its own column onto the page's edge fails here twice
4877 // over — once on the ghost pixel, once on the column it would sample.
4878 let scene = band_scene(&[
4879 (Rect::new(0.0, 0.0, 40.0, 8.0), RED),
4880 // Straddles a band edge on a tile that is only partly covered, so
4881 // the case exercises an alpha-sampled instance cut in two, not
4882 // just a solid one.
4883 (Rect::new(49.0, 0.0, 99.0, 8.0), BLUE),
4884 (Rect::new(160.0, 0.0, 200.0, 8.0), RED),
4885 ]);
4886
4887 let (one_page, one_page_rounds) = band_plan_of(&scene, &PageConfig::default());
4888 let (banded, banded_rounds) = band_plan_of(&scene, &BAND_PAGES);
4889
4890 assert_eq!(
4891 band_bounds(&one_page_rounds).len(),
4892 1,
4893 "the reference plan really does hold the layer on one page"
4894 );
4895 assert!(
4896 band_bounds(&banded_rounds).len() > 1,
4897 "and the case only means anything while the other one is banded"
4898 );
4899
4900 assert_eq!(
4901 painted_pixels(&banded, &banded_rounds),
4902 painted_pixels(&one_page, &one_page_rounds),
4903 "a banded layer paints what one whole-layer page would have"
4904 );
4905 }
4906
4907 #[test]
4908 fn a_band_holds_no_instance_of_a_strip_outside_its_own_column() {
4909 // The counterexample the equality above generalizes: content confined
4910 // to the outer columns, and nothing at all in the middle. A band whose
4911 // column the scene never drew in must render nothing — under a
4912 // saturating shift it would render the leftmost content clamped onto
4913 // its own edge instead.
4914 let scene = band_scene(&[
4915 (Rect::new(0.0, 0.0, 40.0, 8.0), RED),
4916 (Rect::new(160.0, 0.0, 200.0, 8.0), BLUE),
4917 ]);
4918 let (scratch, rounds) = band_plan_of(&scene, &BAND_PAGES);
4919 let bands = band_bounds(&rounds);
4920 assert!(bands.len() > 2, "the layer bands: {bands:?}");
4921
4922 let mut empty = 0_usize;
4923 for (plan, band) in scratch
4924 .rounds
4925 .iter()
4926 .zip(&rounds)
4927 .filter_map(|(plan, round)| round.page().map(|page| (plan, page.bounds)))
4928 {
4929 let strips = plan_strips(&scratch, plan);
4930 if strips.is_empty() {
4931 empty += 1;
4932 }
4933 for span in strips {
4934 let end = span.x + span.width;
4935 assert!(
4936 end <= band.width(),
4937 "an instance of {span:?} reaches past the {band:?} band's own \
4938 column, which its composite never samples"
4939 );
4940 // Mapped back to the scene, every instance lands where one of
4941 // the two rectangles was actually drawn.
4942 let x = band.x0 + span.x;
4943 assert!(
4944 x < 44 || band.x0 + end > 160,
4945 "an instance covers device x {x}..{}, which neither \
4946 rectangle reaches",
4947 band.x0 + end
4948 );
4949 }
4950 }
4951
4952 assert!(
4953 empty > 0,
4954 "a band whose column holds nothing renders nothing: {bands:?}"
4955 );
4956 }
4957
4958 #[test]
4959 fn a_strip_straddling_a_band_edge_advances_its_alpha_column_with_its_geometry() {
4960 // A pixel-aligned rectangle's left edge lands on a tile of its own, so
4961 // a rectangle starting at 49 puts an alpha-sampled instance across
4962 // 48..52 — and a band edge falls inside it. The two pieces together
4963 // have to read the same coverage the whole instance would: neighbouring
4964 // pixels of one strip row sampling neighbouring alpha columns, with no
4965 // column repeated. The rectangles either side of it are what carries
4966 // the layer past the ceiling, so the middle one is banded at all.
4967 let scene = band_scene(&[
4968 (Rect::new(0.0, 0.0, 40.0, 8.0), RED),
4969 (Rect::new(49.0, 0.0, 99.0, 8.0), BLUE),
4970 (Rect::new(160.0, 0.0, 200.0, 8.0), RED),
4971 ]);
4972 let (scratch, rounds) = band_plan_of(&scene, &BAND_PAGES);
4973 let edges: Vec<u16> = band_bounds(&rounds)
4974 .iter()
4975 .map(|bounds| bounds.x0)
4976 .collect();
4977
4978 // Every alpha-sampled pixel of every band, as (strip row, device x,
4979 // the alpha column it samples) — a row at a time, because two rows of
4980 // one draw sample different columns at the same x by construction.
4981 let mut dense: Vec<(u16, u16, u32)> = Vec::new();
4982 for (plan, band) in scratch
4983 .rounds
4984 .iter()
4985 .zip(&rounds)
4986 .filter_map(|(plan, round)| round.page().map(|page| (plan, page.bounds)))
4987 {
4988 for span in plan_strips(&scratch, plan) {
4989 for offset in 0..span.dense_width_or_rect_height {
4990 dense.push((
4991 band.y0 + span.y,
4992 band.x0 + span.x + offset,
4993 span.col_idx_or_rect_frac + u32::from(offset),
4994 ));
4995 }
4996 }
4997 }
4998 dense.sort_unstable();
4999
5000 let neighbours = || {
5001 dense
5002 .windows(2)
5003 .map(|pair| (pair[0], pair[1]))
5004 .filter(|((row, x, _), (next_row, next_x, _))| row == next_row && x + 1 == *next_x)
5005 };
5006 assert!(
5007 neighbours().any(|(_, (_, x, _))| edges.contains(&x)),
5008 "the case only means anything while an alpha-sampled instance \
5009 really is cut by a band edge {edges:?}: {dense:?}"
5010 );
5011
5012 for ((_, _, column), (_, x, next_column)) in neighbours() {
5013 assert_eq!(
5014 next_column,
5015 column + 1,
5016 "neighbouring pixels of one strip row sample neighbouring alpha \
5017 columns, band edge at {x} or not: {dense:?}"
5018 );
5019 }
5020 }
5021
5022 #[test]
5023 fn a_translucent_layer_over_a_clear_is_composited_after_the_punch() {
5024 let (scratch, rounds) = plan_of(&plan_scene(|builder| {
5025 plan_backdrop(builder);
5026 builder.clear_rect(PLAN_SLOT);
5027 builder.push_layer(PLAN_CHIP, 0.5);
5028 builder.fill_rect(PLAN_CHIP, Brush::Solid(BLUE));
5029 builder.pop_layer();
5030 }));
5031
5032 assert_eq!(
5033 scratch.rounds.len(),
5034 rounds.len() + 1,
5035 "the layer's page round, then the surface round cut in two"
5036 );
5037 let cut = scratch
5038 .rounds
5039 .iter()
5040 .position(|plan| plan.punch.1 > 0)
5041 .expect("the cut plan carries the punch");
5042 // The composite is what the cut has to fall before: a layer recorded
5043 // after the clear carries a deeper index than the punch, and the
5044 // composite writes no depth for the test to save it by.
5045 let composited = scratch.segments[scratch.rounds[cut + 1].segments.clone()]
5046 .iter()
5047 .any(|segment| matches!(segment, Segment::Composite(..)));
5048 assert!(
5049 composited,
5050 "the layer composites in the plan AFTER the punch, not before it"
5051 );
5052 assert!(
5053 !scratch.segments[scratch.rounds[cut].segments.clone()]
5054 .iter()
5055 .any(|segment| matches!(segment, Segment::Composite(..))),
5056 "and nothing composites into the plan the punch closes"
5057 );
5058 }
5059}