Expand description
The per-frame bump arena every small GPU upload goes through:
HostBuffer, the BufferSlice it hands back, and the
BufferUploader seam that keeps both host-testable.
A frame produces many small pieces of data a shader has to read —
per-draw uniforms, a handful of vertices, a transform block. Giving each
one its own wgpu::Buffer means an allocation, a bind group and a
separate upload per draw. The arena replaces that with one buffer: every
HostBuffer::alloc appends to a CPU-side staging vector and returns
the (offset, size) slice the caller will bind at, one
HostBuffer::flush uploads the whole frame’s bytes with a single
queue.write_buffer, and HostBuffer::reset rewinds the bump pointer
at the start of the next frame.
§Alignment
Every slice is aligned to at least the adapter’s
TierCaps::min_uniform_buffer_offset_alignment, so any slice is
legal as a uniform binding offset — including a dynamic one. That value
is 256 under the GLES-3.0/WebGL2 profile and on the iOS Simulator (which
misreports its own alignment; crate::context forces the device
request back up to 256 there), which is the strictest target this
workspace has, so an arena built against those caps satisfies every
looser adapter as well. A caller may ask for more alignment for a
specific slice — a vertex stride, say — and never gets less.
§Grow-only, with a high-water mark
The GPU buffer is created on the first flush that has bytes and then
reused. It grows geometrically when a frame outgrows it and never
shrinks, so the steady state after a few frames is zero buffer
allocations per frame; HostBuffer::high_water_mark reports the
largest frame seen, which is what a host tunes an initial size against.
§Why there is no 4-deep ring
A hand-rolled Vulkan/Metal arena double- or quadruple-buffers, because the CPU writes into persistently mapped memory the GPU may still be reading from an in-flight frame; the ring is what keeps this frame’s writes off last frame’s bytes.
wgpu::Queue::write_buffer has no such hazard to guard. The bytes are
copied out of the caller’s slice into a queue-owned staging allocation at
call time, and the actual device-side copy is recorded ahead of the
command buffers submitted after it, in queue order — so a write issued
for frame N lands after frame N-1’s commands have already run, and
wgpu tracks and reclaims the staging memory itself once the GPU is done
with it. The caller never holds a pointer into memory the GPU is reading,
which is the only thing a ring exists to arrange. Growing is safe for the
same reason: wgpu refcounts a replaced buffer until every submission
referencing it has retired.
What the caller must still respect is ordering within its own frame:
rewind with HostBuffer::reset at frame start, allocate, flush once
before submitting the frame’s commands. Rewinding a frame whose commands
are already recorded but not yet flushed would overwrite the bytes those
commands are going to read.
Structs§
- Buffer
Slice - Where a
HostBuffer::alloccall’s bytes ended up: a byte range within the arena’s single GPU buffer. - Host
Buffer - One growable GPU buffer plus the CPU-side staging vector a frame bump- allocates into.
Constants§
- ARENA_
USAGE - The usages every arena buffer is created with: the two binding kinds a
frame’s transient data is read as, plus the copy destination
write_bufferneeds. - MIN_
ARENA_ CAPACITY - The smallest GPU buffer the arena ever creates. A first frame with one 64-byte uniform in it should not lead to a second allocation on the second frame.
Traits§
- Buffer
Uploader - Creates and writes the arena’s GPU buffer.