pub trait SyncWaiter: MaybeSendSync + 'static {
// Required methods
fn wait(&self, timeout: Duration) -> Result<(), Error>;
fn backend(&self) -> BackendKind;
fn as_any(&self) -> &dyn Any;
// Provided methods
fn wait_async<'a>(
&'a self,
timeout: Duration,
) -> BoxFuture<'a, Result<(), Error>> { ... }
fn is_signaled(&self) -> Result<bool, Error> { ... }
fn as_cuda_event_waiter(&self) -> Option<&dyn CudaEventWaiter> { ... }
fn rebind_to_value(&self, value: u64) -> Option<Arc<dyn SyncWaiter>> { ... }
}Expand description
Backend-specific wait dispatch + lifetime anchor for a SyncPoint.
Every non-trivial SyncPoint variant carries an Arc<dyn SyncWaiter>
that:
-
Owns the underlying primitive’s lifetime — the timeline semaphore, D3D12 fence,
MTLSharedEvent,cl_event,GLsync, etc. Cloning theSyncPointclones theArc; the primitive lives as long as any clone outlives. This rules out the “consumer holds a*mut c_voidafter the producer dropped” footgun. -
Provides the wait implementation —
SyncPoint::wait/is_signaled/backenddispatch through this trait, with no global registries and noOnceLock<fn>runtime callbacks.
Implementations live in the producer crate (typically an interop
layer’s per-backend module) and are constructed at the same site that
mints the SyncPoint.
Required Methods§
Sourcefn wait(&self, timeout: Duration) -> Result<(), Error>
fn wait(&self, timeout: Duration) -> Result<(), Error>
Block until the GPU work this sync point represents has
completed, or until timeout elapses.
Duration::MAX means “wait forever” — implementations should
translate this to the platform’s native “infinite” sentinel
(UINT64_MAX for vkWaitSemaphores, INFINITE for
WaitForSingleObject, etc.).
§Timeout granularity
Native wait APIs vary in granularity, and implementations preserve the caller’s intent at the expense of requested (not realised) latency on sub-API-tick timeouts:
- Metal (
MTLSharedEvent::waitUntilSignaledValue:timeoutMS:) accepts millisecond-granularu64. Sub-millisecond positive timeouts (e.g.Duration::from_micros(100)) round up to 1 ms — truncation to 0 would silently behave as “do not wait”. A high-frequency progress poll usingDuration::from_micros(100)therefore observes ~10× the requested latency on a not-yet-signaled event. For non-blocking probes useSyncWaiter::is_signaled(which passesDuration::ZEROand short-circuits on thesignaledValue() >= valueaccessor — driver-side, no timer). - Win32 (
WaitForSingleObject) is also millisecond- granular; the same round-up applies. - Vulkan (
vkWaitSemaphores), D3D12 fence, CUDA external semaphore, OpenCL are nanosecond-granular and honour the caller’sDurationexactly.
Callers that need both a deterministic poll cadence under
1 ms and a real wait fallback should compose the two —
e.g. is_signaled in a hot loop with their own Instant
budget, then a single coarse wait(Duration::from_millis(N))
when the budget is exhausted.
Sourcefn backend(&self) -> BackendKind
fn backend(&self) -> BackendKind
Backend identity for routing decisions on the consumer side.
Sourcefn as_any(&self) -> &dyn Any
fn as_any(&self) -> &dyn Any
Downcast hook for cross-API bridges that need access to the
concrete waiter type — e.g. a Vulkan→CUDA bridge that wants to
pull the VkDevice out of the VulkanWaiter to issue its own
cuImportExternalSemaphore. Most callers use the per-variant
raw fields on SyncPoint directly and never need this.
Provided Methods§
Sourcefn wait_async<'a>(
&'a self,
timeout: Duration,
) -> BoxFuture<'a, Result<(), Error>>
fn wait_async<'a>( &'a self, timeout: Duration, ) -> BoxFuture<'a, Result<(), Error>>
Async sibling of wait. Returns a boxed future so
the trait stays object-safe through Arc<dyn SyncWaiter> — native
async-fn-in-trait is RPITIT and dyn-incompatible by default; callers
dispatch through Arc<dyn SyncWaiter> everywhere, so the explicit
Pin<Box<...>> is mandatory.
§Default implementation
Backends without a hand-tuned implementation get a two-phase
hybrid: a short cooperative-yield spin (catches sub-ms waits
without burning a thread) followed by an iteration-bounded
is_signaled probe loop. Per-backend impls SHOULD override with
their native blocking-with-timeout primitive on a dedicated
waiter thread for waits the spin loop didn’t catch.
The default implementation is correct (resolves when signalled,
returns Error::Timeout after timeout, never blocks the
executor) but not optimal: it busy-polls inside the spin loop
and falls back to a Duration::ZERO probe loop after that. For
production paths, override with a backend-specific impl.
Sourcefn is_signaled(&self) -> Result<bool, Error>
fn is_signaled(&self) -> Result<bool, Error>
Non-blocking probe.
Ok(true)— the sync point has been reached.Ok(false)— work is still in flight (the underlying poll timed out at zero).Err(_)— driver-level failure (device lost / TDR /wgpu::PollError::WrongSubmissionIndex/ equivalent on other backends). Callers polling in a loop must break onErr— folding errors intoOk(false)would loop forever on a TDR’d device.
Default implementation calls wait(Duration::ZERO) and maps
Err(Error::Timeout) → Ok(false); every other Err is
propagated. Backends override when they have a more direct
“is this primitive currently signaled” probe (e.g.
vkGetSemaphoreCounterValue, ID3D12Fence::GetCompletedValue)
that distinguishes “not signaled” from real errors without
going through the wait path.
Sourcefn as_cuda_event_waiter(&self) -> Option<&dyn CudaEventWaiter>
fn as_cuda_event_waiter(&self) -> Option<&dyn CudaEventWaiter>
CudaEventWaiter view of this waiter, when the concrete type
implements it. Default None.
This is the trait-object route to the CUDA-only
CudaEventWaiter::wait_on_foreign_stream extension: Any can
only downcast to concrete types, which consumers in other
crates cannot name — so producers whose waiter implements
CudaEventWaiter override this with Some(self) and
consumers (e.g. a CUDA import path gating its private copy stream
on a producer’s SyncPoint::CudaEvent) reach the extension method
without knowing the concrete type.
Sourcefn rebind_to_value(&self, value: u64) -> Option<Arc<dyn SyncWaiter>>
fn rebind_to_value(&self, value: u64) -> Option<Arc<dyn SyncWaiter>>
Rebuild this waiter bound to value instead of the value it was
constructed with, reusing the same underlying primitive (and the
same lifetime anchors).
§Why this exists
wait / is_signaled take no
value argument — the waiter embeds the value it resolves at,
captured when it was built. So a SyncPoint whose value
field was substituted while its waiter was cloned verbatim has
two surfaces that disagree: the GPU side (a consumer reading
SyncPoint::*.value to stage a queue wait) waits the new value,
while the CPU side silently resolves at the old one — reporting
“already signalled” for a signal the producer has not emitted.
Any code that mints a SyncPoint at a value other than the one a
template was built with MUST route the waiter through this method
rather than cloning it, so the two surfaces cannot diverge.
§Contract
Some(w)—wwaitsvalueon the same primitive this waiter waits on, and holds the same keep-alive chain. Implementations MUST NOT carry over state that is only valid for the original value (e.g. a captured submission index that retires at it).None(the default) — the underlying primitive carries no value, so there is nothing to rebind: binary semaphores, fences,GLsync,cl_semaphore_khr. ANonehere is not a failure; it means the CPU surface has no value to disagree about.
Dyn Compatibility§
This trait is dyn compatible.
In older versions of Rust, dyn compatibility was called "object safety".