#[non_exhaustive]#[repr(u8)]pub enum TensorEvent {
Load0 = 0,
Load1 = 1,
LoadL2_0 = 2,
LoadL2_1 = 3,
Prefetch0 = 4,
Prefetch1 = 5,
CacheOp = 6,
Fma = 7,
Store = 8,
TensorReduce = 9,
TensorQuant = 10,
}Expand description
Tensor co-processor synchronisation events for tensor_wait.
The four-bit EVENT field in the TensorWait xs register selects which
outstanding operation the hart waits for before the instruction retires.
(PRM Table 9-2.)
This enum is #[non_exhaustive]: match arms outside this crate must
include a wildcard arm.
Variants (Non-exhaustive)§
This enum is marked as non-exhaustive
Load0 = 0
Completion of all TensorLoad operations issued with ID = 0 (event 0).
Load1 = 1
Completion of all TensorLoad operations issued with ID = 1 (event 1).
LoadL2_0 = 2
Completion of a TensorLoadL2Scp issued with ID = 0 (event 2).
Use this after tensor_load_l2 with id = false.
Not the same as CacheOp (event 6).
LoadL2_1 = 3
Completion of a TensorLoadL2Scp issued with ID = 1 (event 3).
Use this after tensor_load_l2 with id = true.
Prefetch0 = 4
Completion of L2/L3 prefetch operations with ID = 0 (event 4).
Prefetch1 = 5
Completion of L2/L3 prefetch operations with ID = 1 (event 5).
CacheOp = 6
Completion of all preceding L1 cache management operations: EvictVA
and FlushVA (event 6). Required after [cache::cache_writeback],
[cache::cache_invalidate], or [cache::cache_flush] before issuing
memory accesses to the affected cache lines.
Note: TensorLoadL2Scp requires LoadL2_0/LoadL2_1 (events 2/3),
not this event. L2/L3 prefetch requires Prefetch0/Prefetch1 (events
4/5). The cache op functions already issue this wait internally; use this
variant directly only when batching cache ops and deferring the wait.
Fma = 7
Completion of all preceding TensorFMA operations (event 7). The FP register file holds the final accumulated C tile and may be read or stored.
Store = 8
Completion of all preceding TensorStore DMA transfers (event 8).
Drains only the tensor store DMA; prefer this over a full
fence rw, rw when only tensor-store ordering is required.
TensorReduce = 9
Completion of all preceding TensorSend/TensorRecv operations (event 9).
Required after tensor_recv before reading the FP registers updated
by the receive.
TensorQuant = 10
Completion of all preceding TensorQuant operations (event 10).
Trait Implementations§
Source§impl Clone for TensorEvent
impl Clone for TensorEvent
Source§fn clone(&self) -> TensorEvent
fn clone(&self) -> TensorEvent
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read more