Expand description
Collectives over logical warps within a native device subgroup.
All lanes of a logical warp participate in each call. width is positive,
does not exceed the native subgroup width, and groups must not straddle a
native subgroup. No NVIDIA warp width is assumed by these algorithms.
Modules§
- batched
- broadcast
- exchange
- exclusive_
scan - exclusive_
unseeded - head_
segmented_ reduce - inclusive_
scan - io
- logical_
lane_ id - logical_
warp_ base_ id - logical_
warp_ id - record
- reduce
- scan
- sort
- tail_
segmented_ reduce
Functions§
- broadcast
- Broadcast from a logical lane, not a native subgroup lane.
- exclusive_
scan - Exclusive scan seeded by
initial, with the same participation contract. - exclusive_
unseeded - Unseeded exclusive scan; the first logical lane’s output is unspecified.
- head_
segmented_ reduce - Segmented reduction whose result is valid at each segment’s head lane. The first lane implicitly starts a segment.
- inclusive_
scan - Inclusive scan in lane order. Values beyond
valid_lanesare unspecified.valid_lanesis uniform within a logical warp and lies in1..=width. - logical_
lane_ id - logical_
warp_ base_ id - logical_
warp_ id - reduce
- Ordered reduction, broadcast to all lanes of the logical warp.
- scan
- Compute both scan forms with one scan and return the unseeded aggregate to every participating logical lane.
- tail_
segmented_ reduce - Segmented reduction whose result is valid at each segment’s head lane.
A true
tailmarks the last lane of a segment; the final lane is implicit.