Skip to main content

Module warp

Module warp 

Source
Expand description

Collectives over logical warps within a native device subgroup.

All lanes of a logical warp participate in each call. width is positive, does not exceed the native subgroup width, and groups must not straddle a native subgroup. No NVIDIA warp width is assumed by these algorithms.

Modules§

batched
broadcast
exchange
exclusive_scan
exclusive_unseeded
head_segmented_reduce
inclusive_scan
io
logical_lane_id
logical_warp_base_id
logical_warp_id
record
reduce
scan
sort
tail_segmented_reduce

Functions§

broadcast
Broadcast from a logical lane, not a native subgroup lane.
exclusive_scan
Exclusive scan seeded by initial, with the same participation contract.
exclusive_unseeded
Unseeded exclusive scan; the first logical lane’s output is unspecified.
head_segmented_reduce
Segmented reduction whose result is valid at each segment’s head lane. The first lane implicitly starts a segment.
inclusive_scan
Inclusive scan in lane order. Values beyond valid_lanes are unspecified. valid_lanes is uniform within a logical warp and lies in 1..=width.
logical_lane_id
logical_warp_base_id
logical_warp_id
reduce
Ordered reduction, broadcast to all lanes of the logical warp.
scan
Compute both scan forms with one scan and return the unseeded aggregate to every participating logical lane.
tail_segmented_reduce
Segmented reduction whose result is valid at each segment’s head lane. A true tail marks the last lane of a segment; the final lane is implicit.