1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
// Copyright 2026 Ryan Daum
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
// Demonstrates Phase 1 of the GPU benchmarking work
// (book/src/gpu-sharp-edges.md): declaring a benchmark group as
// `MeasurementDomain::Gpu` so the runner suppresses CPU-PMU bottleneck
// diagnostics and relabels the PMU scheduling byline as
// "host PMU (orchestration)".
//
// The kernel here is a stand-in for one synchronized device operation
// (e.g. `cublasLtMatmul()` + `cudaDeviceSynchronize()`). On a real GPU
// benchmark the host thread is mostly idle, so CPU PMU counters describe
// launch/sync orchestration, not the device kernel. A previous run of
// this example with `MeasurementDomain::Cpu` produced:
//
// possible bottlenecks:
// - Likely data-side memory latency: backend stall is 89.99% with
// cache pressure at 0.00% miss rate and 0.0000 misses/op
//
// which is misleading for GPU work. With `MeasurementDomain::Gpu`:
//
// - the "possible bottlenecks:" section is suppressed entirely
// - the PMU scheduling byline reads:
// host PMU (orchestration): scheduled=100.0% ...
//
// Run with:
// cargo run --example gpu_domain --release
use ;
benchmark_main!;