pub trait Benchmark {
type Input: Clone;
type Output;
// Required methods
fn prepare(&self) -> Self::Input;
fn execute(&self, input: Self::Input) -> Result<Self::Output, String>;
fn name(&self) -> String;
fn sync(&self);
// Provided methods
fn warmup_budget(&self) -> Duration { ... }
fn num_samples(&self) -> usize { ... }
fn options(&self) -> Option<String> { ... }
fn shapes(&self) -> Vec<Vec<usize>> { ... }
fn work(&self) -> Option<Work> { ... }
fn profile(&self, args: Self::Input) -> Result<ProfileDuration, String> { ... }
fn profile_full(&self, args: Self::Input) -> Result<ProfileDuration, String> { ... }
fn run(
&self,
timing_method: TimingMethod,
) -> Result<BenchmarkDurations, String> { ... }
}Expand description
Benchmark trait.
Required Associated Types§
Required Methods§
Sourcefn prepare(&self) -> Self::Input
fn prepare(&self) -> Self::Input
Prepare the benchmark, run anything that is essential for the benchmark, but shouldn’t count as included in the duration.
§Notes
This should not include warmup, the benchmark will be run at least one time without measuring the execution time.
Sourcefn execute(&self, input: Self::Input) -> Result<Self::Output, String>
fn execute(&self, input: Self::Input) -> Result<Self::Output, String>
Execute the benchmark and returns the logical output of the task executed.
It is important to return the output since otherwise deadcode optimization might optimize away code that should be benchmarked.
Provided Methods§
Sourcefn warmup_budget(&self) -> Duration
fn warmup_budget(&self) -> Duration
Wall clock the warmup holds the device for before sampling starts.
A device takes hundreds of milliseconds to reach the clocks it sustains,
and under half a second a GPU still samples a step low. BENCH_WARMUP_MS
buys the wall clock back and pays in accuracy.
Sourcefn num_samples(&self) -> usize
fn num_samples(&self) -> usize
Number of samples per run required to have a statistical significance.
Sourcefn work(&self) -> Option<Work>
fn work(&self) -> Option<Work>
The work one execution performs, for scoring the run against measured peak
throughput. None when the benchmark has no such figure to report.
One figure rather than per-resource bounds because this crate cannot name a
throughput key. A bound builder that can, such as roofline_bounds in
cubecl-std, splits it.
Sourcefn profile(&self, args: Self::Input) -> Result<ProfileDuration, String>
Available on crate feature std only.
fn profile(&self, args: Self::Input) -> Result<ProfileDuration, String>
std only.Start measuring the computation duration.
Sourcefn profile_full(&self, args: Self::Input) -> Result<ProfileDuration, String>
Available on crate feature std only.
fn profile_full(&self, args: Self::Input) -> Result<ProfileDuration, String>
std only.Start measuring the computation duration. Use the full duration irregardless of whether device duration is available or not.
Sourcefn run(&self, timing_method: TimingMethod) -> Result<BenchmarkDurations, String>
fn run(&self, timing_method: TimingMethod) -> Result<BenchmarkDurations, String>
Run the benchmark a number of times.
Dyn Compatibility§
This trait is dyn compatible.
In older versions of Rust, dyn compatibility was called "object safety".