Skip to main content

Crate ying_profiler

Crate ying_profiler 

Source
Expand description

Ying is a native Rust sampling memory profiler which tracks retained memory and allocations. It is designed for production usage in asynchronous Rust programs. Ying(鷹) is Chinese for eagle. 🦅🦅🦅

I started this experiment because existing solutions I looked at were either consuming too much resources, or wrote profiling files that were too large or too cumbersome to consume, or were very bad at producing useful stack traces especially for Rust async programs, or did not support tracking retained memory in its profiling.

Features:

  • Sampling profiler, so it uses little enough resources to be useful in production
  • Track retained memory, including reallocs, as well as total allocations
  • Targets async Rust programs (especially Tokio apps that use tracing for instrumentation), so you know the stack traces will be useful
    • Skips inlined ::poll:: frames in expanded stack traces for clarity
  • Support for detecting leaks or large amounts of allocated memory that has not been freed
    • Tracks realloc() calls as single long-lived allocation
  • Automatic and easy flamegraph generation
  • Allocation lifetime/length histogram
  • Get top stack traces by total allocation
  • Get top traces by retained allocation
  • ProfilerRunner – utility to spin up thread to dump out reports and optionally flamegraphs every N minutes when total memory usage changes significantly
  • Catch and prevent single allocations greater than say 64GB, dump out giant allocation stack trace

To see an example of a long-running toy service that generates the periodic reports and flamegraphs:

cargo run --profile bench --example ying_example

The example runs a cache workload plus a simulated leak for ~12 minutes — long enough for the reporting thread to write a report and flamegraph under ying-profiles/ at every check. For a quicker demo:

YING_EXAMPLE_INTERVAL_SECS=10 YING_EXAMPLE_RUNTIME_SECS=45 cargo run --profile bench --example ying_example

§How to use

Set Ying as the global allocator, then turn on reporting with one line:

use ying_profiler::YingProfiler;

#[global_allocator]
static YING_ALLOC: YingProfiler = YingProfiler::default();

fn main() {
    YING_ALLOC.start_profiling();
    // ... the rest of your program
}

There are three layers here, and it is worth knowing which one you are using.

§1. The allocator: collects, never reports

The #[global_allocator] declaration is what makes Ying sample allocations at all. On its own it is complete and valid: Ying accumulates stats in memory and defers to the System allocator, but writes nothing anywhere. Nothing is scheduled and no thread is spawned.

§2. ProfilerRunner: the primitive that decides how stats get reported

ProfilerRunner is the piece to reach for whenever the defaults do not fit. It owns every decision about turning collected stats into output — how often to look, how much of a change is worth reporting, what to measure, where output goes, and in what form:

SettingMeaning
check_interval_secshow often the background thread wakes up to compare memory use
report_pct_change_triggerhow much retained memory must move before a report is written
reporting_pathdirectory for reports and flamegraphs; created if missing
measure_allocated_not_retainedrank stacks by total allocated bytes instead of retained
gen_flamegraphsalso write an SVG flamegraph alongside each text report
expand_framesexpand inlined symbols within each stack frame in reports

Build one with ProfilerRunnerBuilder, which defaults anything you leave out, then hand it the allocator static to start its thread:

use ying_profiler::{utils::ProfilerRunnerBuilder, YingProfiler};

#[global_allocator]
static YING_ALLOC: YingProfiler = YingProfiler::default();

fn main() {
    let runner = ProfilerRunnerBuilder::default()
        .check_interval_secs(60usize)
        .report_pct_change_trigger(5usize)
        .reporting_path("/var/log/ying")
        .gen_flamegraphs(true)
        .build()
        .unwrap();
    runner.spawn(&YING_ALLOC);
}

YingProfiler::start_profiling is not a separate mechanism: it builds exactly one of these with a specific set of values (5 minute interval, 10% trigger, flamegraphs on, retained memory, output to ying-profiles) and spawns it. Anything it can do, a ProfilerRunner you build yourself can do too.

§3. Reading stats directly

You do not need a runner at all if you would rather decide when to look. YingProfiler exposes the stats as data — YingProfiler::top_k_stacks_by_retained, YingProfiler::top_k_stacks_by_allocated, YingProfiler::total_retained_bytes — and utils::gen_flamegraph writes a one-off flamegraph on demand. This is the route to take when reports should be triggered by your own signals, such as an HTTP endpoint or a health check noticing memory growth.

§Why a new memory profiler?

Rust as an ecosystem is lacking in good memory profiling tools. Bytehound is quite good but has a large CPU impact and writes out huge profiling files as it measures every allocation. Jemalloc/Jeprof does sampling, so it’s great for production use, but its output is difficult to interpret, and it does not track retained memory, which is quite critical for debugging memory issues in production. Both of the above tools are written with generic C/C++ malloc/preload ABI in mind, so the backtraces that one gets from their use are really limited, especially for profiling Rust binaries built for an optimized/release target, and especially async code. The output also often has trouble with mangled symbols.

If we use the backtrace crate and analyze release/bench backtraces in detail, we can see why that is. The Rust compiler does a good job of inlining function calls - even ones across async/await boundaries - in release code. Thus, for a single instruction pointer (IP) in the stack trace, it might correspond to many different places in the code. This is from examples/ying_example.rs:

Some(ying_example::insert_one::{{closure}}::h7eddb5f8ebb3289b)
 > Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::h7a53098577c44da0)
 > Some(ying_example::cache_update_loop::{{closure}}::h38556c7e7ae06bfa)
 > Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::hd319a0f603a1d426)
 > Some(ying_example::main::{{closure}}::h33aa63760e836e2f)
 > Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::hb2fd3cb904946c24)

A generic tool which just examines the IP and tries to figure out a single symbol would miss out on all of the inlined symbols. Some tools can expand on symbols, but the results still aren’t very good.

§Feature Flags

  • profile_spans - records tracing-span information in stacks. NOTE: this feature is experimental and incomplete.

Modules§

callstack
histogram
utils

Structs§

YingProfiler
Ying is a memory profiling Allocator wrapper. Ying is the Chinese word for an eagle.