ying_profiler/lib.rs
1//! Ying is a native Rust sampling memory profiler which tracks retained memory and allocations. It is designed for
2//! production usage in asynchronous Rust programs. Ying(้ทน) is Chinese for eagle. ๐ฆ
๐ฆ
๐ฆ
3//!
4//! I started this experiment because existing solutions I looked at were either consuming too much resources, or wrote
5//! profiling files that were too large or too cumbersome to consume, or were very bad at producing useful stack traces
6//! especially for Rust async programs, or did not support tracking retained memory in its profiling.
7//!
8//! Features:
9//! * Sampling profiler, so it uses little enough resources to be useful in production
10//! * Track retained memory, including reallocs, as well as total allocations
11//! * Targets async Rust programs (especially Tokio apps that use tracing for instrumentation), so you know the stack
12//! traces will be useful
13//! - Skips inlined `::poll::` frames in expanded stack traces for clarity
14//! * Support for detecting leaks or large amounts of allocated memory that has not been freed
15//! - Tracks realloc() calls as single long-lived allocation
16//! * Automatic and easy flamegraph generation
17//! * Allocation lifetime/length histogram
18//! * Get top stack traces by total allocation
19//! * Get top traces by retained allocation
20//! * `ProfilerRunner` -- utility to spin up thread to dump out reports and optionally flamegraphs every N minutes when
21//! total memory usage changes significantly
22//! * Catch and prevent single allocations greater than say 64GB, dump out giant allocation stack trace
23//!
24//! To see an example of a long-running toy service that generates the periodic reports and flamegraphs:
25//!
26//! `cargo run --profile bench --example ying_example`
27//!
28//! The example runs a cache workload plus a simulated leak for ~12 minutes โ long enough for the
29//! reporting thread to write a report and flamegraph under `ying-profiles/` at every check. For a
30//! quicker demo:
31//!
32//! `YING_EXAMPLE_INTERVAL_SECS=10 YING_EXAMPLE_RUNTIME_SECS=45 cargo run --profile bench --example ying_example`
33//!
34//! ## How to use
35//!
36//! Set Ying as the global allocator, then turn on reporting with one line:
37//!
38//! ```no_run
39//! use ying_profiler::YingProfiler;
40//!
41//! #[global_allocator]
42//! static YING_ALLOC: YingProfiler = YingProfiler::default();
43//!
44//! fn main() {
45//! YING_ALLOC.start_profiling();
46//! // ... the rest of your program
47//! }
48//! ```
49//!
50//! There are three layers here, and it is worth knowing which one you are using.
51//!
52//! ### 1. The allocator: collects, never reports
53//!
54//! The `#[global_allocator]` declaration is what makes Ying sample allocations at all. On its own
55//! it is complete and valid: Ying accumulates stats in memory and defers to the System allocator,
56//! but writes nothing anywhere. Nothing is scheduled and no thread is spawned.
57//!
58//! ### 2. [`ProfilerRunner`]: the primitive that decides how stats get reported
59//!
60//! [`ProfilerRunner`] is the piece to reach for whenever the defaults do not fit. It owns every
61//! decision about turning collected stats into output โ how often to look, how much of a change is
62//! worth reporting, what to measure, where output goes, and in what form:
63//!
64//! | Setting | Meaning |
65//! |---|---|
66//! | `check_interval_secs` | how often the background thread wakes up to compare memory use |
67//! | `report_pct_change_trigger` | how much retained memory must move before a report is written |
68//! | `reporting_path` | directory for reports and flamegraphs; created if missing |
69//! | `measure_allocated_not_retained` | rank stacks by total allocated bytes instead of retained |
70//! | `gen_flamegraphs` | also write an SVG flamegraph alongside each text report |
71//! | `expand_frames` | expand inlined symbols within each stack frame in reports |
72//!
73//! Build one with [`ProfilerRunnerBuilder`](utils::ProfilerRunnerBuilder), which defaults
74//! anything you leave out, then hand it
75//! the allocator static to start its thread:
76//!
77//! ```no_run
78//! use ying_profiler::{utils::ProfilerRunnerBuilder, YingProfiler};
79//!
80//! #[global_allocator]
81//! static YING_ALLOC: YingProfiler = YingProfiler::default();
82//!
83//! fn main() {
84//! let runner = ProfilerRunnerBuilder::default()
85//! .check_interval_secs(60usize)
86//! .report_pct_change_trigger(5usize)
87//! .reporting_path("/var/log/ying")
88//! .gen_flamegraphs(true)
89//! .build()
90//! .unwrap();
91//! runner.spawn(&YING_ALLOC);
92//! }
93//! ```
94//!
95//! [`YingProfiler::start_profiling`] is not a separate mechanism: it builds exactly one of these
96//! with a specific set of values (5 minute interval, 10% trigger, flamegraphs on, retained memory,
97//! output to `ying-profiles`) and spawns it. Anything it can do, a [`ProfilerRunner`] you build
98//! yourself can do too.
99//!
100//! ### 3. Reading stats directly
101//!
102//! You do not need a runner at all if you would rather decide when to look. [`YingProfiler`]
103//! exposes the stats as data โ [`YingProfiler::top_k_stacks_by_retained`],
104//! [`YingProfiler::top_k_stacks_by_allocated`], [`YingProfiler::total_retained_bytes`] โ and
105//! [`utils::gen_flamegraph`] writes a one-off flamegraph on demand. This is the route to take when
106//! reports should be triggered by your own signals, such as an HTTP endpoint or a health check
107//! noticing memory growth.
108//!
109//! ## Why a new memory profiler?
110//!
111//! Rust as an ecosystem is lacking in good memory profiling tools. [Bytehound](https://github.com/koute/bytehound)
112//! is quite good but has a large CPU impact and writes out huge profiling files as it measures every allocation.
113//! Jemalloc/Jeprof does sampling, so it's great for production use, but its output is difficult to interpret, and
114//! it does not track retained memory, which is quite critical for debugging memory issues in production.
115//! Both of the above tools are written with generic C/C++ malloc/preload ABI in mind, so the backtraces
116//! that one gets from their use are really limited, especially for profiling Rust binaries built for
117//! an optimized/release target, and especially async code. The output also often has trouble with mangled symbols.
118//!
119//! If we use the backtrace crate and analyze release/bench backtraces in detail, we can see why that is.
120//! The Rust compiler does a good job of inlining function calls - even ones across async/await boundaries -
121//! in release code. Thus, for a single instruction pointer (IP) in the stack trace, it might correspond to
122//! many different places in the code. This is from `examples/ying_example.rs`:
123//!
124//! ```bash
125//! Some(ying_example::insert_one::{{closure}}::h7eddb5f8ebb3289b)
126//! > Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::h7a53098577c44da0)
127//! > Some(ying_example::cache_update_loop::{{closure}}::h38556c7e7ae06bfa)
128//! > Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::hd319a0f603a1d426)
129//! > Some(ying_example::main::{{closure}}::h33aa63760e836e2f)
130//! > Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::hb2fd3cb904946c24)
131//! ```
132//!
133//! A generic tool which just examines the IP and tries to figure out a single symbol would miss out on all of the
134//! inlined symbols. Some tools can expand on symbols, but the results still aren't very good.
135//!
136//! ## Feature Flags
137//!
138//! - `profile_spans` - records tracing-span information in stacks. NOTE: this feature is
139//! experimental and incomplete.
140//!
141use std::alloc::{GlobalAlloc, Layout, System};
142use std::cell::Cell;
143use std::cmp::min;
144use std::env::var;
145use std::fmt::Write;
146use std::mem::needs_drop;
147use std::ptr::{copy_nonoverlapping, null_mut};
148use std::sync::atomic::{AtomicUsize, Ordering::Relaxed, Ordering::SeqCst};
149
150use backtrace::Backtrace;
151use coarsetime::Clock;
152use once_cell::sync::OnceCell;
153
154pub mod callstack;
155pub mod histogram;
156mod sharded_map;
157pub mod utils;
158use callstack::{FriendlySymbol, StackStats, StdCallstack};
159use sharded_map::ShardedMap;
160use utils::{
161 ProfilerRunner, DEFAULT_CHECK_INTERVAL_SECS, DEFAULT_PCT_CHANGE_TRIGGER, DEFAULT_REPORTING_PATH,
162};
163
164/// The number of frames at the top of the stack to skip. Most of these have to do with
165/// backtrace and this profiler infrastructure. This number needs to be adjusted
166/// depending on the implementation.
167/// For release/bench builds with debug = 1 / strip = none, this should be 4.
168/// For debug builds, this is about 9.
169const TOP_FRAMES_TO_SKIP: usize = 3;
170
171const DEFAULT_GIANT_ALLOC_LIMIT: usize = 64 * 1024 * 1024 * 1024;
172
173/// Testing knob: overrides the initial capacity of every profiler map. Not part of the public
174/// API, and read exactly once when profiler state is first initialized.
175const MAP_CAPACITY_ENV_VAR: &str = "YING_TEST_MAP_CAPACITY";
176
177// A map for caching symbols in backtraces so we can mostly store u64's
178pub(crate) type SymbolMap = ShardedMap<Vec<FriendlySymbol>>;
179
180/// Ying is a memory profiling Allocator wrapper.
181/// Ying is the Chinese word for an eagle.
182pub struct YingProfiler {
183 /// Allocation sampling ratio. Eg: 500 means 1 in 500 allocations are sampled.
184 sampling_ratio: u32,
185 /// Prevent and dump stack trace for giant single allocations beyond a certain size
186 single_alloc_limit: usize,
187 /// Global thread local state cache
188 /// Statistics... lazily initialized later
189 state: OnceCell<YingState>,
190}
191
192static TOTAL_RETAINED: AtomicUsize = AtomicUsize::new(0);
193static PROFILED_ALLOCATED: AtomicUsize = AtomicUsize::new(0);
194static PROFILED_RETAINED: AtomicUsize = AtomicUsize::new(0);
195
196impl YingProfiler {
197 /// sampling_ratio: number of allocations for every sampled allocation
198 pub const fn new(sampling_ratio: u32, single_alloc_limit: usize) -> Self {
199 Self {
200 sampling_ratio,
201 single_alloc_limit,
202 state: OnceCell::new(),
203 }
204 }
205
206 pub const fn default() -> Self {
207 Self {
208 sampling_ratio: 500,
209 single_alloc_limit: DEFAULT_GIANT_ALLOC_LIMIT,
210 state: OnceCell::new(),
211 }
212 }
213
214 /// Starts periodic memory reporting on a background thread and returns the [`ProfilerRunner`]
215 /// that was spawned. This is the one-line way to turn Ying on:
216 ///
217 /// ```no_run
218 /// use ying_profiler::YingProfiler;
219 ///
220 /// #[global_allocator]
221 /// static YING_ALLOC: YingProfiler = YingProfiler::default();
222 ///
223 /// fn main() {
224 /// YING_ALLOC.start_profiling();
225 /// }
226 /// ```
227 ///
228 /// This is a convenience only, and adds no capability of its own: it builds a
229 /// [`ProfilerRunner`] with a fixed set of values and calls [`ProfilerRunner::spawn`]. Those
230 /// values are a 5 minute [`DEFAULT_CHECK_INTERVAL_SECS`] check interval, a 10%
231 /// [`DEFAULT_PCT_CHANGE_TRIGGER`] change before a report is written, output to
232 /// [`DEFAULT_REPORTING_PATH`] (created if it does not exist), flamegraphs on, retained rather
233 /// than allocated memory, and inlined frames left unexpanded.
234 ///
235 /// To change any of those, build the [`ProfilerRunner`] yourself with
236 /// [`ProfilerRunnerBuilder`](utils::ProfilerRunnerBuilder) and call
237 /// [`spawn`](ProfilerRunner::spawn) on it instead. The returned runner records the settings in
238 /// use, but the reporting thread has already been started with them; changing the returned
239 /// value has no effect.
240 pub fn start_profiling(&'static self) -> ProfilerRunner {
241 let runner = ProfilerRunner::new(
242 DEFAULT_CHECK_INTERVAL_SECS,
243 DEFAULT_PCT_CHANGE_TRIGGER,
244 DEFAULT_REPORTING_PATH,
245 false,
246 true,
247 false,
248 );
249 runner.spawn(self);
250 runner
251 }
252
253 /// Total outstanding retained bytes (not just sampled but all allocations)
254 #[inline]
255 pub fn total_retained_bytes() -> usize {
256 TOTAL_RETAINED.load(Relaxed)
257 }
258
259 /// Total bytes allocated for profiled allocations
260 #[inline]
261 pub fn profiled_bytes_allocated() -> usize {
262 PROFILED_ALLOCATED.load(Relaxed)
263 }
264
265 /// Profiled bytes retained - retained memory usage amongst profiled allocations
266 #[inline]
267 pub fn profiled_bytes_retained() -> usize {
268 PROFILED_RETAINED.load(Relaxed)
269 }
270
271 #[inline]
272 pub fn symbol_map_size(&self) -> usize {
273 self.lock_out_profiler(|| self.get_state().symbol_map.len())
274 }
275
276 /// Number of entries for outstanding sampled allocations map
277 #[inline]
278 pub fn num_outstanding_allocs(&self) -> usize {
279 self.lock_out_profiler(|| self.get_state().outstanding_allocs.len())
280 }
281
282 /// Number of distinct stack traces that have been sampled
283 #[inline]
284 pub fn num_stack_traces(&self) -> usize {
285 self.lock_out_profiler(|| self.get_state().stack_stats.len())
286 }
287
288 /// Get the top k stack traces by total profiled bytes allocated, in descending order.
289 /// Note that "profiled bytes" refers to the bytes allocated during sampling by this profiler.
290 pub fn top_k_stacks_by_allocated(&self, k: usize) -> Vec<StackStats> {
291 self.lock_out_profiler(|| {
292 let stacks_by_alloc = self.stack_list_allocated_bytes_desc();
293 stacks_by_alloc
294 .iter()
295 .take(k)
296 .filter_map(|&(stack_hash, _bytes_allocated)| {
297 self.get_stats_for_stack_hash(stack_hash)
298 })
299 .collect()
300 })
301 }
302
303 /// Get the top k stack traces by retained sampled memory, in descending order.
304 pub fn top_k_stacks_by_retained(&self, k: usize) -> Vec<StackStats> {
305 self.lock_out_profiler(|| {
306 let stacks_by_retained = self.stack_list_retained_bytes_desc();
307 stacks_by_retained
308 .iter()
309 .take(k)
310 .filter_map(|&(stack_hash, _bytes_retained)| {
311 self.get_stats_for_stack_hash(stack_hash)
312 })
313 .collect()
314 })
315 }
316
317 /// Returns a list of stack IDs (stack_hash, bytes_allocated) in order from highest
318 /// number of bytes allocated to lowest
319 fn stack_list_allocated_bytes_desc(&self) -> Vec<(u64, u64)> {
320 // TODO: filter away entries with minimal allocations, say <1% or some threshold
321 let mut items = self
322 .get_state()
323 .stack_stats
324 .map_to_vec(|stack_hash, stats| (stack_hash, stats.allocated_bytes));
325 items.sort_unstable_by_key(|&(_, bytes)| std::cmp::Reverse(bytes));
326 items
327 }
328
329 /// Returns a list of stack IDs (stack_hash, bytes_retained) in order from highest
330 /// number of bytes retained to lowest
331 fn stack_list_retained_bytes_desc(&self) -> Vec<(u64, u64)> {
332 // TODO: filter away entries with minimal retained allocations, say <1% or some threshold
333 let mut items = self
334 .get_state()
335 .stack_stats
336 .map_to_vec(|stack_hash, stats| (stack_hash, stats.retained_profiled_bytes()));
337 items.sort_unstable_by_key(|&(_, bytes)| std::cmp::Reverse(bytes));
338 items
339 }
340
341 pub fn reset_state_for_testing_only(&self) {
342 self.lock_out_profiler(|| {
343 let state = self.get_state();
344 state.stack_stats.clear();
345 state.outstanding_allocs.clear();
346 })
347 }
348
349 pub fn testing_only_guarantee_next_sample(&self) {
350 THREAD_STATE.with(YingThreadLocal::test_only_reset_sampling_counter)
351 }
352
353 #[inline]
354 fn check_and_deny_giant_allocations(&self, ptr: *mut u8, layout: Layout) -> *mut u8 {
355 // Sorry there is an edge case where this check cannot happen if YING is not initialized
356 if layout.size() >= self.single_alloc_limit && self.state.get().is_some() {
357 // Prevent allocation sampling while we are telling the world who did this
358 self.lock_out_profiler(|| {
359 println!(
360 "WARNING: Huge memory allocation of {} bytes denied by Ying profiler",
361 layout.size()
362 );
363
364 let mut bt = Backtrace::new_unresolved();
365
366 // 2. Create a Callstack, check if there is a similar stack
367 let stack = StdCallstack::from_backtrace_unresolved(&bt);
368 let state = self.get_state();
369 stack.populate_symbol_map(&mut bt, &state.symbol_map);
370 println!(
371 "Stack trace:\n{}",
372 stack.with_symbols_and_filename(&state.symbol_map, true)
373 );
374 });
375 null_mut::<u8>()
376 } else {
377 ptr
378 }
379 }
380
381 #[inline]
382 fn get_state(&self) -> &YingState {
383 // We need to lock out the profiler here, to ensure no tracking of allocations or messes
384 self.lock_out_profiler(|| self.state.get_or_init(YingState::new))
385 }
386
387 #[inline]
388 fn get_stats_for_stack_hash(&self, stack_hash: u64) -> Option<StackStats> {
389 self.get_state().stack_stats.get_cloned(stack_hash)
390 }
391
392 /// Locks the profiler flag so that allocations are not profiled.
393 /// This is for non-profiler code such as debug prints that has to access the Dashmap or state
394 /// and could potentially cause deadlock problems with Dashmap for example.
395 #[inline]
396 fn lock_out_profiler<R>(&self, func: impl FnOnce() -> R) -> R {
397 // Deliberately three short thread local accesses rather than running `func` inside a
398 // `with`: `func` frequently calls back in here (get_state does), and keeping the borrow
399 // scopes disjoint means nesting needs no reasoning about re-entrant `with`.
400 THREAD_STATE.with(YingThreadLocal::set_allocator_lock);
401 let return_val = func();
402 exit_profiler();
403 return_val
404 }
405}
406
407// Private state. We can't put this in the main YingProfiler struct as that one has to be const static
408struct YingState {
409 symbol_map: SymbolMap,
410 // Main map of stack hash to StackStats
411 stack_stats: ShardedMap<StackStats>,
412 // Map of outstanding sampled allocations. Used to figure out amount of outstanding allocations and
413 // statistics about how long lived outstanding allocations are.
414 // (*ptr as u64 -> (stack hash, start_timestamp_epoch_millis))
415 outstanding_allocs: ShardedMap<(u64, u64)>,
416}
417
418impl YingState {
419 pub fn new() -> Self {
420 // Shrinking the maps makes them resize constantly, which is how the re-entrancy deadlocks
421 // were shaken out; the stress tests set this so those paths get exercised on every run.
422 let capacity_scale = var(MAP_CAPACITY_ENV_VAR)
423 .ok()
424 .and_then(|s| s.parse::<usize>().ok());
425 let symbol_map = SymbolMap::with_capacity(capacity_scale.unwrap_or(1000));
426 let stack_stats = ShardedMap::with_capacity(capacity_scale.unwrap_or(1000));
427 let outstanding_allocs = ShardedMap::with_capacity(capacity_scale.unwrap_or(5000));
428 Self {
429 symbol_map,
430 stack_stats,
431 outstanding_allocs,
432 }
433 }
434}
435
436thread_local! {
437 /// Per-thread profiler state, in a thread local rather than a slot in a fixed-size table.
438 ///
439 /// An earlier design hashed the thread id into a 1024 entry array and handed out `&mut` to the
440 /// slot, on the reasoning that a `thread_local!` cannot be used from a `GlobalAlloc` because
441 /// setting up TLS may allocate and re-enter the allocator. That reasoning was sound for a lazily
442 /// initialized thread local, but two threads whose ids collide then share one slot, which is both
443 /// aliasing UB and a lost re-entrancy guard: a non-atomic increment of `alloc_lock` can be dropped,
444 /// the guard falls to zero while a thread is still inside its critical section, and its next
445 /// internal free re-enters the maps and blocks on a shard lock it already holds. That deadlock was
446 /// observed, and the captured backtrace showed exactly that shape.
447 ///
448 /// The `const {}` initializer plus a type that needs no `Drop` is what makes this safe: it skips
449 /// lazy initialization and destructor registration, so no Rust allocation happens on access. On
450 /// ELF targets linked into an executable this compiles to a single thread-pointer-relative load.
451 /// macOS still routes through a resolver call that allocates the thread's TLV block on first touch,
452 /// but that allocation is libc `malloc` inside dyld, which `#[global_allocator]` does not
453 /// interpose, so it cannot re-enter Ying. Measured: 21,225 accesses from inside `alloc` and
454 /// `dealloc` across 50 freshly spawned threads and their teardown, with zero re-entry.
455 ///
456 /// Keep both invariants if you touch this. Adding a field that needs dropping, or dropping the
457 /// `const {}`, reintroduces an allocating first access and with it the deadlock.
458 static THREAD_STATE: YingThreadLocal = const { YingThreadLocal::new() };
459}
460
461/// Enforces the invariant that [`THREAD_STATE`] documents. A field that needs dropping would make
462/// the thread local register a destructor, which allocates on first access and re-enters the
463/// allocator, so this is a compile error rather than a comment nobody reads.
464const _: () = assert!(
465 !needs_drop::<YingThreadLocal>(),
466 "YingThreadLocal must not need Drop, or thread local setup will allocate inside the allocator"
467);
468
469/// A struct to provide a better API around the lock out profiler flag/re-entrancy plus sampling.
470/// Interior mutability via [`Cell`] rather than `&mut`, so that no field needs `Drop` and the
471/// thread local stays on its cheapest code path.
472struct YingThreadLocal {
473 // Counts up for every time we enter a no-allocator critical section (ie where we have to touch
474 // allocator state or cause an allocation within profiling-related code and don't want sampling
475 // of re-entrant allocations done). Nonzero prevents allocator from sampling.
476 alloc_lock: Cell<u32>,
477 sample_count: Cell<u32>,
478}
479
480impl YingThreadLocal {
481 const fn new() -> Self {
482 Self {
483 alloc_lock: Cell::new(0),
484 sample_count: Cell::new(0),
485 }
486 }
487
488 #[inline]
489 fn is_allocator_locked(&self) -> bool {
490 self.alloc_lock.get() > 0
491 }
492
493 #[inline]
494 fn set_allocator_lock(&self) {
495 self.alloc_lock.set(self.alloc_lock.get().saturating_add(1));
496 }
497
498 #[inline]
499 fn release_allocator_lock(&self) {
500 self.alloc_lock.set(self.alloc_lock.get().saturating_sub(1));
501 }
502
503 /// Obtains the counter, checks for sampling ratio, and updates counter in one go
504 #[inline]
505 fn should_sample(&self, ratio: u32) -> bool {
506 let count = self.sample_count.get().wrapping_add(1);
507 self.sample_count.set(count);
508 count % ratio == 0
509 }
510
511 // Resets counter to 0 to guarantee next call to alloc() will sample. TESTING ONLY
512 #[inline]
513 fn test_only_reset_sampling_counter(&self) {
514 self.sample_count.set(0);
515 }
516}
517
518/// Takes the re-entrancy guard if this thread is not already inside the profiler, reporting whether
519/// the caller now owns it and must release it.
520///
521/// Deliberately one thread local access rather than several: this runs on every allocation, and the
522/// common answer is "do not profile this one".
523#[inline]
524fn try_enter_profiler(sampling_ratio: Option<u32>) -> bool {
525 THREAD_STATE.with(|tl| {
526 if tl.is_allocator_locked() {
527 return false;
528 }
529 // Only consume a sampling tick when we were not already locked out, matching the original
530 // short-circuiting behaviour
531 if let Some(ratio) = sampling_ratio {
532 if !tl.should_sample(ratio) {
533 return false;
534 }
535 }
536 tl.set_allocator_lock();
537 true
538 })
539}
540
541#[inline]
542fn exit_profiler() {
543 THREAD_STATE.with(YingThreadLocal::release_allocator_lock)
544}
545
546unsafe impl GlobalAlloc for YingProfiler {
547 unsafe fn alloc(&self, layout: Layout) -> *mut u8 {
548 // NOTE: the code between here and the state.0 = true must be re-entrant
549 // and therefore not allocate, otherwise there will be an infinite loop.
550 let alloc_ptr = self.check_and_deny_giant_allocations(System.alloc(layout), layout);
551 if !alloc_ptr.is_null() {
552 TOTAL_RETAINED.fetch_add(layout.size(), SeqCst);
553
554 // Now, sample allocation - if it falls below threshold, then profile
555 // Also, we set a ThreadLocal to avoid re-entry: ie the code below might allocate,
556 // and we avoid profiling if we are already in the loop below. Avoids cycles.
557 if try_enter_profiler(Some(self.sampling_ratio)) {
558 PROFILED_ALLOCATED.fetch_add(layout.size(), SeqCst);
559 PROFILED_RETAINED.fetch_add(layout.size(), SeqCst);
560
561 // -- Beginning of section that may allocate
562 // 1. Get unresolved backtrace for speed
563 let mut bt = Backtrace::new_unresolved();
564
565 // 2. Create a Callstack, check if there is a similar stack
566 let stack = StdCallstack::from_backtrace_unresolved(&bt);
567 let stack_hash = stack.compute_hash();
568 let state = self.get_state();
569 state.stack_stats.update_or_insert_with(
570 stack_hash,
571 |stats| {
572 // 4. Update stats
573 stats.num_allocations += 1;
574 stats.allocated_bytes += layout.size() as u64;
575 },
576 || {
577 // 3. Resolve symbols if needed (new stack entry).
578 // NOTE: this runs while the stack_stats shard is locked, but it only ever
579 // touches symbol_map, which is a separate map, so it cannot self-deadlock.
580 stack.populate_symbol_map(&mut bt, &state.symbol_map);
581 StackStats::new(stack, Some(layout.size() as u64))
582 },
583 );
584
585 // 4. Record allocation so we can track outstanding vs transient allocs
586 state.outstanding_allocs.update_or_insert_with(
587 alloc_ptr as u64,
588 |_existing| {},
589 || (stack_hash, Clock::recent_since_epoch().as_millis()),
590 );
591
592 // -- End of core profiling section, no more allocations --
593 exit_profiler();
594 }
595 }
596 alloc_ptr
597 }
598
599 unsafe fn dealloc(&self, ptr: *mut u8, layout: Layout) {
600 System.dealloc(ptr, layout);
601 TOTAL_RETAINED.fetch_sub(layout.size(), SeqCst);
602
603 // Return immediately and skip rest of this if YING_STATE is not initialized. It could cause
604 // an infinite loop because during initialization of YING_STATE, dealloc() could be then called
605 if self.state.get().is_none() {
606 return;
607 }
608
609 // IMPORTANT: the re-entrancy check has to happen before we touch any of the profiler maps.
610 // The maps allocate and free internally (eg when a shard resizes), and those frees land back
611 // here while that shard's write lock is held. Touching the map first would then try to take
612 // a read lock the same thread already holds exclusively, which self-deadlocks.
613 if !try_enter_profiler(None) {
614 return;
615 }
616
617 // -- Beginning of section that may allocate
618 // If the allocation was recorded in outstanding_allocs, then remove it and update stats
619 // about number of bytes freed etc.
620 let state = self.get_state();
621 if state.outstanding_allocs.contains_key(ptr as u64) {
622 if let Some((stack_hash, alloc_ts)) = state.outstanding_allocs.remove(ptr as u64) {
623 PROFILED_RETAINED.fetch_sub(layout.size(), SeqCst);
624 let alloc_time_ms = Clock::recent_since_epoch()
625 .as_millis()
626 .saturating_sub(alloc_ts);
627
628 // Update memory profiling freed bytes stats
629 state.stack_stats.update(stack_hash, |stats| {
630 stats.update_free_stats(layout.size() as u64, alloc_time_ms)
631 });
632 }
633 }
634
635 // -- End of core profiling section, no more allocations --
636 exit_profiler();
637 }
638
639 // We implement a custom realloc(). We must count reallocs as the same allocation, but need to do
640 // the following: - update original allocated bytes (but not allocations); move outstanding_allocs
641 // because the pointer moved, but preserve original starting timestamp.
642 // The above also saves us cycles from having to call alloc() and dealloc() separately.
643 unsafe fn realloc(&self, ptr: *mut u8, layout: Layout, new_size: usize) -> *mut u8 {
644 let old_size = layout.size();
645 // SAFETY: the caller must ensure that the `new_size` does not overflow.
646 // `layout.align()` comes from a `Layout` and is thus guaranteed to be valid.
647 let new_layout = Layout::from_size_align_unchecked(new_size, layout.align());
648 // SAFETY: the caller must ensure that `new_layout` is greater than zero.
649 let new_ptr = self.check_and_deny_giant_allocations(System.alloc(new_layout), new_layout);
650 if !new_ptr.is_null() {
651 // SAFETY: the previously allocated block cannot overlap the newly allocated block.
652 // The safety contract for `dealloc` must be upheld by the caller.
653 copy_nonoverlapping(ptr, new_ptr, min(old_size, new_size));
654 System.dealloc(ptr, layout);
655
656 // 1. Update global statistics
657 if new_size > old_size {
658 TOTAL_RETAINED.fetch_add(new_size - old_size, SeqCst);
659 } else {
660 TOTAL_RETAINED.fetch_sub(old_size - new_size, SeqCst);
661 }
662
663 // 2. IF the old pointer was in outstanding_allocs, move it and make a new entry,
664 // keeping the old starting timestamp. Also update stack stats.
665 // But only if state is alredy initialized - otherwise any state initialization that
666 // results in a realloc() could cause this to infinite loop
667 if self.state.get().is_none() {
668 return new_ptr;
669 }
670
671 // As in dealloc(), the re-entrancy check must come before any map access, or a map's own
672 // internal realloc re-enters here while holding that shard's write lock and deadlocks.
673 if !try_enter_profiler(None) {
674 return new_ptr;
675 }
676
677 // -- Beginning of section that may allocate
678 let state = self.get_state();
679 if state.outstanding_allocs.contains_key(ptr as u64) {
680 if let Some((stack_hash, alloc_ts)) = state.outstanding_allocs.remove(ptr as u64) {
681 if new_size > old_size {
682 PROFILED_RETAINED.fetch_add(new_size - old_size, SeqCst);
683 } else {
684 PROFILED_RETAINED.fetch_sub(old_size - new_size, SeqCst);
685 }
686
687 state
688 .outstanding_allocs
689 .insert(new_ptr as u64, (stack_hash, alloc_ts));
690
691 // Update memory profiling freed bytes stats
692 state.stack_stats.update(stack_hash, |stats| {
693 if new_size > old_size {
694 stats.allocated_bytes += (new_size - old_size) as u64;
695 } else {
696 stats.allocated_bytes -= (old_size - new_size) as u64;
697 }
698 // Don't change number of allocations or frees
699 });
700 }
701 }
702
703 // -- End of core profiling section, no more allocations --
704 exit_profiler();
705 }
706 new_ptr
707 }
708}