Skip to main content

ying_profiler/
lib.rs

1//! Ying is a native Rust sampling memory profiler which tracks retained memory and allocations.  It is designed for
2//! production usage in asynchronous Rust programs.  Ying(้ทน) is Chinese for eagle.  ๐Ÿฆ…๐Ÿฆ…๐Ÿฆ…
3//!
4//! I started this experiment because existing solutions I looked at were either consuming too much resources, or wrote
5//! profiling files that were too large or too cumbersome to consume, or were very bad at producing useful stack traces
6//! especially for Rust async programs, or did not support tracking retained memory in its profiling.
7//!
8//! Features:
9//! * Sampling profiler, so it uses little enough resources to be useful in production
10//! * Track retained memory, including reallocs, as well as total allocations
11//! * Targets async Rust programs (especially Tokio apps that use tracing for instrumentation), so you know the stack
12//!   traces will be useful
13//!   - Skips inlined `::poll::` frames in expanded stack traces for clarity
14//! * Support for detecting leaks or large amounts of allocated memory that has not been freed
15//!   - Tracks realloc() calls as single long-lived allocation
16//! * Automatic and easy flamegraph generation
17//! * Allocation lifetime/length histogram
18//! * Get top stack traces by total allocation
19//! * Get top traces by retained allocation
20//! * `ProfilerRunner` -- utility to spin up thread to dump out reports and optionally flamegraphs every N minutes when
21//!   total memory usage changes significantly
22//! * Catch and prevent single allocations greater than say 64GB, dump out giant allocation stack trace
23//!
24//! To see an example of a long-running toy service that generates the periodic reports and flamegraphs:
25//!
26//! `cargo run --profile bench --example ying_example`
27//!
28//! The example runs a cache workload plus a simulated leak for ~12 minutes โ€” long enough for the
29//! reporting thread to write a report and flamegraph under `ying-profiles/` at every check.  For a
30//! quicker demo:
31//!
32//! `YING_EXAMPLE_INTERVAL_SECS=10 YING_EXAMPLE_RUNTIME_SECS=45 cargo run --profile bench --example ying_example`
33//!
34//! ## How to use
35//!
36//! Set Ying as the global allocator, then turn on reporting with one line:
37//!
38//! ```no_run
39//! use ying_profiler::YingProfiler;
40//!
41//! #[global_allocator]
42//! static YING_ALLOC: YingProfiler = YingProfiler::default();
43//!
44//! fn main() {
45//!     YING_ALLOC.start_profiling();
46//!     // ... the rest of your program
47//! }
48//! ```
49//!
50//! There are three layers here, and it is worth knowing which one you are using.
51//!
52//! ### 1. The allocator: collects, never reports
53//!
54//! The `#[global_allocator]` declaration is what makes Ying sample allocations at all.  On its own
55//! it is complete and valid: Ying accumulates stats in memory and defers to the System allocator,
56//! but writes nothing anywhere.  Nothing is scheduled and no thread is spawned.
57//!
58//! ### 2. [`ProfilerRunner`]: the primitive that decides how stats get reported
59//!
60//! [`ProfilerRunner`] is the piece to reach for whenever the defaults do not fit.  It owns every
61//! decision about turning collected stats into output โ€” how often to look, how much of a change is
62//! worth reporting, what to measure, where output goes, and in what form:
63//!
64//! | Setting | Meaning |
65//! |---|---|
66//! | `check_interval_secs` | how often the background thread wakes up to compare memory use |
67//! | `report_pct_change_trigger` | how much retained memory must move before a report is written |
68//! | `reporting_path` | directory for reports and flamegraphs; created if missing |
69//! | `measure_allocated_not_retained` | rank stacks by total allocated bytes instead of retained |
70//! | `gen_flamegraphs` | also write an SVG flamegraph alongside each text report |
71//! | `expand_frames` | expand inlined symbols within each stack frame in reports |
72//!
73//! Build one with [`ProfilerRunnerBuilder`](utils::ProfilerRunnerBuilder), which defaults
74//! anything you leave out, then hand it
75//! the allocator static to start its thread:
76//!
77//! ```no_run
78//! use ying_profiler::{utils::ProfilerRunnerBuilder, YingProfiler};
79//!
80//! #[global_allocator]
81//! static YING_ALLOC: YingProfiler = YingProfiler::default();
82//!
83//! fn main() {
84//!     let runner = ProfilerRunnerBuilder::default()
85//!         .check_interval_secs(60usize)
86//!         .report_pct_change_trigger(5usize)
87//!         .reporting_path("/var/log/ying")
88//!         .gen_flamegraphs(true)
89//!         .build()
90//!         .unwrap();
91//!     runner.spawn(&YING_ALLOC);
92//! }
93//! ```
94//!
95//! [`YingProfiler::start_profiling`] is not a separate mechanism: it builds exactly one of these
96//! with a specific set of values (5 minute interval, 10% trigger, flamegraphs on, retained memory,
97//! output to `ying-profiles`) and spawns it.  Anything it can do, a [`ProfilerRunner`] you build
98//! yourself can do too.
99//!
100//! ### 3. Reading stats directly
101//!
102//! You do not need a runner at all if you would rather decide when to look.  [`YingProfiler`]
103//! exposes the stats as data โ€” [`YingProfiler::top_k_stacks_by_retained`],
104//! [`YingProfiler::top_k_stacks_by_allocated`], [`YingProfiler::total_retained_bytes`] โ€” and
105//! [`utils::gen_flamegraph`] writes a one-off flamegraph on demand.  This is the route to take when
106//! reports should be triggered by your own signals, such as an HTTP endpoint or a health check
107//! noticing memory growth.
108//!
109//! ## Why a new memory profiler?
110//!
111//! Rust as an ecosystem is lacking in good memory profiling tools.  [Bytehound](https://github.com/koute/bytehound)
112//! is quite good but has a large CPU impact and writes out huge profiling files as it measures every allocation.
113//! Jemalloc/Jeprof does sampling, so it's great for production use, but its output is difficult to interpret, and
114//! it does not track retained memory, which is quite critical for debugging memory issues in production.
115//! Both of the above tools are written with generic C/C++ malloc/preload ABI in mind, so the backtraces
116//! that one gets from their use are really limited, especially for profiling Rust binaries built for
117//! an optimized/release target, and especially async code.  The output also often has trouble with mangled symbols.
118//!
119//! If we use the backtrace crate and analyze release/bench backtraces in detail, we can see why that is.
120//! The Rust compiler does a good job of inlining function calls - even ones across async/await boundaries -
121//! in release code.  Thus, for a single instruction pointer (IP) in the stack trace, it might correspond to
122//! many different places in the code.  This is from `examples/ying_example.rs`:
123//!
124//! ```bash
125//! Some(ying_example::insert_one::{{closure}}::h7eddb5f8ebb3289b)
126//!  > Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::h7a53098577c44da0)
127//!  > Some(ying_example::cache_update_loop::{{closure}}::h38556c7e7ae06bfa)
128//!  > Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::hd319a0f603a1d426)
129//!  > Some(ying_example::main::{{closure}}::h33aa63760e836e2f)
130//!  > Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::hb2fd3cb904946c24)
131//! ```
132//!
133//! A generic tool which just examines the IP and tries to figure out a single symbol would miss out on all of the
134//! inlined symbols.  Some tools can expand on symbols, but the results still aren't very good.
135//!
136//! ## Feature Flags
137//!
138//! - `profile_spans` - records tracing-span information in stacks.  NOTE: this feature is
139//!   experimental and incomplete.
140//!
141use std::alloc::{GlobalAlloc, Layout, System};
142use std::cell::Cell;
143use std::cmp::min;
144use std::env::var;
145use std::fmt::Write;
146use std::mem::needs_drop;
147use std::ptr::{copy_nonoverlapping, null_mut};
148use std::sync::atomic::{AtomicUsize, Ordering::Relaxed, Ordering::SeqCst};
149
150use backtrace::Backtrace;
151use coarsetime::Clock;
152use once_cell::sync::OnceCell;
153
154pub mod callstack;
155pub mod histogram;
156mod sharded_map;
157pub mod utils;
158use callstack::{FriendlySymbol, StackStats, StdCallstack};
159use sharded_map::ShardedMap;
160use utils::{
161    ProfilerRunner, DEFAULT_CHECK_INTERVAL_SECS, DEFAULT_PCT_CHANGE_TRIGGER, DEFAULT_REPORTING_PATH,
162};
163
164/// The number of frames at the top of the stack to skip.  Most of these have to do with
165/// backtrace and this profiler infrastructure.  This number needs to be adjusted
166/// depending on the implementation.
167/// For release/bench builds with debug = 1 / strip = none, this should be 4.
168/// For debug builds, this is about 9.
169const TOP_FRAMES_TO_SKIP: usize = 3;
170
171const DEFAULT_GIANT_ALLOC_LIMIT: usize = 64 * 1024 * 1024 * 1024;
172
173/// Testing knob: overrides the initial capacity of every profiler map.  Not part of the public
174/// API, and read exactly once when profiler state is first initialized.
175const MAP_CAPACITY_ENV_VAR: &str = "YING_TEST_MAP_CAPACITY";
176
177// A map for caching symbols in backtraces so we can mostly store u64's
178pub(crate) type SymbolMap = ShardedMap<Vec<FriendlySymbol>>;
179
180/// Ying is a memory profiling Allocator wrapper.
181/// Ying is the Chinese word for an eagle.
182pub struct YingProfiler {
183    /// Allocation sampling ratio.  Eg: 500 means 1 in 500 allocations are sampled.
184    sampling_ratio: u32,
185    /// Prevent and dump stack trace for giant single allocations beyond a certain size
186    single_alloc_limit: usize,
187    /// Global thread local state cache
188    /// Statistics... lazily initialized later
189    state: OnceCell<YingState>,
190}
191
192static TOTAL_RETAINED: AtomicUsize = AtomicUsize::new(0);
193static PROFILED_ALLOCATED: AtomicUsize = AtomicUsize::new(0);
194static PROFILED_RETAINED: AtomicUsize = AtomicUsize::new(0);
195
196impl YingProfiler {
197    /// sampling_ratio: number of allocations for every sampled allocation
198    pub const fn new(sampling_ratio: u32, single_alloc_limit: usize) -> Self {
199        Self {
200            sampling_ratio,
201            single_alloc_limit,
202            state: OnceCell::new(),
203        }
204    }
205
206    pub const fn default() -> Self {
207        Self {
208            sampling_ratio: 500,
209            single_alloc_limit: DEFAULT_GIANT_ALLOC_LIMIT,
210            state: OnceCell::new(),
211        }
212    }
213
214    /// Starts periodic memory reporting on a background thread and returns the [`ProfilerRunner`]
215    /// that was spawned.  This is the one-line way to turn Ying on:
216    ///
217    /// ```no_run
218    ///     use ying_profiler::YingProfiler;
219    ///
220    ///     #[global_allocator]
221    ///     static YING_ALLOC: YingProfiler = YingProfiler::default();
222    ///
223    ///     fn main() {
224    ///         YING_ALLOC.start_profiling();
225    ///     }
226    /// ```
227    ///
228    /// This is a convenience only, and adds no capability of its own: it builds a
229    /// [`ProfilerRunner`] with a fixed set of values and calls [`ProfilerRunner::spawn`].  Those
230    /// values are a 5 minute [`DEFAULT_CHECK_INTERVAL_SECS`] check interval, a 10%
231    /// [`DEFAULT_PCT_CHANGE_TRIGGER`] change before a report is written, output to
232    /// [`DEFAULT_REPORTING_PATH`] (created if it does not exist), flamegraphs on, retained rather
233    /// than allocated memory, and inlined frames left unexpanded.
234    ///
235    /// To change any of those, build the [`ProfilerRunner`] yourself with
236    /// [`ProfilerRunnerBuilder`](utils::ProfilerRunnerBuilder) and call
237    /// [`spawn`](ProfilerRunner::spawn) on it instead.  The returned runner records the settings in
238    /// use, but the reporting thread has already been started with them; changing the returned
239    /// value has no effect.
240    pub fn start_profiling(&'static self) -> ProfilerRunner {
241        let runner = ProfilerRunner::new(
242            DEFAULT_CHECK_INTERVAL_SECS,
243            DEFAULT_PCT_CHANGE_TRIGGER,
244            DEFAULT_REPORTING_PATH,
245            false,
246            true,
247            false,
248        );
249        runner.spawn(self);
250        runner
251    }
252
253    /// Total outstanding retained bytes (not just sampled but all allocations)
254    #[inline]
255    pub fn total_retained_bytes() -> usize {
256        TOTAL_RETAINED.load(Relaxed)
257    }
258
259    /// Total bytes allocated for profiled allocations
260    #[inline]
261    pub fn profiled_bytes_allocated() -> usize {
262        PROFILED_ALLOCATED.load(Relaxed)
263    }
264
265    /// Profiled bytes retained - retained memory usage amongst profiled allocations
266    #[inline]
267    pub fn profiled_bytes_retained() -> usize {
268        PROFILED_RETAINED.load(Relaxed)
269    }
270
271    #[inline]
272    pub fn symbol_map_size(&self) -> usize {
273        self.lock_out_profiler(|| self.get_state().symbol_map.len())
274    }
275
276    /// Number of entries for outstanding sampled allocations map
277    #[inline]
278    pub fn num_outstanding_allocs(&self) -> usize {
279        self.lock_out_profiler(|| self.get_state().outstanding_allocs.len())
280    }
281
282    /// Number of distinct stack traces that have been sampled
283    #[inline]
284    pub fn num_stack_traces(&self) -> usize {
285        self.lock_out_profiler(|| self.get_state().stack_stats.len())
286    }
287
288    /// Get the top k stack traces by total profiled bytes allocated, in descending order.
289    /// Note that "profiled bytes" refers to the bytes allocated during sampling by this profiler.
290    pub fn top_k_stacks_by_allocated(&self, k: usize) -> Vec<StackStats> {
291        self.lock_out_profiler(|| {
292            let stacks_by_alloc = self.stack_list_allocated_bytes_desc();
293            stacks_by_alloc
294                .iter()
295                .take(k)
296                .filter_map(|&(stack_hash, _bytes_allocated)| {
297                    self.get_stats_for_stack_hash(stack_hash)
298                })
299                .collect()
300        })
301    }
302
303    /// Get the top k stack traces by retained sampled memory, in descending order.
304    pub fn top_k_stacks_by_retained(&self, k: usize) -> Vec<StackStats> {
305        self.lock_out_profiler(|| {
306            let stacks_by_retained = self.stack_list_retained_bytes_desc();
307            stacks_by_retained
308                .iter()
309                .take(k)
310                .filter_map(|&(stack_hash, _bytes_retained)| {
311                    self.get_stats_for_stack_hash(stack_hash)
312                })
313                .collect()
314        })
315    }
316
317    /// Returns a list of stack IDs (stack_hash, bytes_allocated) in order from highest
318    /// number of bytes allocated to lowest
319    fn stack_list_allocated_bytes_desc(&self) -> Vec<(u64, u64)> {
320        // TODO: filter away entries with minimal allocations, say <1% or some threshold
321        let mut items = self
322            .get_state()
323            .stack_stats
324            .map_to_vec(|stack_hash, stats| (stack_hash, stats.allocated_bytes));
325        items.sort_unstable_by_key(|&(_, bytes)| std::cmp::Reverse(bytes));
326        items
327    }
328
329    /// Returns a list of stack IDs (stack_hash, bytes_retained) in order from highest
330    /// number of bytes retained to lowest
331    fn stack_list_retained_bytes_desc(&self) -> Vec<(u64, u64)> {
332        // TODO: filter away entries with minimal retained allocations, say <1% or some threshold
333        let mut items = self
334            .get_state()
335            .stack_stats
336            .map_to_vec(|stack_hash, stats| (stack_hash, stats.retained_profiled_bytes()));
337        items.sort_unstable_by_key(|&(_, bytes)| std::cmp::Reverse(bytes));
338        items
339    }
340
341    pub fn reset_state_for_testing_only(&self) {
342        self.lock_out_profiler(|| {
343            let state = self.get_state();
344            state.stack_stats.clear();
345            state.outstanding_allocs.clear();
346        })
347    }
348
349    pub fn testing_only_guarantee_next_sample(&self) {
350        THREAD_STATE.with(YingThreadLocal::test_only_reset_sampling_counter)
351    }
352
353    #[inline]
354    fn check_and_deny_giant_allocations(&self, ptr: *mut u8, layout: Layout) -> *mut u8 {
355        // Sorry there is an edge case where this check cannot happen if YING is not initialized
356        if layout.size() >= self.single_alloc_limit && self.state.get().is_some() {
357            // Prevent allocation sampling while we are telling the world who did this
358            self.lock_out_profiler(|| {
359                println!(
360                    "WARNING: Huge memory allocation of {} bytes denied by Ying profiler",
361                    layout.size()
362                );
363
364                let mut bt = Backtrace::new_unresolved();
365
366                // 2. Create a Callstack, check if there is a similar stack
367                let stack = StdCallstack::from_backtrace_unresolved(&bt);
368                let state = self.get_state();
369                stack.populate_symbol_map(&mut bt, &state.symbol_map);
370                println!(
371                    "Stack trace:\n{}",
372                    stack.with_symbols_and_filename(&state.symbol_map, true)
373                );
374            });
375            null_mut::<u8>()
376        } else {
377            ptr
378        }
379    }
380
381    #[inline]
382    fn get_state(&self) -> &YingState {
383        // We need to lock out the profiler here, to ensure no tracking of allocations or messes
384        self.lock_out_profiler(|| self.state.get_or_init(YingState::new))
385    }
386
387    #[inline]
388    fn get_stats_for_stack_hash(&self, stack_hash: u64) -> Option<StackStats> {
389        self.get_state().stack_stats.get_cloned(stack_hash)
390    }
391
392    /// Locks the profiler flag so that allocations are not profiled.
393    /// This is for non-profiler code such as debug prints that has to access the Dashmap or state
394    /// and could potentially cause deadlock problems with Dashmap for example.
395    #[inline]
396    fn lock_out_profiler<R>(&self, func: impl FnOnce() -> R) -> R {
397        // Deliberately three short thread local accesses rather than running `func` inside a
398        // `with`: `func` frequently calls back in here (get_state does), and keeping the borrow
399        // scopes disjoint means nesting needs no reasoning about re-entrant `with`.
400        THREAD_STATE.with(YingThreadLocal::set_allocator_lock);
401        let return_val = func();
402        exit_profiler();
403        return_val
404    }
405}
406
407// Private state.  We can't put this in the main YingProfiler struct as that one has to be const static
408struct YingState {
409    symbol_map: SymbolMap,
410    // Main map of stack hash to StackStats
411    stack_stats: ShardedMap<StackStats>,
412    // Map of outstanding sampled allocations.  Used to figure out amount of outstanding allocations and
413    // statistics about how long lived outstanding allocations are.
414    // (*ptr as u64 -> (stack hash, start_timestamp_epoch_millis))
415    outstanding_allocs: ShardedMap<(u64, u64)>,
416}
417
418impl YingState {
419    pub fn new() -> Self {
420        // Shrinking the maps makes them resize constantly, which is how the re-entrancy deadlocks
421        // were shaken out; the stress tests set this so those paths get exercised on every run.
422        let capacity_scale = var(MAP_CAPACITY_ENV_VAR)
423            .ok()
424            .and_then(|s| s.parse::<usize>().ok());
425        let symbol_map = SymbolMap::with_capacity(capacity_scale.unwrap_or(1000));
426        let stack_stats = ShardedMap::with_capacity(capacity_scale.unwrap_or(1000));
427        let outstanding_allocs = ShardedMap::with_capacity(capacity_scale.unwrap_or(5000));
428        Self {
429            symbol_map,
430            stack_stats,
431            outstanding_allocs,
432        }
433    }
434}
435
436thread_local! {
437    /// Per-thread profiler state, in a thread local rather than a slot in a fixed-size table.
438    ///
439    /// An earlier design hashed the thread id into a 1024 entry array and handed out `&mut` to the
440    /// slot, on the reasoning that a `thread_local!` cannot be used from a `GlobalAlloc` because
441    /// setting up TLS may allocate and re-enter the allocator.  That reasoning was sound for a lazily
442    /// initialized thread local, but two threads whose ids collide then share one slot, which is both
443    /// aliasing UB and a lost re-entrancy guard: a non-atomic increment of `alloc_lock` can be dropped,
444    /// the guard falls to zero while a thread is still inside its critical section, and its next
445    /// internal free re-enters the maps and blocks on a shard lock it already holds.  That deadlock was
446    /// observed, and the captured backtrace showed exactly that shape.
447    ///
448    /// The `const {}` initializer plus a type that needs no `Drop` is what makes this safe: it skips
449    /// lazy initialization and destructor registration, so no Rust allocation happens on access.  On
450    /// ELF targets linked into an executable this compiles to a single thread-pointer-relative load.
451    /// macOS still routes through a resolver call that allocates the thread's TLV block on first touch,
452    /// but that allocation is libc `malloc` inside dyld, which `#[global_allocator]` does not
453    /// interpose, so it cannot re-enter Ying.  Measured: 21,225 accesses from inside `alloc` and
454    /// `dealloc` across 50 freshly spawned threads and their teardown, with zero re-entry.
455    ///
456    /// Keep both invariants if you touch this.  Adding a field that needs dropping, or dropping the
457    /// `const {}`, reintroduces an allocating first access and with it the deadlock.
458    static THREAD_STATE: YingThreadLocal = const { YingThreadLocal::new() };
459}
460
461/// Enforces the invariant that [`THREAD_STATE`] documents.  A field that needs dropping would make
462/// the thread local register a destructor, which allocates on first access and re-enters the
463/// allocator, so this is a compile error rather than a comment nobody reads.
464const _: () = assert!(
465    !needs_drop::<YingThreadLocal>(),
466    "YingThreadLocal must not need Drop, or thread local setup will allocate inside the allocator"
467);
468
469/// A struct to provide a better API around the lock out profiler flag/re-entrancy plus sampling.
470/// Interior mutability via [`Cell`] rather than `&mut`, so that no field needs `Drop` and the
471/// thread local stays on its cheapest code path.
472struct YingThreadLocal {
473    // Counts up for every time we enter a no-allocator critical section (ie where we have to touch
474    // allocator state or cause an allocation within profiling-related code and don't want sampling
475    // of re-entrant allocations done).  Nonzero prevents allocator from sampling.
476    alloc_lock: Cell<u32>,
477    sample_count: Cell<u32>,
478}
479
480impl YingThreadLocal {
481    const fn new() -> Self {
482        Self {
483            alloc_lock: Cell::new(0),
484            sample_count: Cell::new(0),
485        }
486    }
487
488    #[inline]
489    fn is_allocator_locked(&self) -> bool {
490        self.alloc_lock.get() > 0
491    }
492
493    #[inline]
494    fn set_allocator_lock(&self) {
495        self.alloc_lock.set(self.alloc_lock.get().saturating_add(1));
496    }
497
498    #[inline]
499    fn release_allocator_lock(&self) {
500        self.alloc_lock.set(self.alloc_lock.get().saturating_sub(1));
501    }
502
503    /// Obtains the counter, checks for sampling ratio, and updates counter in one go
504    #[inline]
505    fn should_sample(&self, ratio: u32) -> bool {
506        let count = self.sample_count.get().wrapping_add(1);
507        self.sample_count.set(count);
508        count % ratio == 0
509    }
510
511    // Resets counter to 0 to guarantee next call to alloc() will sample.  TESTING ONLY
512    #[inline]
513    fn test_only_reset_sampling_counter(&self) {
514        self.sample_count.set(0);
515    }
516}
517
518/// Takes the re-entrancy guard if this thread is not already inside the profiler, reporting whether
519/// the caller now owns it and must release it.
520///
521/// Deliberately one thread local access rather than several: this runs on every allocation, and the
522/// common answer is "do not profile this one".
523#[inline]
524fn try_enter_profiler(sampling_ratio: Option<u32>) -> bool {
525    THREAD_STATE.with(|tl| {
526        if tl.is_allocator_locked() {
527            return false;
528        }
529        // Only consume a sampling tick when we were not already locked out, matching the original
530        // short-circuiting behaviour
531        if let Some(ratio) = sampling_ratio {
532            if !tl.should_sample(ratio) {
533                return false;
534            }
535        }
536        tl.set_allocator_lock();
537        true
538    })
539}
540
541#[inline]
542fn exit_profiler() {
543    THREAD_STATE.with(YingThreadLocal::release_allocator_lock)
544}
545
546unsafe impl GlobalAlloc for YingProfiler {
547    unsafe fn alloc(&self, layout: Layout) -> *mut u8 {
548        // NOTE: the code between here and the state.0 = true must be re-entrant
549        // and therefore not allocate, otherwise there will be an infinite loop.
550        let alloc_ptr = self.check_and_deny_giant_allocations(System.alloc(layout), layout);
551        if !alloc_ptr.is_null() {
552            TOTAL_RETAINED.fetch_add(layout.size(), SeqCst);
553
554            // Now, sample allocation - if it falls below threshold, then profile
555            // Also, we set a ThreadLocal to avoid re-entry: ie the code below might allocate,
556            // and we avoid profiling if we are already in the loop below.  Avoids cycles.
557            if try_enter_profiler(Some(self.sampling_ratio)) {
558                PROFILED_ALLOCATED.fetch_add(layout.size(), SeqCst);
559                PROFILED_RETAINED.fetch_add(layout.size(), SeqCst);
560
561                // -- Beginning of section that may allocate
562                // 1. Get unresolved backtrace for speed
563                let mut bt = Backtrace::new_unresolved();
564
565                // 2. Create a Callstack, check if there is a similar stack
566                let stack = StdCallstack::from_backtrace_unresolved(&bt);
567                let stack_hash = stack.compute_hash();
568                let state = self.get_state();
569                state.stack_stats.update_or_insert_with(
570                    stack_hash,
571                    |stats| {
572                        // 4. Update stats
573                        stats.num_allocations += 1;
574                        stats.allocated_bytes += layout.size() as u64;
575                    },
576                    || {
577                        // 3. Resolve symbols if needed (new stack entry).
578                        // NOTE: this runs while the stack_stats shard is locked, but it only ever
579                        // touches symbol_map, which is a separate map, so it cannot self-deadlock.
580                        stack.populate_symbol_map(&mut bt, &state.symbol_map);
581                        StackStats::new(stack, Some(layout.size() as u64))
582                    },
583                );
584
585                // 4. Record allocation so we can track outstanding vs transient allocs
586                state.outstanding_allocs.update_or_insert_with(
587                    alloc_ptr as u64,
588                    |_existing| {},
589                    || (stack_hash, Clock::recent_since_epoch().as_millis()),
590                );
591
592                // -- End of core profiling section, no more allocations --
593                exit_profiler();
594            }
595        }
596        alloc_ptr
597    }
598
599    unsafe fn dealloc(&self, ptr: *mut u8, layout: Layout) {
600        System.dealloc(ptr, layout);
601        TOTAL_RETAINED.fetch_sub(layout.size(), SeqCst);
602
603        // Return immediately and skip rest of this if YING_STATE is not initialized.  It could cause
604        // an infinite loop because during initialization of YING_STATE, dealloc() could be then called
605        if self.state.get().is_none() {
606            return;
607        }
608
609        // IMPORTANT: the re-entrancy check has to happen before we touch any of the profiler maps.
610        // The maps allocate and free internally (eg when a shard resizes), and those frees land back
611        // here while that shard's write lock is held.  Touching the map first would then try to take
612        // a read lock the same thread already holds exclusively, which self-deadlocks.
613        if !try_enter_profiler(None) {
614            return;
615        }
616
617        // -- Beginning of section that may allocate
618        // If the allocation was recorded in outstanding_allocs, then remove it and update stats
619        // about number of bytes freed etc.
620        let state = self.get_state();
621        if state.outstanding_allocs.contains_key(ptr as u64) {
622            if let Some((stack_hash, alloc_ts)) = state.outstanding_allocs.remove(ptr as u64) {
623                PROFILED_RETAINED.fetch_sub(layout.size(), SeqCst);
624                let alloc_time_ms = Clock::recent_since_epoch()
625                    .as_millis()
626                    .saturating_sub(alloc_ts);
627
628                // Update memory profiling freed bytes stats
629                state.stack_stats.update(stack_hash, |stats| {
630                    stats.update_free_stats(layout.size() as u64, alloc_time_ms)
631                });
632            }
633        }
634
635        // -- End of core profiling section, no more allocations --
636        exit_profiler();
637    }
638
639    // We implement a custom realloc().  We must count reallocs as the same allocation, but need to do
640    // the following: - update original allocated bytes (but not allocations); move outstanding_allocs
641    // because the pointer moved, but preserve original starting timestamp.
642    // The above also saves us cycles from having to call alloc() and dealloc() separately.
643    unsafe fn realloc(&self, ptr: *mut u8, layout: Layout, new_size: usize) -> *mut u8 {
644        let old_size = layout.size();
645        // SAFETY: the caller must ensure that the `new_size` does not overflow.
646        // `layout.align()` comes from a `Layout` and is thus guaranteed to be valid.
647        let new_layout = Layout::from_size_align_unchecked(new_size, layout.align());
648        // SAFETY: the caller must ensure that `new_layout` is greater than zero.
649        let new_ptr = self.check_and_deny_giant_allocations(System.alloc(new_layout), new_layout);
650        if !new_ptr.is_null() {
651            // SAFETY: the previously allocated block cannot overlap the newly allocated block.
652            // The safety contract for `dealloc` must be upheld by the caller.
653            copy_nonoverlapping(ptr, new_ptr, min(old_size, new_size));
654            System.dealloc(ptr, layout);
655
656            // 1. Update global statistics
657            if new_size > old_size {
658                TOTAL_RETAINED.fetch_add(new_size - old_size, SeqCst);
659            } else {
660                TOTAL_RETAINED.fetch_sub(old_size - new_size, SeqCst);
661            }
662
663            // 2. IF the old pointer was in outstanding_allocs, move it and make a new entry,
664            //    keeping the old starting timestamp.  Also update stack stats.
665            //    But only if state is alredy initialized - otherwise any state initialization that
666            //    results in a realloc() could cause this to infinite loop
667            if self.state.get().is_none() {
668                return new_ptr;
669            }
670
671            // As in dealloc(), the re-entrancy check must come before any map access, or a map's own
672            // internal realloc re-enters here while holding that shard's write lock and deadlocks.
673            if !try_enter_profiler(None) {
674                return new_ptr;
675            }
676
677            // -- Beginning of section that may allocate
678            let state = self.get_state();
679            if state.outstanding_allocs.contains_key(ptr as u64) {
680                if let Some((stack_hash, alloc_ts)) = state.outstanding_allocs.remove(ptr as u64) {
681                    if new_size > old_size {
682                        PROFILED_RETAINED.fetch_add(new_size - old_size, SeqCst);
683                    } else {
684                        PROFILED_RETAINED.fetch_sub(old_size - new_size, SeqCst);
685                    }
686
687                    state
688                        .outstanding_allocs
689                        .insert(new_ptr as u64, (stack_hash, alloc_ts));
690
691                    // Update memory profiling freed bytes stats
692                    state.stack_stats.update(stack_hash, |stats| {
693                        if new_size > old_size {
694                            stats.allocated_bytes += (new_size - old_size) as u64;
695                        } else {
696                            stats.allocated_bytes -= (old_size - new_size) as u64;
697                        }
698                        // Don't change number of allocations or frees
699                    });
700                }
701            }
702
703            // -- End of core profiling section, no more allocations --
704            exit_profiler();
705        }
706        new_ptr
707    }
708}