1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
//! Self-reported peak resident-set-size (RSS) for the `xberg extract --format json` envelope.
//!
//! The benchmark harness (`tools/benchmark-harness`) compares memory usage across frameworks by
//! reading a `_peak_memory_bytes` field that every competitor's Python wrapper self-reports via
//! `resource.getrusage(resource.RUSAGE_SELF).ru_maxrss` — a kernel-tracked high-water mark that
//! cannot miss a transient allocation spike. The xberg CLI previously emitted no such field, so
//! the harness fell back entirely to its own `sysinfo`-based sampler (a 1-10ms polling loop, see
//! `tools/benchmark-harness/src/monitoring.rs`), which can miss allocations shorter-lived than the
//! sampling interval. That asymmetry biased the published memory comparison in xberg's favor: the
//! same class of transient spike was always visible to competitors and sometimes invisible for
//! xberg. This module closes the gap by reporting the same kernel-tracked `ru_maxrss` value the
//! competitors already report, under the same measurement method.
/// Pure conversion from a raw `ru_maxrss` value to bytes, parameterized explicitly by whether the
/// value came from a Linux `getrusage` call.
///
/// Kept separate from [`ru_maxrss_to_bytes`] (which pins `is_linux` to the actual build target via
/// `cfg!`) so both unit conventions below can be exercised by unit tests on any host platform,
/// without needing `#[cfg(target_os = ...)]` on the tests themselves.
///
/// Gated on `cfg(any(test, unix))` rather than plain `cfg(unix)`: the only non-test caller is
/// [`ru_maxrss_to_bytes`], itself only called from the `cfg(unix)` `peak_memory_bytes`, so on a
/// non-test, non-unix build (e.g. Windows release) this would otherwise be genuinely unused. The
/// `test` half of the predicate keeps it compiled for unit tests on every host platform, per the
/// doc comment above.
/// Converts a raw `ru_maxrss` value (as read directly from `getrusage(2)`) to bytes.
///
/// # Platform units (the actual trap)
/// - **Linux**: `ru_maxrss` is reported in **kibibytes** — must be multiplied by 1024.
/// - **macOS / other BSD-derived libc**: `ru_maxrss` is reported in **bytes** already — must NOT
/// be scaled.
///
/// Mirrors the equivalent conversion in the benchmark harness's Python competitor wrappers (see
/// `tools/benchmark-harness/scripts/docling_extract.py::_get_peak_memory_bytes`), so both sides of
/// a memory comparison agree on units instead of one side silently being off by 1024x.
///
/// Gated on `cfg(any(test, unix))` for the same reason as [`convert_ru_maxrss`]: its only
/// non-test caller is the `cfg(unix)` `peak_memory_bytes`, so it is genuinely unused on a
/// non-test, non-unix (e.g. Windows release) build.
/// Returns this process's peak RSS in bytes via `getrusage(RUSAGE_SELF)`, or `None` if the
/// syscall failed (should not happen in practice on Unix).
///
/// The returned value covers the whole process's lifetime up to the call site, matching the
/// semantics of the competitors' `ru_maxrss` self-report described in the module docs.
/// Non-Unix platforms have no `getrusage`; report "unavailable" rather than guessing.