pub fn reserve_aligned_huge(size: usize, align: usize) -> Option<Reservation>huge-pages only.Expand description
Reserve size bytes aligned to align, requesting OS large / huge
pages (Linux/Android MAP_HUGETLB, Windows MEM_LARGE_PAGES).
Currently a no-op on macOS and other Unix that is neither Linux nor
Android — it falls back to
an ordinary reservation, identical to reserve_aligned.
Transparent-huge-page hinting (Linux/Android MADV_HUGEPAGE) is not used: it
cannot affect an already-explicitly-huge MAP_HUGETLB mapping (the pages
are already huge), so issuing it would be a wasted syscall. This crate’s
strategy is the explicit hugetlbfs path only.
Large pages reduce TLB pressure for big allocator segments. The request is best-effort: if the OS refuses large pages (none configured, no privilege), the reservation transparently falls back to ordinary pages, so this never fails purely because huge pages are unavailable — it fails only on a genuine reservation error (OOM) or a contract violation.
To detect whether huge pages were actually granted (as opposed to having
fallen back to ordinary pages), use the returned Reservation::is_huge
method.
Base/align/size contract is otherwise identical to reserve_aligned,
except on Linux AND Android with huge-pages enabled (the check’s
cfg is any(target_os = "linux", target_os = "android") + feature = "huge-pages"): size and align must BOTH
additionally be multiples of the huge-page size (2 MiB) — a request
that only satisfies reserve_aligned’s own weaker PAGE-multiple contract
is rejected with VmemError::invalid_argument() before any syscall runs,
even though such a request could previously succeed there via the documented
ordinary-page fallback. For the failure cause use
try_reserve_aligned_huge.
Windows limitation: on Windows, this function returns a reservation with
Reservation::is_huge == true only when ALL of the following hold:
- The fast-path condition
align <= GetLargePageMinimum()is satisfied (typicallyalign <= 2 MiBon x86_64) sizeis a multiple of the system’s large-page minimum- The calling process has
SeLockMemoryPrivilegegranted AND has enabled it viaAdjustTokenPrivileges(the crate does not do this for you — a process with the privilege granted but not enabled fails exactly like an unprivileged one and silently falls back to ordinary pages)
NOTE: The widened fast-path condition (II-3, 2026-08-16 audit finding) expanded
the single-call ATTEMPT window from align <= 64 KiB to align <= GetLargePageMinimum(),
but on an unprivileged host the actual paths that SUCCEED (pass the post-call alignment
check) are typically still limited. When large pages are NOT granted (unprivileged),
VirtualAlloc’s alignment guarantee is only 64 KiB; in practice it typically does NOT
happen to land on the requested alignment, so the post-call check fails and the fast
path falls through to the two-call path. Practically, this means is_huge() == true only
for shapes where large pages are actually granted, which requires all three conditions
above to hold.
Extra-syscall cost on unprivileged hosts: For the widened align range
(64 KiB < align <= GetLargePageMinimum()), when large pages are requested but
not granted (e.g., unprivileged process, or SeLockMemoryPrivilege not enabled),
the code attempts VirtualAlloc with MEM_LARGE_PAGES (fails), retries without
it (succeeds with ordinary pages), and if that retry’s base doesn’t happen to
satisfy the requested alignment, the whole thing is released and falls through to
the two-call path. This means an unprivileged reservation in this align range
can cost up to 2 extra VirtualAlloc calls + 1 VirtualFree before reaching
the two-call path, versus before the II-3 change (which would have gone straight
to the two-call path for align > 64 KiB). This is a real, measurable behavior
change, not a correctness bug — the widening genuinely expands the single-call
attempt window, and unprivileged processes pay the extra-syscall cost for shapes
that now attempt but fail the fast path.
If any of these conditions fail, the function falls back to ordinary
pages and returns a reservation with Reservation::is_huge == false.
On Windows, large pages (MEM_LARGE_PAGES) are only ever requested and
possibly granted via the single-call fast path; the two-call path never requests
large pages, so the result never has
Reservation::is_huge == true.
Decommit incompatibility (corrected task #1140): on Windows,
decommit/decommit_lazy never work on huge-page
reservations — VirtualFree with MEM_DECOMMIT unconditionally fails on
large-page regions. decommit_lazy never works on a huge-page
reservation on ANY platform — MADV_FREE (its Linux/Android backend) has
no documented HugeTLB support, unlike MADV_DONTNEED below.
decommit on Linux/Android is more nuanced: it depends on BOTH the
requested range and the running kernel. madvise(2) documents that
MADV_DONTNEED gained HugeTLB support in Linux 5.18, requiring
[base+start, base+end) to be aligned to the mapping’s huge page size (2
MiB) at both endpoints — the same alignment this function already requires
of size/align themselves on Linux/Android, so decommitting an entire
huge reservation, or any 2-MiB-granular sub-range of it, is exactly such an
eligible range on a >= 5.18 kernel. A page_size()-granular (e.g. 4 KiB)
but not 2-MiB-granular offset still gets EINVAL and does nothing, as does
EVERY range on a pre-5.18 kernel. Reservation::decommit/
Reservation::try_decommit (the safe methods) consult both
Reservation::is_huge and the requested range to skip the ineligible
case before issuing the syscall; the free decommit function has no
is_huge() to consult and issues the syscall unconditionally — see that
function’s own doc for the precise split. Either way — an ineligible
range, or any range on a pre-5.18 kernel — the effect is indistinguishable
from a silent no-op: the caller’s RSS does not decrease, and subsequent
reads return the old (stale) data rather than zeroed pages.
Documented per the madvise(2) man page, and — since task #1152 (F1) —
empirically exercised under a real hugetlb pool by this crate’s own CI:
the aligned-vmem-hugetlb-real job (.github/workflows/ci.yml) hard-
asserts (via a path-activation oracle) that this function actually
received a MAP_HUGETLB grant, then drives a huge-page-eligible range
through Reservation::decommit’s eligible-huge branch. What that
job proves, stated precisely (task #1160/F1 correction of an earlier
overclaim; strengthened tasks #1164 and #1174): the eligible-range/post-5.18-kernel
case genuinely REACHES the real madvise(2)/MADV_DONTNEED backend call
— AND, since task #1164’s
ci_hugetlb_real_pool_kernel_actually_accepts_eligible_madvise
(tests/decommit_capability.rs), that the kernel itself returned 0
(accepted) for that call, not -1 (rejected): under bench-internals,
libc_madvise (src/os/unix.rs) records the syscall’s own return value
into a counter pair, and that job hard-asserts it increased for this
eligible-range case — AND, since task #1174’s
ci_hugetlb_real_pool_decommit_actually_zeroes_memory_on_reaccess
(tests/decommit_capability.rs), that the decommitted range reads back
zero on re-access: that test writes a non-zero byte pattern across the
whole eligible range, calls Reservation::decommit,
then reads every byte back and hard-asserts each one is zero —
zero-fill-on-readback is proven for this eligible-range case on a Linux
runner (the code path is gated Linux and Android as a pair; the
Android half is inherited from that shared cfg, not separately executed
by any CI job). What this still does NOT prove: that the kernel’s
acceptance actually corresponds to reclaiming the physical backing — the
job logs HugePages_Free around that test as an observation only, never
a pass/fail gate, because it is a kernel-global counter shared with the
job’s other huge-page reservations. On
builds WITHOUT bench-internals, libc_madvise still discards the
return value entirely (task #719) — the kernel-response proof above is
scoped to the one CI job that enables the counters. Three things remain
reasoned-from-spec rather than empirically verified, named explicitly:
(1) the free decommit entry point, reached only through the safe
methods in CI, never called directly by a test; (2) the ineligible-range
case (still a documented no-op) and every range on a pre-5.18 kernel —
CI’s runner image kernel version is not pinned by this crate; and (3)
physical reclaim of the decommitted backing to the OS/hugetlb pool —
the job’s HugePages_Free log is an observation only, not a gate.
Zero-fill on next access is no longer part of this list: task #1174’s
content test above reads every byte back and hard-asserts zero.
Use reserve_aligned instead if you need decommit to work
unconditionally, regardless of range shape, kernel version, or platform.
Linux/Android hugetlb pool over-reserve (align > 2 MiB): when huge pages
are actually granted through the over-reserve path — which is every
granted align > 2 MiB request (the exact-size fast path exists only
for align == LINUX_HUGE_PAGE_SIZE, 2 MiB), and an align == 2 MiB
request only when that fast path misses — the whole size + align-byte
MAP_HUGETLB mapping is kept for the reservation’s lifetime, and the
Linux kernel reserves pool pages for a private hugetlb mapping’s entire
length at
mmap time (no MAP_NORESERVE is passed). The exactly align bytes
of never-touched slack are therefore charged against the bounded
nr_hugepages pool until the reservation is released: a
size == align == 4 MiB workload consumes 4 pool pages per segment
for 2 needed (2×), reaching pool exhaustion — and the silent
ordinary-page fallback — with half the segments an exact charge would
allow. Workloads bounding the hugetlb pool should prefer
align == 2 MiB shapes, which the exact-size fast path serves with
zero over-reserve whenever it hits (and whose miss cost is at most
align == 2 MiB of slack, not align > 2 MiB). This cost is
REASONED-FROM-SPEC (documented kernel
reservation semantics; updated task #1160/F4: a hugetlb-configured host
now exists in this crate’s CI (aligned-vmem-hugetlb-real,
.github/workflows/ci.yml), but that job does not measure pool-page
consumption before/after a reservation, so this specific over-reserve
cost remains unmeasured on any host available to this project) and, when
huge pages are granted via
this over-reserve path, is deliberately not trimmed away: the
over-reserved mapping is then kept whole as one soundness-driven
unit (a single munmap at the mapping base), and the pool
trade-off has no measurable host available.