Skip to main content

reserve_aligned_huge

Function reserve_aligned_huge 

Source
pub fn reserve_aligned_huge(size: usize, align: usize) -> Option<Reservation>
Available on crate feature huge-pages only.
Expand description

Reserve size bytes aligned to align, requesting OS large / huge pages (Linux/Android MAP_HUGETLB, Windows MEM_LARGE_PAGES). Currently a no-op on macOS and other Unix that is neither Linux nor Android — it falls back to an ordinary reservation, identical to reserve_aligned.

Transparent-huge-page hinting (Linux/Android MADV_HUGEPAGE) is not used: it cannot affect an already-explicitly-huge MAP_HUGETLB mapping (the pages are already huge), so issuing it would be a wasted syscall. This crate’s strategy is the explicit hugetlbfs path only.

Large pages reduce TLB pressure for big allocator segments. The request is best-effort: if the OS refuses large pages (none configured, no privilege), the reservation transparently falls back to ordinary pages, so this never fails purely because huge pages are unavailable — it fails only on a genuine reservation error (OOM) or a contract violation.

To detect whether huge pages were actually granted (as opposed to having fallen back to ordinary pages), use the returned Reservation::is_huge method.

Base/align/size contract is otherwise identical to reserve_aligned, except on Linux AND Android with huge-pages enabled (the check’s cfg is any(target_os = "linux", target_os = "android") + feature = "huge-pages"): size and align must BOTH additionally be multiples of the huge-page size (2 MiB) — a request that only satisfies reserve_aligned’s own weaker PAGE-multiple contract is rejected with VmemError::invalid_argument() before any syscall runs, even though such a request could previously succeed there via the documented ordinary-page fallback. For the failure cause use try_reserve_aligned_huge.

Windows limitation: on Windows, this function returns a reservation with Reservation::is_huge == true only when ALL of the following hold:

  1. The fast-path condition align <= GetLargePageMinimum() is satisfied (typically align <= 2 MiB on x86_64)
  2. size is a multiple of the system’s large-page minimum
  3. The calling process has SeLockMemoryPrivilege granted AND has enabled it via AdjustTokenPrivileges (the crate does not do this for you — a process with the privilege granted but not enabled fails exactly like an unprivileged one and silently falls back to ordinary pages)

NOTE: The widened fast-path condition (II-3, 2026-08-16 audit finding) expanded the single-call ATTEMPT window from align <= 64 KiB to align <= GetLargePageMinimum(), but on an unprivileged host the actual paths that SUCCEED (pass the post-call alignment check) are typically still limited. When large pages are NOT granted (unprivileged), VirtualAlloc’s alignment guarantee is only 64 KiB; in practice it typically does NOT happen to land on the requested alignment, so the post-call check fails and the fast path falls through to the two-call path. Practically, this means is_huge() == true only for shapes where large pages are actually granted, which requires all three conditions above to hold.

Extra-syscall cost on unprivileged hosts: For the widened align range (64 KiB < align <= GetLargePageMinimum()), when large pages are requested but not granted (e.g., unprivileged process, or SeLockMemoryPrivilege not enabled), the code attempts VirtualAlloc with MEM_LARGE_PAGES (fails), retries without it (succeeds with ordinary pages), and if that retry’s base doesn’t happen to satisfy the requested alignment, the whole thing is released and falls through to the two-call path. This means an unprivileged reservation in this align range can cost up to 2 extra VirtualAlloc calls + 1 VirtualFree before reaching the two-call path, versus before the II-3 change (which would have gone straight to the two-call path for align > 64 KiB). This is a real, measurable behavior change, not a correctness bug — the widening genuinely expands the single-call attempt window, and unprivileged processes pay the extra-syscall cost for shapes that now attempt but fail the fast path.

If any of these conditions fail, the function falls back to ordinary pages and returns a reservation with Reservation::is_huge == false. On Windows, large pages (MEM_LARGE_PAGES) are only ever requested and possibly granted via the single-call fast path; the two-call path never requests large pages, so the result never has Reservation::is_huge == true.

Decommit incompatibility (corrected task #1140): on Windows, decommit/decommit_lazy never work on huge-page reservations — VirtualFree with MEM_DECOMMIT unconditionally fails on large-page regions. decommit_lazy never works on a huge-page reservation on ANY platform — MADV_FREE (its Linux/Android backend) has no documented HugeTLB support, unlike MADV_DONTNEED below.

decommit on Linux/Android is more nuanced: it depends on BOTH the requested range and the running kernel. madvise(2) documents that MADV_DONTNEED gained HugeTLB support in Linux 5.18, requiring [base+start, base+end) to be aligned to the mapping’s huge page size (2 MiB) at both endpoints — the same alignment this function already requires of size/align themselves on Linux/Android, so decommitting an entire huge reservation, or any 2-MiB-granular sub-range of it, is exactly such an eligible range on a >= 5.18 kernel. A page_size()-granular (e.g. 4 KiB) but not 2-MiB-granular offset still gets EINVAL and does nothing, as does EVERY range on a pre-5.18 kernel. Reservation::decommit/ Reservation::try_decommit (the safe methods) consult both Reservation::is_huge and the requested range to skip the ineligible case before issuing the syscall; the free decommit function has no is_huge() to consult and issues the syscall unconditionally — see that function’s own doc for the precise split. Either way — an ineligible range, or any range on a pre-5.18 kernel — the effect is indistinguishable from a silent no-op: the caller’s RSS does not decrease, and subsequent reads return the old (stale) data rather than zeroed pages.

Documented per the madvise(2) man page, and — since task #1152 (F1) — empirically exercised under a real hugetlb pool by this crate’s own CI: the aligned-vmem-hugetlb-real job (.github/workflows/ci.yml) hard- asserts (via a path-activation oracle) that this function actually received a MAP_HUGETLB grant, then drives a huge-page-eligible range through Reservation::decommit’s eligible-huge branch. What that job proves, stated precisely (task #1160/F1 correction of an earlier overclaim; strengthened tasks #1164 and #1174): the eligible-range/post-5.18-kernel case genuinely REACHES the real madvise(2)/MADV_DONTNEED backend call — AND, since task #1164’s ci_hugetlb_real_pool_kernel_actually_accepts_eligible_madvise (tests/decommit_capability.rs), that the kernel itself returned 0 (accepted) for that call, not -1 (rejected): under bench-internals, libc_madvise (src/os/unix.rs) records the syscall’s own return value into a counter pair, and that job hard-asserts it increased for this eligible-range case — AND, since task #1174’s ci_hugetlb_real_pool_decommit_actually_zeroes_memory_on_reaccess (tests/decommit_capability.rs), that the decommitted range reads back zero on re-access: that test writes a non-zero byte pattern across the whole eligible range, calls Reservation::decommit, then reads every byte back and hard-asserts each one is zero — zero-fill-on-readback is proven for this eligible-range case on a Linux runner (the code path is gated Linux and Android as a pair; the Android half is inherited from that shared cfg, not separately executed by any CI job). What this still does NOT prove: that the kernel’s acceptance actually corresponds to reclaiming the physical backing — the job logs HugePages_Free around that test as an observation only, never a pass/fail gate, because it is a kernel-global counter shared with the job’s other huge-page reservations. On builds WITHOUT bench-internals, libc_madvise still discards the return value entirely (task #719) — the kernel-response proof above is scoped to the one CI job that enables the counters. Three things remain reasoned-from-spec rather than empirically verified, named explicitly: (1) the free decommit entry point, reached only through the safe methods in CI, never called directly by a test; (2) the ineligible-range case (still a documented no-op) and every range on a pre-5.18 kernel — CI’s runner image kernel version is not pinned by this crate; and (3) physical reclaim of the decommitted backing to the OS/hugetlb pool — the job’s HugePages_Free log is an observation only, not a gate. Zero-fill on next access is no longer part of this list: task #1174’s content test above reads every byte back and hard-asserts zero.

Use reserve_aligned instead if you need decommit to work unconditionally, regardless of range shape, kernel version, or platform.

Linux/Android hugetlb pool over-reserve (align > 2 MiB): when huge pages are actually granted through the over-reserve path — which is every granted align > 2 MiB request (the exact-size fast path exists only for align == LINUX_HUGE_PAGE_SIZE, 2 MiB), and an align == 2 MiB request only when that fast path misses — the whole size + align-byte MAP_HUGETLB mapping is kept for the reservation’s lifetime, and the Linux kernel reserves pool pages for a private hugetlb mapping’s entire length at mmap time (no MAP_NORESERVE is passed). The exactly align bytes of never-touched slack are therefore charged against the bounded nr_hugepages pool until the reservation is released: a size == align == 4 MiB workload consumes 4 pool pages per segment for 2 needed (2×), reaching pool exhaustion — and the silent ordinary-page fallback — with half the segments an exact charge would allow. Workloads bounding the hugetlb pool should prefer align == 2 MiB shapes, which the exact-size fast path serves with zero over-reserve whenever it hits (and whose miss cost is at most align == 2 MiB of slack, not align > 2 MiB). This cost is REASONED-FROM-SPEC (documented kernel reservation semantics; updated task #1160/F4: a hugetlb-configured host now exists in this crate’s CI (aligned-vmem-hugetlb-real, .github/workflows/ci.yml), but that job does not measure pool-page consumption before/after a reservation, so this specific over-reserve cost remains unmeasured on any host available to this project) and, when huge pages are granted via this over-reserve path, is deliberately not trimmed away: the over-reserved mapping is then kept whole as one soundness-driven unit (a single munmap at the mapping base), and the pool trade-off has no measurable host available.