Linux · intermediate
Users report that jobs requesting 100 GB of RAM get OOM-killed on a node that 'has 500 GB'. Nothing big shows in top. /proc/meminfo shows the excerpt below. What is going on?
MemTotal: 528167328 kB
MemFree: 21504212 kB
MemAvailable: 23811640 kB
HugePages_Total: 480
HugePages_Free: 480
HugePages_Rsvd: 0
HugePages_Surp: 0
Hugepagesize: 1048576 kB
The options
- A480 x 1 GiB hugepages (~480 GiB) are reserved in the hugepage pool; that memory is carved out of general RAM and cannot back normal 4K allocations even though the pages are 'Free' — shrink vm.nr_hugepages or make workloads actually use hugetlbfs
- BMemFree is low because a slow kernel-side memory leak has consumed roughly 480 GiB that no process owns, which is exactly why nothing large shows up in top; a rolling reboot of the node will hand that memory back and the OOM kills will stop.
- CMemAvailable deliberately ignores hugepage reservations, so the node genuinely does have roughly 500 GiB available for normal allocations; the OOM kills must therefore be coming from per-job cgroup memory limits applied by the scheduler rather than from any real shortage.
- DThe node needs swap enabled to cover the gap between MemTotal and MemAvailable: with no swap device configured the kernel has no overflow area to fall back on, so a 100 GB request is refused outright and the job is killed instead of being paged out.
The answer
A. 480 x 1 GiB hugepages (~480 GiB) are reserved in the hugepage pool; that memory is carved out of general RAM and cannot back normal 4K allocations even though the pages are 'Free' — shrink vm.nr_hugepages or make workloads actually use hugetlbfs
Why
Hugepagesize is 1048576 kB (1 GiB) and HugePages_Total is 480, so ~480 GiB of the node's 503 GiB is locked in the explicit hugepage pool — that memory is unusable for ordinary anonymous allocations even while HugePages_Free shows the pool is untouched, which is why MemAvailable is only ~23 GiB. This is a classic leftover from a DPDK/database tuning or a mis-applied sysctl. It is not a leak (no process owns it, so a reboot with the same boot/sysctl config brings it right back), MemAvailable is correctly excluding the pool rather than ignoring it, and swap would just let 100 GB jobs thrash instead of fail fast.
More Linux questions
This is 1 of 10 free questions. The full bank is 150 questions and 10 incident labs against a simulated 4-node HGX cluster you can break and repair — €7.99. All free questions · Field notes
← Back to ClusterDrill