cluster/drill

Linux · intermediate

Users report that jobs requesting 100 GB of RAM get OOM-killed on a node that 'has 500 GB'. Nothing big shows in top. /proc/meminfo shows the excerpt below. What is going on?

MemTotal:       528167328 kB
MemFree:         21504212 kB
MemAvailable:    23811640 kB
HugePages_Total:     480
HugePages_Free:      480
HugePages_Rsvd:        0
HugePages_Surp:        0
Hugepagesize:    1048576 kB

The options

The answer

A. 480 x 1 GiB hugepages (~480 GiB) are reserved in the hugepage pool; that memory is carved out of general RAM and cannot back normal 4K allocations even though the pages are 'Free' — shrink vm.nr_hugepages or make workloads actually use hugetlbfs

Why

Hugepagesize is 1048576 kB (1 GiB) and HugePages_Total is 480, so ~480 GiB of the node's 503 GiB is locked in the explicit hugepage pool — that memory is unusable for ordinary anonymous allocations even while HugePages_Free shows the pool is untouched, which is why MemAvailable is only ~23 GiB. This is a classic leftover from a DPDK/database tuning or a mis-applied sysctl. It is not a leak (no process owns it, so a reboot with the same boot/sysctl config brings it right back), MemAvailable is correctly excluding the pool rather than ignoring it, and swap would just let 100 GB jobs thrash instead of fail fast.

More Linux questions

This is 1 of 10 free questions. The full bank is 150 questions and 10 incident labs against a simulated 4-node HGX cluster you can break and repair — €7.99. All free questions · Field notes


← Back to ClusterDrill