cluster/drill

Linux · basic

A user opens a ticket: 'Node cn117 has almost no free memory even though nothing is running — please reboot it before my job starts.' You run free -g and see the output below. What is the correct response?

$ free -g
              total        used        free      shared  buff/cache   available
Mem:            503          41           7           2         454         456
Swap:             0           0           0

The options

The answer

B. The node is healthy: the 454 GiB in buff/cache is page cache the kernel reclaims automatically when applications need memory, and 'available' (456 GiB) is what a new job can actually allocate — no reboot needed

Why

Linux deliberately uses otherwise-idle RAM as page cache to speed up file I/O; that memory is counted in buff/cache and is reclaimed on demand, which is exactly what the 'available' column estimates (456 of 503 GiB). Low 'free' on a long-running Linux box is normal and healthy, not a leak. Adding swap to a compute node is generally undesirable (it turns OOM conditions into node-crippling thrash), and routinely dropping caches only forces the next job to re-read data from disk.

More Linux questions

This is 1 of 10 free questions. The full bank is 150 questions and 10 incident labs against a simulated 4-node HGX cluster you can break and repair — €7.99. All free questions · Field notes


← Back to ClusterDrill