cluster/drill

NCCL · basic

A 16-node run dies mid-epoch. Rank 37's log is below, and similar 'unhandled system error' traces appear on dozens of other ranks. What is the most likely root cause and your first move?

node05:22314:22398 [1] NCCL WARN socketProgress: Connection closed by remote peer node09.cluster<34211>
node05:22314:22398 [1] NCCL INFO transport/net_socket.cc:573 -> 6
[E ProcessGroupNCCL.cpp:915] [Rank 37] NCCL watchdog thread terminated with exception:
NCCL error: unhandled system error, NCCL version 2.19.3
ncclSystemError: System call (e.g. socket, malloc) or external library call failed or device error.

The options

The answer

B. A rank on node09 died first (OOM kill, segfault, preemption); every surviving rank then reports secondary connection errors - find the earliest-failing rank and check node09's job log and dmesg

Why

The key line is 'Connection closed by remote peer node09' - it names the side that vanished. When one rank dies, every peer that had an open connection to it throws its own error, so the loudest or most numerous traces are almost never the root cause; the earliest exit is. dmesg on node09 will often show the OOM killer or a segfault. Raising the watchdog timeout only delays the report of an already-dead peer, and there is no evidence against node05's own NIC since it was the receiving end of a closed connection.

More NCCL questions

This is 1 of 10 free questions. The full bank is 150 questions and 10 incident labs against a simulated 4-node HGX cluster you can break and repair — €7.99. All free questions · Field notes


← Back to ClusterDrill