GPU · intermediate
Mid-shift, a node's jobs all fail at once, nvidia-smi hangs, and dmesg shows the lines below. What is the appropriate operator response?
[43121.008812] NVRM: GPU at PCI:0000:af:00: GPU-9d1a44f2-6c1e-11ee-8c90-0242ac120002
[43121.008833] NVRM: Xid (PCI:0000:af:00): 79, pid=52210, GPU has fallen off the bus.
[43121.008841] NVRM: GPU 0000:af:00.0: GPU has fallen off the bus.
The options
- ADrain the node; XID 79 means the GPU dropped off PCIe, so investigate power delivery, PCIe seating/risers, and cooling, and expect a power cycle before it returns
- BRequeue the failed jobs on the same node; XID 79 is transient and clears by itself as soon as the next CUDA context initializes on that GPU, so no drain or hardware work is needed
- CTreat it as an application fault in the same family as XID 31, and ask the user to fix the code that pushed the device into this state before they resubmit the job to the queue
- DReload the NVIDIA kernel modules so that the driver re-enumerates the missing device, which brings the GPU back into service without any disruptive reboot of the whole node
The answer
A. Drain the node; XID 79 means the GPU dropped off PCIe, so investigate power delivery, PCIe seating/risers, and cooling, and expect a power cycle before it returns
Why
XID 79 means the driver lost contact with a GPU that vanished from the PCIe bus, which is a hardware-level event: common culprits are power delivery (cables, VRMs), PCIe signal/seating problems, or overheating. The node must come out of the scheduler and usually needs a full power cycle; if the GPU keeps falling off the bus, it is an RMA candidate. It is not transient in the way a requeue assumes, it is not an application bug, and a module reload cannot restore a device that is no longer enumerating, which is also why lspci is a useful confirmation step.
More GPU questions
This is 1 of 10 free questions. The full bank is 150 questions and 10 incident labs against a simulated 4-node HGX cluster you can break and repair — €7.99. All free questions · Field notes
← Back to ClusterDrill