cluster/drill

GPU · intermediate

Mid-shift, a node's jobs all fail at once, nvidia-smi hangs, and dmesg shows the lines below. What is the appropriate operator response?

[43121.008812] NVRM: GPU at PCI:0000:af:00: GPU-9d1a44f2-6c1e-11ee-8c90-0242ac120002
[43121.008833] NVRM: Xid (PCI:0000:af:00): 79, pid=52210, GPU has fallen off the bus.
[43121.008841] NVRM: GPU 0000:af:00.0: GPU has fallen off the bus.

The options

The answer

A. Drain the node; XID 79 means the GPU dropped off PCIe, so investigate power delivery, PCIe seating/risers, and cooling, and expect a power cycle before it returns

Why

XID 79 means the driver lost contact with a GPU that vanished from the PCIe bus, which is a hardware-level event: common culprits are power delivery (cables, VRMs), PCIe signal/seating problems, or overheating. The node must come out of the scheduler and usually needs a full power cycle; if the GPU keeps falling off the bus, it is an RMA candidate. It is not transient in the way a requeue assumes, it is not an application bug, and a module reload cannot restore a device that is no longer enumerating, which is also why lspci is a useful confirmation step.

More GPU questions

This is 1 of 10 free questions. The full bank is 150 questions and 10 incident labs against a simulated 4-node HGX cluster you can break and repair — €7.99. All free questions · Field notes


← Back to ClusterDrill