cluster/drill

GPU · basic

Right after your config-management system pushed a new NVIDIA driver package to a live node, users report CUDA failures and you see the output below. What is the correct fix?

$ nvidia-smi
Failed to initialize NVML: Driver/library version mismatch
NVML library version: 550.54

The options

The answer

B. Reboot the node (or unload and reload the NVIDIA kernel modules) so the running kernel module matches the new userspace libraries

Why

The package update replaced the userspace NVML library (550.54) on disk, but the old driver kernel module is still loaded in memory, so the two disagree. Rebooting, or stopping all GPU clients and doing rmmod/modprobe of the nvidia modules, brings the loaded module in line with the libraries. The CUDA toolkit is unrelated to NVML versioning, a fallen-off-the-bus GPU produces XID 79 and missing devices rather than this message, and nouveau conflicts prevent the NVIDIA module from loading at all instead of causing a version mismatch.

More GPU questions

This is 1 of 10 free questions. The full bank is 150 questions and 10 incident labs against a simulated 4-node HGX cluster you can break and repair — €7.99. All free questions · Field notes


← Back to ClusterDrill