Fix Blackwell GPU driver: open kernel modules, bump pin to 595
The gpu role pinned nvidia-driver-580-server (proprietary kernel modules). Those install and load cleanly on the RTX PRO 6000, but the GPU never initialises: NVRM fails at RmInitAdapter (0x22:0x56:897) and nvidia-smi reports "No devices were found". Blackwell (10de:2bb5) requires the open kernel modules. - defaults/main.yml: nvidia-driver-580-server -> 595-server-open. The "-open" part is the actual fix and is version independent. 595 is the newest -server branch in the 24.04 archive (610 is desktop only); projektplan §5.3 asked for the branch to be re-checked and pinned at install time, which had not happened yet. Caveat noted in the file: 595 ships from multiverse, 580/590 from restricted. - tasks/main.yml: install_recommends: false, keeps nvidia-settings and its GTK chain off a headless server (~65 packages). - manuals/20260714-nvidia-driver-install.md: same command, documents the open-module requirement and the failure signature. - SETUP.md: tick off the completed OS install steps. Verified on phy-z-srv-gpu01: driver 595.71.05, CUDA 13.2, RTX PRO 6000 Blackwell at 00000001:5C:00.0 with 97887MiB visible, no NVRM errors in dmesg, single DKMS version registered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -17,6 +17,7 @@
|
||||
apt:
|
||||
name: "{{ nvidia_driver_package }}"
|
||||
state: present
|
||||
install_recommends: false
|
||||
register: nvidia_driver_install
|
||||
|
||||
- name: Reboot after fresh driver install
|
||||
|
||||
Reference in New Issue
Block a user