From e7ef29c5d5ab9241af5557116b6614dcd1afd333 Mon Sep 17 00:00:00 2001 From: Petar Cubela Date: Mon, 31 Aug 2026 09:30:02 +0200 Subject: [PATCH] Fix Blackwell GPU driver: open kernel modules, bump pin to 595 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The gpu role pinned nvidia-driver-580-server (proprietary kernel modules). Those install and load cleanly on the RTX PRO 6000, but the GPU never initialises: NVRM fails at RmInitAdapter (0x22:0x56:897) and nvidia-smi reports "No devices were found". Blackwell (10de:2bb5) requires the open kernel modules. - defaults/main.yml: nvidia-driver-580-server -> 595-server-open. The "-open" part is the actual fix and is version independent. 595 is the newest -server branch in the 24.04 archive (610 is desktop only); projektplan §5.3 asked for the branch to be re-checked and pinned at install time, which had not happened yet. Caveat noted in the file: 595 ships from multiverse, 580/590 from restricted. - tasks/main.yml: install_recommends: false, keeps nvidia-settings and its GTK chain off a headless server (~65 packages). - manuals/20260714-nvidia-driver-install.md: same command, documents the open-module requirement and the failure signature. - SETUP.md: tick off the completed OS install steps. Verified on phy-z-srv-gpu01: driver 595.71.05, CUDA 13.2, RTX PRO 6000 Blackwell at 00000001:5C:00.0 with 97887MiB visible, no NVRM errors in dmesg, single DKMS version registered. Co-Authored-By: Claude Opus 5 --- SETUP.md | 12 +++++++----- ansible/roles/nvidia_gpu/defaults/main.yml | 10 +++++++--- ansible/roles/nvidia_gpu/tasks/main.yml | 1 + .../manuals/20260714-nvidia-driver-install.md | 14 ++++++++++---- 4 files changed, 25 insertions(+), 12 deletions(-) diff --git a/SETUP.md b/SETUP.md index 03501ff..dc139ac 100644 --- a/SETUP.md +++ b/SETUP.md @@ -14,8 +14,10 @@ Document all the steps done for the initial generic setup of the server. - [x] iLO setup: Hostname, snmp (v1 disable, configure v3), license was already inserted, - [x] RAID setup of server in BIOS - [ ] OS Install - - [ ] new local user (sbxadmin) - - [ ] Hostname - - [ ] Network settings - - [ ] filesystem setup - - [ ] select root password + - [x] new local user (sbxadmin) + - [x] Hostname + - [x] Network settings + - [x] filesystem setup + - [x] select root password + - [ ] base hardening (ssh, updaetes,..) + - [x] remove snap (ubuntu only) diff --git a/ansible/roles/nvidia_gpu/defaults/main.yml b/ansible/roles/nvidia_gpu/defaults/main.yml index 3360c16..98a1461 100644 --- a/ansible/roles/nvidia_gpu/defaults/main.yml +++ b/ansible/roles/nvidia_gpu/defaults/main.yml @@ -1,7 +1,11 @@ --- -# Driver pinned explicitly — RTX PRO 6000 (Blackwell) needs >= 580. -# Re-check for a newer branch at install time (projektplan §5). -nvidia_driver_package: nvidia-driver-580-server +# Driver pinned explicitly (projektplan §5: re-check the branch at install time). +# 595 is the newest "-server" branch in the Ubuntu 24.04 archive; 610 is desktop +# only. MUST be the "-open" variant: Blackwell (PCI ID 10de:2bb5) is not +# supported by the proprietary kernel modules — the closed package installs and +# loads fine but fails at RmInitAdapter, so nvidia-smi reports "No devices were +# found". Note 595 lives in multiverse, while 580/590 are in restricted. +nvidia_driver_package: nvidia-driver-595-server-open # CUDA apt repository keyring (Ubuntu 24.04) nvidia_cuda_keyring_url: https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb diff --git a/ansible/roles/nvidia_gpu/tasks/main.yml b/ansible/roles/nvidia_gpu/tasks/main.yml index 69e9e60..e459f85 100644 --- a/ansible/roles/nvidia_gpu/tasks/main.yml +++ b/ansible/roles/nvidia_gpu/tasks/main.yml @@ -17,6 +17,7 @@ apt: name: "{{ nvidia_driver_package }}" state: present + install_recommends: false register: nvidia_driver_install - name: Reboot after fresh driver install diff --git a/server/phy-z-srv-gpu01/manuals/20260714-nvidia-driver-install.md b/server/phy-z-srv-gpu01/manuals/20260714-nvidia-driver-install.md index 8f1c248..88cbea4 100644 --- a/server/phy-z-srv-gpu01/manuals/20260714-nvidia-driver-install.md +++ b/server/phy-z-srv-gpu01/manuals/20260714-nvidia-driver-install.md @@ -18,13 +18,19 @@ Search for available drivers for your GPUs: sudo ubuntu-drivers devices ``` -Install the driver pinned. The RTX PRO 6000 (Blackwell) needs **driver >= 580**; -use the `-server` variant and prefer a pinned install over `ubuntu-drivers autoinstall` -so the choice is explicit and reproducible (re-check for a newer branch at install time): +Install the driver pinned. The RTX PRO 6000 (Blackwell) needs **driver >= 580** +and **only works with the open kernel modules** — the proprietary ones load but +fail at `RmInitAdapter`, leaving `nvidia-smi` with "No devices were found". +Use the `-server-open` variant and prefer a pinned install over +`ubuntu-drivers autoinstall` so the choice is explicit and reproducible +(re-check for a newer branch at install time): ```bash -sudo apt install -y nvidia-driver-580-server +sudo apt install -y --no-install-recommends nvidia-driver-595-server-open ``` +`--no-install-recommends` keeps `nvidia-settings` and its GTK dependency chain +off this headless server. + > Note: do **not** additionally install `cuda-drivers` from the NVIDIA repo — > that would mix the Ubuntu-archive driver with the NVIDIA-repo driver and the > two can conflict. Pick one source; we use the Ubuntu archive.