Files
infra-phytron/ansible/roles/nvidia_gpu/tasks/main.yml
T
CubelaPetarandClaude Opus 5 e7ef29c5d5 Fix Blackwell GPU driver: open kernel modules, bump pin to 595
The gpu role pinned nvidia-driver-580-server (proprietary kernel
modules). Those install and load cleanly on the RTX PRO 6000, but the
GPU never initialises: NVRM fails at RmInitAdapter (0x22:0x56:897) and
nvidia-smi reports "No devices were found". Blackwell (10de:2bb5)
requires the open kernel modules.

- defaults/main.yml: nvidia-driver-580-server -> 595-server-open.
  The "-open" part is the actual fix and is version independent. 595 is
  the newest -server branch in the 24.04 archive (610 is desktop only);
  projektplan §5.3 asked for the branch to be re-checked and pinned at
  install time, which had not happened yet. Caveat noted in the file:
  595 ships from multiverse, 580/590 from restricted.
- tasks/main.yml: install_recommends: false, keeps nvidia-settings and
  its GTK chain off a headless server (~65 packages).
- manuals/20260714-nvidia-driver-install.md: same command, documents
  the open-module requirement and the failure signature.
- SETUP.md: tick off the completed OS install steps.

Verified on phy-z-srv-gpu01: driver 595.71.05, CUDA 13.2, RTX PRO 6000
Blackwell at 00000001:5C:00.0 with 97887MiB visible, no NVRM errors in
dmesg, single DKMS version registered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 09:30:02 +02:00

72 lines
2.1 KiB
YAML

---
# Automates server/phy-z-srv-gpu01/manuals/20260714-nvidia-driver-install.md
# Requires Docker (geerlingguy.docker) for the container toolkit part.
- name: Check Secure Boot state
command: mokutil --sb-state
register: nvidia_sb_state
changed_when: false
failed_when: false
- name: Fail if Secure Boot is enabled
fail:
msg: "Secure Boot is enabled — disable it in the BIOS before installing the NVIDIA driver."
when: "'SecureBoot enabled' in nvidia_sb_state.stdout"
- name: Install NVIDIA driver
apt:
name: "{{ nvidia_driver_package }}"
state: present
install_recommends: false
register: nvidia_driver_install
- name: Reboot after fresh driver install
reboot:
reboot_timeout: 600
when: nvidia_driver_install.changed
- name: Install NVIDIA CUDA repository keyring
apt:
deb: "{{ nvidia_cuda_keyring_url }}"
- name: Install CUDA toolkit (optional, see defaults)
apt:
name: "{{ nvidia_cuda_toolkit_package }}"
state: present
update_cache: true
when: nvidia_install_cuda_toolkit
- name: Add NVIDIA container toolkit repository key
shell: >
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey
| gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
args:
creates: /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
- name: Add NVIDIA container toolkit repository
copy:
dest: /etc/apt/sources.list.d/nvidia-container-toolkit.list
content: "deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://nvidia.github.io/libnvidia-container/stable/deb/$(ARCH) /\n"
mode: "0644"
- name: Install NVIDIA container toolkit
apt:
name: nvidia-container-toolkit
state: present
update_cache: true
- name: Check whether Docker already uses the NVIDIA runtime
command: grep -q nvidia /etc/docker/daemon.json
register: nvidia_docker_runtime
changed_when: false
failed_when: false
- name: Configure Docker to use the NVIDIA runtime
command: nvidia-ctk runtime configure --runtime=docker
when: nvidia_docker_runtime.rc != 0
notify: restart docker
- name: Verify the driver works
command: nvidia-smi
changed_when: false