The gpu role pinned nvidia-driver-580-server (proprietary kernel modules). Those install and load cleanly on the RTX PRO 6000, but the GPU never initialises: NVRM fails at RmInitAdapter (0x22:0x56:897) and nvidia-smi reports "No devices were found". Blackwell (10de:2bb5) requires the open kernel modules. - defaults/main.yml: nvidia-driver-580-server -> 595-server-open. The "-open" part is the actual fix and is version independent. 595 is the newest -server branch in the 24.04 archive (610 is desktop only); projektplan §5.3 asked for the branch to be re-checked and pinned at install time, which had not happened yet. Caveat noted in the file: 595 ships from multiverse, 580/590 from restricted. - tasks/main.yml: install_recommends: false, keeps nvidia-settings and its GTK chain off a headless server (~65 packages). - manuals/20260714-nvidia-driver-install.md: same command, documents the open-module requirement and the failure signature. - SETUP.md: tick off the completed OS install steps. Verified on phy-z-srv-gpu01: driver 595.71.05, CUDA 13.2, RTX PRO 6000 Blackwell at 00000001:5C:00.0 with 97887MiB visible, no NVRM errors in dmesg, single DKMS version registered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
72 lines
2.1 KiB
YAML
72 lines
2.1 KiB
YAML
---
|
|
# Automates server/phy-z-srv-gpu01/manuals/20260714-nvidia-driver-install.md
|
|
# Requires Docker (geerlingguy.docker) for the container toolkit part.
|
|
|
|
- name: Check Secure Boot state
|
|
command: mokutil --sb-state
|
|
register: nvidia_sb_state
|
|
changed_when: false
|
|
failed_when: false
|
|
|
|
- name: Fail if Secure Boot is enabled
|
|
fail:
|
|
msg: "Secure Boot is enabled — disable it in the BIOS before installing the NVIDIA driver."
|
|
when: "'SecureBoot enabled' in nvidia_sb_state.stdout"
|
|
|
|
- name: Install NVIDIA driver
|
|
apt:
|
|
name: "{{ nvidia_driver_package }}"
|
|
state: present
|
|
install_recommends: false
|
|
register: nvidia_driver_install
|
|
|
|
- name: Reboot after fresh driver install
|
|
reboot:
|
|
reboot_timeout: 600
|
|
when: nvidia_driver_install.changed
|
|
|
|
- name: Install NVIDIA CUDA repository keyring
|
|
apt:
|
|
deb: "{{ nvidia_cuda_keyring_url }}"
|
|
|
|
- name: Install CUDA toolkit (optional, see defaults)
|
|
apt:
|
|
name: "{{ nvidia_cuda_toolkit_package }}"
|
|
state: present
|
|
update_cache: true
|
|
when: nvidia_install_cuda_toolkit
|
|
|
|
- name: Add NVIDIA container toolkit repository key
|
|
shell: >
|
|
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey
|
|
| gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
|
|
args:
|
|
creates: /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
|
|
|
|
- name: Add NVIDIA container toolkit repository
|
|
copy:
|
|
dest: /etc/apt/sources.list.d/nvidia-container-toolkit.list
|
|
content: "deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://nvidia.github.io/libnvidia-container/stable/deb/$(ARCH) /\n"
|
|
mode: "0644"
|
|
|
|
- name: Install NVIDIA container toolkit
|
|
apt:
|
|
name: nvidia-container-toolkit
|
|
state: present
|
|
update_cache: true
|
|
|
|
- name: Check whether Docker already uses the NVIDIA runtime
|
|
command: grep -q nvidia /etc/docker/daemon.json
|
|
register: nvidia_docker_runtime
|
|
changed_when: false
|
|
failed_when: false
|
|
|
|
- name: Configure Docker to use the NVIDIA runtime
|
|
command: nvidia-ctk runtime configure --runtime=docker
|
|
when: nvidia_docker_runtime.rc != 0
|
|
notify: restart docker
|
|
|
|
- name: Verify the driver works
|
|
command: nvidia-smi
|
|
changed_when: false
|