Files
infra-phytron/server/phy-z-srv-gpu01/manuals/20260714-nvidia-driver-install.md
T
CubelaPetarandClaude Opus 5 e7ef29c5d5 Fix Blackwell GPU driver: open kernel modules, bump pin to 595
The gpu role pinned nvidia-driver-580-server (proprietary kernel
modules). Those install and load cleanly on the RTX PRO 6000, but the
GPU never initialises: NVRM fails at RmInitAdapter (0x22:0x56:897) and
nvidia-smi reports "No devices were found". Blackwell (10de:2bb5)
requires the open kernel modules.

- defaults/main.yml: nvidia-driver-580-server -> 595-server-open.
  The "-open" part is the actual fix and is version independent. 595 is
  the newest -server branch in the 24.04 archive (610 is desktop only);
  projektplan §5.3 asked for the branch to be re-checked and pinned at
  install time, which had not happened yet. Caveat noted in the file:
  595 ships from multiverse, 580/590 from restricted.
- tasks/main.yml: install_recommends: false, keeps nvidia-settings and
  its GTK chain off a headless server (~65 packages).
- manuals/20260714-nvidia-driver-install.md: same command, documents
  the open-module requirement and the failure signature.
- SETUP.md: tick off the completed OS install steps.

Verified on phy-z-srv-gpu01: driver 595.71.05, CUDA 13.2, RTX PRO 6000
Blackwell at 00000001:5C:00.0 with 97887MiB visible, no NVRM errors in
dmesg, single DKMS version registered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 09:30:02 +02:00

88 lines
2.8 KiB
Markdown

# Manual nvidia driver, cuda repo and container toolkit
Automated by the Ansible role `ansible/roles/nvidia_gpu` — this manual documents the steps.
**Disable Secure Boot in BIOS first** (unsigned kernel modules won't load otherwise).
## NVIDIA driver
Check if GPUs are recognized by the base OS:
```bash
sudo lspci | grep -i nvidia
```
Which should show some output if it finds nvidia devices.
Search for available drivers for your GPUs:
```bash
sudo ubuntu-drivers devices
```
Install the driver pinned. The RTX PRO 6000 (Blackwell) needs **driver >= 580**
and **only works with the open kernel modules** — the proprietary ones load but
fail at `RmInitAdapter`, leaving `nvidia-smi` with "No devices were found".
Use the `-server-open` variant and prefer a pinned install over
`ubuntu-drivers autoinstall` so the choice is explicit and reproducible
(re-check for a newer branch at install time):
```bash
sudo apt install -y --no-install-recommends nvidia-driver-595-server-open
```
`--no-install-recommends` keeps `nvidia-settings` and its GTK dependency chain
off this headless server.
> Note: do **not** additionally install `cuda-drivers` from the NVIDIA repo —
> that would mix the Ubuntu-archive driver with the NVIDIA-repo driver and the
> two can conflict. Pick one source; we use the Ubuntu archive.
Reboot the system for changes to take effect:
```bash
sudo reboot
```
Show GPU stats with:
```bash
nvidia-smi
```
## CUDA repository (toolkit optional)
Add the NVIDIA CUDA apt repository:
```bash
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
```
The full CUDA toolkit is **not needed** for Docker-based workloads (vLLM etc. —
the driver plus container toolkit suffice). Only if compiling on the host:
```bash
sudo apt install -y cuda-toolkit # meta package, pulls the current release
```
## Container toolkit
Install the Nvidia Container toolkit:
```bash
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg \
&& curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
sudo apt install -y nvidia-container-toolkit
```
Configure Docker to use the NVIDIA runtime (writes `/etc/docker/daemon.json`) and restart it —
without this step `docker run --gpus all` fails:
```bash
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
```
Test a simple cuda container and nvidia-smi command inside:
```bash
docker run --rm --gpus all nvidia/cuda:13.0.0-base-ubuntu24.04 nvidia-smi
```