Files
infra-phytron/server/phy-srv-gpu01/manuals/20260714-nvidia-driver-install.md
CubelaPetarandClaude Opus 4.8 75296ce58c Rename phy-z-srv-gpu01 to phy-srv-gpu01; fix share analysis scripts
- rename host/group/folder everywhere to match the server's actual
  hostname and physical label
- share-analysis.ps1: sanitize the output filename prefix — '-Paths "D:"'
  produced 'D:-file-types.csv' and Export-Csv failed with 'path format
  not supported', so no CSVs were written
- share-analysis.ps1: new -FolderDepth so folders can be aggregated at
  D:\Abteilungen\<Share> level, which matches the share layout
- list-shares.ps1: -WithSize walks local paths instead of UNC when run on
  the server itself (UNC was orders of magnitude slower and looked stuck)
  and prints progress every 50k files

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 14:47:07 +02:00

2.8 KiB

Manual nvidia driver, cuda repo and container toolkit

Automated by the Ansible role ansible/roles/nvidia_gpu — this manual documents the steps.

Disable Secure Boot in BIOS first (unsigned kernel modules won't load otherwise).

NVIDIA driver

Check if GPUs are recognized by the base OS:

sudo lspci | grep -i nvidia

Which should show some output if it finds nvidia devices.

Search for available drivers for your GPUs:

sudo ubuntu-drivers devices

Install the driver pinned. The RTX PRO 6000 (Blackwell) needs driver >= 580 and only works with the open kernel modules — the proprietary ones load but fail at RmInitAdapter, leaving nvidia-smi with "No devices were found". Use the -server-open variant and prefer a pinned install over ubuntu-drivers autoinstall so the choice is explicit and reproducible (re-check for a newer branch at install time):

sudo apt install -y --no-install-recommends nvidia-driver-595-server-open

--no-install-recommends keeps nvidia-settings and its GTK dependency chain off this headless server.

Note: do not additionally install cuda-drivers from the NVIDIA repo — that would mix the Ubuntu-archive driver with the NVIDIA-repo driver and the two can conflict. Pick one source; we use the Ubuntu archive.

Reboot the system for changes to take effect:

sudo reboot

Show GPU stats with:

nvidia-smi

CUDA repository (toolkit optional)

Add the NVIDIA CUDA apt repository:

wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update

The full CUDA toolkit is not needed for Docker-based workloads (vLLM etc. — the driver plus container toolkit suffice). Only if compiling on the host:

sudo apt install -y cuda-toolkit    # meta package, pulls the current release

Container toolkit

Install the Nvidia Container toolkit:

curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg \
  && curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
    sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
    sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
sudo apt install -y nvidia-container-toolkit

Configure Docker to use the NVIDIA runtime (writes /etc/docker/daemon.json) and restart it — without this step docker run --gpus all fails:

sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

Test a simple cuda container and nvidia-smi command inside:

docker run --rm --gpus all nvidia/cuda:13.0.0-base-ubuntu24.04 nvidia-smi