e7ef29c5d5ab9241af5557116b6614dcd1afd333
The gpu role pinned nvidia-driver-580-server (proprietary kernel modules). Those install and load cleanly on the RTX PRO 6000, but the GPU never initialises: NVRM fails at RmInitAdapter (0x22:0x56:897) and nvidia-smi reports "No devices were found". Blackwell (10de:2bb5) requires the open kernel modules. - defaults/main.yml: nvidia-driver-580-server -> 595-server-open. The "-open" part is the actual fix and is version independent. 595 is the newest -server branch in the 24.04 archive (610 is desktop only); projektplan §5.3 asked for the branch to be re-checked and pinned at install time, which had not happened yet. Caveat noted in the file: 595 ships from multiverse, 580/590 from restricted. - tasks/main.yml: install_recommends: false, keeps nvidia-settings and its GTK chain off a headless server (~65 packages). - manuals/20260714-nvidia-driver-install.md: same command, documents the open-module requirement and the failure signature. - SETUP.md: tick off the completed OS install steps. Verified on phy-z-srv-gpu01: driver 595.71.05, CUDA 13.2, RTX PRO 6000 Blackwell at 00000001:5C:00.0 with 97887MiB visible, no NVRM errors in dmesg, single DKMS version registered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
infra-phytron
Infrastructure as Code and runbooks for Phytron's server infrastructure.
ansible/— global configuration for all servers (inventory, group_vars, roles, playbooks). See ansible/README.md.server/<hostname>/— per-server documentation, runbooks, scripts, and files.
All servers are listed with their IPs in ansible/hosts.ini; each has a README under server/.
Languages
PowerShell
94.1%
Just
3.1%
Jinja
2.8%