The gpu role pinned nvidia-driver-580-server (proprietary kernel
modules). Those install and load cleanly on the RTX PRO 6000, but the
GPU never initialises: NVRM fails at RmInitAdapter (0x22:0x56:897) and
nvidia-smi reports "No devices were found". Blackwell (10de:2bb5)
requires the open kernel modules.
- defaults/main.yml: nvidia-driver-580-server -> 595-server-open.
The "-open" part is the actual fix and is version independent. 595 is
the newest -server branch in the 24.04 archive (610 is desktop only);
projektplan §5.3 asked for the branch to be re-checked and pinned at
install time, which had not happened yet. Caveat noted in the file:
595 ships from multiverse, 580/590 from restricted.
- tasks/main.yml: install_recommends: false, keeps nvidia-settings and
its GTK chain off a headless server (~65 packages).
- manuals/20260714-nvidia-driver-install.md: same command, documents
the open-module requirement and the failure signature.
- SETUP.md: tick off the completed OS install steps.
Verified on phy-z-srv-gpu01: driver 595.71.05, CUDA 13.2, RTX PRO 6000
Blackwell at 00000001:5C:00.0 with 97887MiB visible, no NVRM errors in
dmesg, single DKMS version registered.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- ansible/README.md: drop removed shutdown playbook mention
- CLAUDE.md: list nvidia_gpu under custom roles
- sftp README: link the config notes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- run.yml: base plays (geerlingguy.security) for jira and git; gpu01
play with security + docker + nvidia_gpu
- roles/nvidia_gpu: driver pinned >=580 (Blackwell), CUDA repo,
container toolkit incl. the nvidia-ctk runtime configure step
- manuals/20260714-nvidia-driver-install.md: dated per convention,
corrected (pinned -server driver instead of autoinstall+cuda-drivers
mix, toolkit optional, added missing nvidia-ctk/docker restart step)
- gpu01 folder: planning docs under notes/, runbooks under manuals/,
scripts/; convention documented in CLAUDE.md
- scripts/share-analysis.ps1: read-only SMB share analysis for the
Windows server (projektplan §2.1)
- TODO.md: Phase-0 pre-work items from the projektplan
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- sftp IP 192.168.99.68 in CLAUDE.md and server README (DMZ move)
- ansible.cfg: ansible_python_interpreter is a host var, moved to
group_vars/all.yml; [ssh_connections] -> [ssh_connection] so
pipelining actually applies
- remove playbooks/shutdown.yml (targeted nonexistent k3s_cluster group)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- hosts.ini: underscore group names, hostname aliases with ansible_host
- move ansible_user/ansible_port to group_vars/all.yml
- rename group_vars files to match underscore group names
- trim sftp group_vars to its only override (password auth off)
- run.yml: load moved secrets file (group_vars/secrets.yml)
- untrack .DS_Store, extend .gitignore
- prefill root/ansible/server READMEs, add jira + cloud server folders
- update CLAUDE.md to match
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Rename inventory groups and group_vars to hostnames (phy-z-*)
- Add group_vars for all servers incl. planned phy-z-srv-gpu01
- Add server/ docs: sftp01, git, gpu01 (README, HW specs, assessments)
- Add rules set to CLAUDE.md
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>