Commit Graph
4 Commits
Author SHA1 Message Date
CubelaPetarandClaude Opus 4.8 75296ce58c Rename phy-z-srv-gpu01 to phy-srv-gpu01; fix share analysis scripts
- rename host/group/folder everywhere to match the server's actual
  hostname and physical label
- share-analysis.ps1: sanitize the output filename prefix — '-Paths "D:"'
  produced 'D:-file-types.csv' and Export-Csv failed with 'path format
  not supported', so no CSVs were written
- share-analysis.ps1: new -FolderDepth so folders can be aggregated at
  D:\Abteilungen\<Share> level, which matches the share layout
- list-shares.ps1: -WithSize walks local paths instead of UNC when run on
  the server itself (UNC was orders of magnitude slower and looked stuck)
  and prints progress every 50k files

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 14:47:07 +02:00
CubelaPetarandClaude Opus 5 e7ef29c5d5 Fix Blackwell GPU driver: open kernel modules, bump pin to 595
The gpu role pinned nvidia-driver-580-server (proprietary kernel
modules). Those install and load cleanly on the RTX PRO 6000, but the
GPU never initialises: NVRM fails at RmInitAdapter (0x22:0x56:897) and
nvidia-smi reports "No devices were found". Blackwell (10de:2bb5)
requires the open kernel modules.

- defaults/main.yml: nvidia-driver-580-server -> 595-server-open.
  The "-open" part is the actual fix and is version independent. 595 is
  the newest -server branch in the 24.04 archive (610 is desktop only);
  projektplan §5.3 asked for the branch to be re-checked and pinned at
  install time, which had not happened yet. Caveat noted in the file:
  595 ships from multiverse, 580/590 from restricted.
- tasks/main.yml: install_recommends: false, keeps nvidia-settings and
  its GTK chain off a headless server (~65 packages).
- manuals/20260714-nvidia-driver-install.md: same command, documents
  the open-module requirement and the failure signature.
- SETUP.md: tick off the completed OS install steps.

Verified on phy-z-srv-gpu01: driver 595.71.05, CUDA 13.2, RTX PRO 6000
Blackwell at 00000001:5C:00.0 with 97887MiB visible, no NVRM errors in
dmesg, single DKMS version registered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 09:30:02 +02:00
CubelaPetarandClaude Fable 5 43054f32b8 Add base plays for all hosts, nvidia_gpu role, GPU pre-work items
- run.yml: base plays (geerlingguy.security) for jira and git; gpu01
  play with security + docker + nvidia_gpu
- roles/nvidia_gpu: driver pinned >=580 (Blackwell), CUDA repo,
  container toolkit incl. the nvidia-ctk runtime configure step
- manuals/20260714-nvidia-driver-install.md: dated per convention,
  corrected (pinned -server driver instead of autoinstall+cuda-drivers
  mix, toolkit optional, added missing nvidia-ctk/docker restart step)
- gpu01 folder: planning docs under notes/, runbooks under manuals/,
  scripts/; convention documented in CLAUDE.md
- scripts/share-analysis.ps1: read-only SMB share analysis for the
  Windows server (projektplan §2.1)
- TODO.md: Phase-0 pre-work items from the projektplan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 10:30:43 +02:00
CubelaPetar 652fba50c1 first commit 2026-07-10 09:57:37 +02:00