Files
infra-phytron/server/phy-z-srv-gpu01/README.md
T
CubelaPetarandClaude Fable 5 43054f32b8 Add base plays for all hosts, nvidia_gpu role, GPU pre-work items
- run.yml: base plays (geerlingguy.security) for jira and git; gpu01
  play with security + docker + nvidia_gpu
- roles/nvidia_gpu: driver pinned >=580 (Blackwell), CUDA repo,
  container toolkit incl. the nvidia-ctk runtime configure step
- manuals/20260714-nvidia-driver-install.md: dated per convention,
  corrected (pinned -server driver instead of autoinstall+cuda-drivers
  mix, toolkit optional, added missing nvidia-ctk/docker restart step)
- gpu01 folder: planning docs under notes/, runbooks under manuals/,
  scripts/; convention documented in CLAUDE.md
- scripts/share-analysis.ps1: read-only SMB share analysis for the
  Windows server (projektplan §2.1)
- TODO.md: Phase-0 pre-work items from the projektplan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 10:30:43 +02:00

1.2 KiB
Raw Blame History

phy-z-srv-gpu01

GPU server for AI/ML workloads. Hardware is ordered/assessed; OS setup and configuration are the next step.

IP 192.168.66.69 (planned)
Ansible group phy_z_srv_gpu01
Status planned — not yet configured

Hardware

HPE ProLiant DL380 Gen12, 2× Intel Xeon 6714P (8-core, 4.0 GHz), 128 GB RAM, NVIDIA RTX PRO 6000 96 GB, 2× 960 GB NVMe SSD, redundant PSU — full BOM in HW.md.

Planning notes

Open pre-work items are tracked in the repo-root TODO.md.

Runbooks

Scripts

  • share-analysis.ps1 — read-only SMB share analysis (projektplan §2.1); run on a Windows machine with read access to the shares