Files
infra-phytron/server/phy-srv-gpu01
CubelaPetarandClaude Opus 4.8 b6b242196d Add LLM compose stack for phy-srv-gpu01 (not deployed yet)
Follows the homelab pattern: ironicbadger.docker_compose_generator v2
renders services/<host>/NN-<stack>/compose.yml templates into
~/docker/compose.yaml on the host.

- 01-vllm: chat model, fixed --gpu-memory-utilization
- 02-embeddings: second vLLM instance (--task embed) rather than a
  separate toolchain, so SM120 support only has to be solved once
- 03-openwebui: Open WebUI + pgvector (not chroma — corpus size)
- 99-network: shared bridge; leading comment keeps networks: top-level
- pin docker_compose_generator to 2.0.1 — galaxy tags mix v1/v2 formats
- group_vars: stack config incl. LDAP placeholders still to be filled

The role only writes the compose file; starting the stack stays manual.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 14:53:59 +02:00
..

phy-srv-gpu01

GPU server for AI/ML workloads. Hardware is ordered/assessed; OS setup and configuration are the next step.

IP 192.168.66.69 (planned)
Ansible group phy_srv_gpu01
Status planned — not yet configured

Hardware

HPE ProLiant DL380 Gen12, 2× Intel Xeon 6714P (8-core, 4.0 GHz), 128 GB RAM, NVIDIA RTX PRO 6000 96 GB, 2× 960 GB NVMe SSD, redundant PSU — full BOM in HW.md.

Planning notes

Open pre-work items are tracked in the repo-root TODO.md.

Runbooks

Scripts

  • share-analysis.ps1 — read-only SMB share analysis (projektplan §2.1); run on a Windows machine with read access to the shares