Commit Graph
20 Commits
Author SHA1 Message Date
CubelaPetarandClaude Opus 4.8 80a862efaf Document compose-template approach in projektplan 2.4
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 14:55:54 +02:00
CubelaPetarandClaude Opus 4.8 922352ae61 Update TODO after compose stack and rename
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 14:54:34 +02:00
CubelaPetarandClaude Opus 4.8 b6b242196d Add LLM compose stack for phy-srv-gpu01 (not deployed yet)
Follows the homelab pattern: ironicbadger.docker_compose_generator v2
renders services/<host>/NN-<stack>/compose.yml templates into
~/docker/compose.yaml on the host.

- 01-vllm: chat model, fixed --gpu-memory-utilization
- 02-embeddings: second vLLM instance (--task embed) rather than a
  separate toolchain, so SM120 support only has to be solved once
- 03-openwebui: Open WebUI + pgvector (not chroma — corpus size)
- 99-network: shared bridge; leading comment keeps networks: top-level
- pin docker_compose_generator to 2.0.1 — galaxy tags mix v1/v2 formats
- group_vars: stack config incl. LDAP placeholders still to be filled

The role only writes the compose file; starting the stack stays manual.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 14:53:59 +02:00
CubelaPetarandClaude Opus 4.8 75296ce58c Rename phy-z-srv-gpu01 to phy-srv-gpu01; fix share analysis scripts
- rename host/group/folder everywhere to match the server's actual
  hostname and physical label
- share-analysis.ps1: sanitize the output filename prefix — '-Paths "D:"'
  produced 'D:-file-types.csv' and Export-Csv failed with 'path format
  not supported', so no CSVs were written
- share-analysis.ps1: new -FolderDepth so folders can be aggregated at
  D:\Abteilungen\<Share> level, which matches the share layout
- list-shares.ps1: -WithSize walks local paths instead of UNC when run on
  the server itself (UNC was orders of magnitude slower and looked stuck)
  and prints progress every 50k files

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 14:47:07 +02:00
CubelaPetarandClaude Opus 4.8 9c38c60783 Record model decision (Qwen3-32B FP8) and vLLM SM120 caveat
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 14:13:11 +02:00
CubelaPetarandClaude Opus 4.8 be6efa7b34 Update projektplan and TODO with share analysis results
- projektplan: status section (server delivered, driver done), share
  analysis results in §2.1, customer checklist answers in §2.2, storage
  resolved in §2.3, new §2.6 on scoping the corpus, updated risks
- scripts/list-shares.ps1: enumerate SMB shares incl. paths, permissions
  and DFS namespaces on the file server
- TODO.md: restructured into blocking/server/done

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 14:10:14 +02:00
CubelaPetarandClaude Opus 5 e7ef29c5d5 Fix Blackwell GPU driver: open kernel modules, bump pin to 595
The gpu role pinned nvidia-driver-580-server (proprietary kernel
modules). Those install and load cleanly on the RTX PRO 6000, but the
GPU never initialises: NVRM fails at RmInitAdapter (0x22:0x56:897) and
nvidia-smi reports "No devices were found". Blackwell (10de:2bb5)
requires the open kernel modules.

- defaults/main.yml: nvidia-driver-580-server -> 595-server-open.
  The "-open" part is the actual fix and is version independent. 595 is
  the newest -server branch in the 24.04 archive (610 is desktop only);
  projektplan §5.3 asked for the branch to be re-checked and pinned at
  install time, which had not happened yet. Caveat noted in the file:
  595 ships from multiverse, 580/590 from restricted.
- tasks/main.yml: install_recommends: false, keeps nvidia-settings and
  its GTK chain off a headless server (~65 packages).
- manuals/20260714-nvidia-driver-install.md: same command, documents
  the open-module requirement and the failure signature.
- SETUP.md: tick off the completed OS install steps.

Verified on phy-z-srv-gpu01: driver 595.71.05, CUDA 13.2, RTX PRO 6000
Blackwell at 00000001:5C:00.0 with 97887MiB visible, no NVRM errors in
dmesg, single DKMS version registered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 09:30:02 +02:00
CubelaPetar 3160c6ca28 add new steps while setting up the os initally 2026-08-27 10:54:22 +02:00
CubelaPetar 01f67cac89 added new step 2026-08-26 16:25:14 +02:00
CubelaPetar f4ea864841 create file to document the setup in the server for a future SOP 2026-08-26 15:30:48 +02:00
CubelaPetarandClaude Fable 5 37c5f8c9de Sync docs after playbook/structure changes
- ansible/README.md: drop removed shutdown playbook mention
- CLAUDE.md: list nvidia_gpu under custom roles
- sftp README: link the config notes

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 10:43:20 +02:00
CubelaPetarandClaude Fable 5 43054f32b8 Add base plays for all hosts, nvidia_gpu role, GPU pre-work items
- run.yml: base plays (geerlingguy.security) for jira and git; gpu01
  play with security + docker + nvidia_gpu
- roles/nvidia_gpu: driver pinned >=580 (Blackwell), CUDA repo,
  container toolkit incl. the nvidia-ctk runtime configure step
- manuals/20260714-nvidia-driver-install.md: dated per convention,
  corrected (pinned -server driver instead of autoinstall+cuda-drivers
  mix, toolkit optional, added missing nvidia-ctk/docker restart step)
- gpu01 folder: planning docs under notes/, runbooks under manuals/,
  scripts/; convention documented in CLAUDE.md
- scripts/share-analysis.ps1: read-only SMB share analysis for the
  Windows server (projektplan §2.1)
- TODO.md: Phase-0 pre-work items from the projektplan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 10:30:43 +02:00
CubelaPetarandClaude Fable 5 277f378fb8 Fix stale sftp IP, ansible.cfg keys, drop dead shutdown playbook
- sftp IP 192.168.99.68 in CLAUDE.md and server README (DMZ move)
- ansible.cfg: ansible_python_interpreter is a host var, moved to
  group_vars/all.yml; [ssh_connections] -> [ssh_connection] so
  pipelining actually applies
- remove playbooks/shutdown.yml (targeted nonexistent k3s_cluster group)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 10:30:43 +02:00
CubelaPetarandClaude Fable 5 b60bc32391 Add sftp config notes and GPU driver install notes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 09:54:30 +02:00
CubelaPetar 4e3594f8f8 changed sftp ip in hosts file as it lives in dmz and not in srv vlan 2026-07-14 09:47:46 +02:00
CubelaPetarandClaude Fable 5 e685bb6e99 Add project plan for phy-z-srv-gpu01 setup
Open WebUI over Onyx (verified: Onyx file connector is manual-upload
only, SSO/RBAC are EE-paid), share analysis as mandatory pre-work,
phased setup via Ansible, re-entry checklist for when the server
arrives in a few months.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 16:41:33 +02:00
CubelaPetarandClaude Fable 5 fbafe9670c Clean up inventory, group_vars, and prefill READMEs
- hosts.ini: underscore group names, hostname aliases with ansible_host
- move ansible_user/ansible_port to group_vars/all.yml
- rename group_vars files to match underscore group names
- trim sftp group_vars to its only override (password auth off)
- run.yml: load moved secrets file (group_vars/secrets.yml)
- untrack .DS_Store, extend .gitignore
- prefill root/ansible/server READMEs, add jira + cloud server folders
- update CLAUDE.md to match

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 14:40:00 +02:00
CubelaPetarandClaude Fable 5 2bc2c476c4 Document repo structure and conventions in CLAUDE.md
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 14:01:54 +02:00
CubelaPetarandClaude Fable 5 82e10aa31f Restructure inventory per host, add server docs
- Rename inventory groups and group_vars to hostnames (phy-z-*)
- Add group_vars for all servers incl. planned phy-z-srv-gpu01
- Add server/ docs: sftp01, git, gpu01 (README, HW specs, assessments)
- Add rules set to CLAUDE.md

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 14:00:57 +02:00
CubelaPetar 652fba50c1 first commit 2026-07-10 09:57:37 +02:00