30 shares, 2128 GB, 3.11M files. Corrects the earlier EDV assessment
(446 GB / 64k files = installers, not documents) and records that
Zeichnungen measured empty. Excluding the obvious shares still leaves
2.08M files, so filtering must continue at subfolder level.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Follows the homelab pattern: ironicbadger.docker_compose_generator v2
renders services/<host>/NN-<stack>/compose.yml templates into
~/docker/compose.yaml on the host.
- 01-vllm: chat model, fixed --gpu-memory-utilization
- 02-embeddings: second vLLM instance (--task embed) rather than a
separate toolchain, so SM120 support only has to be solved once
- 03-openwebui: Open WebUI + pgvector (not chroma — corpus size)
- 99-network: shared bridge; leading comment keeps networks: top-level
- pin docker_compose_generator to 2.0.1 — galaxy tags mix v1/v2 formats
- group_vars: stack config incl. LDAP placeholders still to be filled
The role only writes the compose file; starting the stack stays manual.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- rename host/group/folder everywhere to match the server's actual
hostname and physical label
- share-analysis.ps1: sanitize the output filename prefix — '-Paths "D:"'
produced 'D:-file-types.csv' and Export-Csv failed with 'path format
not supported', so no CSVs were written
- share-analysis.ps1: new -FolderDepth so folders can be aggregated at
D:\Abteilungen\<Share> level, which matches the share layout
- list-shares.ps1: -WithSize walks local paths instead of UNC when run on
the server itself (UNC was orders of magnitude slower and looked stuck)
and prints progress every 50k files
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- projektplan: status section (server delivered, driver done), share
analysis results in §2.1, customer checklist answers in §2.2, storage
resolved in §2.3, new §2.6 on scoping the corpus, updated risks
- scripts/list-shares.ps1: enumerate SMB shares incl. paths, permissions
and DFS namespaces on the file server
- TODO.md: restructured into blocking/server/done
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The gpu role pinned nvidia-driver-580-server (proprietary kernel
modules). Those install and load cleanly on the RTX PRO 6000, but the
GPU never initialises: NVRM fails at RmInitAdapter (0x22:0x56:897) and
nvidia-smi reports "No devices were found". Blackwell (10de:2bb5)
requires the open kernel modules.
- defaults/main.yml: nvidia-driver-580-server -> 595-server-open.
The "-open" part is the actual fix and is version independent. 595 is
the newest -server branch in the 24.04 archive (610 is desktop only);
projektplan §5.3 asked for the branch to be re-checked and pinned at
install time, which had not happened yet. Caveat noted in the file:
595 ships from multiverse, 580/590 from restricted.
- tasks/main.yml: install_recommends: false, keeps nvidia-settings and
its GTK chain off a headless server (~65 packages).
- manuals/20260714-nvidia-driver-install.md: same command, documents
the open-module requirement and the failure signature.
- SETUP.md: tick off the completed OS install steps.
Verified on phy-z-srv-gpu01: driver 595.71.05, CUDA 13.2, RTX PRO 6000
Blackwell at 00000001:5C:00.0 with 97887MiB visible, no NVRM errors in
dmesg, single DKMS version registered.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- ansible/README.md: drop removed shutdown playbook mention
- CLAUDE.md: list nvidia_gpu under custom roles
- sftp README: link the config notes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- run.yml: base plays (geerlingguy.security) for jira and git; gpu01
play with security + docker + nvidia_gpu
- roles/nvidia_gpu: driver pinned >=580 (Blackwell), CUDA repo,
container toolkit incl. the nvidia-ctk runtime configure step
- manuals/20260714-nvidia-driver-install.md: dated per convention,
corrected (pinned -server driver instead of autoinstall+cuda-drivers
mix, toolkit optional, added missing nvidia-ctk/docker restart step)
- gpu01 folder: planning docs under notes/, runbooks under manuals/,
scripts/; convention documented in CLAUDE.md
- scripts/share-analysis.ps1: read-only SMB share analysis for the
Windows server (projektplan §2.1)
- TODO.md: Phase-0 pre-work items from the projektplan
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- sftp IP 192.168.99.68 in CLAUDE.md and server README (DMZ move)
- ansible.cfg: ansible_python_interpreter is a host var, moved to
group_vars/all.yml; [ssh_connections] -> [ssh_connection] so
pipelining actually applies
- remove playbooks/shutdown.yml (targeted nonexistent k3s_cluster group)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Open WebUI over Onyx (verified: Onyx file connector is manual-upload
only, SSO/RBAC are EE-paid), share analysis as mandatory pre-work,
phased setup via Ansible, re-entry checklist for when the server
arrives in a few months.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- hosts.ini: underscore group names, hostname aliases with ansible_host
- move ansible_user/ansible_port to group_vars/all.yml
- rename group_vars files to match underscore group names
- trim sftp group_vars to its only override (password auth off)
- run.yml: load moved secrets file (group_vars/secrets.yml)
- untrack .DS_Store, extend .gitignore
- prefill root/ansible/server READMEs, add jira + cloud server folders
- update CLAUDE.md to match
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Rename inventory groups and group_vars to hostnames (phy-z-*)
- Add group_vars for all servers incl. planned phy-z-srv-gpu01
- Add server/ docs: sftp01, git, gpu01 (README, HW specs, assessments)
- Add rules set to CLAUDE.md
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>