Files
infra-phytron/TODO.md
T
2026-09-03 14:54:34 +02:00

46 lines
2.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TODO
## phy-srv-gpu01
Per [projektplan](server/phy-srv-gpu01/notes/20260710-projektplan.md) (§ references below).
Server is delivered; base setup + NVIDIA driver are done (see `SETUP.md`).
### Blocking the knowledge base (customer conversation)
- [ ] **Scope the corpus (§2.6)** — share analysis found 1.26M extractable documents,
~12× the planning threshold. Decide with the customer which shares/folders
actually form the knowledge base. Everything in Phase 5 depends on this.
- [ ] Evaluate the two CSVs still sitting on the file server: `D-toplevel-folders.csv`
(basis for the include/exclude list) and `D-file-types.csv` (resolve the
1.58M unidentified `other` files)
- [ ] Run `scripts/list-shares.ps1` on Z-FILESERVER — map shares → local paths,
so the analysis of `D:` can be tied to actual shares
- [ ] Data protection / works council: are chat logs stored? (§2.2) — before rollout
- [ ] AD details still needed: bind service account + base DN (group `llm_users` and
the read-only SMB account are agreed)
- [ ] German eval set (§2.5): 2030 Q&A with the customer — acceptance criterion
### Server / Ansible
- [ ] Add the three secrets to `group_vars/secrets.yml` (`just vault edit`):
`vault_owui_secret_key`, `vault_owui_db_password`, `vault_ldap_bind_password`
- [ ] Fill the real LDAP bind DN / base DN in `group_vars/phy_srv_gpu01.yml`
- [ ] Pin `vllm_image` to a build verified on SM120, then `just compose phy_srv_gpu01`
and start the stack manually (see `ansible/services/phy-srv-gpu01/README.md`)
- [ ] Role `cifs_mounts` — read-only mounts, credentials from `group_vars/secrets.yml`
- [ ] oikb sync (share → knowledge base) once the corpus scope is decided
- [ ] Smoke-test a vLLM container on the server (official images may not support SM120)
— pin the working image, then build `llm_stack`. Start model: Qwen3-32B FP8
- [ ] OCR pipeline for ~166k scanned PDFs (§2.1) — own work block, competes with vLLM for the GPU
- [ ] TLS: currently plain HTTP on `chat.phytron.local`; retrofit an internal CA
certificate (AD passwords travel in clear text until then)
- [ ] Finish base hardening in `SETUP.md` (ssh, updates) via `just run phy_srv_gpu01`
### Done
- [x] Renamed host/group/folder to `phy-srv-gpu01` everywhere
- [x] LLM compose stack written (`ansible/services/phy-srv-gpu01/`) — not deployed
- [x] Share analysis script → `scripts/share-analysis.ps1`, run on 2026-07-14 (§2.1)
- [x] Storage decision (§2.3) — 876 GB usable, no extra disks needed
- [x] NVIDIA driver via role `nvidia_gpu` (595 open kernel modules)