Rename phy-z-srv-gpu01 to phy-srv-gpu01; fix share analysis scripts
- rename host/group/folder everywhere to match the server's actual hostname and physical label - share-analysis.ps1: sanitize the output filename prefix — '-Paths "D:"' produced 'D:-file-types.csv' and Export-Csv failed with 'path format not supported', so no CSVs were written - share-analysis.ps1: new -FolderDepth so folders can be aggregated at D:\Abteilungen\<Share> level, which matches the share layout - list-shares.ps1: -WithSize walks local paths instead of UNC when run on the server itself (UNC was orders of magnitude slower and looked stuck) and prints progress every 50k files Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,87 @@
|
||||
# Manual nvidia driver, cuda repo and container toolkit
|
||||
|
||||
Automated by the Ansible role `ansible/roles/nvidia_gpu` — this manual documents the steps.
|
||||
|
||||
**Disable Secure Boot in BIOS first** (unsigned kernel modules won't load otherwise).
|
||||
|
||||
## NVIDIA driver
|
||||
|
||||
Check if GPUs are recognized by the base OS:
|
||||
```bash
|
||||
sudo lspci | grep -i nvidia
|
||||
```
|
||||
|
||||
Which should show some output if it finds nvidia devices.
|
||||
|
||||
Search for available drivers for your GPUs:
|
||||
```bash
|
||||
sudo ubuntu-drivers devices
|
||||
```
|
||||
|
||||
Install the driver pinned. The RTX PRO 6000 (Blackwell) needs **driver >= 580**
|
||||
and **only works with the open kernel modules** — the proprietary ones load but
|
||||
fail at `RmInitAdapter`, leaving `nvidia-smi` with "No devices were found".
|
||||
Use the `-server-open` variant and prefer a pinned install over
|
||||
`ubuntu-drivers autoinstall` so the choice is explicit and reproducible
|
||||
(re-check for a newer branch at install time):
|
||||
```bash
|
||||
sudo apt install -y --no-install-recommends nvidia-driver-595-server-open
|
||||
```
|
||||
|
||||
`--no-install-recommends` keeps `nvidia-settings` and its GTK dependency chain
|
||||
off this headless server.
|
||||
|
||||
> Note: do **not** additionally install `cuda-drivers` from the NVIDIA repo —
|
||||
> that would mix the Ubuntu-archive driver with the NVIDIA-repo driver and the
|
||||
> two can conflict. Pick one source; we use the Ubuntu archive.
|
||||
|
||||
Reboot the system for changes to take effect:
|
||||
```bash
|
||||
sudo reboot
|
||||
```
|
||||
|
||||
Show GPU stats with:
|
||||
```bash
|
||||
nvidia-smi
|
||||
```
|
||||
|
||||
## CUDA repository (toolkit optional)
|
||||
|
||||
Add the NVIDIA CUDA apt repository:
|
||||
|
||||
```bash
|
||||
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
|
||||
sudo dpkg -i cuda-keyring_1.1-1_all.deb
|
||||
sudo apt update
|
||||
```
|
||||
|
||||
The full CUDA toolkit is **not needed** for Docker-based workloads (vLLM etc. —
|
||||
the driver plus container toolkit suffice). Only if compiling on the host:
|
||||
```bash
|
||||
sudo apt install -y cuda-toolkit # meta package, pulls the current release
|
||||
```
|
||||
|
||||
## Container toolkit
|
||||
|
||||
Install the Nvidia Container toolkit:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg \
|
||||
&& curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
|
||||
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
|
||||
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
|
||||
sudo apt update
|
||||
sudo apt install -y nvidia-container-toolkit
|
||||
```
|
||||
|
||||
Configure Docker to use the NVIDIA runtime (writes `/etc/docker/daemon.json`) and restart it —
|
||||
without this step `docker run --gpus all` fails:
|
||||
```bash
|
||||
sudo nvidia-ctk runtime configure --runtime=docker
|
||||
sudo systemctl restart docker
|
||||
```
|
||||
|
||||
Test a simple cuda container and nvidia-smi command inside:
|
||||
```bash
|
||||
docker run --rm --gpus all nvidia/cuda:13.0.0-base-ubuntu24.04 nvidia-smi
|
||||
```
|
||||
Reference in New Issue
Block a user