Files
CubelaPetarandClaude Opus 4.8 b6b242196d Add LLM compose stack for phy-srv-gpu01 (not deployed yet)
Follows the homelab pattern: ironicbadger.docker_compose_generator v2
renders services/<host>/NN-<stack>/compose.yml templates into
~/docker/compose.yaml on the host.

- 01-vllm: chat model, fixed --gpu-memory-utilization
- 02-embeddings: second vLLM instance (--task embed) rather than a
  separate toolchain, so SM120 support only has to be solved once
- 03-openwebui: Open WebUI + pgvector (not chroma — corpus size)
- 99-network: shared bridge; leading comment keeps networks: top-level
- pin docker_compose_generator to 2.0.1 — galaxy tags mix v1/v2 formats
- group_vars: stack config incl. LDAP placeholders still to be filled

The role only writes the compose file; starting the stack stays manual.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 14:53:59 +02:00

37 lines
1.1 KiB
YAML

services:
vllm:
image: "{{ vllm_image }}"
container_name: vllm
networks:
- llmnet
ports:
- "8000:8000"
volumes:
# model cache on local disk — tens of GB per model
- "{{ appdata_path }}/models/huggingface:/root/.cache/huggingface"
environment:
- NVIDIA_VISIBLE_DEVICES=all
- NVIDIA_DRIVER_CAPABILITIES=compute,utility
- "HUGGING_FACE_HUB_TOKEN={{ hf_token | default('') }}"
command:
- --model
- "{{ vllm_model }}"
- --served-model-name
- "{{ vllm_served_model_name }}"
# fixed VRAM share so the embedding server keeps its slice (§0: "it just works")
- --gpu-memory-utilization
- "{{ vllm_gpu_memory_utilization }}"
- --max-model-len
- "{{ vllm_max_model_len }}"
# vLLM needs a large shared-memory segment; without this it dies on startup
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
runtime: nvidia
restart: unless-stopped