homelab-codex-ws/services/ai-cluster
oskar b124e54e66 feat(ai-cluster): manifest + zrodla pod SOLARIA (autoring, cutover NIE wykonany)
Stack ai-cluster (openclaw, codex/planner/service-ops workery, redis, mosquitto)
biega na VPS spoza GitOps, z /home/dockeruser/docker/ai-cluster/. Branch
feat/vps-service-migration (862c04a) probowal wciagnac go do GitOps *w miejscu*,
na VPS. Ten etap pomijamy: workloady sa compute'owe, VPS ma 4 GiB bez swapu i rolę
ingressu, SOLARIA ma 64 GiB i GPU. Migrujemy od razu na SOLARIA, VPS zostaje samym
NPM-em proxujacym po Tailscale.

Zero deployu. Kontenery na VPS nietkniete — recon byl read-only (katalog nalezy do
dockeruser, odczyt przez efemeryczny kontener z mountem :ro, bez dotykania
dzialajacych kontenerow). Plan przelaczenia: services/ai-cluster/CUTOVER.md.

Zrodla wciagniete do services/ai-cluster/src/ (openclaw, worker, mosquitto).
Pominiete: .env i mosquitto/passwd (sekrety), start-codex.sh* (smiec spoza stacku),
__pycache__.

Ustalenia z reconu, ktore zmienily manifest z 862c04a:

- codex-worker/ na VPS to NIE osobny build context ani duplikat worker/ — to
  martwy 609-bajtowy prototyp na `requests` (biblioteka nie wystepuje w zadnym
  requirements.txt), bez Dockerfile'a i bez referencji w compose. Realny
  codex-worker buduje sie z ./worker (Dockerfile default). Nie przenoszony.
- image: ai-cluster-* -> build: z src/openclaw i src/worker (3 Dockerfile'e nad
  jednym kontekstem: default/planner/service-ops).
- openclaw: usuniete publish 0.0.0.0:8000 -> "100.100.231.104:8000:8000";
  usunieta siec npm_default (nie istnieje na SOLARIA).
- mosquitto: bind 100.95.58.48 -> 100.100.231.104.
- mem_limity urealnione (256m workery/openclaw, 64m redis/mosquitto) i przeniesione
  do hosts/solaria/runtime/. Zmierzone RSS na VPS: workery 7-10 MiB, openclaw
  24.6 MiB, redis 4.8, mosquitto 4.2 — limity to zapas ~10x, nie reakcja na presje.
- healthcheck zostaje na python/urllib: obraz openclaw (python:3.12-slim) nie ma
  ani wget, ani curl — sprawdzone na zywym obrazie.
- GATEWAY_BASE_URL http://piha:8080 -> http://100.108.208.3:8080. Tailscale DNS
  jest WYLACZONY na SOLARIA (`tailscale dns status` -> disabled), wiec zadna nazwa
  MagicDNS sie nie rozwiazuje. Na VPS bare `piha` tez nie rozwiazuje sie (rc=2),
  czyli ten default byl martwy juz tam.
- telegram-bot nie przenoszony (na VPS nie biega zaden kontener). telegram_bot.py
  zostaje w src/, bo openclaw/Dockerfile go COPY-uje — bez niego build sie wywala.
- sekrety wylacznie przez env_file /opt/homelab/config/ai-cluster/.env, zero
  interpolacji ${VAR:?} — stack nie zalezy od cwd deployu. service-ops-worker
  montuje ten katalog pod ta sama sciezka absolutna, bo compose waliduje env_file
  na etapie parsowania i inaczej `docker compose ps` w tym kontenerze padnie.
- AGENT_ID zostaja `vps-*`: to cele routingu MQTT, nie hostnamey. Zmiana nazw
  lamie producentow adresujacych konkretnego agenta — osobne, skoordynowane zadanie.

Konsumenci mosquitto do aktualizacji: w repo BRAK. chelsty (z2m, frigate,
stability-agent) uzywaja wlasnego, host-networkowego brokera i musza zostac
offline-first. Lista i sposob potwierdzenia klientow spoza repo — CUTOVER.md krok 5.

DoD (docker build + smoke + pytest) swiadomie odroczone: zadanie zabrania budowania
obrazow. Pokrywaja to kroki 1-3 CUTOVER.md, do wykonania w sesji cutoveru.
Zweryfikowane tutaj: `docker compose config` z overrideem (rc=0, bindy host_ip,
mem_limity, merge volumes), skladnia YAML, `bash -n healthcheck.sh`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 18:30:00 +02:00
..
src feat(ai-cluster): manifest + zrodla pod SOLARIA (autoring, cutover NIE wykonany) 2026-07-27 18:30:00 +02:00
CUTOVER.md feat(ai-cluster): manifest + zrodla pod SOLARIA (autoring, cutover NIE wykonany) 2026-07-27 18:30:00 +02:00
docker-compose.yml feat(ai-cluster): manifest + zrodla pod SOLARIA (autoring, cutover NIE wykonany) 2026-07-27 18:30:00 +02:00
env.example feat(ai-cluster): manifest + zrodla pod SOLARIA (autoring, cutover NIE wykonany) 2026-07-27 18:30:00 +02:00
healthcheck.sh feat(ai-cluster): manifest + zrodla pod SOLARIA (autoring, cutover NIE wykonany) 2026-07-27 18:30:00 +02:00
README.md feat(ai-cluster): manifest + zrodla pod SOLARIA (autoring, cutover NIE wykonany) 2026-07-27 18:30:00 +02:00
service.yaml feat(ai-cluster): manifest + zrodla pod SOLARIA (autoring, cutover NIE wykonany) 2026-07-27 18:30:00 +02:00

ai-cluster

Small MQTT/Redis task pipeline: an API front-end (openclaw) hands tasks to role-scoped Python workers.

Status: authored, not deployed. The stack currently runs on the VPS, outside GitOps, from /home/dockeruser/docker/ai-cluster/. This directory is the target definition for SOLARIA. Nothing here has been built or started yet — see CUTOVER.md.

Why SOLARIA

These are compute workloads pointed at an LLM gateway. The VPS is a 4 GiB public-ingress box with no swap; SOLARIA is the 64 GiB compute/GPU node. The earlier plan (branch feat/vps-service-migration) was to bring the stack into GitOps in place on the VPS. That intermediate step is skipped: the stack moves straight to SOLARIA, and the VPS keeps only the ingress role via NPM.

Architecture

client ──HTTP──> openclaw ──┬── Redis  (primary queue, result store)
                            └── MQTT   (secondary transport, codex/tasks)
                                  │
                                  ├── codex-worker        (ROLE=dev)
                                  ├── planner-worker      (ROLE=planner)
                                  └── service-ops-worker  (ROLE=service-ops)

Tasks are routed by a target field: role:<ROLE> or an exact AGENT_ID. Results go back to Redis (result:{id}, read by GET /result/{id}) and to the codex/results MQTT topic.

Layout

Path What
src/openclaw/ FastAPI app + build context for the openclaw image
src/worker/ Three Dockerfiles over one context: Dockerfile (codex-worker), Dockerfile.planner, Dockerfile.service-ops
src/mosquitto/ mosquitto.conf + acl templates — passwd is generated on the node, never committed
docker-compose.yml Base stack; builds from src/, no registry pulls
../../hosts/solaria/runtime/ai-cluster/docker-compose.override.yml mem limits + the repo-path bind mount

Networking

What Bind Reachable from
openclaw API 100.100.231.104:8000 Tailscale mesh only
mosquitto 100.100.231.104:1883 Tailscale mesh only, auth required

Neither is published on 0.0.0.0. Public access to openclaw goes through NPM on the VPS, which proxies over Tailscale to 100.100.231.104:8000.

There is no loopback bind on 8000 — this deviates from the ollama compose on the same node, which binds both 127.0.0.1 and the Tailscale IP. healthcheck.sh therefore probes the Tailscale IP. If you'd rather have the loopback convenience bind, add it in the host override.

Gotcha: piha does not resolve

GATEWAY_BASE_URL defaults to http://100.108.208.3:8080 (PIHA's Tailscale IP), not http://piha:8080. Tailscale DNS is disabled on SOLARIA:

$ tailscale dns status
Tailscale DNS: disabled.

so no MagicDNS name resolves — not the short name, not the FQDN. The same check on the VPS shows piha never resolved there either, so the original default was already dead. Use IPs in this stack.

Secrets

Everything sensitive comes from /opt/homelab/config/ai-cluster/.env via env_file: — there is no ${VAR:?} interpolation in the compose file, so the stack does not care which directory you run docker compose from. Start from env.example.

Mosquitto's passwd is generated on the node and lives next to the config at /opt/homelab/config/ai-cluster/mosquitto/passwd.

Health

./healthcheck.sh          # containers + openclaw /health + mosquitto port

openclaw also has a container-level healthcheck. It uses python -c urllib... rather than wget/curl because the python:3.12-slim base has neither — verified against the running image.