fix(ollama): restore GPU reservation on SOLARIA, close 07-15 cutover docs
Driver fixed 2026-07-16 (nvidia-driver-595-open from distro repo, old graphics-drivers PPA for jammy disabled) — RTX 4070 Ti SUPER 16GB, CUDA 13.2, nvidia-container-toolkit already present from the 07-15 cutover prerequisite install. Uncomments deploy.resources.reservations (nvidia gpu) in services/ollama/docker-compose.yml, restoring clean formatting. Docs close out the loose ends from the 07-15 declarative cutover: - ollama-solaria-cutover runbook gets a "Wykonanie" section documenting what actually happened (mv instead of rsync for the model store, the missing nvidia-container-toolkit prerequisite, the driver-missing discovery, CPU-only cutover, then the 07-16 GPU fix) plus a note on the container disappearing after the 07-15 evening reboot (one-off, boots fine now, root cause not established). - hosts/solaria/services.yaml: ollama role comment now reflects actual GPU-backed state instead of the previously-aspirational wording. - hosts/solaria/README.md: drop stale Open WebUI mention (not in repo). - docs/backlog.md: close the NVIDIA driver item; leave two follow-ups (Ollama call batching before the mail phase, UNIQUE(envelope_id, chunk_index) schema change for multi-model embeddings at phase 3). - jobs/documents-ingest/README.md: timing section gets a GPU placeholder line next to the existing 0.79s/chunk CPU baseline, to be filled in after the live GPU benchmark. Live recreate + GPU-vs-CPU embedding benchmark deliberately left for the operator to run from the main checkout after merge, per worktree-aware discipline — this worktree only owns the declarative fix and the docs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
parent
542bba4693
commit
5e920dd7c1
|
|
@ -740,3 +740,33 @@ compose. Docker cache'uje obraz po tagu, nie po zawartości src/.
|
|||
`docker compose up -d --build` (przebuduje gdy src się zmienił; no-op gdy nie). Ewentualnie
|
||||
`--force-recreate` gdy zmienił się env-file. Bez tego każdy code-only deploy wymaga
|
||||
ręcznego rebuild — łatwo przeoczyć (deploy mówi green).
|
||||
|
||||
## Ollama SOLARIA: brak sterownika NVIDII — ZAMKNIĘTE (2026-07-16)
|
||||
|
||||
**Kontekst.** Cutover 2026-07-15 (`docs/infra/ollama-solaria-cutover-2026-07-15.md`)
|
||||
odkrył, że SOLARIA nie miała zainstalowanego żadnego sterownika NVIDII —
|
||||
`nvidia-smi` nie istniał na hoście. `hosts/solaria/services.yaml` opisywał
|
||||
ollama jako "GPU-backed" od dawna, ale to było aspiracyjne — Ollama zawsze
|
||||
szła CPU-only. GPU reservation zakomentowana w
|
||||
`services/ollama/docker-compose.yml` (`f57a01a`); item trafił do backlogu
|
||||
jako blokujący fazę mailową embeddingów (moduł 5).
|
||||
|
||||
**Fix (2026-07-16).** Zainstalowany `nvidia-driver-595-open` z repo dystrybucji
|
||||
(nie stary PPA `graphics-drivers` dla jammy — zdezaktywowany przez rename na
|
||||
`.disabled`). RTX 4070 Ti SUPER 16GB, CUDA 13.2, `nvidia-smi` działa na
|
||||
hoście. `nvidia-container-toolkit` był już obecny (doinstalowany jako
|
||||
prerequisite przy cutoverze 07-15). GPU reservation przywrócona w compose.
|
||||
Pomiar throughput GPU vs CPU baseline (0.79s/chunk) — patrz
|
||||
`jobs/documents-ingest/README.md`, sekcja timing.
|
||||
|
||||
**Status:** ZAMKNIĘTE.
|
||||
|
||||
**Follow-upy pozostawione (osobne taski):**
|
||||
- **Batching wywołań Ollamy** — przed fazą mailową (225k kopert). Sekwencyjne
|
||||
wywołania `/api/embeddings` (nawet na GPU) będą wąskim gardłem przy takiej
|
||||
skali; ocenić równoległość/batch API Ollamy.
|
||||
- **`UNIQUE(envelope_id, chunk_index)` bez `model`** w `document_chunk`
|
||||
(`services/kb-postgres/init/002_chunks.sql`) — re-embedding innym modelem
|
||||
cicho no-opuje się przez istniejący constraint. Schema change do zrobienia
|
||||
przy fazie 3 (patrz `jobs/documents-ingest/README.md`, sekcja "Idempotency"
|
||||
kroku 6 embed).
|
||||
|
|
|
|||
|
|
@ -1,9 +1,8 @@
|
|||
# Ollama SOLARIA: manual → declarative cutover runbook
|
||||
|
||||
Date: 2026-07-15
|
||||
Status: repo changes done (this doc), **live cutover not yet executed** — no SSH
|
||||
access to SOLARIA from the worktree this was prepared in. Run on SATURN or by
|
||||
the operator directly on SOLARIA.
|
||||
Status: **executed** 2026-07-15 (CPU-only, see "Wykonanie" below); GPU restored
|
||||
2026-07-16 once the missing NVIDIA driver was fixed.
|
||||
|
||||
## Background
|
||||
|
||||
|
|
@ -110,3 +109,52 @@ sudo systemctl enable --now ollama
|
|||
```
|
||||
The native install's model directory was copied (not moved) in step 3, so it
|
||||
is untouched and the manual instance comes back exactly as it was.
|
||||
|
||||
*(See "Wykonanie" below — the actual run used `mv`, not `rsync`, so this
|
||||
rollback description no longer matches what happened live.)*
|
||||
|
||||
## Wykonanie (2026-07-15, GPU restored 2026-07-16)
|
||||
|
||||
Deviations from the plan above, discovered while executing it live on
|
||||
SOLARIA:
|
||||
|
||||
- **Step 3 (relocate model store) used `mv`, not `rsync -a`.** The model
|
||||
store (36G) was moved in one shot rather than copied. Consequence: the
|
||||
rollback described above ("source untouched since copied") does not apply
|
||||
as executed — a real rollback would need the reverse `mv`
|
||||
(`/opt/homelab/data/ollama/* → /usr/share/ollama/.ollama/`), not just
|
||||
re-enabling the native systemd service against an already-relocated
|
||||
directory.
|
||||
- **Step 2 pre-check (`nvidia-container-toolkit`) failed the first time**: it
|
||||
was not installed. Installed as a prerequisite before continuing — this
|
||||
wasn't previously flagged as something that could be missing.
|
||||
- **Step 2 pre-check (driver) failed harder than expected**: `nvidia-smi` did
|
||||
not exist on the host at all — no NVIDIA driver installed on SOLARIA.
|
||||
`hosts/solaria/services.yaml` had described ollama as GPU-backed for a
|
||||
while; that was aspirational, not actual — Ollama had been running CPU-only
|
||||
the entire time regardless of what the manifest said.
|
||||
- Given the missing driver, the cutover proceeded **in CPU-only mode**: the
|
||||
`deploy.resources` GPU reservation was commented out in
|
||||
`services/ollama/docker-compose.yml` (commit `f57a01a`), and the driver fix
|
||||
was filed as a backlog item (see `docs/backlog.md`) blocking the module 5
|
||||
mail-embedding phase.
|
||||
- **2026-07-16: driver fixed.** Installed `nvidia-driver-595-open` from the
|
||||
distro repository — not the old `ppa:graphics-drivers/ppa` (jammy), which
|
||||
was deactivated by renaming its `sources.list.d` entry to `.disabled`.
|
||||
Result: RTX 4070 Ti SUPER 16GB, CUDA 13.2, `nvidia-smi` working on the
|
||||
host. `nvidia-container-toolkit` from the earlier prerequisite install was
|
||||
already in place (`nvidia` runtime registered in `daemon.json`), so no
|
||||
further toolkit work was needed. The GPU reservation was restored in
|
||||
`docker-compose.yml` and the container recreated.
|
||||
- **GPU vs CPU embedding throughput**: GPU — TBD after measurement
|
||||
2026-07-16 (CPU baseline: ≈0.79s/chunk, see
|
||||
`jobs/documents-ingest/README.md` timing section, updated with the GPU
|
||||
number once measured).
|
||||
|
||||
## Nota operacyjna: kontener zniknął po reboocie (2026-07-15 wieczór)
|
||||
|
||||
Po reboocie hosta wieczorem 2026-07-15 kontener `ollama` **zniknął**
|
||||
(`docker ps -a` nie pokazywał go w ogóle — nie `Exited`), mimo
|
||||
`restart: unless-stopped`. Jednorazowe zdarzenie: po boocie 2026-07-16
|
||||
kontener wstał poprawnie i automatycznie. Przyczyna nie została odtworzona —
|
||||
flagowane jako obserwacja, nie bug do śledzenia, chyba że się powtórzy.
|
||||
|
|
|
|||
|
|
@ -6,7 +6,6 @@
|
|||
|
||||
## Configured Services
|
||||
- Ollama
|
||||
- Open WebUI
|
||||
|
||||
## Runtime Data
|
||||
- `/opt/homelab/data/ollama`
|
||||
|
|
|
|||
|
|
@ -15,7 +15,7 @@ services:
|
|||
logs_path: /opt/homelab/events
|
||||
|
||||
ollama:
|
||||
role: llm-inference # GPU-backed inference: embeddings (bge-m3) + coder models
|
||||
role: llm-inference # GPU-backed inference (RTX 4070 Ti SUPER, driver restored 2026-07-16): embeddings (bge-m3) + coder models
|
||||
deployment_model: docker-compose
|
||||
exposure: private # Tailscale-only bind (TAILSCALE_BIND_IP) + loopback; consumed by llm-gateway@PIHA over mesh
|
||||
offline_required: false
|
||||
|
|
|
|||
|
|
@ -464,8 +464,8 @@ kb-postgres@PIHA:
|
|||
average across the 160 chunked documents, ≈35 minutes wall-clock for the
|
||||
full 186-document pilot. This is the real-world input for scaling this
|
||||
pipeline to the much larger mail corpus later (plan §7 assumed GPU-based
|
||||
"minutes for the whole pilot"; SOLARIA's Ollama currently runs CPU-only
|
||||
per the recent GPU-reservation-disabled fix). The 186-document pilot's
|
||||
"minutes for the whole pilot"; SOLARIA's Ollama ran CPU-only for this pilot
|
||||
per the then-disabled GPU reservation). The 186-document pilot's
|
||||
≈13.2s/document average is dominated by Paperless' long OCR text (≈22k
|
||||
chars/doc average, per plan §1.2) — 225 030 mail envelopes will have a
|
||||
very different, likely much shorter, per-envelope chunk count (email
|
||||
|
|
@ -475,3 +475,10 @@ kb-postgres@PIHA:
|
|||
average chunk count per item will need either a GPU driver fix,
|
||||
concurrent/batched Ollama calls, or both, before a full mail-corpus run
|
||||
is practical — flagged for whoever picks up the mail-indexer phase.
|
||||
- **GPU (RTX 4070 Ti SUPER, driver 595-open, restored 2026-07-16)**: TBD po
|
||||
pomiarze 2026-07-16 — sterownik NVIDII naprawiony i GPU reservation
|
||||
przywrócona w `services/ollama/docker-compose.yml`
|
||||
(`docs/backlog.md`, "Ollama SOLARIA: brak sterownika NVIDII"), ale
|
||||
sekwencyjny embedding benchmark vs. ≈0.79s/chunk CPU baseline jeszcze nie
|
||||
zmierzony w tym repo — liczba do uzupełnienia osobnym commitem po
|
||||
pomiarze na żywym Ollamie.
|
||||
|
|
|
|||
|
|
@ -13,12 +13,13 @@ services:
|
|||
- "${TAILSCALE_BIND_IP}:11434:11434"
|
||||
volumes:
|
||||
- /opt/homelab/data/ollama:/root/.ollama
|
||||
# GPU disabled 2026-07-15: brak sterownika NVIDII na hoście (nvidia-smi not found).
|
||||
# Przywrócić po naprawie sterownika — task w backlogu.
|
||||
# deploy:
|
||||
# resources:
|
||||
# reservations:
|
||||
# devices:
|
||||
# - driver: nvidia
|
||||
# count: all
|
||||
# capabilities: [gpu]
|
||||
# GPU przywrócone 2026-07-16: sterownik nvidia-driver-595-open (repo distro)
|
||||
# zainstalowany, CUDA 13.2, nvidia-container-toolkit już obecny. Patrz
|
||||
# docs/infra/ollama-solaria-cutover-2026-07-15.md dla historii cutoveru.
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
count: all
|
||||
capabilities: [gpu]
|
||||
|
|
|
|||
Loading…
Reference in a new issue