2026-08-04 15:00:19 +02:00
|
|
|
|
---
|
|
|
|
|
|
okf: "0.1"
|
|
|
|
|
|
type: runbook
|
|
|
|
|
|
visibility: private
|
|
|
|
|
|
status: active
|
|
|
|
|
|
updated: 2026-07-16
|
|
|
|
|
|
links: []
|
|
|
|
|
|
---
|
|
|
|
|
|
|
2026-07-15 15:51:28 +02:00
|
|
|
|
# Ollama SOLARIA: manual → declarative cutover runbook
|
|
|
|
|
|
|
|
|
|
|
|
Date: 2026-07-15
|
fix(ollama): restore GPU reservation on SOLARIA, close 07-15 cutover docs
Driver fixed 2026-07-16 (nvidia-driver-595-open from distro repo, old
graphics-drivers PPA for jammy disabled) — RTX 4070 Ti SUPER 16GB, CUDA
13.2, nvidia-container-toolkit already present from the 07-15 cutover
prerequisite install. Uncomments deploy.resources.reservations (nvidia
gpu) in services/ollama/docker-compose.yml, restoring clean formatting.
Docs close out the loose ends from the 07-15 declarative cutover:
- ollama-solaria-cutover runbook gets a "Wykonanie" section documenting
what actually happened (mv instead of rsync for the model store, the
missing nvidia-container-toolkit prerequisite, the driver-missing
discovery, CPU-only cutover, then the 07-16 GPU fix) plus a note on
the container disappearing after the 07-15 evening reboot (one-off,
boots fine now, root cause not established).
- hosts/solaria/services.yaml: ollama role comment now reflects actual
GPU-backed state instead of the previously-aspirational wording.
- hosts/solaria/README.md: drop stale Open WebUI mention (not in repo).
- docs/backlog.md: close the NVIDIA driver item; leave two follow-ups
(Ollama call batching before the mail phase, UNIQUE(envelope_id,
chunk_index) schema change for multi-model embeddings at phase 3).
- jobs/documents-ingest/README.md: timing section gets a GPU placeholder
line next to the existing 0.79s/chunk CPU baseline, to be filled in
after the live GPU benchmark.
Live recreate + GPU-vs-CPU embedding benchmark deliberately left for the
operator to run from the main checkout after merge, per worktree-aware
discipline — this worktree only owns the declarative fix and the docs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-16 14:20:16 +02:00
|
|
|
|
Status: **executed** 2026-07-15 (CPU-only, see "Wykonanie" below); GPU restored
|
|
|
|
|
|
2026-07-16 once the missing NVIDIA driver was fixed.
|
2026-07-15 15:51:28 +02:00
|
|
|
|
|
|
|
|
|
|
## Background
|
|
|
|
|
|
|
|
|
|
|
|
`services/ollama/service.yaml` declared `owner_node: solaria` but was never
|
|
|
|
|
|
added to `hosts/solaria/services.yaml` — the file `deploy-node.sh` actually
|
|
|
|
|
|
reads to decide what to bring up on a host. That's the whole gap; Ollama has
|
|
|
|
|
|
been running natively/manually (systemd) the entire time instead. This change
|
|
|
|
|
|
adds the missing `hosts/solaria/services.yaml` entry and fixes the compose
|
|
|
|
|
|
port bind (see below) so `deploy-node.sh` can manage it going forward.
|
|
|
|
|
|
|
|
|
|
|
|
The running native instance already has models pulled, including large coder
|
|
|
|
|
|
models (`qwen2.5-coder:14b`, `qwen3-coder:30b`, `deepseek-coder:latest`,
|
|
|
|
|
|
`deepcoder:14b`) actively used by `llm-gateway@PIHA`, plus `bge-m3` (pulled
|
|
|
|
|
|
manually ahead of this task for the KB embeddings job). **Do not re-pull —
|
|
|
|
|
|
bind-mount the existing model directory.**
|
|
|
|
|
|
|
|
|
|
|
|
## Pre-cutover checks (read-only, on SOLARIA)
|
|
|
|
|
|
|
|
|
|
|
|
1. Confirm current model store path and owning user:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
sudo systemctl show ollama -p Environment # look for OLLAMA_MODELS override
|
|
|
|
|
|
sudo ls -la /usr/share/ollama/.ollama/models 2>&1 # default for systemd-installed ollama
|
|
|
|
|
|
ls -la ~/.ollama/models 2>&1 # default for a user-run instance
|
|
|
|
|
|
```
|
|
|
|
|
|
2. Confirm `nvidia-container-toolkit` is installed and Docker sees the GPU runtime:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
docker info | grep -i nvidia
|
|
|
|
|
|
nvidia-smi
|
|
|
|
|
|
```
|
|
|
|
|
|
3. Confirm the current OLLAMA_HOST bind (how llm-gateway@PIHA is reaching it today):
|
|
|
|
|
|
```bash
|
|
|
|
|
|
sudo systemctl show ollama -p Environment | grep -o 'OLLAMA_HOST=[^ ]*'
|
|
|
|
|
|
```
|
|
|
|
|
|
4. **Timing: do this when llm-gateway@PIHA is not actively serving a request.**
|
|
|
|
|
|
The cutover has a brief availability gap between disabling the native
|
|
|
|
|
|
service and the container passing its healthcheck — pick a quiet window,
|
|
|
|
|
|
don't run it mid-request.
|
|
|
|
|
|
|
|
|
|
|
|
## Cutover steps
|
|
|
|
|
|
|
|
|
|
|
|
1. `git pull` on SOLARIA (picks up this branch once merged to master).
|
|
|
|
|
|
2. Populate `services/ollama/.env` from `services/ollama/env.example`
|
|
|
|
|
|
(`TAILSCALE_BIND_IP=100.100.231.104` — verify with `tailscale ip -4`).
|
|
|
|
|
|
3. **Relocate the model store BEFORE the first `docker compose up`:**
|
|
|
|
|
|
```bash
|
|
|
|
|
|
sudo mkdir -p /opt/homelab/data/ollama
|
|
|
|
|
|
# Adjust source path per the pre-check above — this assumes the systemd
|
|
|
|
|
|
# default. Use rsync (not mv) so nothing is lost if the path guess is wrong;
|
|
|
|
|
|
# delete the source only after the post-start verification below passes.
|
|
|
|
|
|
sudo rsync -a /usr/share/ollama/.ollama/ /opt/homelab/data/ollama/
|
|
|
|
|
|
sudo chown -R "$(id -u):$(id -g)" /opt/homelab/data/ollama # container runs as root by default; adjust if the image user differs
|
|
|
|
|
|
```
|
|
|
|
|
|
An empty bind mount here is a silent failure mode: the container starts
|
|
|
|
|
|
fine, `ollama list` comes back empty, and llm-gateway's `/api/generate`
|
|
|
|
|
|
calls start 404-ing on model names that used to work — no crash, no loud
|
|
|
|
|
|
error, just wrong answers for whoever's calling it next.
|
|
|
|
|
|
4. Disable (not just stop) the native service so it can't come back and fight
|
|
|
|
|
|
the container for port 11434:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
sudo systemctl disable --now ollama
|
|
|
|
|
|
```
|
|
|
|
|
|
5. Bring up the declarative stack:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
cd ~/homelab-codex-ws
|
|
|
|
|
|
./scripts/deploy/deploy-node.sh
|
|
|
|
|
|
# or directly: docker compose -f services/ollama/docker-compose.yml --env-file services/ollama/.env up -d
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
## Post-cutover verification
|
|
|
|
|
|
|
|
|
|
|
|
1. **Model completeness — do this before anything else calls the new instance:**
|
|
|
|
|
|
```bash
|
|
|
|
|
|
docker exec ollama ollama list
|
|
|
|
|
|
```
|
|
|
|
|
|
Confirm it lists ALL of: `qwen2.5-coder:14b`, `qwen3-coder:30b`,
|
|
|
|
|
|
`deepseek-coder:latest`, `deepcoder:14b`, `bge-m3` — not just a subset. If
|
|
|
|
|
|
any are missing, the bind mount points at the wrong source directory; stop
|
|
|
|
|
|
and fix before letting llm-gateway traffic resume (it depends on these
|
|
|
|
|
|
coder models being present).
|
|
|
|
|
|
2. Container is on GPU, not CPU fallback:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
nvidia-smi # ollama process should appear once a model is loaded
|
|
|
|
|
|
docker exec ollama ollama ps
|
|
|
|
|
|
```
|
|
|
|
|
|
3. Embeddings endpoint + vector dimension (deferred check from
|
|
|
|
|
|
`docs/kb/modules/05-faza2-plan.md` §6 step 2):
|
|
|
|
|
|
```bash
|
|
|
|
|
|
curl -s http://localhost:11434/api/embeddings -d '{"model":"bge-m3","prompt":"test"}' \
|
|
|
|
|
|
| python3 -c "import json,sys; v=json.load(sys.stdin)['embedding']; print(len(v))"
|
|
|
|
|
|
# expect: 1024
|
|
|
|
|
|
```
|
|
|
|
|
|
4. Reachability from PIHA (llm-gateway's actual path):
|
|
|
|
|
|
```bash
|
|
|
|
|
|
curl -sf http://solaria:11434/api/tags # from PIHA, over Tailscale
|
|
|
|
|
|
```
|
|
|
|
|
|
5. `services/ollama/healthcheck.sh` passes on SOLARIA.
|
|
|
|
|
|
|
|
|
|
|
|
## Rollback
|
|
|
|
|
|
|
|
|
|
|
|
If the container fails to come up healthy or the model list is incomplete:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
docker compose -f services/ollama/docker-compose.yml down
|
|
|
|
|
|
sudo systemctl enable --now ollama
|
|
|
|
|
|
```
|
|
|
|
|
|
The native install's model directory was copied (not moved) in step 3, so it
|
|
|
|
|
|
is untouched and the manual instance comes back exactly as it was.
|
fix(ollama): restore GPU reservation on SOLARIA, close 07-15 cutover docs
Driver fixed 2026-07-16 (nvidia-driver-595-open from distro repo, old
graphics-drivers PPA for jammy disabled) — RTX 4070 Ti SUPER 16GB, CUDA
13.2, nvidia-container-toolkit already present from the 07-15 cutover
prerequisite install. Uncomments deploy.resources.reservations (nvidia
gpu) in services/ollama/docker-compose.yml, restoring clean formatting.
Docs close out the loose ends from the 07-15 declarative cutover:
- ollama-solaria-cutover runbook gets a "Wykonanie" section documenting
what actually happened (mv instead of rsync for the model store, the
missing nvidia-container-toolkit prerequisite, the driver-missing
discovery, CPU-only cutover, then the 07-16 GPU fix) plus a note on
the container disappearing after the 07-15 evening reboot (one-off,
boots fine now, root cause not established).
- hosts/solaria/services.yaml: ollama role comment now reflects actual
GPU-backed state instead of the previously-aspirational wording.
- hosts/solaria/README.md: drop stale Open WebUI mention (not in repo).
- docs/backlog.md: close the NVIDIA driver item; leave two follow-ups
(Ollama call batching before the mail phase, UNIQUE(envelope_id,
chunk_index) schema change for multi-model embeddings at phase 3).
- jobs/documents-ingest/README.md: timing section gets a GPU placeholder
line next to the existing 0.79s/chunk CPU baseline, to be filled in
after the live GPU benchmark.
Live recreate + GPU-vs-CPU embedding benchmark deliberately left for the
operator to run from the main checkout after merge, per worktree-aware
discipline — this worktree only owns the declarative fix and the docs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-16 14:20:16 +02:00
|
|
|
|
|
|
|
|
|
|
*(See "Wykonanie" below — the actual run used `mv`, not `rsync`, so this
|
|
|
|
|
|
rollback description no longer matches what happened live.)*
|
|
|
|
|
|
|
|
|
|
|
|
## Wykonanie (2026-07-15, GPU restored 2026-07-16)
|
|
|
|
|
|
|
|
|
|
|
|
Deviations from the plan above, discovered while executing it live on
|
|
|
|
|
|
SOLARIA:
|
|
|
|
|
|
|
|
|
|
|
|
- **Step 3 (relocate model store) used `mv`, not `rsync -a`.** The model
|
|
|
|
|
|
store (36G) was moved in one shot rather than copied. Consequence: the
|
|
|
|
|
|
rollback described above ("source untouched since copied") does not apply
|
|
|
|
|
|
as executed — a real rollback would need the reverse `mv`
|
|
|
|
|
|
(`/opt/homelab/data/ollama/* → /usr/share/ollama/.ollama/`), not just
|
|
|
|
|
|
re-enabling the native systemd service against an already-relocated
|
|
|
|
|
|
directory.
|
|
|
|
|
|
- **Step 2 pre-check (`nvidia-container-toolkit`) failed the first time**: it
|
|
|
|
|
|
was not installed. Installed as a prerequisite before continuing — this
|
|
|
|
|
|
wasn't previously flagged as something that could be missing.
|
|
|
|
|
|
- **Step 2 pre-check (driver) failed harder than expected**: `nvidia-smi` did
|
|
|
|
|
|
not exist on the host at all — no NVIDIA driver installed on SOLARIA.
|
|
|
|
|
|
`hosts/solaria/services.yaml` had described ollama as GPU-backed for a
|
|
|
|
|
|
while; that was aspirational, not actual — Ollama had been running CPU-only
|
|
|
|
|
|
the entire time regardless of what the manifest said.
|
|
|
|
|
|
- Given the missing driver, the cutover proceeded **in CPU-only mode**: the
|
|
|
|
|
|
`deploy.resources` GPU reservation was commented out in
|
|
|
|
|
|
`services/ollama/docker-compose.yml` (commit `f57a01a`), and the driver fix
|
|
|
|
|
|
was filed as a backlog item (see `docs/backlog.md`) blocking the module 5
|
|
|
|
|
|
mail-embedding phase.
|
|
|
|
|
|
- **2026-07-16: driver fixed.** Installed `nvidia-driver-595-open` from the
|
|
|
|
|
|
distro repository — not the old `ppa:graphics-drivers/ppa` (jammy), which
|
|
|
|
|
|
was deactivated by renaming its `sources.list.d` entry to `.disabled`.
|
|
|
|
|
|
Result: RTX 4070 Ti SUPER 16GB, CUDA 13.2, `nvidia-smi` working on the
|
|
|
|
|
|
host. `nvidia-container-toolkit` from the earlier prerequisite install was
|
|
|
|
|
|
already in place (`nvidia` runtime registered in `daemon.json`), so no
|
|
|
|
|
|
further toolkit work was needed. The GPU reservation was restored in
|
|
|
|
|
|
`docker-compose.yml` and the container recreated.
|
2026-07-16 15:02:07 +02:00
|
|
|
|
- **GPU vs CPU embedding throughput** (measured 2026-07-16): GPU 207ms/embed
|
|
|
|
|
|
vs CPU ≈790ms — ~3.8× faster sequential (50 calls, ~600-tok prompt, bge-m3
|
|
|
|
|
|
at 100% GPU per `ollama ps`). Single-request overhead (HTTP/tokenization)
|
|
|
|
|
|
dominates; batching remains the real lever (backlog).
|
fix(ollama): restore GPU reservation on SOLARIA, close 07-15 cutover docs
Driver fixed 2026-07-16 (nvidia-driver-595-open from distro repo, old
graphics-drivers PPA for jammy disabled) — RTX 4070 Ti SUPER 16GB, CUDA
13.2, nvidia-container-toolkit already present from the 07-15 cutover
prerequisite install. Uncomments deploy.resources.reservations (nvidia
gpu) in services/ollama/docker-compose.yml, restoring clean formatting.
Docs close out the loose ends from the 07-15 declarative cutover:
- ollama-solaria-cutover runbook gets a "Wykonanie" section documenting
what actually happened (mv instead of rsync for the model store, the
missing nvidia-container-toolkit prerequisite, the driver-missing
discovery, CPU-only cutover, then the 07-16 GPU fix) plus a note on
the container disappearing after the 07-15 evening reboot (one-off,
boots fine now, root cause not established).
- hosts/solaria/services.yaml: ollama role comment now reflects actual
GPU-backed state instead of the previously-aspirational wording.
- hosts/solaria/README.md: drop stale Open WebUI mention (not in repo).
- docs/backlog.md: close the NVIDIA driver item; leave two follow-ups
(Ollama call batching before the mail phase, UNIQUE(envelope_id,
chunk_index) schema change for multi-model embeddings at phase 3).
- jobs/documents-ingest/README.md: timing section gets a GPU placeholder
line next to the existing 0.79s/chunk CPU baseline, to be filled in
after the live GPU benchmark.
Live recreate + GPU-vs-CPU embedding benchmark deliberately left for the
operator to run from the main checkout after merge, per worktree-aware
discipline — this worktree only owns the declarative fix and the docs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-16 14:20:16 +02:00
|
|
|
|
|
|
|
|
|
|
## Nota operacyjna: kontener zniknął po reboocie (2026-07-15 wieczór)
|
|
|
|
|
|
|
|
|
|
|
|
Po reboocie hosta wieczorem 2026-07-15 kontener `ollama` **zniknął**
|
|
|
|
|
|
(`docker ps -a` nie pokazywał go w ogóle — nie `Exited`), mimo
|
|
|
|
|
|
`restart: unless-stopped`. Jednorazowe zdarzenie: po boocie 2026-07-16
|
|
|
|
|
|
kontener wstał poprawnie i automatycznie. Przyczyna nie została odtworzona —
|
|
|
|
|
|
flagowane jako obserwacja, nie bug do śledzenia, chyba że się powtórzy.
|