homelab-codex-ws/services/ollama-piha/docker-compose.yml
oskar 3d4ee3818d feat(kb-query): active embed fallback SOLARIA→PIHA (module 5 phase 4, plan §2/§5)
Last missing core piece of KB phase 4: kb-query no longer hard-fails /search
when Ollama@SOLARIA is unreachable. app/fallback.py implements the plan's
circuit-breaker exactly (30s cached health probe, 3s hard embed timeout on
SOLARIA, one-shot same-request switch to a new local ollama-piha@PIHA
container on timeout/error). sol_status in /healthz and /search now reflects
the real breaker state instead of a hardcoded "up".

New services/ollama-piha (bge-m3, OLLAMA_KEEP_ALIVE=0, arm64/no-GPU) is the
local fallback leg. Live calibration on PIHA (2026-07-27, normal load):
embed latency 4.2-5.2s, RAM peak ~983MiB against a 2.5GiB ceiling -- both
inside the plan's go-bar, so the fallback is enabled by default rather than
gated behind a flag. Calibration also surfaced and disabled (not removed) a
previously-undocumented orphaned native ollama.service on PIHA that had been
conflicting with the container's port.

The embed-model invariant (query embedding == document_chunk.model) still
enforces once at startup, since both fallback legs share one EMBED_MODEL
constant by construction; a redundant per-request DB check was deliberately
skipped and the invariant is instead proven structurally by test.

retrieval_eval.py gains --transport http (plan §2 decision 6/§9), previously
unimplemented. Verified live: HTTP transport is bit-identical to direct
transport against the same live SOLARIA (0 mismatches), and a live sol-down
simulation (kb-query's own OLLAMA_URL pointed at a dead address, no other
Ollama consumer touched) shows the PIHA fallback answering with the same
hit@3 gate outcome and dist within ~3e-4 of the SOLARIA baseline.

Zero changes to DB schema or kb_retrieval's retrieval logic -- only the
embed + health layer, per task constraints.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-27 22:23:53 +02:00

27 lines
1.5 KiB
YAML

services:
ollama-piha:
image: ollama/ollama:latest
container_name: ollama-piha
restart: unless-stopped
ports:
# Loopback: healthcheck.sh curls localhost directly on the node. LAN IP: kb-query@PIHA
# reaches this over the host's LAN interface -- kb-query runs in its own Docker network
# (separate compose project), same reasoning as kb-query's KB_DSN reaching kb-postgres.
# Requires .env (from env.example) next to this file at deploy.
- "127.0.0.1:11434:11434"
- "${LAN_BIND_IP}:11434:11434"
environment:
# Module 5 phase 4 plan §2 decision 2 / §5 requirement: the model is loaded only for the
# duration of a request and released immediately after -- RPi5 has no GPU and limited RAM
# (hosts/piha/capabilities.yaml: arm64, 4 cores, no acceleration), so this is a short burst
# spike (idle Ollama binary ~100 MB) rather than a permanent ~1.5-2 GB resident cost. This
# is a fallback-only path (kb-query only reaches this when SOLARIA is down/times out), not
# the default hot path, so paying a cold-load per request here is the correct trade-off.
- OLLAMA_KEEP_ALIVE=0
volumes:
- /opt/homelab/data/ollama-piha:/root/.ollama
# No GPU reservation -- PIHA is arm64 with no acceleration (hosts/piha/capabilities.yaml),
# unlike services/ollama@SOLARIA. CPU-only inference here is expected to be slower; that is
# exactly what the plan §5 live calibration step measures before this is trusted as a
# default fallback (see README.md "Calibration status").