homelab-codex-ws/services/kb-query/app/main.py
oskar e7625cd322 feat(kb): aktywny fallback embeddingów SOLARIA→PIHA dla kb-query (faza 4 Krok 2)
Ostatni krok fazy 4 KB (plan §2 Decyzja 2, §5): kb-query przestaje być martwe
przez ~16 h/dobę, gdy SOLARIA (GPU) śpi — zapytania embeduje wtedy lokalna
Ollama CPU na PIHA (wolniej: ~790 ms+ vs ~207 ms na GPU, ale działa).

Nowy serwis services/ollama-piha (GitOps, owner_node: piha):
- ollama/ollama:latest (arm64 natywnie), OLLAMA_KEEP_ALIVE=0 — model zwalnia
  RAM natychmiast po każdym wywołaniu (spike, nie rezydent; PIHA dzieli 8 GB z HA)
- bind wyłącznie 127.0.0.1 + LAN_BIND_IP (192.168.31.5), nigdy 0.0.0.0/Tailscale
- named volume ollama_piha_models (NVMe data-root) zamiast bind-mounta — obraz
  biega jako root w kontenerze i bind łamałby wzorzec uid PIHA (oskar=1004,
  kontenery uid 1000, setgid pi)
- override hosts/piha/runtime/ollama-piha: mem_limit 2560m (wartość startowa
  z planu, do potwierdzenia kalibracją na żywo), świadomie bez mem_reservation
- pull bge-m3 to jawny, ręczny krok deployu (README) — obraz nie ma modeli

kb-query — maszyna stanów fallbacku (app/embed_router.py):
- health-check SOLARII (GET /api/tags, timeout 1.5 s) z cache 30 s — zero
  sondowania per request; po powrocie SOLARII ruch wraca na GPU w ≤30 s
- primary up → embed na SOLARII z twardym timeoutem 3 s; błąd W TRAKCIE
  zapytania = jednorazowe przełączenie (krok 3b planu): status down na 30 s
  i TO SAMO zapytanie leci na fallback — user nie widzi błędu SOLARII
- primary down → embed prosto na ollama-piha (bez twardego timeoutu: CPU +
  zimny load modelu to legalnie pojedyncze sekundy)
- 503 tylko gdy oba backendy padłe (lub fallback nieskonfigurowany)
- inwariant modelu, druga połowa: każdy backend weryfikowany raz, leniwie przy
  pierwszym użyciu, że /api/tags zawiera EMBED_MODEL (bge-m3 — ta sama wartość
  co startowy check przeciw document_chunk.model/document_summary.embedding_model);
  niezgodność = ERROR log + 500, nigdy ciche liczenie dystansów między
  różnymi przestrzeniami embeddingów; leniwie, bo śpiąca SOLARIA nie może
  blokować startu serwisu
- odpowiedź /search: nowe pole embed_backend ("solaria"|"piha") + sol_status
  wg realnego świata routera (UI już renderuje down jako "offline (fallback
  embed)"); log INFO backend=... elapsed_ms=... per zapytanie
- /healthz: sol_status przez cache routera (spójny widok z routingiem) +
  fallback_status (żywa, tania sonda /api/tags)

Konfiguracja spójnie przez env (compose + env.example + service.yaml + README):
EMBED_PRIMARY_URL (zastępuje OLLAMA_URL), EMBED_FALLBACK_URL (pusty = brak
fallbacku, zachowanie sprzed kroku 2), EMBED_{PRIMARY,FALLBACK}_NAME,
EMBED_HEALTH_TTL_S/EMBED_HEALTH_TIMEOUT_S/EMBED_PRIMARY_TIMEOUT_S.

Testy: 39 pass (14 nowych w test_embed_router.py: cache TTL, failover w trakcie
zapytania, powrót po TTL, oba padłe, mismatch modelu na primary i fallbacku,
tag "bge-m3:latest" vs "bge-m3"); docker build + smoke (importy + uvicorn do
guardu KB_DSN) OK; compose config OK dla obu stacków.

Deploy (Oskar, na PIHA z mastera po merge):
  cd ~/homelab-codex-ws && git pull
  # 1. ollama-piha
  cp services/ollama-piha/env.example services/ollama-piha/.env
  docker compose -f services/ollama-piha/docker-compose.yml \
    -f hosts/piha/runtime/ollama-piha/docker-compose.override.yml \
    --env-file services/ollama-piha/.env up -d
  docker exec ollama-piha ollama pull bge-m3     # ręczny krok, obowiązkowy
  services/ollama-piha/healthcheck.sh
  # 2. kb-query (dopisać fallback do istniejącego .env)
  echo 'EMBED_FALLBACK_URL=http://192.168.31.5:11434' >> services/kb-query/.env
  docker compose -f services/kb-query/docker-compose.yml \
    -f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build
  services/kb-query/healthcheck.sh
  # (deploy-node.sh też podniesie oba serwisy z hosts/piha/services.yaml,
  #  ale pull bge-m3 i .env pozostają ręczne)
Weryfikacja: testy A/B/C w services/kb-query/README.md (backend=solaria przy
SOLARII online; backend=piha przy symulacji offline; powrót na GPU w ≤30 s).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 19:01:29 +02:00

131 lines
5.9 KiB
Python

"""kb-query -- module 5 phase 4 (docs/kb/modules/05-faza4-plan.md §4): first user-facing HTTP
entry point to the KB. Wraps `kb_retrieval` retrieval (module 5 phase 3, gated PASS --
docs/sessions/2026-07-21.md) in FastAPI. This is a search API, not chat: no answer synthesis,
no LLM call over the results (that is phase 5, out of scope here).
Embed path (Krok 2, plan §2 decision 2 / §5 -- the active fallback): queries embed through
`app.embed_router.EmbedRouter` -- SOLARIA's GPU Ollama (EMBED_PRIMARY_URL) when its cached
~30 s health-check says up, the local CPU `ollama-piha` (EMBED_FALLBACK_URL) when SOLARIA
sleeps. A mid-query primary failure fails over the SAME request instead of surfacing an
error. Only when BOTH backends are unavailable does `/search` return 503; a backend serving
the wrong model (vs the bge-m3 space `document_chunk` is indexed in) is a loud 500, never a
silent cross-space distance computation.
`GET /` (Krok 4, plan §7) serves the search UI from this same FastAPI process -- one image, one
container (plan §2 decision 4): a Jinja2 shell + a static vanilla-JS file, no node build step.
`/` and `/static/*` need no DB/Ollama, so they stay reachable even while `/search` is 503ing.
`mode=hybrid` (faza mailowa, plan Krok 3, docs/kb/modules/05-faza-mailowa-plan.md §6) is
available explicitly starting here, but the default stays `cascade` until the quality gate
(plan §8) PASSes on the full mail corpus -- flipping the default is a separate, later change.
"""
from __future__ import annotations
import logging
import os
import pathlib
from contextlib import asynccontextmanager
import aiohttp
from fastapi import FastAPI, HTTPException, Query, Request
from fastapi.responses import HTMLResponse
from fastapi.staticfiles import StaticFiles
from fastapi.templating import Jinja2Templates
from app.db import create_pool
from app.embed_router import EmbedBackendError, EmbedRouter, ModelMismatchError
from app.search import run_search
from app.startup import validate_embed_model
# uvicorn only configures its own loggers; without this the router's
# backend=solaria/piha lines (kb-query.embed) would never reach docker logs.
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(name)s %(message)s")
BASE_DIR = pathlib.Path(__file__).resolve().parent
KB_DSN = os.environ.get("KB_DSN")
# EMBED_PRIMARY_URL replaced OLLAMA_URL in Krok 2 (aktywny fallback) -- one naming scheme
# for both legs. Same default upstream as before: Ollama @ SOLARIA over Tailscale.
EMBED_PRIMARY_URL = os.environ.get("EMBED_PRIMARY_URL", "http://solaria:11434")
# Unset/empty -> no fallback leg: /search 503s when SOLARIA is down (pre-Krok-2 behaviour).
EMBED_FALLBACK_URL = os.environ.get("EMBED_FALLBACK_URL") or None
EMBED_PRIMARY_NAME = os.environ.get("EMBED_PRIMARY_NAME", "solaria")
EMBED_FALLBACK_NAME = os.environ.get("EMBED_FALLBACK_NAME", "piha")
EMBED_MODEL = os.environ.get("EMBED_MODEL", "bge-m3")
SUMMARY_MODEL = os.environ.get("SUMMARY_MODEL", "claude-haiku-4-5")
# State-machine parameters (plan §2 D2 table; timeouts rationale in app/embed_router.py).
EMBED_HEALTH_TTL_S = float(os.environ.get("EMBED_HEALTH_TTL_S", "30"))
EMBED_HEALTH_TIMEOUT_S = float(os.environ.get("EMBED_HEALTH_TIMEOUT_S", "1.5"))
EMBED_PRIMARY_TIMEOUT_S = float(os.environ.get("EMBED_PRIMARY_TIMEOUT_S", "3"))
@asynccontextmanager
async def lifespan(app: FastAPI):
if not KB_DSN:
raise RuntimeError("KB_DSN is required (see env.example)")
pool = await create_pool(KB_DSN)
async with pool.acquire() as conn:
# Hard invariant (plan §2 decision 2): refuse to start rather than silently serve
# queries against a mismatched embedding space. The per-backend half of the same
# invariant (does the Ollama actually serve EMBED_MODEL?) lives in EmbedRouter and
# runs lazily at each backend's first use -- a sleeping SOLARIA must not block boot.
await validate_embed_model(conn, EMBED_MODEL)
app.state.pool = pool
app.state.http = aiohttp.ClientSession()
app.state.embed_router = EmbedRouter(
EMBED_PRIMARY_URL,
EMBED_FALLBACK_URL,
embed_model=EMBED_MODEL,
primary_name=EMBED_PRIMARY_NAME,
fallback_name=EMBED_FALLBACK_NAME,
health_ttl_s=EMBED_HEALTH_TTL_S,
health_timeout_s=EMBED_HEALTH_TIMEOUT_S,
primary_embed_timeout_s=EMBED_PRIMARY_TIMEOUT_S,
)
try:
yield
finally:
await app.state.http.close()
await pool.close()
app = FastAPI(lifespan=lifespan)
app.mount("/static", StaticFiles(directory=BASE_DIR / "static"), name="static")
templates = Jinja2Templates(directory=BASE_DIR / "templates")
@app.get("/", response_class=HTMLResponse)
async def index(request: Request):
return templates.TemplateResponse(request, "index.html")
@app.get("/healthz")
async def healthz() -> dict:
# sol_status goes through the router's 30 s cache -- /healthz and /search routing
# deliberately share one world view (and monitoring no longer re-probes a sleeping
# SOLARIA on every scrape). fallback_status is a live cheap /api/tags probe.
router = app.state.embed_router
return {
"status": "ok",
"sol_status": await router.primary_status(app.state.http),
"fallback_status": await router.fallback_status(app.state.http),
}
@app.get("/search")
async def search(
q: str = Query(..., min_length=1),
mode: str = Query("cascade", pattern="^(cascade|flat|hybrid)$"),
) -> dict:
try:
async with app.state.pool.acquire() as conn:
return await run_search(
conn, app.state.http, app.state.embed_router, q, mode, SUMMARY_MODEL
)
except ModelMismatchError as exc:
# Config/ops error (backend up but wrong model set) -- 500, deliberately loud.
raise HTTPException(status_code=500, detail=str(exc)) from exc
except (EmbedBackendError, aiohttp.ClientError) as exc:
raise HTTPException(status_code=503, detail=f"embed backend unavailable: {exc}") from exc