homelab-codex-ws/services/kb-query/README.md
oskar e7625cd322 feat(kb): aktywny fallback embeddingów SOLARIA→PIHA dla kb-query (faza 4 Krok 2)
Ostatni krok fazy 4 KB (plan §2 Decyzja 2, §5): kb-query przestaje być martwe
przez ~16 h/dobę, gdy SOLARIA (GPU) śpi — zapytania embeduje wtedy lokalna
Ollama CPU na PIHA (wolniej: ~790 ms+ vs ~207 ms na GPU, ale działa).

Nowy serwis services/ollama-piha (GitOps, owner_node: piha):
- ollama/ollama:latest (arm64 natywnie), OLLAMA_KEEP_ALIVE=0 — model zwalnia
  RAM natychmiast po każdym wywołaniu (spike, nie rezydent; PIHA dzieli 8 GB z HA)
- bind wyłącznie 127.0.0.1 + LAN_BIND_IP (192.168.31.5), nigdy 0.0.0.0/Tailscale
- named volume ollama_piha_models (NVMe data-root) zamiast bind-mounta — obraz
  biega jako root w kontenerze i bind łamałby wzorzec uid PIHA (oskar=1004,
  kontenery uid 1000, setgid pi)
- override hosts/piha/runtime/ollama-piha: mem_limit 2560m (wartość startowa
  z planu, do potwierdzenia kalibracją na żywo), świadomie bez mem_reservation
- pull bge-m3 to jawny, ręczny krok deployu (README) — obraz nie ma modeli

kb-query — maszyna stanów fallbacku (app/embed_router.py):
- health-check SOLARII (GET /api/tags, timeout 1.5 s) z cache 30 s — zero
  sondowania per request; po powrocie SOLARII ruch wraca na GPU w ≤30 s
- primary up → embed na SOLARII z twardym timeoutem 3 s; błąd W TRAKCIE
  zapytania = jednorazowe przełączenie (krok 3b planu): status down na 30 s
  i TO SAMO zapytanie leci na fallback — user nie widzi błędu SOLARII
- primary down → embed prosto na ollama-piha (bez twardego timeoutu: CPU +
  zimny load modelu to legalnie pojedyncze sekundy)
- 503 tylko gdy oba backendy padłe (lub fallback nieskonfigurowany)
- inwariant modelu, druga połowa: każdy backend weryfikowany raz, leniwie przy
  pierwszym użyciu, że /api/tags zawiera EMBED_MODEL (bge-m3 — ta sama wartość
  co startowy check przeciw document_chunk.model/document_summary.embedding_model);
  niezgodność = ERROR log + 500, nigdy ciche liczenie dystansów między
  różnymi przestrzeniami embeddingów; leniwie, bo śpiąca SOLARIA nie może
  blokować startu serwisu
- odpowiedź /search: nowe pole embed_backend ("solaria"|"piha") + sol_status
  wg realnego świata routera (UI już renderuje down jako "offline (fallback
  embed)"); log INFO backend=... elapsed_ms=... per zapytanie
- /healthz: sol_status przez cache routera (spójny widok z routingiem) +
  fallback_status (żywa, tania sonda /api/tags)

Konfiguracja spójnie przez env (compose + env.example + service.yaml + README):
EMBED_PRIMARY_URL (zastępuje OLLAMA_URL), EMBED_FALLBACK_URL (pusty = brak
fallbacku, zachowanie sprzed kroku 2), EMBED_{PRIMARY,FALLBACK}_NAME,
EMBED_HEALTH_TTL_S/EMBED_HEALTH_TIMEOUT_S/EMBED_PRIMARY_TIMEOUT_S.

Testy: 39 pass (14 nowych w test_embed_router.py: cache TTL, failover w trakcie
zapytania, powrót po TTL, oba padłe, mismatch modelu na primary i fallbacku,
tag "bge-m3:latest" vs "bge-m3"); docker build + smoke (importy + uvicorn do
guardu KB_DSN) OK; compose config OK dla obu stacków.

Deploy (Oskar, na PIHA z mastera po merge):
  cd ~/homelab-codex-ws && git pull
  # 1. ollama-piha
  cp services/ollama-piha/env.example services/ollama-piha/.env
  docker compose -f services/ollama-piha/docker-compose.yml \
    -f hosts/piha/runtime/ollama-piha/docker-compose.override.yml \
    --env-file services/ollama-piha/.env up -d
  docker exec ollama-piha ollama pull bge-m3     # ręczny krok, obowiązkowy
  services/ollama-piha/healthcheck.sh
  # 2. kb-query (dopisać fallback do istniejącego .env)
  echo 'EMBED_FALLBACK_URL=http://192.168.31.5:11434' >> services/kb-query/.env
  docker compose -f services/kb-query/docker-compose.yml \
    -f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build
  services/kb-query/healthcheck.sh
  # (deploy-node.sh też podniesie oba serwisy z hosts/piha/services.yaml,
  #  ale pull bge-m3 i .env pozostają ręczne)
Weryfikacja: testy A/B/C w services/kb-query/README.md (backend=solaria przy
SOLARII online; backend=piha przy symulacji offline; powrót na GPU w ≤30 s).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 19:01:29 +02:00

11 KiB
Raw Blame History

kb-query

FastAPI search API in front of the module-5 KB retrieval engine (packages/kb-retrieval/). Runs on PIHA, bound to PIHA's LAN IP only (exposure: private, same class as paperless/nextcloud — no public ingress yet). This is a search API, not chat: no answer synthesis over results, that's phase 5.

Endpoints

Endpoint Method Purpose
/ GET Search UI (Jinja2 shell + /static/app.js, no login yet — plan §8 OIDC is a later step)
/static/* GET UI assets (app.js, style.css)
/healthz GET {"status": "ok", "sol_status": "up"|"down", "fallback_status": "up"|"down"|"unconfigured"}sol_status comes from the embed router's ~30 s-cached SOLARIA probe (same world view /search routes by), fallback_status is a live cheap probe of ollama-piha; no auth required (monitoring must reach it)
/search?q=<text>&mode=cascade|flat GET query_text -> embed -> cascade_query/flat_query -> results, mode defaults to cascade

/search response shape (module 5 phase 4 plan §4, summary/summary_tags added in Krok 4 for the UI's per-envelope result header — additive, does not change any field the plan §4 shape already defined):

{
  "query": "...", "mode": "cascade", "sol_status": "up", "embed_backend": "solaria",
  "results": [
    {"envelope_id": "paperless:119", "source": "paperless", "dist": 0.34,
     "chunk_index": 2, "text": "...", "link": "https://paper.kapala.org/documents/119/details",
     "summary": "...", "summary_tags": ["..."]},
    {"envelope_id": "<Message-ID>", "source": "gmail", "dist": 0.44,
     "chunk_index": 0, "text": "...", "subject": "...", "from": "...", "date": "...",
     "link": null, "mail_ui_url": null, "summary": null, "summary_tags": []}
  ]
}

dist is never filtered server-side — the 0.45/0.55 colour thresholds (below) are a frontend concern, not an API contract. summary/summary_tags come from document_summary for SUMMARY_MODEL; null/[] when the envelope has no summary yet.

Frontend (Krok 4, plan §7)

One page, served from this same FastAPI process — no separate frontend container, no node build step (plan §2 decision 4): app/templates/index.html (Jinja2 shell) + app/static/app.js (vanilla JS, fetch() to /search) + app/static/style.css. Wszystko po polsku.

  • Pole zapytania + submit (Enter lub przycisk), przełącznik trybu kaskada/flat (domyślnie kaskada — checkbox "tryb flat (debug)").
  • Wyniki grupowane po envelope_id (dokument): nagłówek trafienia to streszczenie dokumentu (summary, tor haiku) gdy dostępne, w przeciwnym razie envelope_id; chunki są rozwijanymi fragmentami (<details>) pod nagłówkiem, posortowane po dist.
  • Kolorowanie progów (fazy 3, zweryfikowane bramką): dist < 0.45 zielony, 0.450.55 żółty (nadal renderowany, z wizualnym ostrzeżeniem), > 0.55 nigdy nie renderowany jako pojedynczy wynik. Jeśli po tym filtrze żadna grupa nie zostaje nic do pokazania (wszystkie trafienia > 0.55, albo brak trafień w ogóle), całość zastępuje komunikat "Brak odpowiedzi w KB dla tego zapytania" z najlepszym (najniższym) zaobserwowanym dist w nawiasie.
  • Źródło: Paperless → link "Otwórz w Paperless" (link); Gmail → metadane (subject/from/date) + przycisk "Kopiuj Message-ID" (envelope_id jest Message-ID, plan §2 decyzja 3) — nie ma dokąd linkować, więc kopiowalny identyfikator zamiast martwego linku.
  • Stopka pokazuje sol_status dyskretnie (odświeżane z /healthz przy starcie strony i po każdym wyszukiwaniu).

Embed path — active SOLARIA→PIHA fallback (Krok 2, plan §2 D2/§5)

Query embeddings go through app/embed_router.py (EmbedRouter, one instance per process):

  1. SOLARIA's health verdict (GET /api/tags, 1.5 s timeout) is cached for 30 s — no per-request probing.
  2. Verdict up → embed on EMBED_PRIMARY_URL (SOLARIA GPU, ~207 ms) under a hard 3 s timeout.
  3. A failure during a real embed flips the verdict to down for one TTL window and the same request is served from EMBED_FALLBACK_URL (ollama-piha, CPU, OLLAMA_KEEP_ALIVE=0 — slower, single seconds, but alive) — the user never sees a SOLARIA error while a fallback exists.
  4. Verdict down → straight to the fallback until the TTL expires; when SOLARIA wakes up, traffic returns to the GPU within ≤30 s.

Every /search response and log line says which backend embedded the query (embed_backend: "solaria"|"piha", log backend=… in kb-query.embed) — needed to debug result quality per backend. sol_status in the response is the router's world view; the UI footer renders down as „SOLARIA: offline (fallback embed)".

Model invariant, both halves: at startup kb-query pins EMBED_MODEL (bge-m3) to document_chunk.model/document_summary.embedding_model (below); additionally each backend is verified once, at its first use, that its /api/tags actually lists EMBED_MODEL. A backend serving the wrong model is a loud ERROR + 500 — never a silent distance computation across two different embedding spaces. Verification is lazy because SOLARIA may be asleep at boot and must not block startup.

/search returns 503 only when BOTH backends are unreachable (or EMBED_FALLBACK_URL is unset — then the pre-fallback behaviour applies); /healthz still answers, same tolerance pattern as llm-gateway.

Startup invariant (hard-fail)

At startup, kb-query queries document_chunk.model and document_summary.embedding_model for the set of models behind active embeddings, and refuses to start (crash-loop, visible via container restarts) if the configured EMBED_MODEL (default bge-m3) isn't in both sets. This guards against querying with an embedding space that doesn't match what's actually indexed — see app/startup.py for why the check reads document_summary.embedding_model and not .model (the latter is the LLM that wrote the summary, e.g. claude-haiku-4-5, not the embedder).

Configuration

.envgitignored, copy from env.example. Required: LAN_BIND_IP, KB_DSN. On PIHA also set EMBED_FALLBACK_URL=http://192.168.31.5:11434 (ollama-piha). Optional (defaults in parentheses): EMBED_PRIMARY_URL (http://solaria:11434), EMBED_PRIMARY_NAME/EMBED_FALLBACK_NAME (solaria/piha), EMBED_HEALTH_TTL_S (30), EMBED_HEALTH_TIMEOUT_S (1.5), EMBED_PRIMARY_TIMEOUT_S (3), EMBED_MODEL (bge-m3), SUMMARY_MODEL (claude-haiku-4-5). OLLAMA_URL was renamed to EMBED_PRIMARY_URL in Krok 2 — if an old .env sets OLLAMA_URL, it is ignored.

Deploy (PIHA)

  1. Prerequisite: ollama-piha deployed and bge-m3 pulled — see services/ollama-piha/README.md (the pull is a manual deploy step).
  2. git pull on PIHA (~/homelab-codex-ws).
  3. cp services/kb-query/env.example services/kb-query/.env and fill in the real KB_DSN password (the template already sets EMBED_FALLBACK_URL). On an existing install: add EMBED_FALLBACK_URL=http://192.168.31.5:11434 to the existing .env.
  4. docker compose -f services/kb-query/docker-compose.yml \
      -f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build
    
  5. Verify: services/kb-query/healthcheck.sh, then from PIHA: curl "http://192.168.31.5:8230/search?q=test" and open http://192.168.31.5:8230/ in a browser.

Fallback verification (execution: operator, after deploy)

  • Test A — SOLARIA online: query via UI/curl; response has "embed_backend": "solaria", docker logs kb-query shows backend=solaria, latency ~sub-second.
  • Test B — SOLARIA offline: either wait for the nightly power-off, or simulate: set EMBED_PRIMARY_URL=http://192.0.2.1:11434 (TEST-NET, always unreachable) in .env and docker compose … up -d again. Query still works; response has "embed_backend": "piha", log shows backend=piha plus a circuit open for 30s warning on the first hit; latency visibly higher (CPU + cold model load each time, OLLAMA_KEEP_ALIVE=0). Revert .env afterwards if simulated.
  • Test C — SOLARIA returns: after it is back up, within ≤30 s (one health-cache TTL) responses show "embed_backend": "solaria" again, no restart needed.

Tests

pip install -e packages/kb-retrieval/
cd services/kb-query && pip install -r requirements.txt pytest pytest-asyncio && pytest

Unit tests mock the DB connection and Ollama HTTP session (no live DB/Ollama required) — same style as packages/kb-retrieval/tests/. tests/test_frontend.py drives GET ///static/* through FastAPI's TestClient without entering it as a context manager, so the DB-requiring lifespan never runs.

Frontend JS has its own pure-function tests (query-URL encoding, threshold colouring, envelope grouping), run without a browser via Node's built-in test runner: node --test services/kb-query/tests/frontend/.

Ingress (kb.kapala.org, plan §8)

Wired up 2026-07-23 (docs/sessions/2026-07-23-kb-f4-ingress.md), no code change in this service — pure infra step:

  • npm@PIHA proxy host #35: kb.kapala.orghttp://192.168.31.5:8230, cert #49 (*.kapala.org wildcard, DNS-01 via Cloudflare, expires 2026-09-28) — same pattern as paper./vikunja./ha.kapala.org.
  • Pi-hole Local DNS (/etc/pihole/custom.list on PIHA, runtime, not in Git): kb.kapala.org192.168.31.5. This is the first kapala.org entry in that file — every other kapala.org vhost (paper/ha/immich/vikunja/forgejo) has no LAN override and resolves via the public Cloudflare record (Tailscale IP) even from LAN, a hairpin the plan assumed was already avoided for those too. Not fixed here (out of this task's scope — no other vhosts touched); worth a follow-up if it matters for those services.
  • Cloudflare A record kb.kapala.org100.108.208.3 (Tailscale PIHA, DNS only): added manually by the operator (no CF API token available in the environment that did the rest of this step), verified against Cloudflare's own authoritative NS and 8.8.8.8/1.1.1.1 — resolves everywhere now, both LAN (via Pi-hole override) and Tailscale/public (via this record).

No auth. OIDC (plan §8: authlib, /login, /auth/callback) is explicitly not implemented — confirmed no forward-auth/reverse-proxy-level auth pattern exists anywhere in this repo (NPM community edition doesn't support it either); the three precedents (paperless/nextcloud/vikunja) all do OIDC inside the app. Building that is real service code (authlib dependency, session middleware, Forgejo OAuth2 app registration) — deliberately deferred to a separate session, decision confirmed with the operator 2026-07-23. Until then kb.kapala.org is reachable by anyone on the LAN/tailnet with no login, same as before this vhost existed.

Out of scope for this step

  • OIDC login (see above) — separate session, needs authlib + Forgejo OAuth2 app.
  • Calibration verdict for the fallback (plan §5 steps 45: live PIHA measurements → keep/tune/degrade decision) — the mechanism is built and default-on when EMBED_FALLBACK_URL is set; the measurements are the operator's post-deploy step (see services/ollama-piha/README.md).