homelab-codex-ws/services/kb-query/README.md
oskar e7625cd322 feat(kb): aktywny fallback embeddingów SOLARIA→PIHA dla kb-query (faza 4 Krok 2)
Ostatni krok fazy 4 KB (plan §2 Decyzja 2, §5): kb-query przestaje być martwe
przez ~16 h/dobę, gdy SOLARIA (GPU) śpi — zapytania embeduje wtedy lokalna
Ollama CPU na PIHA (wolniej: ~790 ms+ vs ~207 ms na GPU, ale działa).

Nowy serwis services/ollama-piha (GitOps, owner_node: piha):
- ollama/ollama:latest (arm64 natywnie), OLLAMA_KEEP_ALIVE=0 — model zwalnia
  RAM natychmiast po każdym wywołaniu (spike, nie rezydent; PIHA dzieli 8 GB z HA)
- bind wyłącznie 127.0.0.1 + LAN_BIND_IP (192.168.31.5), nigdy 0.0.0.0/Tailscale
- named volume ollama_piha_models (NVMe data-root) zamiast bind-mounta — obraz
  biega jako root w kontenerze i bind łamałby wzorzec uid PIHA (oskar=1004,
  kontenery uid 1000, setgid pi)
- override hosts/piha/runtime/ollama-piha: mem_limit 2560m (wartość startowa
  z planu, do potwierdzenia kalibracją na żywo), świadomie bez mem_reservation
- pull bge-m3 to jawny, ręczny krok deployu (README) — obraz nie ma modeli

kb-query — maszyna stanów fallbacku (app/embed_router.py):
- health-check SOLARII (GET /api/tags, timeout 1.5 s) z cache 30 s — zero
  sondowania per request; po powrocie SOLARII ruch wraca na GPU w ≤30 s
- primary up → embed na SOLARII z twardym timeoutem 3 s; błąd W TRAKCIE
  zapytania = jednorazowe przełączenie (krok 3b planu): status down na 30 s
  i TO SAMO zapytanie leci na fallback — user nie widzi błędu SOLARII
- primary down → embed prosto na ollama-piha (bez twardego timeoutu: CPU +
  zimny load modelu to legalnie pojedyncze sekundy)
- 503 tylko gdy oba backendy padłe (lub fallback nieskonfigurowany)
- inwariant modelu, druga połowa: każdy backend weryfikowany raz, leniwie przy
  pierwszym użyciu, że /api/tags zawiera EMBED_MODEL (bge-m3 — ta sama wartość
  co startowy check przeciw document_chunk.model/document_summary.embedding_model);
  niezgodność = ERROR log + 500, nigdy ciche liczenie dystansów między
  różnymi przestrzeniami embeddingów; leniwie, bo śpiąca SOLARIA nie może
  blokować startu serwisu
- odpowiedź /search: nowe pole embed_backend ("solaria"|"piha") + sol_status
  wg realnego świata routera (UI już renderuje down jako "offline (fallback
  embed)"); log INFO backend=... elapsed_ms=... per zapytanie
- /healthz: sol_status przez cache routera (spójny widok z routingiem) +
  fallback_status (żywa, tania sonda /api/tags)

Konfiguracja spójnie przez env (compose + env.example + service.yaml + README):
EMBED_PRIMARY_URL (zastępuje OLLAMA_URL), EMBED_FALLBACK_URL (pusty = brak
fallbacku, zachowanie sprzed kroku 2), EMBED_{PRIMARY,FALLBACK}_NAME,
EMBED_HEALTH_TTL_S/EMBED_HEALTH_TIMEOUT_S/EMBED_PRIMARY_TIMEOUT_S.

Testy: 39 pass (14 nowych w test_embed_router.py: cache TTL, failover w trakcie
zapytania, powrót po TTL, oba padłe, mismatch modelu na primary i fallbacku,
tag "bge-m3:latest" vs "bge-m3"); docker build + smoke (importy + uvicorn do
guardu KB_DSN) OK; compose config OK dla obu stacków.

Deploy (Oskar, na PIHA z mastera po merge):
  cd ~/homelab-codex-ws && git pull
  # 1. ollama-piha
  cp services/ollama-piha/env.example services/ollama-piha/.env
  docker compose -f services/ollama-piha/docker-compose.yml \
    -f hosts/piha/runtime/ollama-piha/docker-compose.override.yml \
    --env-file services/ollama-piha/.env up -d
  docker exec ollama-piha ollama pull bge-m3     # ręczny krok, obowiązkowy
  services/ollama-piha/healthcheck.sh
  # 2. kb-query (dopisać fallback do istniejącego .env)
  echo 'EMBED_FALLBACK_URL=http://192.168.31.5:11434' >> services/kb-query/.env
  docker compose -f services/kb-query/docker-compose.yml \
    -f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build
  services/kb-query/healthcheck.sh
  # (deploy-node.sh też podniesie oba serwisy z hosts/piha/services.yaml,
  #  ale pull bge-m3 i .env pozostają ręczne)
Weryfikacja: testy A/B/C w services/kb-query/README.md (backend=solaria przy
SOLARII online; backend=piha przy symulacji offline; powrót na GPU w ≤30 s).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 19:01:29 +02:00

210 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# kb-query
FastAPI search API in front of the module-5 KB retrieval engine
(`packages/kb-retrieval/`). Runs on **PIHA**, bound to PIHA's LAN IP only
(`exposure: private`, same class as paperless/nextcloud — no public ingress
yet). This is a **search API, not chat**: no answer synthesis over results,
that's phase 5.
## Endpoints
| Endpoint | Method | Purpose |
|---|---|---|
| `/` | GET | Search UI (Jinja2 shell + `/static/app.js`, no login yet — plan §8 OIDC is a later step) |
| `/static/*` | GET | UI assets (`app.js`, `style.css`) |
| `/healthz` | GET | `{"status": "ok", "sol_status": "up"\|"down", "fallback_status": "up"\|"down"\|"unconfigured"}``sol_status` comes from the embed router's ~30 s-cached SOLARIA probe (same world view `/search` routes by), `fallback_status` is a live cheap probe of ollama-piha; no auth required (monitoring must reach it) |
| `/search?q=<text>&mode=cascade\|flat` | GET | `query_text -> embed -> cascade_query/flat_query -> results`, `mode` defaults to `cascade` |
`/search` response shape (module 5 phase 4 plan §4, `summary`/`summary_tags`
added in Krok 4 for the UI's per-envelope result header — additive, does not
change any field the plan §4 shape already defined):
```json
{
"query": "...", "mode": "cascade", "sol_status": "up", "embed_backend": "solaria",
"results": [
{"envelope_id": "paperless:119", "source": "paperless", "dist": 0.34,
"chunk_index": 2, "text": "...", "link": "https://paper.kapala.org/documents/119/details",
"summary": "...", "summary_tags": ["..."]},
{"envelope_id": "<Message-ID>", "source": "gmail", "dist": 0.44,
"chunk_index": 0, "text": "...", "subject": "...", "from": "...", "date": "...",
"link": null, "mail_ui_url": null, "summary": null, "summary_tags": []}
]
}
```
`dist` is never filtered server-side — the 0.45/0.55 colour thresholds (below)
are a frontend concern, not an API contract. `summary`/`summary_tags` come
from `document_summary` for `SUMMARY_MODEL`; `null`/`[]` when the envelope
has no summary yet.
## Frontend (Krok 4, plan §7)
One page, served from this same FastAPI process — no separate frontend
container, no node build step (plan §2 decision 4): `app/templates/index.html`
(Jinja2 shell) + `app/static/app.js` (vanilla JS, `fetch()` to `/search`) +
`app/static/style.css`. Wszystko po polsku.
- Pole zapytania + submit (Enter lub przycisk), przełącznik trybu
kaskada/flat (domyślnie kaskada — checkbox "tryb flat (debug)").
- Wyniki grupowane po `envelope_id` (dokument): nagłówek trafienia to
streszczenie dokumentu (`summary`, tor haiku) gdy dostępne, w przeciwnym
razie `envelope_id`; chunki są rozwijanymi fragmentami (`<details>`) pod
nagłówkiem, posortowane po `dist`.
- Kolorowanie progów (fazy 3, zweryfikowane bramką): `dist < 0.45` zielony,
`0.450.55` żółty (nadal renderowany, z wizualnym ostrzeżeniem), `> 0.55`
nigdy nie renderowany jako pojedynczy wynik. Jeśli po tym filtrze żadna
grupa nie zostaje nic do pokazania (wszystkie trafienia > 0.55, albo brak
trafień w ogóle), całość zastępuje komunikat "Brak odpowiedzi w KB dla
tego zapytania" z najlepszym (najniższym) zaobserwowanym `dist` w nawiasie.
- Źródło: Paperless → link "Otwórz w Paperless" (`link`); Gmail → metadane
(`subject`/`from`/`date`) + przycisk "Kopiuj Message-ID" (`envelope_id`
**jest** Message-ID, plan §2 decyzja 3) — nie ma dokąd linkować, więc
kopiowalny identyfikator zamiast martwego linku.
- Stopka pokazuje `sol_status` dyskretnie (odświeżane z `/healthz` przy
starcie strony i po każdym wyszukiwaniu).
## Embed path — active SOLARIA→PIHA fallback (Krok 2, plan §2 D2/§5)
Query embeddings go through `app/embed_router.py` (`EmbedRouter`, one instance
per process):
1. SOLARIA's health verdict (`GET /api/tags`, 1.5 s timeout) is **cached for
30 s** — no per-request probing.
2. Verdict `up` → embed on `EMBED_PRIMARY_URL` (SOLARIA GPU, ~207 ms) under a
hard 3 s timeout.
3. A failure **during a real embed** flips the verdict to `down` for one TTL
window and the **same request** is served from `EMBED_FALLBACK_URL`
(`ollama-piha`, CPU, `OLLAMA_KEEP_ALIVE=0` — slower, single seconds, but
alive) — the user never sees a SOLARIA error while a fallback exists.
4. Verdict `down` → straight to the fallback until the TTL expires; when
SOLARIA wakes up, traffic returns to the GPU within ≤30 s.
Every `/search` response and log line says which backend embedded the query
(`embed_backend: "solaria"|"piha"`, log `backend=…` in `kb-query.embed`) —
needed to debug result quality per backend. `sol_status` in the response is
the router's world view; the UI footer renders `down` as
„SOLARIA: offline (fallback embed)".
**Model invariant, both halves:** at startup kb-query pins `EMBED_MODEL`
(bge-m3) to `document_chunk.model`/`document_summary.embedding_model` (below);
additionally each backend is verified once, at its first use, that its
`/api/tags` actually lists `EMBED_MODEL`. A backend serving the wrong model is
a **loud ERROR + 500** — never a silent distance computation across two
different embedding spaces. Verification is lazy because SOLARIA may be asleep
at boot and must not block startup.
`/search` returns **503** only when BOTH backends are unreachable (or
`EMBED_FALLBACK_URL` is unset — then the pre-fallback behaviour applies);
`/healthz` still answers, same tolerance pattern as `llm-gateway`.
## Startup invariant (hard-fail)
At startup, kb-query queries `document_chunk.model` and
`document_summary.embedding_model` for the set of models behind *active*
embeddings, and refuses to start (crash-loop, visible via container restarts)
if the configured `EMBED_MODEL` (default `bge-m3`) isn't in both sets. This
guards against querying with an embedding space that doesn't match what's
actually indexed — see `app/startup.py` for why the check reads
`document_summary.embedding_model` and not `.model` (the latter is the LLM
that *wrote* the summary, e.g. `claude-haiku-4-5`, not the embedder).
## Configuration
`.env`**gitignored**, copy from `env.example`. Required: `LAN_BIND_IP`,
`KB_DSN`. On PIHA also set `EMBED_FALLBACK_URL=http://192.168.31.5:11434`
(ollama-piha). Optional (defaults in parentheses): `EMBED_PRIMARY_URL`
(`http://solaria:11434`), `EMBED_PRIMARY_NAME`/`EMBED_FALLBACK_NAME`
(`solaria`/`piha`), `EMBED_HEALTH_TTL_S` (30), `EMBED_HEALTH_TIMEOUT_S` (1.5),
`EMBED_PRIMARY_TIMEOUT_S` (3), `EMBED_MODEL` (`bge-m3`), `SUMMARY_MODEL`
(`claude-haiku-4-5`). `OLLAMA_URL` was **renamed** to `EMBED_PRIMARY_URL` in
Krok 2 — if an old `.env` sets `OLLAMA_URL`, it is ignored.
## Deploy (PIHA)
0. Prerequisite: `ollama-piha` deployed and `bge-m3` pulled — see
`services/ollama-piha/README.md` (the pull is a **manual** deploy step).
1. `git pull` on PIHA (`~/homelab-codex-ws`).
2. `cp services/kb-query/env.example services/kb-query/.env` and fill in the
real `KB_DSN` password (the template already sets `EMBED_FALLBACK_URL`).
On an existing install: add `EMBED_FALLBACK_URL=http://192.168.31.5:11434`
to the existing `.env`.
3. ```
docker compose -f services/kb-query/docker-compose.yml \
-f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build
```
4. Verify: `services/kb-query/healthcheck.sh`, then from PIHA:
`curl "http://192.168.31.5:8230/search?q=test"` and open
`http://192.168.31.5:8230/` in a browser.
## Fallback verification (execution: operator, after deploy)
- **Test A — SOLARIA online**: query via UI/`curl`; response has
`"embed_backend": "solaria"`, `docker logs kb-query` shows
`backend=solaria`, latency ~sub-second.
- **Test B — SOLARIA offline**: either wait for the nightly power-off, or
simulate: set `EMBED_PRIMARY_URL=http://192.0.2.1:11434` (TEST-NET, always
unreachable) in `.env` and `docker compose … up -d` again. Query still
works; response has `"embed_backend": "piha"`, log shows `backend=piha`
plus a `circuit open for 30s` warning on the first hit; latency visibly
higher (CPU + cold model load each time, `OLLAMA_KEEP_ALIVE=0`). Revert
`.env` afterwards if simulated.
- **Test C — SOLARIA returns**: after it is back up, within ≤30 s (one
health-cache TTL) responses show `"embed_backend": "solaria"` again, no
restart needed.
## Tests
```
pip install -e packages/kb-retrieval/
cd services/kb-query && pip install -r requirements.txt pytest pytest-asyncio && pytest
```
Unit tests mock the DB connection and Ollama HTTP session (no live DB/Ollama
required) — same style as `packages/kb-retrieval/tests/`. `tests/test_frontend.py`
drives `GET /`/`/static/*` through FastAPI's `TestClient` without entering it
as a context manager, so the DB-requiring `lifespan` never runs.
Frontend JS has its own pure-function tests (query-URL encoding, threshold
colouring, envelope grouping), run without a browser via Node's built-in
test runner: `node --test services/kb-query/tests/frontend/`.
## Ingress (`kb.kapala.org`, plan §8)
Wired up 2026-07-23 (`docs/sessions/2026-07-23-kb-f4-ingress.md`), **no code
change in this service** — pure infra step:
- npm@PIHA proxy host #35: `kb.kapala.org``http://192.168.31.5:8230`,
cert #49 (`*.kapala.org` wildcard, DNS-01 via Cloudflare, expires
2026-09-28) — same pattern as `paper.`/`vikunja.`/`ha.kapala.org`.
- Pi-hole Local DNS (`/etc/pihole/custom.list` on PIHA, runtime, not in Git):
`kb.kapala.org``192.168.31.5`. **This is the first `kapala.org` entry in
that file** — every other `kapala.org` vhost (paper/ha/immich/vikunja/forgejo)
has no LAN override and resolves via the public Cloudflare record
(Tailscale IP) even from LAN, a hairpin the plan assumed was already avoided
for those too. Not fixed here (out of this task's scope — no other vhosts
touched); worth a follow-up if it matters for those services.
- Cloudflare A record `kb.kapala.org``100.108.208.3` (Tailscale PIHA, DNS
only): added manually by the operator (no CF API token available in the
environment that did the rest of this step), verified against Cloudflare's
own authoritative NS and 8.8.8.8/1.1.1.1 — resolves everywhere now, both
LAN (via Pi-hole override) and Tailscale/public (via this record).
**No auth.** OIDC (plan §8: `authlib`, `/login`, `/auth/callback`) is explicitly
**not implemented** — confirmed no forward-auth/reverse-proxy-level auth pattern
exists anywhere in this repo (NPM community edition doesn't support it either);
the three precedents (paperless/nextcloud/vikunja) all do OIDC inside the app.
Building that is real service code (`authlib` dependency, session middleware,
Forgejo OAuth2 app registration) — deliberately deferred to a separate session,
decision confirmed with the operator 2026-07-23. Until then `kb.kapala.org` is
reachable by anyone on the LAN/tailnet with no login, same as before this vhost
existed.
## Out of scope for this step
- OIDC login (see above) — separate session, needs `authlib` + Forgejo OAuth2 app.
- Calibration verdict for the fallback (plan §5 steps 45: live PIHA
measurements → keep/tune/degrade decision) — the mechanism is built and
default-on when `EMBED_FALLBACK_URL` is set; the measurements are the
operator's post-deploy step (see `services/ollama-piha/README.md`).