Ostatni krok fazy 4 KB (plan §2 Decyzja 2, §5): kb-query przestaje być martwe
przez ~16 h/dobę, gdy SOLARIA (GPU) śpi — zapytania embeduje wtedy lokalna
Ollama CPU na PIHA (wolniej: ~790 ms+ vs ~207 ms na GPU, ale działa).
Nowy serwis services/ollama-piha (GitOps, owner_node: piha):
- ollama/ollama:latest (arm64 natywnie), OLLAMA_KEEP_ALIVE=0 — model zwalnia
RAM natychmiast po każdym wywołaniu (spike, nie rezydent; PIHA dzieli 8 GB z HA)
- bind wyłącznie 127.0.0.1 + LAN_BIND_IP (192.168.31.5), nigdy 0.0.0.0/Tailscale
- named volume ollama_piha_models (NVMe data-root) zamiast bind-mounta — obraz
biega jako root w kontenerze i bind łamałby wzorzec uid PIHA (oskar=1004,
kontenery uid 1000, setgid pi)
- override hosts/piha/runtime/ollama-piha: mem_limit 2560m (wartość startowa
z planu, do potwierdzenia kalibracją na żywo), świadomie bez mem_reservation
- pull bge-m3 to jawny, ręczny krok deployu (README) — obraz nie ma modeli
kb-query — maszyna stanów fallbacku (app/embed_router.py):
- health-check SOLARII (GET /api/tags, timeout 1.5 s) z cache 30 s — zero
sondowania per request; po powrocie SOLARII ruch wraca na GPU w ≤30 s
- primary up → embed na SOLARII z twardym timeoutem 3 s; błąd W TRAKCIE
zapytania = jednorazowe przełączenie (krok 3b planu): status down na 30 s
i TO SAMO zapytanie leci na fallback — user nie widzi błędu SOLARII
- primary down → embed prosto na ollama-piha (bez twardego timeoutu: CPU +
zimny load modelu to legalnie pojedyncze sekundy)
- 503 tylko gdy oba backendy padłe (lub fallback nieskonfigurowany)
- inwariant modelu, druga połowa: każdy backend weryfikowany raz, leniwie przy
pierwszym użyciu, że /api/tags zawiera EMBED_MODEL (bge-m3 — ta sama wartość
co startowy check przeciw document_chunk.model/document_summary.embedding_model);
niezgodność = ERROR log + 500, nigdy ciche liczenie dystansów między
różnymi przestrzeniami embeddingów; leniwie, bo śpiąca SOLARIA nie może
blokować startu serwisu
- odpowiedź /search: nowe pole embed_backend ("solaria"|"piha") + sol_status
wg realnego świata routera (UI już renderuje down jako "offline (fallback
embed)"); log INFO backend=... elapsed_ms=... per zapytanie
- /healthz: sol_status przez cache routera (spójny widok z routingiem) +
fallback_status (żywa, tania sonda /api/tags)
Konfiguracja spójnie przez env (compose + env.example + service.yaml + README):
EMBED_PRIMARY_URL (zastępuje OLLAMA_URL), EMBED_FALLBACK_URL (pusty = brak
fallbacku, zachowanie sprzed kroku 2), EMBED_{PRIMARY,FALLBACK}_NAME,
EMBED_HEALTH_TTL_S/EMBED_HEALTH_TIMEOUT_S/EMBED_PRIMARY_TIMEOUT_S.
Testy: 39 pass (14 nowych w test_embed_router.py: cache TTL, failover w trakcie
zapytania, powrót po TTL, oba padłe, mismatch modelu na primary i fallbacku,
tag "bge-m3:latest" vs "bge-m3"); docker build + smoke (importy + uvicorn do
guardu KB_DSN) OK; compose config OK dla obu stacków.
Deploy (Oskar, na PIHA z mastera po merge):
cd ~/homelab-codex-ws && git pull
# 1. ollama-piha
cp services/ollama-piha/env.example services/ollama-piha/.env
docker compose -f services/ollama-piha/docker-compose.yml \
-f hosts/piha/runtime/ollama-piha/docker-compose.override.yml \
--env-file services/ollama-piha/.env up -d
docker exec ollama-piha ollama pull bge-m3 # ręczny krok, obowiązkowy
services/ollama-piha/healthcheck.sh
# 2. kb-query (dopisać fallback do istniejącego .env)
echo 'EMBED_FALLBACK_URL=http://192.168.31.5:11434' >> services/kb-query/.env
docker compose -f services/kb-query/docker-compose.yml \
-f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build
services/kb-query/healthcheck.sh
# (deploy-node.sh też podniesie oba serwisy z hosts/piha/services.yaml,
# ale pull bge-m3 i .env pozostają ręczne)
Weryfikacja: testy A/B/C w services/kb-query/README.md (backend=solaria przy
SOLARII online; backend=piha przy symulacji offline; powrót na GPU w ≤30 s).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|---|---|---|
| .. | ||
| app | ||
| tests | ||
| docker-compose.yml | ||
| Dockerfile | ||
| Dockerfile.dockerignore | ||
| env.example | ||
| healthcheck.sh | ||
| pytest.ini | ||
| README.md | ||
| requirements.txt | ||
| service.yaml | ||
kb-query
FastAPI search API in front of the module-5 KB retrieval engine
(packages/kb-retrieval/). Runs on PIHA, bound to PIHA's LAN IP only
(exposure: private, same class as paperless/nextcloud — no public ingress
yet). This is a search API, not chat: no answer synthesis over results,
that's phase 5.
Endpoints
| Endpoint | Method | Purpose |
|---|---|---|
/ |
GET | Search UI (Jinja2 shell + /static/app.js, no login yet — plan §8 OIDC is a later step) |
/static/* |
GET | UI assets (app.js, style.css) |
/healthz |
GET | {"status": "ok", "sol_status": "up"|"down", "fallback_status": "up"|"down"|"unconfigured"} — sol_status comes from the embed router's ~30 s-cached SOLARIA probe (same world view /search routes by), fallback_status is a live cheap probe of ollama-piha; no auth required (monitoring must reach it) |
/search?q=<text>&mode=cascade|flat |
GET | query_text -> embed -> cascade_query/flat_query -> results, mode defaults to cascade |
/search response shape (module 5 phase 4 plan §4, summary/summary_tags
added in Krok 4 for the UI's per-envelope result header — additive, does not
change any field the plan §4 shape already defined):
{
"query": "...", "mode": "cascade", "sol_status": "up", "embed_backend": "solaria",
"results": [
{"envelope_id": "paperless:119", "source": "paperless", "dist": 0.34,
"chunk_index": 2, "text": "...", "link": "https://paper.kapala.org/documents/119/details",
"summary": "...", "summary_tags": ["..."]},
{"envelope_id": "<Message-ID>", "source": "gmail", "dist": 0.44,
"chunk_index": 0, "text": "...", "subject": "...", "from": "...", "date": "...",
"link": null, "mail_ui_url": null, "summary": null, "summary_tags": []}
]
}
dist is never filtered server-side — the 0.45/0.55 colour thresholds (below)
are a frontend concern, not an API contract. summary/summary_tags come
from document_summary for SUMMARY_MODEL; null/[] when the envelope
has no summary yet.
Frontend (Krok 4, plan §7)
One page, served from this same FastAPI process — no separate frontend
container, no node build step (plan §2 decision 4): app/templates/index.html
(Jinja2 shell) + app/static/app.js (vanilla JS, fetch() to /search) +
app/static/style.css. Wszystko po polsku.
- Pole zapytania + submit (Enter lub przycisk), przełącznik trybu kaskada/flat (domyślnie kaskada — checkbox "tryb flat (debug)").
- Wyniki grupowane po
envelope_id(dokument): nagłówek trafienia to streszczenie dokumentu (summary, tor haiku) gdy dostępne, w przeciwnym razieenvelope_id; chunki są rozwijanymi fragmentami (<details>) pod nagłówkiem, posortowane podist. - Kolorowanie progów (fazy 3, zweryfikowane bramką):
dist < 0.45zielony,0.45–0.55żółty (nadal renderowany, z wizualnym ostrzeżeniem),> 0.55nigdy nie renderowany jako pojedynczy wynik. Jeśli po tym filtrze żadna grupa nie zostaje nic do pokazania (wszystkie trafienia > 0.55, albo brak trafień w ogóle), całość zastępuje komunikat "Brak odpowiedzi w KB dla tego zapytania" z najlepszym (najniższym) zaobserwowanymdistw nawiasie. - Źródło: Paperless → link "Otwórz w Paperless" (
link); Gmail → metadane (subject/from/date) + przycisk "Kopiuj Message-ID" (envelope_idjest Message-ID, plan §2 decyzja 3) — nie ma dokąd linkować, więc kopiowalny identyfikator zamiast martwego linku. - Stopka pokazuje
sol_statusdyskretnie (odświeżane z/healthzprzy starcie strony i po każdym wyszukiwaniu).
Embed path — active SOLARIA→PIHA fallback (Krok 2, plan §2 D2/§5)
Query embeddings go through app/embed_router.py (EmbedRouter, one instance
per process):
- SOLARIA's health verdict (
GET /api/tags, 1.5 s timeout) is cached for 30 s — no per-request probing. - Verdict
up→ embed onEMBED_PRIMARY_URL(SOLARIA GPU, ~207 ms) under a hard 3 s timeout. - A failure during a real embed flips the verdict to
downfor one TTL window and the same request is served fromEMBED_FALLBACK_URL(ollama-piha, CPU,OLLAMA_KEEP_ALIVE=0— slower, single seconds, but alive) — the user never sees a SOLARIA error while a fallback exists. - Verdict
down→ straight to the fallback until the TTL expires; when SOLARIA wakes up, traffic returns to the GPU within ≤30 s.
Every /search response and log line says which backend embedded the query
(embed_backend: "solaria"|"piha", log backend=… in kb-query.embed) —
needed to debug result quality per backend. sol_status in the response is
the router's world view; the UI footer renders down as
„SOLARIA: offline (fallback embed)".
Model invariant, both halves: at startup kb-query pins EMBED_MODEL
(bge-m3) to document_chunk.model/document_summary.embedding_model (below);
additionally each backend is verified once, at its first use, that its
/api/tags actually lists EMBED_MODEL. A backend serving the wrong model is
a loud ERROR + 500 — never a silent distance computation across two
different embedding spaces. Verification is lazy because SOLARIA may be asleep
at boot and must not block startup.
/search returns 503 only when BOTH backends are unreachable (or
EMBED_FALLBACK_URL is unset — then the pre-fallback behaviour applies);
/healthz still answers, same tolerance pattern as llm-gateway.
Startup invariant (hard-fail)
At startup, kb-query queries document_chunk.model and
document_summary.embedding_model for the set of models behind active
embeddings, and refuses to start (crash-loop, visible via container restarts)
if the configured EMBED_MODEL (default bge-m3) isn't in both sets. This
guards against querying with an embedding space that doesn't match what's
actually indexed — see app/startup.py for why the check reads
document_summary.embedding_model and not .model (the latter is the LLM
that wrote the summary, e.g. claude-haiku-4-5, not the embedder).
Configuration
.env — gitignored, copy from env.example. Required: LAN_BIND_IP,
KB_DSN. On PIHA also set EMBED_FALLBACK_URL=http://192.168.31.5:11434
(ollama-piha). Optional (defaults in parentheses): EMBED_PRIMARY_URL
(http://solaria:11434), EMBED_PRIMARY_NAME/EMBED_FALLBACK_NAME
(solaria/piha), EMBED_HEALTH_TTL_S (30), EMBED_HEALTH_TIMEOUT_S (1.5),
EMBED_PRIMARY_TIMEOUT_S (3), EMBED_MODEL (bge-m3), SUMMARY_MODEL
(claude-haiku-4-5). OLLAMA_URL was renamed to EMBED_PRIMARY_URL in
Krok 2 — if an old .env sets OLLAMA_URL, it is ignored.
Deploy (PIHA)
- Prerequisite:
ollama-pihadeployed andbge-m3pulled — seeservices/ollama-piha/README.md(the pull is a manual deploy step). git pullon PIHA (~/homelab-codex-ws).cp services/kb-query/env.example services/kb-query/.envand fill in the realKB_DSNpassword (the template already setsEMBED_FALLBACK_URL). On an existing install: addEMBED_FALLBACK_URL=http://192.168.31.5:11434to the existing.env.-
docker compose -f services/kb-query/docker-compose.yml \ -f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build - Verify:
services/kb-query/healthcheck.sh, then from PIHA:curl "http://192.168.31.5:8230/search?q=test"and openhttp://192.168.31.5:8230/in a browser.
Fallback verification (execution: operator, after deploy)
- Test A — SOLARIA online: query via UI/
curl; response has"embed_backend": "solaria",docker logs kb-queryshowsbackend=solaria, latency ~sub-second. - Test B — SOLARIA offline: either wait for the nightly power-off, or
simulate: set
EMBED_PRIMARY_URL=http://192.0.2.1:11434(TEST-NET, always unreachable) in.envanddocker compose … up -dagain. Query still works; response has"embed_backend": "piha", log showsbackend=pihaplus acircuit open for 30swarning on the first hit; latency visibly higher (CPU + cold model load each time,OLLAMA_KEEP_ALIVE=0). Revert.envafterwards if simulated. - Test C — SOLARIA returns: after it is back up, within ≤30 s (one
health-cache TTL) responses show
"embed_backend": "solaria"again, no restart needed.
Tests
pip install -e packages/kb-retrieval/
cd services/kb-query && pip install -r requirements.txt pytest pytest-asyncio && pytest
Unit tests mock the DB connection and Ollama HTTP session (no live DB/Ollama
required) — same style as packages/kb-retrieval/tests/. tests/test_frontend.py
drives GET ///static/* through FastAPI's TestClient without entering it
as a context manager, so the DB-requiring lifespan never runs.
Frontend JS has its own pure-function tests (query-URL encoding, threshold
colouring, envelope grouping), run without a browser via Node's built-in
test runner: node --test services/kb-query/tests/frontend/.
Ingress (kb.kapala.org, plan §8)
Wired up 2026-07-23 (docs/sessions/2026-07-23-kb-f4-ingress.md), no code
change in this service — pure infra step:
- npm@PIHA proxy host #35:
kb.kapala.org→http://192.168.31.5:8230, cert #49 (*.kapala.orgwildcard, DNS-01 via Cloudflare, expires 2026-09-28) — same pattern aspaper./vikunja./ha.kapala.org. - Pi-hole Local DNS (
/etc/pihole/custom.liston PIHA, runtime, not in Git):kb.kapala.org→192.168.31.5. This is the firstkapala.orgentry in that file — every otherkapala.orgvhost (paper/ha/immich/vikunja/forgejo) has no LAN override and resolves via the public Cloudflare record (Tailscale IP) even from LAN, a hairpin the plan assumed was already avoided for those too. Not fixed here (out of this task's scope — no other vhosts touched); worth a follow-up if it matters for those services. - Cloudflare A record
kb.kapala.org→100.108.208.3(Tailscale PIHA, DNS only): added manually by the operator (no CF API token available in the environment that did the rest of this step), verified against Cloudflare's own authoritative NS and 8.8.8.8/1.1.1.1 — resolves everywhere now, both LAN (via Pi-hole override) and Tailscale/public (via this record).
No auth. OIDC (plan §8: authlib, /login, /auth/callback) is explicitly
not implemented — confirmed no forward-auth/reverse-proxy-level auth pattern
exists anywhere in this repo (NPM community edition doesn't support it either);
the three precedents (paperless/nextcloud/vikunja) all do OIDC inside the app.
Building that is real service code (authlib dependency, session middleware,
Forgejo OAuth2 app registration) — deliberately deferred to a separate session,
decision confirmed with the operator 2026-07-23. Until then kb.kapala.org is
reachable by anyone on the LAN/tailnet with no login, same as before this vhost
existed.
Out of scope for this step
- OIDC login (see above) — separate session, needs
authlib+ Forgejo OAuth2 app. - Calibration verdict for the fallback (plan §5 steps 4–5: live PIHA
measurements → keep/tune/degrade decision) — the mechanism is built and
default-on when
EMBED_FALLBACK_URLis set; the measurements are the operator's post-deploy step (seeservices/ollama-piha/README.md).