Krok 4 of the phase-4 plan done ahead of the local-embed-fallback step (Krok 2, deliberately deferred -- embed stays a plain SOLARIA call, per task instruction): one FastAPI process now serves both the /search API and the UI, no separate frontend build (plan §2 decision 4). - GET / renders a Jinja2 shell; app/static/app.js (vanilla, no build) and style.css are the whole client. Query -> /search, results grouped by envelope_id client-side (chunks sorted by dist, <details> fragments). - Colour thresholds per plan §7: dist<0.45 green, 0.45-0.55 yellow (still shown with a warning), >0.55 never rendered as an individual result; if a query ends up with nothing renderable, one "Brak odpowiedzi w KB" message replaces the list, carrying the best observed dist. - Paperless hits link out; gmail hits get a "kopiuj Message-ID" button (there's nothing to link to yet, plan §2 decision 3) plus header metadata. Cascade/flat toggle defaults to cascade. Footer shows sol_status, refreshed from /healthz on load and after each search. - /search gained additive summary/summary_tags fields (document_summary, haiku track) so the UI can show a document summary as each result group's header -- non-breaking, existing response shape untouched. - Tests: app/db.py + app/search.py unit tests (mocked DB/HTTP, no live deps) cover the new summary join; tests/test_frontend.py drives GET / and /static/* via TestClient without running the DB-requiring lifespan; tests/frontend/app.test.js (Node's built-in test runner, no framework) covers query-URL encoding, threshold colouring, and envelope grouping. - Verified live: docker build + container against kb-postgres@PIHA over LAN and Ollama@SOLARIA over Tailscale -- GET / (HTML), /static/app.js, /healthz, and /search (cascade + flat) all round-tripped correctly, including real summary/summary_tags data. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
5.8 KiB
kb-query
FastAPI search API in front of the module-5 KB retrieval engine
(packages/kb-retrieval/). Runs on PIHA, bound to PIHA's LAN IP only
(exposure: private, same class as paperless/nextcloud — no public ingress
yet). This is a search API, not chat: no answer synthesis over results,
that's phase 5.
Endpoints
| Endpoint | Method | Purpose |
|---|---|---|
/ |
GET | Search UI (Jinja2 shell + /static/app.js, no login yet — plan §8 OIDC is a later step) |
/static/* |
GET | UI assets (app.js, style.css) |
/healthz |
GET | {"status": "ok", "sol_status": "up"|"down"} — sol_status is a live probe of Ollama@SOLARIA, no auth required (monitoring must reach it) |
/search?q=<text>&mode=cascade|flat |
GET | query_text -> embed -> cascade_query/flat_query -> results, mode defaults to cascade |
/search response shape (module 5 phase 4 plan §4, summary/summary_tags
added in Krok 4 for the UI's per-envelope result header — additive, does not
change any field the plan §4 shape already defined):
{
"query": "...", "mode": "cascade", "sol_status": "up",
"results": [
{"envelope_id": "paperless:119", "source": "paperless", "dist": 0.34,
"chunk_index": 2, "text": "...", "link": "https://paper.kapala.org/documents/119/details",
"summary": "...", "summary_tags": ["..."]},
{"envelope_id": "<Message-ID>", "source": "gmail", "dist": 0.44,
"chunk_index": 0, "text": "...", "subject": "...", "from": "...", "date": "...",
"link": null, "mail_ui_url": null, "summary": null, "summary_tags": []}
]
}
dist is never filtered server-side — the 0.45/0.55 colour thresholds (below)
are a frontend concern, not an API contract. summary/summary_tags come
from document_summary for SUMMARY_MODEL; null/[] when the envelope
has no summary yet.
Frontend (Krok 4, plan §7)
One page, served from this same FastAPI process — no separate frontend
container, no node build step (plan §2 decision 4): app/templates/index.html
(Jinja2 shell) + app/static/app.js (vanilla JS, fetch() to /search) +
app/static/style.css. Wszystko po polsku.
- Pole zapytania + submit (Enter lub przycisk), przełącznik trybu kaskada/flat (domyślnie kaskada — checkbox "tryb flat (debug)").
- Wyniki grupowane po
envelope_id(dokument): nagłówek trafienia to streszczenie dokumentu (summary, tor haiku) gdy dostępne, w przeciwnym razieenvelope_id; chunki są rozwijanymi fragmentami (<details>) pod nagłówkiem, posortowane podist. - Kolorowanie progów (fazy 3, zweryfikowane bramką):
dist < 0.45zielony,0.45–0.55żółty (nadal renderowany, z wizualnym ostrzeżeniem),> 0.55nigdy nie renderowany jako pojedynczy wynik. Jeśli po tym filtrze żadna grupa nie zostaje nic do pokazania (wszystkie trafienia > 0.55, albo brak trafień w ogóle), całość zastępuje komunikat "Brak odpowiedzi w KB dla tego zapytania" z najlepszym (najniższym) zaobserwowanymdistw nawiasie. - Źródło: Paperless → link "Otwórz w Paperless" (
link); Gmail → metadane (subject/from/date) + przycisk "Kopiuj Message-ID" (envelope_idjest Message-ID, plan §2 decyzja 3) — nie ma dokąd linkować, więc kopiowalny identyfikator zamiast martwego linku. - Stopka pokazuje
sol_statusdyskretnie (odświeżane z/healthzprzy starcie strony i po każdym wyszukiwaniu).
Embed path (current step)
Calls Ollama on SOLARIA directly per request — no cache, no circuit breaker,
no local-PIHA fallback yet (that state machine, plan §2 decision 2/§5, is a
separate later step). If SOLARIA is unreachable, /search returns 503;
/healthz still answers (sol_status: "down"), same tolerance pattern as
llm-gateway.
Startup invariant (hard-fail)
At startup, kb-query queries document_chunk.model and
document_summary.embedding_model for the set of models behind active
embeddings, and refuses to start (crash-loop, visible via container restarts)
if the configured EMBED_MODEL (default bge-m3) isn't in both sets. This
guards against querying with an embedding space that doesn't match what's
actually indexed — see app/startup.py for why the check reads
document_summary.embedding_model and not .model (the latter is the LLM
that wrote the summary, e.g. claude-haiku-4-5, not the embedder).
Configuration
.env — gitignored, copy from env.example. Required: LAN_BIND_IP,
KB_DSN. Optional: OLLAMA_URL, EMBED_MODEL, SUMMARY_MODEL.
Deploy (PIHA)
git pullon PIHA (~/homelab-codex-ws).cp services/kb-query/env.example services/kb-query/.envand fill in the realKB_DSNpassword.-
docker compose -f services/kb-query/docker-compose.yml \ -f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build - Verify:
services/kb-query/healthcheck.sh, then from PIHA:curl "http://192.168.31.5:8230/search?q=test"and openhttp://192.168.31.5:8230/in a browser.
Tests
pip install -e packages/kb-retrieval/
cd services/kb-query && pip install -r requirements.txt pytest pytest-asyncio && pytest
Unit tests mock the DB connection and Ollama HTTP session (no live DB/Ollama
required) — same style as packages/kb-retrieval/tests/. tests/test_frontend.py
drives GET ///static/* through FastAPI's TestClient without entering it
as a context manager, so the DB-requiring lifespan never runs.
Frontend JS has its own pure-function tests (query-URL encoding, threshold
colouring, envelope grouping), run without a browser via Node's built-in
test runner: node --test services/kb-query/tests/frontend/.
Out of scope for this step
- Local-PIHA embed fallback / circuit breaker (plan §2 decision 2, §5).
- npm@PIHA vhost, OIDC login, DNS (plan §8) —
kb.kapala.orgis not wired up yet; reach the API/UI directly over LAN/Tailscale for now, no auth.