homelab-codex-ws/services/kb-query/app/main.py
oskar a73a986bbe feat(kb-query): add search frontend (module 5 phase 4, plan §7, Krok 4)
Krok 4 of the phase-4 plan done ahead of the local-embed-fallback step
(Krok 2, deliberately deferred -- embed stays a plain SOLARIA call, per
task instruction): one FastAPI process now serves both the /search API
and the UI, no separate frontend build (plan §2 decision 4).

- GET / renders a Jinja2 shell; app/static/app.js (vanilla, no build) and
  style.css are the whole client. Query -> /search, results grouped by
  envelope_id client-side (chunks sorted by dist, <details> fragments).
- Colour thresholds per plan §7: dist<0.45 green, 0.45-0.55 yellow (still
  shown with a warning), >0.55 never rendered as an individual result; if
  a query ends up with nothing renderable, one "Brak odpowiedzi w KB"
  message replaces the list, carrying the best observed dist.
- Paperless hits link out; gmail hits get a "kopiuj Message-ID" button
  (there's nothing to link to yet, plan §2 decision 3) plus header
  metadata. Cascade/flat toggle defaults to cascade. Footer shows
  sol_status, refreshed from /healthz on load and after each search.
- /search gained additive summary/summary_tags fields (document_summary,
  haiku track) so the UI can show a document summary as each result
  group's header -- non-breaking, existing response shape untouched.
- Tests: app/db.py + app/search.py unit tests (mocked DB/HTTP, no live
  deps) cover the new summary join; tests/test_frontend.py drives GET /
  and /static/* via TestClient without running the DB-requiring lifespan;
  tests/frontend/app.test.js (Node's built-in test runner, no framework)
  covers query-URL encoding, threshold colouring, and envelope grouping.
- Verified live: docker build + container against kb-postgres@PIHA over
  LAN and Ollama@SOLARIA over Tailscale -- GET / (HTML), /static/app.js,
  /healthz, and /search (cascade + flat) all round-tripped correctly,
  including real summary/summary_tags data.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-22 16:49:36 +02:00

89 lines
3.3 KiB
Python

"""kb-query -- module 5 phase 4 (docs/kb/modules/05-faza4-plan.md §4): first user-facing HTTP
entry point to the KB. Wraps `kb_retrieval.cascade_query`/`flat_query` (module 5 phase 3,
already gated PASS -- docs/sessions/2026-07-21.md) in FastAPI. This is a search API, not chat:
no answer synthesis, no LLM call over the results (that is phase 5, out of scope here).
Embed path is deliberately simple for this step: calls Ollama on SOLARIA directly, no
cache/circuit-breaker/local-PIHA-fallback (plan §2 decision 2, §5) -- that state machine is a
later, separate step. A failed embed call (SOLARIA unreachable) surfaces as 503 to the caller
rather than a bare 500.
`GET /` (Krok 4, plan §7) serves the search UI from this same FastAPI process -- one image, one
container (plan §2 decision 4): a Jinja2 shell + a static vanilla-JS file, no node build step.
`/` and `/static/*` need no DB/Ollama, so they stay reachable even while `/search` is 503ing.
"""
from __future__ import annotations
import os
import pathlib
from contextlib import asynccontextmanager
import aiohttp
from fastapi import FastAPI, HTTPException, Query, Request
from fastapi.responses import HTMLResponse
from fastapi.staticfiles import StaticFiles
from fastapi.templating import Jinja2Templates
from kb_retrieval.embed import check_ollama_health
from app.db import create_pool
from app.search import run_search
from app.startup import validate_embed_model
BASE_DIR = pathlib.Path(__file__).resolve().parent
KB_DSN = os.environ.get("KB_DSN")
OLLAMA_URL = os.environ.get("OLLAMA_URL", "http://solaria:11434")
EMBED_MODEL = os.environ.get("EMBED_MODEL", "bge-m3")
SUMMARY_MODEL = os.environ.get("SUMMARY_MODEL", "claude-haiku-4-5")
OLLAMA_HEALTH_TIMEOUT_S = 3.0
@asynccontextmanager
async def lifespan(app: FastAPI):
if not KB_DSN:
raise RuntimeError("KB_DSN is required (see env.example)")
pool = await create_pool(KB_DSN)
async with pool.acquire() as conn:
# Hard invariant (plan §2 decision 2): refuse to start rather than silently serve
# queries against a mismatched embedding space.
await validate_embed_model(conn, EMBED_MODEL)
app.state.pool = pool
app.state.http = aiohttp.ClientSession()
try:
yield
finally:
await app.state.http.close()
await pool.close()
app = FastAPI(lifespan=lifespan)
app.mount("/static", StaticFiles(directory=BASE_DIR / "static"), name="static")
templates = Jinja2Templates(directory=BASE_DIR / "templates")
@app.get("/", response_class=HTMLResponse)
async def index(request: Request):
return templates.TemplateResponse(request, "index.html")
@app.get("/healthz")
async def healthz() -> dict:
sol_up = await check_ollama_health(app.state.http, OLLAMA_URL, OLLAMA_HEALTH_TIMEOUT_S)
return {"status": "ok", "sol_status": "up" if sol_up else "down"}
@app.get("/search")
async def search(
q: str = Query(..., min_length=1),
mode: str = Query("cascade", pattern="^(cascade|flat)$"),
) -> dict:
try:
async with app.state.pool.acquire() as conn:
return await run_search(
conn, app.state.http, OLLAMA_URL, q, mode, EMBED_MODEL, SUMMARY_MODEL
)
except aiohttp.ClientError as exc:
raise HTTPException(status_code=503, detail=f"embed backend unavailable: {exc}") from exc