homelab-codex-ws/services/kb-query/app/main.py
oskar 3d4ee3818d feat(kb-query): active embed fallback SOLARIA→PIHA (module 5 phase 4, plan §2/§5)
Last missing core piece of KB phase 4: kb-query no longer hard-fails /search
when Ollama@SOLARIA is unreachable. app/fallback.py implements the plan's
circuit-breaker exactly (30s cached health probe, 3s hard embed timeout on
SOLARIA, one-shot same-request switch to a new local ollama-piha@PIHA
container on timeout/error). sol_status in /healthz and /search now reflects
the real breaker state instead of a hardcoded "up".

New services/ollama-piha (bge-m3, OLLAMA_KEEP_ALIVE=0, arm64/no-GPU) is the
local fallback leg. Live calibration on PIHA (2026-07-27, normal load):
embed latency 4.2-5.2s, RAM peak ~983MiB against a 2.5GiB ceiling -- both
inside the plan's go-bar, so the fallback is enabled by default rather than
gated behind a flag. Calibration also surfaced and disabled (not removed) a
previously-undocumented orphaned native ollama.service on PIHA that had been
conflicting with the container's port.

The embed-model invariant (query embedding == document_chunk.model) still
enforces once at startup, since both fallback legs share one EMBED_MODEL
constant by construction; a redundant per-request DB check was deliberately
skipped and the invariant is instead proven structurally by test.

retrieval_eval.py gains --transport http (plan §2 decision 6/§9), previously
unimplemented. Verified live: HTTP transport is bit-identical to direct
transport against the same live SOLARIA (0 mismatches), and a live sol-down
simulation (kb-query's own OLLAMA_URL pointed at a dead address, no other
Ollama consumer touched) shows the PIHA fallback answering with the same
hit@3 gate outcome and dist within ~3e-4 of the SOLARIA baseline.

Zero changes to DB schema or kb_retrieval's retrieval logic -- only the
embed + health layer, per task constraints.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-27 22:23:53 +02:00

103 lines
4.5 KiB
Python

"""kb-query -- module 5 phase 4 (docs/kb/modules/05-faza4-plan.md §4): first user-facing HTTP
entry point to the KB. Wraps `kb_retrieval.cascade_query`/`flat_query` (module 5 phase 3,
already gated PASS -- docs/sessions/2026-07-21.md) in FastAPI. This is a search API, not chat:
no answer synthesis, no LLM call over the results (that is phase 5, out of scope here).
Embed path (plan §2 decision 2, §5): `app/fallback.py`'s `SolCircuitBreaker` health-checks
Ollama@SOLARIA (cached ~30s, ~500ms probe), embeds there with a hard ~3s timeout when up, and
falls through to Ollama@PIHA (`OLLAMA_PIHA_URL`, local fallback container `ollama-piha`) in the
same request on timeout/error -- the user only ever pays one extra timeout, once per cache
window. A failure on *both* legs still surfaces as 503 to the caller rather than a bare 500.
`GET /` (Krok 4, plan §7) serves the search UI from this same FastAPI process -- one image, one
container (plan §2 decision 4): a Jinja2 shell + a static vanilla-JS file, no node build step.
`/` and `/static/*` need no DB/Ollama, so they stay reachable even while `/search` is 503ing.
`mode=hybrid` (faza mailowa, plan Krok 3, docs/kb/modules/05-faza-mailowa-plan.md §6) is
available explicitly starting here, but the default stays `cascade` until the quality gate
(plan §8) PASSes on the full mail corpus -- flipping the default is a separate, later change.
"""
from __future__ import annotations
import os
import pathlib
from contextlib import asynccontextmanager
import aiohttp
from fastapi import FastAPI, HTTPException, Query, Request
from fastapi.responses import HTMLResponse
from fastapi.staticfiles import StaticFiles
from fastapi.templating import Jinja2Templates
from app.db import create_pool
from app.fallback import SolCircuitBreaker, resolve_sol_status
from app.search import run_search
from app.startup import validate_embed_model
BASE_DIR = pathlib.Path(__file__).resolve().parent
KB_DSN = os.environ.get("KB_DSN")
OLLAMA_URL = os.environ.get("OLLAMA_URL", "http://solaria:11434")
# Placeholder until ollama-piha is deployed (plan §5 calibration gate) -- nothing listens on
# localhost:11434 inside this container, so this fails closed (connection refused) exactly like
# SOLARIA-down does today, never a false success.
OLLAMA_PIHA_URL = os.environ.get("OLLAMA_PIHA_URL", "http://localhost:11434")
EMBED_MODEL = os.environ.get("EMBED_MODEL", "bge-m3")
SUMMARY_MODEL = os.environ.get("SUMMARY_MODEL", "claude-haiku-4-5")
@asynccontextmanager
async def lifespan(app: FastAPI):
if not KB_DSN:
raise RuntimeError("KB_DSN is required (see env.example)")
pool = await create_pool(KB_DSN)
async with pool.acquire() as conn:
# Hard invariant (plan §2 decision 2): refuse to start rather than silently serve
# queries against a mismatched embedding space. embed_model is a single constant used
# identically by both the SOLARIA and PIHA embed legs (app/fallback.py) -- this one
# check covers both paths, see app/fallback.py's module docstring for why a second,
# per-request DB check would be redundant.
await validate_embed_model(conn, EMBED_MODEL)
app.state.pool = pool
app.state.http = aiohttp.ClientSession()
app.state.sol_breaker = SolCircuitBreaker()
try:
yield
finally:
await app.state.http.close()
await pool.close()
app = FastAPI(lifespan=lifespan)
app.mount("/static", StaticFiles(directory=BASE_DIR / "static"), name="static")
templates = Jinja2Templates(directory=BASE_DIR / "templates")
@app.get("/", response_class=HTMLResponse)
async def index(request: Request):
return templates.TemplateResponse(request, "index.html")
@app.get("/healthz")
async def healthz() -> dict:
# Shares the same cached breaker /search uses (app/fallback.py) -- healthz and search
# always agree on the current sol_status instead of running independent probes.
sol_status = await resolve_sol_status(app.state.sol_breaker, app.state.http, OLLAMA_URL)
return {"status": "ok", "sol_status": sol_status}
@app.get("/search")
async def search(
q: str = Query(..., min_length=1),
mode: str = Query("cascade", pattern="^(cascade|flat|hybrid)$"),
) -> dict:
try:
async with app.state.pool.acquire() as conn:
return await run_search(
conn, app.state.http, app.state.sol_breaker, OLLAMA_URL, OLLAMA_PIHA_URL,
q, mode, EMBED_MODEL, SUMMARY_MODEL,
)
except aiohttp.ClientError as exc:
raise HTTPException(status_code=503, detail=f"embed backend unavailable: {exc}") from exc