Last missing core piece of KB phase 4: kb-query no longer hard-fails /search when Ollama@SOLARIA is unreachable. app/fallback.py implements the plan's circuit-breaker exactly (30s cached health probe, 3s hard embed timeout on SOLARIA, one-shot same-request switch to a new local ollama-piha@PIHA container on timeout/error). sol_status in /healthz and /search now reflects the real breaker state instead of a hardcoded "up". New services/ollama-piha (bge-m3, OLLAMA_KEEP_ALIVE=0, arm64/no-GPU) is the local fallback leg. Live calibration on PIHA (2026-07-27, normal load): embed latency 4.2-5.2s, RAM peak ~983MiB against a 2.5GiB ceiling -- both inside the plan's go-bar, so the fallback is enabled by default rather than gated behind a flag. Calibration also surfaced and disabled (not removed) a previously-undocumented orphaned native ollama.service on PIHA that had been conflicting with the container's port. The embed-model invariant (query embedding == document_chunk.model) still enforces once at startup, since both fallback legs share one EMBED_MODEL constant by construction; a redundant per-request DB check was deliberately skipped and the invariant is instead proven structurally by test. retrieval_eval.py gains --transport http (plan §2 decision 6/§9), previously unimplemented. Verified live: HTTP transport is bit-identical to direct transport against the same live SOLARIA (0 mismatches), and a live sol-down simulation (kb-query's own OLLAMA_URL pointed at a dead address, no other Ollama consumer touched) shows the PIHA fallback answering with the same hit@3 gate outcome and dist within ~3e-4 of the SOLARIA baseline. Zero changes to DB schema or kb_retrieval's retrieval logic -- only the embed + health layer, per task constraints. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
103 lines
4.5 KiB
Python
103 lines
4.5 KiB
Python
"""kb-query -- module 5 phase 4 (docs/kb/modules/05-faza4-plan.md §4): first user-facing HTTP
|
|
entry point to the KB. Wraps `kb_retrieval.cascade_query`/`flat_query` (module 5 phase 3,
|
|
already gated PASS -- docs/sessions/2026-07-21.md) in FastAPI. This is a search API, not chat:
|
|
no answer synthesis, no LLM call over the results (that is phase 5, out of scope here).
|
|
|
|
Embed path (plan §2 decision 2, §5): `app/fallback.py`'s `SolCircuitBreaker` health-checks
|
|
Ollama@SOLARIA (cached ~30s, ~500ms probe), embeds there with a hard ~3s timeout when up, and
|
|
falls through to Ollama@PIHA (`OLLAMA_PIHA_URL`, local fallback container `ollama-piha`) in the
|
|
same request on timeout/error -- the user only ever pays one extra timeout, once per cache
|
|
window. A failure on *both* legs still surfaces as 503 to the caller rather than a bare 500.
|
|
|
|
`GET /` (Krok 4, plan §7) serves the search UI from this same FastAPI process -- one image, one
|
|
container (plan §2 decision 4): a Jinja2 shell + a static vanilla-JS file, no node build step.
|
|
`/` and `/static/*` need no DB/Ollama, so they stay reachable even while `/search` is 503ing.
|
|
|
|
`mode=hybrid` (faza mailowa, plan Krok 3, docs/kb/modules/05-faza-mailowa-plan.md §6) is
|
|
available explicitly starting here, but the default stays `cascade` until the quality gate
|
|
(plan §8) PASSes on the full mail corpus -- flipping the default is a separate, later change.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
import pathlib
|
|
from contextlib import asynccontextmanager
|
|
|
|
import aiohttp
|
|
from fastapi import FastAPI, HTTPException, Query, Request
|
|
from fastapi.responses import HTMLResponse
|
|
from fastapi.staticfiles import StaticFiles
|
|
from fastapi.templating import Jinja2Templates
|
|
|
|
from app.db import create_pool
|
|
from app.fallback import SolCircuitBreaker, resolve_sol_status
|
|
from app.search import run_search
|
|
from app.startup import validate_embed_model
|
|
|
|
BASE_DIR = pathlib.Path(__file__).resolve().parent
|
|
KB_DSN = os.environ.get("KB_DSN")
|
|
OLLAMA_URL = os.environ.get("OLLAMA_URL", "http://solaria:11434")
|
|
# Placeholder until ollama-piha is deployed (plan §5 calibration gate) -- nothing listens on
|
|
# localhost:11434 inside this container, so this fails closed (connection refused) exactly like
|
|
# SOLARIA-down does today, never a false success.
|
|
OLLAMA_PIHA_URL = os.environ.get("OLLAMA_PIHA_URL", "http://localhost:11434")
|
|
EMBED_MODEL = os.environ.get("EMBED_MODEL", "bge-m3")
|
|
SUMMARY_MODEL = os.environ.get("SUMMARY_MODEL", "claude-haiku-4-5")
|
|
|
|
|
|
@asynccontextmanager
|
|
async def lifespan(app: FastAPI):
|
|
if not KB_DSN:
|
|
raise RuntimeError("KB_DSN is required (see env.example)")
|
|
|
|
pool = await create_pool(KB_DSN)
|
|
async with pool.acquire() as conn:
|
|
# Hard invariant (plan §2 decision 2): refuse to start rather than silently serve
|
|
# queries against a mismatched embedding space. embed_model is a single constant used
|
|
# identically by both the SOLARIA and PIHA embed legs (app/fallback.py) -- this one
|
|
# check covers both paths, see app/fallback.py's module docstring for why a second,
|
|
# per-request DB check would be redundant.
|
|
await validate_embed_model(conn, EMBED_MODEL)
|
|
|
|
app.state.pool = pool
|
|
app.state.http = aiohttp.ClientSession()
|
|
app.state.sol_breaker = SolCircuitBreaker()
|
|
try:
|
|
yield
|
|
finally:
|
|
await app.state.http.close()
|
|
await pool.close()
|
|
|
|
|
|
app = FastAPI(lifespan=lifespan)
|
|
app.mount("/static", StaticFiles(directory=BASE_DIR / "static"), name="static")
|
|
templates = Jinja2Templates(directory=BASE_DIR / "templates")
|
|
|
|
|
|
@app.get("/", response_class=HTMLResponse)
|
|
async def index(request: Request):
|
|
return templates.TemplateResponse(request, "index.html")
|
|
|
|
|
|
@app.get("/healthz")
|
|
async def healthz() -> dict:
|
|
# Shares the same cached breaker /search uses (app/fallback.py) -- healthz and search
|
|
# always agree on the current sol_status instead of running independent probes.
|
|
sol_status = await resolve_sol_status(app.state.sol_breaker, app.state.http, OLLAMA_URL)
|
|
return {"status": "ok", "sol_status": sol_status}
|
|
|
|
|
|
@app.get("/search")
|
|
async def search(
|
|
q: str = Query(..., min_length=1),
|
|
mode: str = Query("cascade", pattern="^(cascade|flat|hybrid)$"),
|
|
) -> dict:
|
|
try:
|
|
async with app.state.pool.acquire() as conn:
|
|
return await run_search(
|
|
conn, app.state.http, app.state.sol_breaker, OLLAMA_URL, OLLAMA_PIHA_URL,
|
|
q, mode, EMBED_MODEL, SUMMARY_MODEL,
|
|
)
|
|
except aiohttp.ClientError as exc:
|
|
raise HTTPException(status_code=503, detail=f"embed backend unavailable: {exc}") from exc
|