homelab-codex-ws/services/kb-query/app/search.py
oskar 3d4ee3818d feat(kb-query): active embed fallback SOLARIA→PIHA (module 5 phase 4, plan §2/§5)
Last missing core piece of KB phase 4: kb-query no longer hard-fails /search
when Ollama@SOLARIA is unreachable. app/fallback.py implements the plan's
circuit-breaker exactly (30s cached health probe, 3s hard embed timeout on
SOLARIA, one-shot same-request switch to a new local ollama-piha@PIHA
container on timeout/error). sol_status in /healthz and /search now reflects
the real breaker state instead of a hardcoded "up".

New services/ollama-piha (bge-m3, OLLAMA_KEEP_ALIVE=0, arm64/no-GPU) is the
local fallback leg. Live calibration on PIHA (2026-07-27, normal load):
embed latency 4.2-5.2s, RAM peak ~983MiB against a 2.5GiB ceiling -- both
inside the plan's go-bar, so the fallback is enabled by default rather than
gated behind a flag. Calibration also surfaced and disabled (not removed) a
previously-undocumented orphaned native ollama.service on PIHA that had been
conflicting with the container's port.

The embed-model invariant (query embedding == document_chunk.model) still
enforces once at startup, since both fallback legs share one EMBED_MODEL
constant by construction; a redundant per-request DB check was deliberately
skipped and the invariant is instead proven structurally by test.

retrieval_eval.py gains --transport http (plan §2 decision 6/§9), previously
unimplemented. Verified live: HTTP transport is bit-identical to direct
transport against the same live SOLARIA (0 mismatches), and a live sol-down
simulation (kb-query's own OLLAMA_URL pointed at a dead address, no other
Ollama consumer touched) shows the PIHA fallback answering with the same
hit@3 gate outcome and dist within ~3e-4 of the SOLARIA baseline.

Zero changes to DB schema or kb_retrieval's retrieval logic -- only the
embed + health layer, per task constraints.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-27 22:23:53 +02:00

76 lines
3.2 KiB
Python

"""`/search` core -- module 5 phase 4 (docs/kb/modules/05-faza4-plan.md §4). Kept decoupled
from FastAPI so it can be unit-tested with fake `conn`/`session` objects, the same style as
`kb_retrieval`'s own tests, instead of needing a live DB/Ollama behind a TestClient.
Response shape (plan §4 exactly): `{"query", "mode", "sol_status", "results": [...]}`, each
result carrying `dist` un-filtered -- the 0.45/0.55 colour thresholds (plan §7) are a frontend
concern (Krok 4, out of this step's scope), never applied server-side.
`summary`/`summary_tags` (document_summary, haiku track) are an additive field added in Krok 4
for the frontend's per-envelope result header (plan §7) -- `None`/`[]` when the envelope has no
summary for `summary_model` yet. Purely additive: does not change any field already covered by
the phase-4 gate's HTTP-equivalence check (plan §9).
Embed fallback (plan §2 decision 2 / §5, `app/fallback.py`): the query embedding is computed
once via `embed_with_fallback` (SOLARIA, or PIHA if SOLARIA is down/times out), then handed to
`kb_retrieval`'s `*_retrieve` functions as a plain vector literal -- `flat_query`/`cascade_query`/
`hybrid_query` (which embed *and* retrieve in one call) are deliberately bypassed here so the
fallback decision lives entirely in this HTTP layer, per this task's constraint of zero changes
to `kb_retrieval`'s retrieval logic. `sol_status` in the response is the real outcome of that
call, not a hardcoded "up".
"""
from __future__ import annotations
import aiohttp
import asyncpg
from kb_retrieval.embed import _vector_literal
from kb_retrieval.retrieval import cascade_retrieve, flat_retrieve, hybrid_retrieve
from app.db import fetch_envelopes, fetch_summaries
from app.fallback import SolCircuitBreaker, embed_with_fallback
from app.links import build_result
async def run_search(
conn: asyncpg.Connection,
session: aiohttp.ClientSession,
breaker: SolCircuitBreaker,
solaria_url: str,
piha_url: str,
query_text: str,
mode: str,
embed_model: str,
summary_model: str,
) -> dict:
embedding, sol_status = await embed_with_fallback(
breaker, session, solaria_url, piha_url, embed_model, query_text
)
query_vector = _vector_literal(embedding)
if mode == "flat":
chunks = await flat_retrieve(conn, query_vector) # flat_retrieve returns a plain list
elif mode == "hybrid":
chunks = (await hybrid_retrieve(conn, query_vector, summary_model=summary_model))["chunks"]
else:
chunks = (await cascade_retrieve(conn, query_vector, summary_model=summary_model))["chunks"]
envelope_ids = sorted({c["envelope_id"] for c in chunks})
envelopes = await fetch_envelopes(conn, envelope_ids)
summaries = await fetch_summaries(conn, envelope_ids, summary_model)
results = []
for chunk in chunks:
result = build_result(chunk, envelopes.get(chunk["envelope_id"]))
summary = summaries.get(chunk["envelope_id"])
result["summary"] = summary["summary"] if summary else None
result["summary_tags"] = summary["tags"] if summary else []
results.append(result)
return {
"query": query_text,
"mode": mode,
"sol_status": sol_status,
"results": results,
}