Last missing core piece of KB phase 4: kb-query no longer hard-fails /search when Ollama@SOLARIA is unreachable. app/fallback.py implements the plan's circuit-breaker exactly (30s cached health probe, 3s hard embed timeout on SOLARIA, one-shot same-request switch to a new local ollama-piha@PIHA container on timeout/error). sol_status in /healthz and /search now reflects the real breaker state instead of a hardcoded "up". New services/ollama-piha (bge-m3, OLLAMA_KEEP_ALIVE=0, arm64/no-GPU) is the local fallback leg. Live calibration on PIHA (2026-07-27, normal load): embed latency 4.2-5.2s, RAM peak ~983MiB against a 2.5GiB ceiling -- both inside the plan's go-bar, so the fallback is enabled by default rather than gated behind a flag. Calibration also surfaced and disabled (not removed) a previously-undocumented orphaned native ollama.service on PIHA that had been conflicting with the container's port. The embed-model invariant (query embedding == document_chunk.model) still enforces once at startup, since both fallback legs share one EMBED_MODEL constant by construction; a redundant per-request DB check was deliberately skipped and the invariant is instead proven structurally by test. retrieval_eval.py gains --transport http (plan §2 decision 6/§9), previously unimplemented. Verified live: HTTP transport is bit-identical to direct transport against the same live SOLARIA (0 mismatches), and a live sol-down simulation (kb-query's own OLLAMA_URL pointed at a dead address, no other Ollama consumer touched) shows the PIHA fallback answering with the same hit@3 gate outcome and dist within ~3e-4 of the SOLARIA baseline. Zero changes to DB schema or kb_retrieval's retrieval logic -- only the embed + health layer, per task constraints. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
76 lines
3.2 KiB
Python
76 lines
3.2 KiB
Python
"""`/search` core -- module 5 phase 4 (docs/kb/modules/05-faza4-plan.md §4). Kept decoupled
|
|
from FastAPI so it can be unit-tested with fake `conn`/`session` objects, the same style as
|
|
`kb_retrieval`'s own tests, instead of needing a live DB/Ollama behind a TestClient.
|
|
|
|
Response shape (plan §4 exactly): `{"query", "mode", "sol_status", "results": [...]}`, each
|
|
result carrying `dist` un-filtered -- the 0.45/0.55 colour thresholds (plan §7) are a frontend
|
|
concern (Krok 4, out of this step's scope), never applied server-side.
|
|
|
|
`summary`/`summary_tags` (document_summary, haiku track) are an additive field added in Krok 4
|
|
for the frontend's per-envelope result header (plan §7) -- `None`/`[]` when the envelope has no
|
|
summary for `summary_model` yet. Purely additive: does not change any field already covered by
|
|
the phase-4 gate's HTTP-equivalence check (plan §9).
|
|
|
|
Embed fallback (plan §2 decision 2 / §5, `app/fallback.py`): the query embedding is computed
|
|
once via `embed_with_fallback` (SOLARIA, or PIHA if SOLARIA is down/times out), then handed to
|
|
`kb_retrieval`'s `*_retrieve` functions as a plain vector literal -- `flat_query`/`cascade_query`/
|
|
`hybrid_query` (which embed *and* retrieve in one call) are deliberately bypassed here so the
|
|
fallback decision lives entirely in this HTTP layer, per this task's constraint of zero changes
|
|
to `kb_retrieval`'s retrieval logic. `sol_status` in the response is the real outcome of that
|
|
call, not a hardcoded "up".
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import aiohttp
|
|
import asyncpg
|
|
|
|
from kb_retrieval.embed import _vector_literal
|
|
from kb_retrieval.retrieval import cascade_retrieve, flat_retrieve, hybrid_retrieve
|
|
|
|
from app.db import fetch_envelopes, fetch_summaries
|
|
from app.fallback import SolCircuitBreaker, embed_with_fallback
|
|
from app.links import build_result
|
|
|
|
|
|
async def run_search(
|
|
conn: asyncpg.Connection,
|
|
session: aiohttp.ClientSession,
|
|
breaker: SolCircuitBreaker,
|
|
solaria_url: str,
|
|
piha_url: str,
|
|
query_text: str,
|
|
mode: str,
|
|
embed_model: str,
|
|
summary_model: str,
|
|
) -> dict:
|
|
embedding, sol_status = await embed_with_fallback(
|
|
breaker, session, solaria_url, piha_url, embed_model, query_text
|
|
)
|
|
query_vector = _vector_literal(embedding)
|
|
|
|
if mode == "flat":
|
|
chunks = await flat_retrieve(conn, query_vector) # flat_retrieve returns a plain list
|
|
elif mode == "hybrid":
|
|
chunks = (await hybrid_retrieve(conn, query_vector, summary_model=summary_model))["chunks"]
|
|
else:
|
|
chunks = (await cascade_retrieve(conn, query_vector, summary_model=summary_model))["chunks"]
|
|
|
|
envelope_ids = sorted({c["envelope_id"] for c in chunks})
|
|
envelopes = await fetch_envelopes(conn, envelope_ids)
|
|
summaries = await fetch_summaries(conn, envelope_ids, summary_model)
|
|
|
|
results = []
|
|
for chunk in chunks:
|
|
result = build_result(chunk, envelopes.get(chunk["envelope_id"]))
|
|
summary = summaries.get(chunk["envelope_id"])
|
|
result["summary"] = summary["summary"] if summary else None
|
|
result["summary_tags"] = summary["tags"] if summary else []
|
|
results.append(result)
|
|
|
|
return {
|
|
"query": query_text,
|
|
"mode": mode,
|
|
"sol_status": sol_status,
|
|
"results": results,
|
|
}
|