Module 5 phase 4 step 1 (docs/kb/modules/05-faza4-plan.md, §4): first user-facing HTTP entry point to the KB. FastAPI wrapping kb_retrieval.cascade_query/flat_query — GET /search (query_text -> embed via Ollama@SOLARIA -> cascade/flat -> envelope join -> JSON with per-source links) and GET /healthz. Search API only, no answer synthesis (phase 5) and no server-side dist filtering — the 0.45/0.55 colour thresholds are a frontend concern (plan §7, a later step). Hard startup invariant (plan §2 decision 2): refuses to start unless the configured EMBED_MODEL is present in both document_chunk.model and document_summary.embedding_model. Note the latter: document_summary.model is the LLM that *wrote* the summary (claude-haiku-4-5/gemma3:12b), not the embedder — checked live against kb-postgres@PIHA before writing this, see app/startup.py's docstring. Verified end-to-end with a live docker run: the invariant crash-loops on a mismatched EMBED_MODEL and passes through to a real /search hit against the live corpus with a correct model. Repo-only: no deploy, no npm/OIDC/DNS wiring (plan §8, later step), no local embed fallback (plan §5, later step) — Ollama@SOLARIA is called directly and a failure surfaces as 503, not a crash. Also: scripts/deploy/deploy.sh's gate now builds each service via `docker compose build` instead of a raw `docker build <svc_dir>`, so a service whose docker-compose.yml declares a repo-root build context (needed here to COPY packages/kb-retrieval/, the packages/ Dockerfile convention already documented in CLAUDE.md) resolves the same way in the gate as it does at real deploy time (deploy-node.sh's `docker compose ... up --build`). No behavior change for existing single-context services — verified against llm-gateway's compose file. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
3.3 KiB
kb-query
FastAPI search API in front of the module-5 KB retrieval engine
(packages/kb-retrieval/). Runs on PIHA, bound to PIHA's LAN IP only
(exposure: private, same class as paperless/nextcloud — no public ingress
yet). This is a search API, not chat: no answer synthesis over results,
that's phase 5.
Endpoints
| Endpoint | Method | Purpose |
|---|---|---|
/healthz |
GET | {"status": "ok", "sol_status": "up"|"down"} — sol_status is a live probe of Ollama@SOLARIA, no auth required (monitoring must reach it) |
/search?q=<text>&mode=cascade|flat |
GET | query_text -> embed -> cascade_query/flat_query -> results, mode defaults to cascade |
/search response shape (module 5 phase 4 plan §4):
{
"query": "...", "mode": "cascade", "sol_status": "up",
"results": [
{"envelope_id": "paperless:119", "source": "paperless", "dist": 0.34,
"chunk_index": 2, "text": "...", "link": "https://paper.kapala.org/documents/119/details"},
{"envelope_id": "<Message-ID>", "source": "gmail", "dist": 0.44,
"chunk_index": 0, "text": "...", "subject": "...", "from": "...", "date": "...",
"link": null, "mail_ui_url": null}
]
}
dist is never filtered server-side — the 0.45/0.55 colour thresholds are a
frontend concern (a later step), not an API contract.
Embed path (current step)
Calls Ollama on SOLARIA directly per request — no cache, no circuit breaker,
no local-PIHA fallback yet (that state machine, plan §2 decision 2/§5, is a
separate later step). If SOLARIA is unreachable, /search returns 503;
/healthz still answers (sol_status: "down"), same tolerance pattern as
llm-gateway.
Startup invariant (hard-fail)
At startup, kb-query queries document_chunk.model and
document_summary.embedding_model for the set of models behind active
embeddings, and refuses to start (crash-loop, visible via container restarts)
if the configured EMBED_MODEL (default bge-m3) isn't in both sets. This
guards against querying with an embedding space that doesn't match what's
actually indexed — see app/startup.py for why the check reads
document_summary.embedding_model and not .model (the latter is the LLM
that wrote the summary, e.g. claude-haiku-4-5, not the embedder).
Configuration
.env — gitignored, copy from env.example. Required: LAN_BIND_IP,
KB_DSN. Optional: OLLAMA_URL, EMBED_MODEL, SUMMARY_MODEL.
Deploy (PIHA)
git pullon PIHA (~/homelab-codex-ws).cp services/kb-query/env.example services/kb-query/.envand fill in the realKB_DSNpassword.-
docker compose -f services/kb-query/docker-compose.yml \ -f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build - Verify:
services/kb-query/healthcheck.sh, then from PIHA:curl "http://192.168.31.5:8230/search?q=test".
Tests
pip install -e packages/kb-retrieval/
cd services/kb-query && pytest
Unit tests mock the DB connection and Ollama HTTP session (no live DB/Ollama
required) — same style as packages/kb-retrieval/tests/.
Out of scope for this step
- Local-PIHA embed fallback / circuit breaker (plan §2 decision 2, §5).
- npm@PIHA vhost, OIDC login, DNS (plan §8) —
kb.kapala.orgis not wired up yet; reach the API directly over LAN/Tailscale for now. - Frontend (plan §7) —
/searchreturns bare JSON.