# kb-query FastAPI search API in front of the module-5 KB retrieval engine (`packages/kb-retrieval/`). Runs on **PIHA**, bound to PIHA's LAN IP only (`exposure: private`, same class as paperless/nextcloud — no public ingress yet). This is a **search API, not chat**: no answer synthesis over results, that's phase 5. ## Endpoints | Endpoint | Method | Purpose | |---|---|---| | `/healthz` | GET | `{"status": "ok", "sol_status": "up"\|"down"}` — `sol_status` is a live probe of Ollama@SOLARIA, no auth required (monitoring must reach it) | | `/search?q=&mode=cascade\|flat` | GET | `query_text -> embed -> cascade_query/flat_query -> results`, `mode` defaults to `cascade` | `/search` response shape (module 5 phase 4 plan §4): ```json { "query": "...", "mode": "cascade", "sol_status": "up", "results": [ {"envelope_id": "paperless:119", "source": "paperless", "dist": 0.34, "chunk_index": 2, "text": "...", "link": "https://paper.kapala.org/documents/119/details"}, {"envelope_id": "", "source": "gmail", "dist": 0.44, "chunk_index": 0, "text": "...", "subject": "...", "from": "...", "date": "...", "link": null, "mail_ui_url": null} ] } ``` `dist` is never filtered server-side — the 0.45/0.55 colour thresholds are a frontend concern (a later step), not an API contract. ## Embed path (current step) Calls Ollama on SOLARIA directly per request — no cache, no circuit breaker, no local-PIHA fallback yet (that state machine, plan §2 decision 2/§5, is a separate later step). If SOLARIA is unreachable, `/search` returns **503**; `/healthz` still answers (`sol_status: "down"`), same tolerance pattern as `llm-gateway`. ## Startup invariant (hard-fail) At startup, kb-query queries `document_chunk.model` and `document_summary.embedding_model` for the set of models behind *active* embeddings, and refuses to start (crash-loop, visible via container restarts) if the configured `EMBED_MODEL` (default `bge-m3`) isn't in both sets. This guards against querying with an embedding space that doesn't match what's actually indexed — see `app/startup.py` for why the check reads `document_summary.embedding_model` and not `.model` (the latter is the LLM that *wrote* the summary, e.g. `claude-haiku-4-5`, not the embedder). ## Configuration `.env` — **gitignored**, copy from `env.example`. Required: `LAN_BIND_IP`, `KB_DSN`. Optional: `OLLAMA_URL`, `EMBED_MODEL`, `SUMMARY_MODEL`. ## Deploy (PIHA) 1. `git pull` on PIHA (`~/homelab-codex-ws`). 2. `cp services/kb-query/env.example services/kb-query/.env` and fill in the real `KB_DSN` password. 3. ``` docker compose -f services/kb-query/docker-compose.yml \ -f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build ``` 4. Verify: `services/kb-query/healthcheck.sh`, then from PIHA: `curl "http://192.168.31.5:8230/search?q=test"`. ## Tests ``` pip install -e packages/kb-retrieval/ cd services/kb-query && pytest ``` Unit tests mock the DB connection and Ollama HTTP session (no live DB/Ollama required) — same style as `packages/kb-retrieval/tests/`. ## Out of scope for this step - Local-PIHA embed fallback / circuit breaker (plan §2 decision 2, §5). - npm@PIHA vhost, OIDC login, DNS (plan §8) — `kb.kapala.org` is not wired up yet; reach the API directly over LAN/Tailscale for now. - Frontend (plan §7) — `/search` returns bare JSON.