homelab-codex-ws/services/kb-query/README.md

86 lines
3.3 KiB
Markdown
Raw Normal View History

feat(kb): add kb-query service skeleton (search API, no ingress yet) Module 5 phase 4 step 1 (docs/kb/modules/05-faza4-plan.md, §4): first user-facing HTTP entry point to the KB. FastAPI wrapping kb_retrieval.cascade_query/flat_query — GET /search (query_text -> embed via Ollama@SOLARIA -> cascade/flat -> envelope join -> JSON with per-source links) and GET /healthz. Search API only, no answer synthesis (phase 5) and no server-side dist filtering — the 0.45/0.55 colour thresholds are a frontend concern (plan §7, a later step). Hard startup invariant (plan §2 decision 2): refuses to start unless the configured EMBED_MODEL is present in both document_chunk.model and document_summary.embedding_model. Note the latter: document_summary.model is the LLM that *wrote* the summary (claude-haiku-4-5/gemma3:12b), not the embedder — checked live against kb-postgres@PIHA before writing this, see app/startup.py's docstring. Verified end-to-end with a live docker run: the invariant crash-loops on a mismatched EMBED_MODEL and passes through to a real /search hit against the live corpus with a correct model. Repo-only: no deploy, no npm/OIDC/DNS wiring (plan §8, later step), no local embed fallback (plan §5, later step) — Ollama@SOLARIA is called directly and a failure surfaces as 503, not a crash. Also: scripts/deploy/deploy.sh's gate now builds each service via `docker compose build` instead of a raw `docker build <svc_dir>`, so a service whose docker-compose.yml declares a repo-root build context (needed here to COPY packages/kb-retrieval/, the packages/ Dockerfile convention already documented in CLAUDE.md) resolves the same way in the gate as it does at real deploy time (deploy-node.sh's `docker compose ... up --build`). No behavior change for existing single-context services — verified against llm-gateway's compose file. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-22 16:06:18 +02:00
# kb-query
FastAPI search API in front of the module-5 KB retrieval engine
(`packages/kb-retrieval/`). Runs on **PIHA**, bound to PIHA's LAN IP only
(`exposure: private`, same class as paperless/nextcloud — no public ingress
yet). This is a **search API, not chat**: no answer synthesis over results,
that's phase 5.
## Endpoints
| Endpoint | Method | Purpose |
|---|---|---|
| `/healthz` | GET | `{"status": "ok", "sol_status": "up"\|"down"}``sol_status` is a live probe of Ollama@SOLARIA, no auth required (monitoring must reach it) |
| `/search?q=<text>&mode=cascade\|flat` | GET | `query_text -> embed -> cascade_query/flat_query -> results`, `mode` defaults to `cascade` |
`/search` response shape (module 5 phase 4 plan §4):
```json
{
"query": "...", "mode": "cascade", "sol_status": "up",
"results": [
{"envelope_id": "paperless:119", "source": "paperless", "dist": 0.34,
"chunk_index": 2, "text": "...", "link": "https://paper.kapala.org/documents/119/details"},
{"envelope_id": "<Message-ID>", "source": "gmail", "dist": 0.44,
"chunk_index": 0, "text": "...", "subject": "...", "from": "...", "date": "...",
"link": null, "mail_ui_url": null}
]
}
```
`dist` is never filtered server-side — the 0.45/0.55 colour thresholds are a
frontend concern (a later step), not an API contract.
## Embed path (current step)
Calls Ollama on SOLARIA directly per request — no cache, no circuit breaker,
no local-PIHA fallback yet (that state machine, plan §2 decision 2/§5, is a
separate later step). If SOLARIA is unreachable, `/search` returns **503**;
`/healthz` still answers (`sol_status: "down"`), same tolerance pattern as
`llm-gateway`.
## Startup invariant (hard-fail)
At startup, kb-query queries `document_chunk.model` and
`document_summary.embedding_model` for the set of models behind *active*
embeddings, and refuses to start (crash-loop, visible via container restarts)
if the configured `EMBED_MODEL` (default `bge-m3`) isn't in both sets. This
guards against querying with an embedding space that doesn't match what's
actually indexed — see `app/startup.py` for why the check reads
`document_summary.embedding_model` and not `.model` (the latter is the LLM
that *wrote* the summary, e.g. `claude-haiku-4-5`, not the embedder).
## Configuration
`.env`**gitignored**, copy from `env.example`. Required: `LAN_BIND_IP`,
`KB_DSN`. Optional: `OLLAMA_URL`, `EMBED_MODEL`, `SUMMARY_MODEL`.
## Deploy (PIHA)
1. `git pull` on PIHA (`~/homelab-codex-ws`).
2. `cp services/kb-query/env.example services/kb-query/.env` and fill in the
real `KB_DSN` password.
3. ```
docker compose -f services/kb-query/docker-compose.yml \
-f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build
```
4. Verify: `services/kb-query/healthcheck.sh`, then from PIHA:
`curl "http://192.168.31.5:8230/search?q=test"`.
## Tests
```
pip install -e packages/kb-retrieval/
cd services/kb-query && pytest
```
Unit tests mock the DB connection and Ollama HTTP session (no live DB/Ollama
required) — same style as `packages/kb-retrieval/tests/`.
## Out of scope for this step
- Local-PIHA embed fallback / circuit breaker (plan §2 decision 2, §5).
- npm@PIHA vhost, OIDC login, DNS (plan §8) — `kb.kapala.org` is not wired up
yet; reach the API directly over LAN/Tailscale for now.
- Frontend (plan §7) — `/search` returns bare JSON.