Module 5 phase 4 step 1 (docs/kb/modules/05-faza4-plan.md, §4): first user-facing HTTP entry point to the KB. FastAPI wrapping kb_retrieval.cascade_query/flat_query — GET /search (query_text -> embed via Ollama@SOLARIA -> cascade/flat -> envelope join -> JSON with per-source links) and GET /healthz. Search API only, no answer synthesis (phase 5) and no server-side dist filtering — the 0.45/0.55 colour thresholds are a frontend concern (plan §7, a later step). Hard startup invariant (plan §2 decision 2): refuses to start unless the configured EMBED_MODEL is present in both document_chunk.model and document_summary.embedding_model. Note the latter: document_summary.model is the LLM that *wrote* the summary (claude-haiku-4-5/gemma3:12b), not the embedder — checked live against kb-postgres@PIHA before writing this, see app/startup.py's docstring. Verified end-to-end with a live docker run: the invariant crash-loops on a mismatched EMBED_MODEL and passes through to a real /search hit against the live corpus with a correct model. Repo-only: no deploy, no npm/OIDC/DNS wiring (plan §8, later step), no local embed fallback (plan §5, later step) — Ollama@SOLARIA is called directly and a failure surfaces as 503, not a crash. Also: scripts/deploy/deploy.sh's gate now builds each service via `docker compose build` instead of a raw `docker build <svc_dir>`, so a service whose docker-compose.yml declares a repo-root build context (needed here to COPY packages/kb-retrieval/, the packages/ Dockerfile convention already documented in CLAUDE.md) resolves the same way in the gate as it does at real deploy time (deploy-node.sh's `docker compose ... up --build`). No behavior change for existing single-context services — verified against llm-gateway's compose file. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
86 lines
3.3 KiB
Markdown
86 lines
3.3 KiB
Markdown
# kb-query
|
|
|
|
FastAPI search API in front of the module-5 KB retrieval engine
|
|
(`packages/kb-retrieval/`). Runs on **PIHA**, bound to PIHA's LAN IP only
|
|
(`exposure: private`, same class as paperless/nextcloud — no public ingress
|
|
yet). This is a **search API, not chat**: no answer synthesis over results,
|
|
that's phase 5.
|
|
|
|
## Endpoints
|
|
|
|
| Endpoint | Method | Purpose |
|
|
|---|---|---|
|
|
| `/healthz` | GET | `{"status": "ok", "sol_status": "up"\|"down"}` — `sol_status` is a live probe of Ollama@SOLARIA, no auth required (monitoring must reach it) |
|
|
| `/search?q=<text>&mode=cascade\|flat` | GET | `query_text -> embed -> cascade_query/flat_query -> results`, `mode` defaults to `cascade` |
|
|
|
|
`/search` response shape (module 5 phase 4 plan §4):
|
|
|
|
```json
|
|
{
|
|
"query": "...", "mode": "cascade", "sol_status": "up",
|
|
"results": [
|
|
{"envelope_id": "paperless:119", "source": "paperless", "dist": 0.34,
|
|
"chunk_index": 2, "text": "...", "link": "https://paper.kapala.org/documents/119/details"},
|
|
{"envelope_id": "<Message-ID>", "source": "gmail", "dist": 0.44,
|
|
"chunk_index": 0, "text": "...", "subject": "...", "from": "...", "date": "...",
|
|
"link": null, "mail_ui_url": null}
|
|
]
|
|
}
|
|
```
|
|
|
|
`dist` is never filtered server-side — the 0.45/0.55 colour thresholds are a
|
|
frontend concern (a later step), not an API contract.
|
|
|
|
## Embed path (current step)
|
|
|
|
Calls Ollama on SOLARIA directly per request — no cache, no circuit breaker,
|
|
no local-PIHA fallback yet (that state machine, plan §2 decision 2/§5, is a
|
|
separate later step). If SOLARIA is unreachable, `/search` returns **503**;
|
|
`/healthz` still answers (`sol_status: "down"`), same tolerance pattern as
|
|
`llm-gateway`.
|
|
|
|
## Startup invariant (hard-fail)
|
|
|
|
At startup, kb-query queries `document_chunk.model` and
|
|
`document_summary.embedding_model` for the set of models behind *active*
|
|
embeddings, and refuses to start (crash-loop, visible via container restarts)
|
|
if the configured `EMBED_MODEL` (default `bge-m3`) isn't in both sets. This
|
|
guards against querying with an embedding space that doesn't match what's
|
|
actually indexed — see `app/startup.py` for why the check reads
|
|
`document_summary.embedding_model` and not `.model` (the latter is the LLM
|
|
that *wrote* the summary, e.g. `claude-haiku-4-5`, not the embedder).
|
|
|
|
## Configuration
|
|
|
|
`.env` — **gitignored**, copy from `env.example`. Required: `LAN_BIND_IP`,
|
|
`KB_DSN`. Optional: `OLLAMA_URL`, `EMBED_MODEL`, `SUMMARY_MODEL`.
|
|
|
|
## Deploy (PIHA)
|
|
|
|
1. `git pull` on PIHA (`~/homelab-codex-ws`).
|
|
2. `cp services/kb-query/env.example services/kb-query/.env` and fill in the
|
|
real `KB_DSN` password.
|
|
3. ```
|
|
docker compose -f services/kb-query/docker-compose.yml \
|
|
-f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build
|
|
```
|
|
4. Verify: `services/kb-query/healthcheck.sh`, then from PIHA:
|
|
`curl "http://192.168.31.5:8230/search?q=test"`.
|
|
|
|
## Tests
|
|
|
|
```
|
|
pip install -e packages/kb-retrieval/
|
|
cd services/kb-query && pytest
|
|
```
|
|
|
|
Unit tests mock the DB connection and Ollama HTTP session (no live DB/Ollama
|
|
required) — same style as `packages/kb-retrieval/tests/`.
|
|
|
|
## Out of scope for this step
|
|
|
|
- Local-PIHA embed fallback / circuit breaker (plan §2 decision 2, §5).
|
|
- npm@PIHA vhost, OIDC login, DNS (plan §8) — `kb.kapala.org` is not wired up
|
|
yet; reach the API directly over LAN/Tailscale for now.
|
|
- Frontend (plan §7) — `/search` returns bare JSON.
|