homelab-codex-ws/services/kb-query/README.md
oskar b2379e3275 feat(kb): add kb-query service skeleton (search API, no ingress yet)
Module 5 phase 4 step 1 (docs/kb/modules/05-faza4-plan.md, §4): first
user-facing HTTP entry point to the KB. FastAPI wrapping
kb_retrieval.cascade_query/flat_query — GET /search (query_text -> embed via
Ollama@SOLARIA -> cascade/flat -> envelope join -> JSON with per-source
links) and GET /healthz. Search API only, no answer synthesis (phase 5) and
no server-side dist filtering — the 0.45/0.55 colour thresholds are a
frontend concern (plan §7, a later step).

Hard startup invariant (plan §2 decision 2): refuses to start unless the
configured EMBED_MODEL is present in both document_chunk.model and
document_summary.embedding_model. Note the latter: document_summary.model is
the LLM that *wrote* the summary (claude-haiku-4-5/gemma3:12b), not the
embedder — checked live against kb-postgres@PIHA before writing this, see
app/startup.py's docstring. Verified end-to-end with a live docker run: the
invariant crash-loops on a mismatched EMBED_MODEL and passes through to a
real /search hit against the live corpus with a correct model.

Repo-only: no deploy, no npm/OIDC/DNS wiring (plan §8, later step), no local
embed fallback (plan §5, later step) — Ollama@SOLARIA is called directly and
a failure surfaces as 503, not a crash.

Also: scripts/deploy/deploy.sh's gate now builds each service via
`docker compose build` instead of a raw `docker build <svc_dir>`, so a
service whose docker-compose.yml declares a repo-root build context (needed
here to COPY packages/kb-retrieval/, the packages/ Dockerfile convention
already documented in CLAUDE.md) resolves the same way in the gate as it
does at real deploy time (deploy-node.sh's `docker compose ... up --build`).
No behavior change for existing single-context services — verified against
llm-gateway's compose file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-22 16:06:18 +02:00

3.3 KiB

kb-query

FastAPI search API in front of the module-5 KB retrieval engine (packages/kb-retrieval/). Runs on PIHA, bound to PIHA's LAN IP only (exposure: private, same class as paperless/nextcloud — no public ingress yet). This is a search API, not chat: no answer synthesis over results, that's phase 5.

Endpoints

Endpoint Method Purpose
/healthz GET {"status": "ok", "sol_status": "up"|"down"}sol_status is a live probe of Ollama@SOLARIA, no auth required (monitoring must reach it)
/search?q=<text>&mode=cascade|flat GET query_text -> embed -> cascade_query/flat_query -> results, mode defaults to cascade

/search response shape (module 5 phase 4 plan §4):

{
  "query": "...", "mode": "cascade", "sol_status": "up",
  "results": [
    {"envelope_id": "paperless:119", "source": "paperless", "dist": 0.34,
     "chunk_index": 2, "text": "...", "link": "https://paper.kapala.org/documents/119/details"},
    {"envelope_id": "<Message-ID>", "source": "gmail", "dist": 0.44,
     "chunk_index": 0, "text": "...", "subject": "...", "from": "...", "date": "...",
     "link": null, "mail_ui_url": null}
  ]
}

dist is never filtered server-side — the 0.45/0.55 colour thresholds are a frontend concern (a later step), not an API contract.

Embed path (current step)

Calls Ollama on SOLARIA directly per request — no cache, no circuit breaker, no local-PIHA fallback yet (that state machine, plan §2 decision 2/§5, is a separate later step). If SOLARIA is unreachable, /search returns 503; /healthz still answers (sol_status: "down"), same tolerance pattern as llm-gateway.

Startup invariant (hard-fail)

At startup, kb-query queries document_chunk.model and document_summary.embedding_model for the set of models behind active embeddings, and refuses to start (crash-loop, visible via container restarts) if the configured EMBED_MODEL (default bge-m3) isn't in both sets. This guards against querying with an embedding space that doesn't match what's actually indexed — see app/startup.py for why the check reads document_summary.embedding_model and not .model (the latter is the LLM that wrote the summary, e.g. claude-haiku-4-5, not the embedder).

Configuration

.envgitignored, copy from env.example. Required: LAN_BIND_IP, KB_DSN. Optional: OLLAMA_URL, EMBED_MODEL, SUMMARY_MODEL.

Deploy (PIHA)

  1. git pull on PIHA (~/homelab-codex-ws).
  2. cp services/kb-query/env.example services/kb-query/.env and fill in the real KB_DSN password.
  3. docker compose -f services/kb-query/docker-compose.yml \
      -f hosts/piha/runtime/kb-query/docker-compose.override.yml up -d --build
    
  4. Verify: services/kb-query/healthcheck.sh, then from PIHA: curl "http://192.168.31.5:8230/search?q=test".

Tests

pip install -e packages/kb-retrieval/
cd services/kb-query && pytest

Unit tests mock the DB connection and Ollama HTTP session (no live DB/Ollama required) — same style as packages/kb-retrieval/tests/.

Out of scope for this step

  • Local-PIHA embed fallback / circuit breaker (plan §2 decision 2, §5).
  • npm@PIHA vhost, OIDC login, DNS (plan §8) — kb.kapala.org is not wired up yet; reach the API directly over LAN/Tailscale for now.
  • Frontend (plan §7) — /search returns bare JSON.