feat(kb): add kb-query service skeleton (search API, no ingress yet)
Module 5 phase 4 step 1 (docs/kb/modules/05-faza4-plan.md, §4): first
user-facing HTTP entry point to the KB. FastAPI wrapping
kb_retrieval.cascade_query/flat_query — GET /search (query_text -> embed via
Ollama@SOLARIA -> cascade/flat -> envelope join -> JSON with per-source
links) and GET /healthz. Search API only, no answer synthesis (phase 5) and
no server-side dist filtering — the 0.45/0.55 colour thresholds are a
frontend concern (plan §7, a later step).
Hard startup invariant (plan §2 decision 2): refuses to start unless the
configured EMBED_MODEL is present in both document_chunk.model and
document_summary.embedding_model. Note the latter: document_summary.model is
the LLM that *wrote* the summary (claude-haiku-4-5/gemma3:12b), not the
embedder — checked live against kb-postgres@PIHA before writing this, see
app/startup.py's docstring. Verified end-to-end with a live docker run: the
invariant crash-loops on a mismatched EMBED_MODEL and passes through to a
real /search hit against the live corpus with a correct model.
Repo-only: no deploy, no npm/OIDC/DNS wiring (plan §8, later step), no local
embed fallback (plan §5, later step) — Ollama@SOLARIA is called directly and
a failure surfaces as 503, not a crash.
Also: scripts/deploy/deploy.sh's gate now builds each service via
`docker compose build` instead of a raw `docker build <svc_dir>`, so a
service whose docker-compose.yml declares a repo-root build context (needed
here to COPY packages/kb-retrieval/, the packages/ Dockerfile convention
already documented in CLAUDE.md) resolves the same way in the gate as it
does at real deploy time (deploy-node.sh's `docker compose ... up --build`).
No behavior change for existing single-context services — verified against
llm-gateway's compose file.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-22 16:06:18 +02:00
|
|
|
# kb-query secrets + host-local binds — copy to .env (gitignored) next to
|
|
|
|
|
# docker-compose.yml and fill in real values. Never commit .env.
|
|
|
|
|
|
|
|
|
|
# LAN IP of PIHA. The published port (8230) binds ONLY to this interface —
|
|
|
|
|
# never 0.0.0.0. Verify after host rebuilds: ip -4 addr.
|
|
|
|
|
LAN_BIND_IP=192.168.31.5
|
|
|
|
|
|
|
|
|
|
# asyncpg DSN for kb-postgres@PIHA. kb-query runs in its own Docker network
|
|
|
|
|
# (separate compose project from kb-postgres), so it reaches kb-postgres's
|
|
|
|
|
# published port over the host's LAN interface, not "localhost" — same
|
|
|
|
|
# reasoning as paperless-worker@SOLARIA reaching paperless@PIHA over LAN.
|
|
|
|
|
KB_DSN=postgresql://kb:CHANGE-ME@192.168.31.5:5433/kb
|
|
|
|
|
|
|
|
|
|
# Ollama upstream (SOLARIA, over Tailscale MagicDNS — same trick as
|
|
|
|
|
# llm-gateway's OLLAMA_URL). Optional: defaults to this value if unset.
|
|
|
|
|
# OLLAMA_URL=http://solaria:11434
|
|
|
|
|
|
feat(kb-query): active embed fallback SOLARIA→PIHA (module 5 phase 4, plan §2/§5)
Last missing core piece of KB phase 4: kb-query no longer hard-fails /search
when Ollama@SOLARIA is unreachable. app/fallback.py implements the plan's
circuit-breaker exactly (30s cached health probe, 3s hard embed timeout on
SOLARIA, one-shot same-request switch to a new local ollama-piha@PIHA
container on timeout/error). sol_status in /healthz and /search now reflects
the real breaker state instead of a hardcoded "up".
New services/ollama-piha (bge-m3, OLLAMA_KEEP_ALIVE=0, arm64/no-GPU) is the
local fallback leg. Live calibration on PIHA (2026-07-27, normal load):
embed latency 4.2-5.2s, RAM peak ~983MiB against a 2.5GiB ceiling -- both
inside the plan's go-bar, so the fallback is enabled by default rather than
gated behind a flag. Calibration also surfaced and disabled (not removed) a
previously-undocumented orphaned native ollama.service on PIHA that had been
conflicting with the container's port.
The embed-model invariant (query embedding == document_chunk.model) still
enforces once at startup, since both fallback legs share one EMBED_MODEL
constant by construction; a redundant per-request DB check was deliberately
skipped and the invariant is instead proven structurally by test.
retrieval_eval.py gains --transport http (plan §2 decision 6/§9), previously
unimplemented. Verified live: HTTP transport is bit-identical to direct
transport against the same live SOLARIA (0 mismatches), and a live sol-down
simulation (kb-query's own OLLAMA_URL pointed at a dead address, no other
Ollama consumer touched) shows the PIHA fallback answering with the same
hit@3 gate outcome and dist within ~3e-4 of the SOLARIA baseline.
Zero changes to DB schema or kb_retrieval's retrieval logic -- only the
embed + health layer, per task constraints.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-27 18:57:46 +02:00
|
|
|
# Local fallback Ollama (ollama-piha@PIHA, plan §2 decision 2 / §5) — used only when SOLARIA
|
|
|
|
|
# is unreachable/times out. kb-query runs in its own Docker network (separate compose project
|
|
|
|
|
# from ollama-piha), so this must be PIHA's LAN IP + published port, not "localhost" — same
|
|
|
|
|
# reasoning as KB_DSN above. Optional: defaults to http://localhost:11434, which fails closed
|
|
|
|
|
# (nothing listens there in this container) until you set the real value below.
|
|
|
|
|
# OLLAMA_PIHA_URL=http://192.168.31.5:11434
|
|
|
|
|
|
feat(kb): add kb-query service skeleton (search API, no ingress yet)
Module 5 phase 4 step 1 (docs/kb/modules/05-faza4-plan.md, §4): first
user-facing HTTP entry point to the KB. FastAPI wrapping
kb_retrieval.cascade_query/flat_query — GET /search (query_text -> embed via
Ollama@SOLARIA -> cascade/flat -> envelope join -> JSON with per-source
links) and GET /healthz. Search API only, no answer synthesis (phase 5) and
no server-side dist filtering — the 0.45/0.55 colour thresholds are a
frontend concern (plan §7, a later step).
Hard startup invariant (plan §2 decision 2): refuses to start unless the
configured EMBED_MODEL is present in both document_chunk.model and
document_summary.embedding_model. Note the latter: document_summary.model is
the LLM that *wrote* the summary (claude-haiku-4-5/gemma3:12b), not the
embedder — checked live against kb-postgres@PIHA before writing this, see
app/startup.py's docstring. Verified end-to-end with a live docker run: the
invariant crash-loops on a mismatched EMBED_MODEL and passes through to a
real /search hit against the live corpus with a correct model.
Repo-only: no deploy, no npm/OIDC/DNS wiring (plan §8, later step), no local
embed fallback (plan §5, later step) — Ollama@SOLARIA is called directly and
a failure surfaces as 503, not a crash.
Also: scripts/deploy/deploy.sh's gate now builds each service via
`docker compose build` instead of a raw `docker build <svc_dir>`, so a
service whose docker-compose.yml declares a repo-root build context (needed
here to COPY packages/kb-retrieval/, the packages/ Dockerfile convention
already documented in CLAUDE.md) resolves the same way in the gate as it
does at real deploy time (deploy-node.sh's `docker compose ... up --build`).
No behavior change for existing single-context services — verified against
llm-gateway's compose file.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-22 16:06:18 +02:00
|
|
|
# Embedding model kb-query enforces as a startup invariant (plan §2 decision
|
|
|
|
|
# 2) against document_chunk.model / document_summary.embedding_model.
|
|
|
|
|
# Optional: defaults to bge-m3.
|
|
|
|
|
# EMBED_MODEL=bge-m3
|
|
|
|
|
|
|
|
|
|
# document_summary.model kb-query's cascade path pre-filters on (the
|
|
|
|
|
# compilation track, plan §2 D3). Optional: defaults to claude-haiku-4-5.
|
|
|
|
|
# SUMMARY_MODEL=claude-haiku-4-5
|