Last missing core piece of KB phase 4: kb-query no longer hard-fails /search when Ollama@SOLARIA is unreachable. app/fallback.py implements the plan's circuit-breaker exactly (30s cached health probe, 3s hard embed timeout on SOLARIA, one-shot same-request switch to a new local ollama-piha@PIHA container on timeout/error). sol_status in /healthz and /search now reflects the real breaker state instead of a hardcoded "up". New services/ollama-piha (bge-m3, OLLAMA_KEEP_ALIVE=0, arm64/no-GPU) is the local fallback leg. Live calibration on PIHA (2026-07-27, normal load): embed latency 4.2-5.2s, RAM peak ~983MiB against a 2.5GiB ceiling -- both inside the plan's go-bar, so the fallback is enabled by default rather than gated behind a flag. Calibration also surfaced and disabled (not removed) a previously-undocumented orphaned native ollama.service on PIHA that had been conflicting with the container's port. The embed-model invariant (query embedding == document_chunk.model) still enforces once at startup, since both fallback legs share one EMBED_MODEL constant by construction; a redundant per-request DB check was deliberately skipped and the invariant is instead proven structurally by test. retrieval_eval.py gains --transport http (plan §2 decision 6/§9), previously unimplemented. Verified live: HTTP transport is bit-identical to direct transport against the same live SOLARIA (0 mismatches), and a live sol-down simulation (kb-query's own OLLAMA_URL pointed at a dead address, no other Ollama consumer touched) shows the PIHA fallback answering with the same hit@3 gate outcome and dist within ~3e-4 of the SOLARIA baseline. Zero changes to DB schema or kb_retrieval's retrieval logic -- only the embed + health layer, per task constraints. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
35 lines
1.9 KiB
YAML
35 lines
1.9 KiB
YAML
service:
|
|
name: kb-query
|
|
owner_node: piha
|
|
role: kb-search-api # module 5 phase 4: first user-facing HTTP entry point to the KB (API + minimal search UI, plan §7)
|
|
exposure: private # LAN/Tailscale only, npm@PIHA vhost (kb.kapala.org) is a later step
|
|
dependencies:
|
|
- kb-postgres
|
|
- forgejo # OIDC identity provider, wired in a later step (plan §8); not yet enforced
|
|
# ollama@SOLARIA is an external, optional runtime dependency (embed calls), not a hard
|
|
# dependency here -- SOLARIA may be offline; /search then fails per-request with 503,
|
|
# /healthz still answers (same tolerance pattern as llm-gateway).
|
|
ports:
|
|
- container: 8080
|
|
host: 8230 # LAN_BIND_IP only, never 0.0.0.0
|
|
protocol: tcp
|
|
healthcheck:
|
|
type: http
|
|
endpoint: http://192.168.31.5:8230/healthz # LAN bind — localhost does not answer
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 5
|
|
restart_policy: unless-stopped
|
|
persistence:
|
|
paths: [] # stateless — kb-postgres holds all state
|
|
runtime:
|
|
config_files:
|
|
- .env # KB_DSN, LAN_BIND_IP (gitignored, from env.example)
|
|
env_vars:
|
|
- LAN_BIND_IP # required — compose port-bind interpolation
|
|
- KB_DSN # required — asyncpg DSN for kb-postgres@PIHA
|
|
- OLLAMA_URL # optional — defaults to http://solaria:11434
|
|
- OLLAMA_PIHA_URL # optional — local fallback (ollama-piha@PIHA), defaults to http://localhost:11434 (fails closed until set)
|
|
- EMBED_MODEL # optional — defaults to bge-m3; startup invariant vs document_chunk/document_summary; same model used on both the SOLARIA and PIHA embed legs
|
|
- SUMMARY_MODEL # optional — defaults to claude-haiku-4-5; cascade_query's stage-1 model
|