Find a file
oskar 17640526ef feat(mail-body-ingest): circuit breaker na martwy backend embed przed Etapem B
Tolerancja pojedynczego nieudanego batcha jest słuszna, przeżycie martwej Ollamy
już nie. Parse archiwum jest jednowątkowy i wyprzedza GPU, więc przy Etapie B
(~212k kopert bez --since) padnięta Ollama przemieliłaby resztę korpusu z
prędkością parse'u, oznaczając każdy chunk jako chunks_errors — bez ani jednego
zapisu, ale kosztem ~2 h przebiegu do powtórzenia. Znany tryb awarii
Ollama@SOLARIA jest totalny (zniknięcie kontenera / network-detach, 4 incydenty,
§1.4/§7), nie częściowy, więc próg z kolejnych porażek trafia w niego od razu.

--max-embed-failures N (domyślnie 5, 0 wyłącza) → EmbedBackendUnavailableError
i exit 2, odrębny od exit 1 (który pełny korpus osiąga legalnie na pojedynczych
parse_errors — §1.5). Licznik zeruje się po udanym batchu, więc kryterium jest
"kolejnych", nie "łącznie". Przy abortcie dopychane są zaległe wpisy
entities[type=threading]: nie zależą od Ollamy, są idempotentne, a ich odtworzenie
oznaczałoby ponowny odczyt tych samych 27 GB. Nowy licznik embed_batch_failures
jest wyłącznie diagnostyczny — równania bilansu bez zmian.

Plan §9: dopisane decyzje operatora do Etapu B (plastry po 50k, breaker, pominięty
dry-run całości, hybrid default poza zakresem) + nota jak czytać exit 1 vs exit 2.

Testy: 4 nowe (trip po N kolejnych, reset po sukcesie, 0 wyłącza, flush threadingu
przy abortcie); 55 passed mail-body-ingest, 25 passed kb-retrieval. Smoke:
--limit 5 dry-run na żywym kb-postgres@PIHA — bilans domknięty, zero zapisów,
zero wywołań Ollamy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 16:53:04 +02:00
.claude/skills docs: session 2026-06-11 — lustro ssh shipping fix + ha-diag-agent piha + backlog/flota-bomba 2026-06-11 14:18:00 +02:00
backups/zigbee Add Zigbee coordinator backup 2026-05-14 18:24:26 +02:00
docs feat(mail-body-ingest): circuit breaker na martwy backend embed przed Etapem B 2026-08-04 16:53:04 +02:00
dotfiles add shared zshrc 2026-05-10 20:52:44 +02:00
hardware/esp/ir-ac-ha-integration feat(gree-ir): wlasny komponent ESPHome dla klimy Gree 2026-07-22 16:46:03 +02:00
hosts fix(vps): M1 mitigation — NODE_TYPE=lte_node disables unfiltered prune 2026-08-04 16:20:02 +02:00
inventory docs(topology): lustro daily duty cycle — nightly power-off, expected liveness cycles 2026-07-30 15:36:13 +02:00
jobs feat(mail-body-ingest): circuit breaker na martwy backend embed przed Etapem B 2026-08-04 16:53:04 +02:00
packages feat(kb-retrieval,kb-query): add hybrid retrieval mode (faza mailowa Krok 3) 2026-07-23 17:06:49 +02:00
scripts fix(control-plane): redeploy wykonywalny — dispatch do host-side deploy-runnera 2026-08-03 18:26:30 +02:00
services fix(node-agent): R1/R2/R3 — stop deleting deliberately stopped containers 2026-08-04 16:20:02 +02:00
.codex Document current homelab state 2026-04-15 17:37:25 +02:00
.gitignore feat(hardware): ESPHome Gree IR blaster (Wemos D1 Mini + Grove IR) 2026-07-22 16:46:03 +02:00
.mcp.json feat(ha-mcp): read-only MCP server (faza 2a) 2026-07-30 16:47:27 +02:00
CLAUDE.md fix(control-plane): redeploy wykonywalny — dispatch do host-side deploy-runnera 2026-08-03 18:26:30 +02:00
codex_context Add session context state 2026-04-20 22:10:39 +02:00
codex_context.yaml add shared context lock 2026-05-05 17:25:50 +02:00
deploy_agent.py Add deploy escalation output 2026-04-22 22:08:26 +02:00
ollama_client.py Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
README.md docs: add link to maintenance plan in README.md 2026-08-03 13:03:58 +02:00
start-aider.sh Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
start-codex.sh Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
sync-context.sh add shared context lock 2026-05-05 17:25:50 +02:00
tech-debt.md docs(tech-debt): cleanup artefaktow paste-error w ~oskar (SOLARIA, niski prio) 2026-06-24 17:13:51 +02:00
update-context.md Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00

Homelab Codex

GitOps-lite orchestration for a distributed homelab environment.

Architecture

The homelab consists of several nodes connected via a Tailscale internal mesh.

Host Role Description
SATURN Primary Node Development, orchestration, and git source of truth (commit node).
SOLARIA Compute Node GPU, inference, and heavy compute workloads.
PIHA Infra Node Core infrastructure services, automation, and monitoring.
VPS Edge Node Public ingress, reverse proxy, and edge services.

Agent System

The homelab uses a multi-agent orchestration model with human-in-the-loop for destructive actions:

Agent Node Role
stability-agent all nodes Per-node watchdog — monitors Docker, disk, Tailscale, MQTT; emits events
node-agent all nodes Publishes container health events to Redis pub/sub
observer VPS Synthesizes world state from events into /opt/homelab/world/*.json
supervisor VPS Detects drift between desired and actual state; writes pending actions
planner-agent SOLARIA LLM-powered diagnosis — listens to Redis, proposes remediation actions
executor VPS Executes actions only after operator approval
operator-ui + telegram-bot VPS / PIHA Operator reviews and approves/rejects pending actions

Action approval flow: pending/ → operator approves → approved/ → executor runs.

Repository Structure

Getting Started

  1. Standardization: Follow the Infrastructure Standards.
  2. Deployment: See Deployment Conventions for how to roll out changes.
  3. SATURN: Remember that SATURN is the only node where commits should be made.

Documentation Index


Note: This repository documents the state of the homelab. Runtime state lives outside the repository in /opt/homelab.