Find a file
oskar 4bfd6c4429 fix(observer): map containers_not_running to unhealthy status + incident — dead/crash-looping containers created no incident, so supervisor never alerted
node-agent (and stability-agent) emit containers_not_running for exited/dead
and crash-looping containers, but the observer's event->status/incident map
only handled service_recovered/service_healthy/service_unhealthy/
healthcheck_failed. containers_not_running fell through: status stayed at its
last "healthy" value and no incident was opened, so the supervisor saw no drift
and generated no action — a dead container produced ZERO operator alerts
(matches the "action queue empty despite failures" symptom).

- containers_not_running now sets status=unhealthy and opens an incident whose
  trigger_type ("containers_not_running") is already in the supervisor's
  CONTAINER_RESTART_TRIGGERS, so remediation (container_restart) fires with no
  supervisor change. Recovery is unchanged: service_healthy resolves the
  incident via the existing svc_key->incident_id link.
- container_restarting / container_state_unexpected (added in 4746ebe) are kept
  intentionally observational — no incident, status not flipped to unhealthy
  (would cause a false redeploy for a transient blip) — but leave a
  last_observation trace so they don't vanish. A real crash-loop still escalates
  via node-agent re-emitting containers_not_running.

Tests: services/control-plane/tests/test_observer_container_events.py — status
+ incident + trigger_type, end-to-end reconcile -> container_restart, recovery
auto-resolve, observational no-incident/trace, idempotent (no incident
multiplication). Full control-plane suite: 117 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 16:28:31 +02:00
.claude/skills docs: session 2026-06-11 — lustro ssh shipping fix + ha-diag-agent piha + backlog/flota-bomba 2026-06-11 14:18:00 +02:00
backups/zigbee Add Zigbee coordinator backup 2026-05-14 18:24:26 +02:00
docs docs(sesja): 2026-07-22 — faza 4 kroki 1-2 done, kb-query LIVE (pierwszy HTTP do KB); wzorzec incydentów Ollama do backlogu 2026-07-22 16:21:10 +02:00
dotfiles add shared zshrc 2026-05-10 20:52:44 +02:00
hosts feat(kb): add kb-query service skeleton (search API, no ingress yet) 2026-07-22 16:06:18 +02:00
inventory feat(kb): faza 3 krok 5 — cykliczny ingest (systemd timer) + alerting 2026-07-17 15:54:13 +02:00
jobs refactor(kb): extract packages/kb-retrieval from documents-ingest 2026-07-22 16:06:02 +02:00
packages refactor(kb): extract packages/kb-retrieval from documents-ingest 2026-07-22 16:06:02 +02:00
scripts fix(observer): map containers_not_running to unhealthy status + incident — dead/crash-looping containers created no incident, so supervisor never alerted 2026-07-22 16:28:31 +02:00
services fix(observer): map containers_not_running to unhealthy status + incident — dead/crash-looping containers created no incident, so supervisor never alerted 2026-07-22 16:28:31 +02:00
.codex Document current homelab state 2026-04-15 17:37:25 +02:00
.gitignore fix(piha): capabilities — realny RAM/NVMe + gitignore build dirs 2026-06-25 17:07:11 +02:00
CLAUDE.md docs(kb): etap 1 fundament done + session log 2026-06-17 2026-06-19 20:02:25 +02:00
codex_context Add session context state 2026-04-20 22:10:39 +02:00
codex_context.yaml add shared context lock 2026-05-05 17:25:50 +02:00
deploy_agent.py Add deploy escalation output 2026-04-22 22:08:26 +02:00
ollama_client.py Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
README.md docs: add planner-agent docs and session summary 2026-05-27 2026-05-27 22:35:59 +02:00
start-aider.sh Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
start-codex.sh Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
sync-context.sh add shared context lock 2026-05-05 17:25:50 +02:00
tech-debt.md docs(tech-debt): cleanup artefaktow paste-error w ~oskar (SOLARIA, niski prio) 2026-06-24 17:13:51 +02:00
update-context.md Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00

Homelab Codex

GitOps-lite orchestration for a distributed homelab environment.

Architecture

The homelab consists of several nodes connected via a Tailscale internal mesh.

Host Role Description
SATURN Primary Node Development, orchestration, and git source of truth (commit node).
SOLARIA Compute Node GPU, inference, and heavy compute workloads.
PIHA Infra Node Core infrastructure services, automation, and monitoring.
VPS Edge Node Public ingress, reverse proxy, and edge services.

Agent System

The homelab uses a multi-agent orchestration model with human-in-the-loop for destructive actions:

Agent Node Role
stability-agent all nodes Per-node watchdog — monitors Docker, disk, Tailscale, MQTT; emits events
node-agent all nodes Publishes container health events to Redis pub/sub
observer VPS Synthesizes world state from events into /opt/homelab/world/*.json
supervisor VPS Detects drift between desired and actual state; writes pending actions
planner-agent SOLARIA LLM-powered diagnosis — listens to Redis, proposes remediation actions
executor VPS Executes actions only after operator approval
operator-ui + telegram-bot VPS / PIHA Operator reviews and approves/rejects pending actions

Action approval flow: pending/ → operator approves → approved/ → executor runs.

Repository Structure

Getting Started

  1. Standardization: Follow the Infrastructure Standards.
  2. Deployment: See Deployment Conventions for how to roll out changes.
  3. SATURN: Remember that SATURN is the only node where commits should be made.

Documentation Index


Note: This repository documents the state of the homelab. Runtime state lives outside the repository in /opt/homelab.