Find a file
oskar 71eaab0025 feat(supervisor): duty-cycle nodes — liveness transitions logged, not actioned
solaria (powered off ~16 h/day by design) and lustro (nightly display
power-off) generated node_offline/node_stale/node_online alerts on every
daily cycle. Six of them have sat in actions/pending/ since 2026-06-17/18,
unapproved. Because an unapproved pending action suppresses its own dedup
ID indefinitely (recon D14, supervisor.py pending/approved/running check),
those stale alerts also meant a *real* future outage on either node would
generate nothing at all.

Suppression is data-driven from inventory/topology.yaml, not a hardcoded
node-name check:

- topology.yaml: new `duty_cycle` (+ `duty_cycle_reason`) on solaria and
  lustro, mirroring the existing dormant/dormant_reason shape. vps and piha
  deliberately do not carry it — an offline 24/7 node is a real incident.
- supervisor: _load_dormant_nodes() -> _load_node_policy(), loading both
  dormant_nodes and duty_cycle_nodes from one topology read. dormant
  behavior is byte-for-byte unchanged.
- supervisor: one guard in _route_node_event. Duty-cycle liveness events
  are logged at INFO and return; no action is written.

duty_cycle is deliberately NOT dormant. A duty-cycle node stays fully
active: its services are still reconciled (missing_service -> redeploy),
its disk pressure still generates disk_cleanup, and its ha_* events still
route. Only the liveness alert is suppressed. Regression tests pin all
three.

Fail-loud: an unreadable topology leaves both sets empty, which disables
suppression and lets alerts through. A broken topology must never silently
mute the fleet.

Accepted trade-off: a genuine permanent outage of solaria or lustro no
longer alerts. It stays visible in the operator UI (which computes liveness
independently at read time) and in the event feed. An "offline longer than
the expected window" escalation is the natural follow-up and needs a
schedule in the topology field rather than a bare marker.

Tests: 169 passed in services/control-plane/tests (was 157; +12).
Runtime deployment is deliberately NOT part of this commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 13:48:29 +02:00
.claude/skills fix(kb): przepiecie wszystkich odwolan wewnetrznych po migracji 2026-08-04 16:58:46 +02:00
backups/zigbee Add Zigbee coordinator backup 2026-05-14 18:24:26 +02:00
docs/sessions chore(kb): usuniecie docs/questions.md 2026-08-04 16:58:46 +02:00
dotfiles add shared zshrc 2026-05-10 20:52:44 +02:00
hardware/esp/ir-ac-ha-integration feat(kb): przenosiny type=service do kb/services/ (13 plikow, bez SPLIT) 2026-08-04 16:58:04 +02:00
hosts docs(kb-site): przepisz nieaktualne kb.okit.pl na kb-e2a24af3.okit.pl 2026-08-05 13:08:13 +02:00
inventory feat(supervisor): duty-cycle nodes — liveness transitions logged, not actioned 2026-08-05 13:48:29 +02:00
jobs fix(kb): 13 pozostalych odwolan do sciezek sprzed migracji 2026-08-04 17:02:12 +02:00
kb fix(kb-site): wycofaj robots.txt — blokowal legalny fetch z tokenem 2026-08-05 13:11:43 +02:00
packages fix(kb): przepiecie wszystkich odwolan wewnetrznych po migracji 2026-08-04 16:58:46 +02:00
scripts feat(kb-site): noindex + robots.txt + obscure subdomain jako domyslny base-url 2026-08-05 13:08:13 +02:00
services feat(supervisor): duty-cycle nodes — liveness transitions logged, not actioned 2026-08-05 13:48:29 +02:00
.codex Document current homelab state 2026-04-15 17:37:25 +02:00
.gitignore feat(kb): generator publicznej warstwy KB (gen_pages.py) 2026-08-04 17:57:15 +02:00
.mcp.json feat(ha-mcp): read-only MCP server (faza 2a) 2026-07-30 16:47:27 +02:00
CLAUDE.md fix(kb): przepiecie wszystkich odwolan wewnetrznych po migracji 2026-08-04 16:58:46 +02:00
codex_context Add session context state 2026-04-20 22:10:39 +02:00
codex_context.yaml add shared context lock 2026-05-05 17:25:50 +02:00
deploy_agent.py Add deploy escalation output 2026-04-22 22:08:26 +02:00
ollama_client.py Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
README.md fix(kb): przepiecie wszystkich odwolan wewnetrznych po migracji 2026-08-04 16:58:46 +02:00
start-aider.sh Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
start-codex.sh Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
sync-context.sh add shared context lock 2026-05-05 17:25:50 +02:00
update-context.md Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00

Homelab Codex

GitOps-lite orchestration for a distributed homelab environment.

Architecture

The homelab consists of several nodes connected via a Tailscale internal mesh.

Host Role Description
SATURN Primary Node Development, orchestration, and git source of truth (commit node).
SOLARIA Compute Node GPU, inference, and heavy compute workloads.
PIHA Infra Node Core infrastructure services, automation, and monitoring.
VPS Edge Node Public ingress, reverse proxy, and edge services.

Agent System

The homelab uses a multi-agent orchestration model with human-in-the-loop for destructive actions:

Agent Node Role
stability-agent all nodes Per-node watchdog — monitors Docker, disk, Tailscale, MQTT; emits events
node-agent all nodes Publishes container health events to Redis pub/sub
observer VPS Synthesizes world state from events into /opt/homelab/world/*.json
supervisor VPS Detects drift between desired and actual state; writes pending actions
planner-agent SOLARIA LLM-powered diagnosis — listens to Redis, proposes remediation actions
executor VPS Executes actions only after operator approval
operator-ui + telegram-bot VPS / PIHA Operator reviews and approves/rejects pending actions

Action approval flow: pending/ → operator approves → approved/ → executor runs.

Repository Structure

Getting Started

  1. Standardization: Follow the Infrastructure Standards.
  2. Deployment: See Deployment Conventions for how to roll out changes.
  3. SATURN: Remember that SATURN is the only node where commits should be made.

Documentation Index


Note: This repository documents the state of the homelab. Runtime state lives outside the repository in /opt/homelab.