Find a file
oskar 17ac070b46 fix(control-plane): unique container_restart action_id, no more history overwrite
_generate_recommendation() built container_restart ids as the bare
container-restart-<node>-<service>. Two DIFFERENT incidents for the
same node+service (e.g. a generic containers_not_running restart,
later followed — after recovery and recurrence — by an unrelated
restart for the same service) produced the identical id. Once the
first action reached cancelled/completed/failed, the second action's
own transition into that same directory silently overwrote the first
one's history file. This is exactly what happened 2026-08-26 to a
shadow-mode HA-websocket restart colliding with an unrelated 08-06
entry (docs/sessions/2026-08-26.md) — worked around by hand-renaming
the file that session.

Fix: suffix the id with the triggering incident's started_at —
container-restart-<node>-<service>-<unixts> — NOT time.time() at
generation call time. reconcile() calls _generate_recommendation() on
every loop iteration while the drift persists, and the pending/
approved/running existence check immediately below is what makes that
idempotent; it only works if repeated calls for the SAME ongoing
incident produce the SAME id. started_at is fixed for an incident's
whole life (observer._handle_incident only bumps
last_occurrence/occurrence_count on repeat occurrences — see
COMMIT-1-adjacent code) and changes only when a genuinely new incident
opens for that service, which is exactly "same id while ongoing,
different id on recurrence". Falls back to time.time() if the
incident record is missing/malformed so a restart is still generated.

Scope: only the generic CONTAINER_RESTART_TRIGGERS path
(_generate_recommendation). Left unchanged, deliberately:
  - redeploy-<node>-<service> ids — no observed collision, out of
    scope for this fix (flagged as a latent follow-up below).
  - The HA-specific container-restart-<node>-homeassistant id used by
    _generate_ha_container_restart / _generate_ha_shadow_alert /
    _cancel_ha_container_restart: these three functions rely on an
    exact-match lookup of that fixed id (cooldown check via
    _ha_action_recently_completed, and the cancel path finding the
    specific pending file to move) — adding a suffix there would
    break both without a broader refactor to prefix-glob lookups.
  - alert-ha-*/alert-node-* ids: _ha_action_recently_completed also
    exact-matches these for cooldown dedup; a suffix would defeat
    cooldown entirely (every occurrence would look "new").

node-agent idempotency gate confirmed unaffected: _already_processed()
in node_agent.py does a full-string action_id match against
processed-actions/<id>.done, guarding against RE-processing the exact
same dispatched action file (e.g. a duplicate rsync delivery) — not
against a new action_id for a new occurrence of the same service. A
suffixed id is legitimately a new action to node-agent, which is the
correct behavior (a genuine new incident should actually restart the
container again).

Tests: new test_supervisor_action_id_uniqueness.py covers (1) repeated
_generate_recommendation() calls for the same ongoing incident produce
the same id and do not duplicate the pending file, (2) a new incident
after the old one completed gets a different id and does not overwrite
the old completed record, (3) fallback to time.time() when the
incident record is missing, (4) redeploy ids stay bare. Updated
test_observer_container_events.py's end-to-end assertion to match by
prefix instead of exact filename. Full control-plane suite: 183
passed; node-agent suite: 70 passed (unchanged, confirming the
idempotency gate needed no code change).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017WDKj5LRY8vdQMx57dfNnu
2026-08-26 21:08:04 +02:00
.claude/skills feat(kb): skill i skrypt do pisania dokumentow kb/ (OKF authoring) 2026-08-26 17:05:04 +02:00
backups/zigbee Add Zigbee coordinator backup 2026-05-14 18:24:26 +02:00
docs/sessions docs(kb): incydent 20 dni ciszy kb-mail-sync + session log 2026-08-26 20:39:45 +02:00
dotfiles add shared zshrc 2026-05-10 20:52:44 +02:00
hardware/esp/ir-ac-ha-integration feat(kb): przenosiny type=service do kb/services/ (13 plikow, bez SPLIT) 2026-08-04 16:58:04 +02:00
hosts feat(mail-sync): scheduler PIHA, takt kb-ingest, runbook i dokumentacja 2026-08-06 15:30:57 +02:00
inventory feat(supervisor): duty-cycle nodes — liveness transitions logged, not actioned 2026-08-05 13:48:29 +02:00
jobs feat(mail-sync): scheduler PIHA, takt kb-ingest, runbook i dokumentacja 2026-08-06 15:30:57 +02:00
kb docs(kb): incydent 20 dni ciszy kb-mail-sync + session log 2026-08-26 20:39:45 +02:00
packages feat(mail-imap-sync): job przyrostowki + wpiecie w tor body-ingest 2026-08-06 15:30:57 +02:00
scripts fix(control-plane): unwedge incidents that never get service_healthy 2026-08-26 21:05:38 +02:00
services fix(control-plane): unique container_restart action_id, no more history overwrite 2026-08-26 21:08:04 +02:00
.codex Document current homelab state 2026-04-15 17:37:25 +02:00
.gitignore feat(kb): generator publicznej warstwy KB (gen_pages.py) 2026-08-04 17:57:15 +02:00
.mcp.json feat(ha-mcp): read-only MCP server (faza 2a) 2026-07-30 16:47:27 +02:00
CLAUDE.md fix(kb): przepiecie wszystkich odwolan wewnetrznych po migracji 2026-08-04 16:58:46 +02:00
codex_context Add session context state 2026-04-20 22:10:39 +02:00
codex_context.yaml add shared context lock 2026-05-05 17:25:50 +02:00
deploy_agent.py Add deploy escalation output 2026-04-22 22:08:26 +02:00
ollama_client.py Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
README.md fix(kb): przepiecie wszystkich odwolan wewnetrznych po migracji 2026-08-04 16:58:46 +02:00
start-aider.sh Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
start-codex.sh Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
sync-context.sh add shared context lock 2026-05-05 17:25:50 +02:00
update-context.md Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00

Homelab Codex

GitOps-lite orchestration for a distributed homelab environment.

Architecture

The homelab consists of several nodes connected via a Tailscale internal mesh.

Host Role Description
SATURN Primary Node Development, orchestration, and git source of truth (commit node).
SOLARIA Compute Node GPU, inference, and heavy compute workloads.
PIHA Infra Node Core infrastructure services, automation, and monitoring.
VPS Edge Node Public ingress, reverse proxy, and edge services.

Agent System

The homelab uses a multi-agent orchestration model with human-in-the-loop for destructive actions:

Agent Node Role
stability-agent all nodes Per-node watchdog — monitors Docker, disk, Tailscale, MQTT; emits events
node-agent all nodes Publishes container health events to Redis pub/sub
observer VPS Synthesizes world state from events into /opt/homelab/world/*.json
supervisor VPS Detects drift between desired and actual state; writes pending actions
planner-agent SOLARIA LLM-powered diagnosis — listens to Redis, proposes remediation actions
executor VPS Executes actions only after operator approval
operator-ui + telegram-bot VPS / PIHA Operator reviews and approves/rejects pending actions

Action approval flow: pending/ → operator approves → approved/ → executor runs.

Repository Structure

Getting Started

  1. Standardization: Follow the Infrastructure Standards.
  2. Deployment: See Deployment Conventions for how to roll out changes.
  3. SATURN: Remember that SATURN is the only node where commits should be made.

Documentation Index


Note: This repository documents the state of the homelab. Runtime state lives outside the repository in /opt/homelab.