Find a file
oskar ffe2a0c35b docs(incidents): root-cause zniknięcia ollama@SOLARIA — node-agent prune bez filtrów co 60 s
Kontener nie padł ani nie zniknął przy starcie: usunął go własny node-agent.
19 s po operatorskim `docker stop ollama` cykl cleanupu wywołał
`containers.prune()` bez filtrów — Docker kasuje każdy kontener nie-running,
ignorując restart policy i labele compose.

Dowód: ten sam cykl logu node-agent, dwie kolejne linie —
  16:45:33,528 WARNING Container exited: ollama (restart=unless-stopped)
  16:45:33,563 INFO    Pruned stopped containers (57 MB reclaimed)
57 MB to jedyna niezerowa wartość SpaceReclaimed w całym dniu (reszta 0 MB).

Mina uzbroiła się dzień wcześniej: prune istnieje od 01b7758, ale na SOLARII
node-agent nie miał dostępu do docker.sock (GID 996 vs 999) do czasu ddae57c.

Wykluczone: remediation pipeline (dispatch/solaria pusty, whitelist tylko
container_restart), stability-agent (zero ścieżek usuwających), config compose
(AutoRemove=false, brak --rm, brak compose down), cron/systemd (brak prune).

To NIE jest ollama-solaria-start-race — tamte dotyczyły startu kontenera.
Tu problem jest w cleanupie node-agenta i dotyczy każdego serwisu: ai_node
(solaria) i standard (vps) prune'ują bez rate-limitu, sd_card raz na 24 h.
`docker stop` jest obecnie na 4 z 6 nodów operacją destrukcyjną.

Fix świadomie niezaimplementowany — rekomendacje R1-R4 + mitygacja M1 w §7.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 20:19:06 +02:00
.claude/skills docs: session 2026-06-11 — lustro ssh shipping fix + ha-diag-agent piha + backlog/flota-bomba 2026-06-11 14:18:00 +02:00
backups/zigbee Add Zigbee coordinator backup 2026-05-14 18:24:26 +02:00
docs docs(incidents): root-cause zniknięcia ollama@SOLARIA — node-agent prune bez filtrów co 60 s 2026-07-30 20:19:06 +02:00
dotfiles add shared zshrc 2026-05-10 20:52:44 +02:00
hardware/esp/ir-ac-ha-integration feat(gree-ir): wlasny komponent ESPHome dla klimy Gree 2026-07-22 16:46:03 +02:00
hosts test(kb-query): luki T1–T3 z test_fallback.py + status kalibracji ollama-piha (salvage S3+S4) 2026-07-30 16:40:24 +02:00
inventory docs(topology): lustro daily duty cycle — nightly power-off, expected liveness cycles 2026-07-30 15:36:13 +02:00
jobs feat(eval): retrieval_eval --transport http — bramka §9 przez żywe /search kb-query (salvage S1) 2026-07-30 16:40:24 +02:00
packages feat(kb-retrieval,kb-query): add hybrid retrieval mode (faza mailowa Krok 3) 2026-07-23 17:06:49 +02:00
scripts fix(supervisor): route healthcheck_failed to container_restart 2026-07-29 19:24:01 +02:00
services feat(ha-mcp): read-only MCP server (faza 2a) 2026-07-30 16:47:27 +02:00
.codex Document current homelab state 2026-04-15 17:37:25 +02:00
.gitignore feat(hardware): ESPHome Gree IR blaster (Wemos D1 Mini + Grove IR) 2026-07-22 16:46:03 +02:00
.mcp.json feat(ha-mcp): read-only MCP server (faza 2a) 2026-07-30 16:47:27 +02:00
CLAUDE.md fix(supervisor): route healthcheck_failed to container_restart 2026-07-29 19:24:01 +02:00
codex_context Add session context state 2026-04-20 22:10:39 +02:00
codex_context.yaml add shared context lock 2026-05-05 17:25:50 +02:00
deploy_agent.py Add deploy escalation output 2026-04-22 22:08:26 +02:00
ollama_client.py Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
README.md docs: add planner-agent docs and session summary 2026-05-27 2026-05-27 22:35:59 +02:00
start-aider.sh Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
start-codex.sh Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00
sync-context.sh add shared context lock 2026-05-05 17:25:50 +02:00
tech-debt.md docs(tech-debt): cleanup artefaktow paste-error w ~oskar (SOLARIA, niski prio) 2026-06-24 17:13:51 +02:00
update-context.md Initial shared homelab agent workspace 2026-05-03 19:37:40 +02:00

Homelab Codex

GitOps-lite orchestration for a distributed homelab environment.

Architecture

The homelab consists of several nodes connected via a Tailscale internal mesh.

Host Role Description
SATURN Primary Node Development, orchestration, and git source of truth (commit node).
SOLARIA Compute Node GPU, inference, and heavy compute workloads.
PIHA Infra Node Core infrastructure services, automation, and monitoring.
VPS Edge Node Public ingress, reverse proxy, and edge services.

Agent System

The homelab uses a multi-agent orchestration model with human-in-the-loop for destructive actions:

Agent Node Role
stability-agent all nodes Per-node watchdog — monitors Docker, disk, Tailscale, MQTT; emits events
node-agent all nodes Publishes container health events to Redis pub/sub
observer VPS Synthesizes world state from events into /opt/homelab/world/*.json
supervisor VPS Detects drift between desired and actual state; writes pending actions
planner-agent SOLARIA LLM-powered diagnosis — listens to Redis, proposes remediation actions
executor VPS Executes actions only after operator approval
operator-ui + telegram-bot VPS / PIHA Operator reviews and approves/rejects pending actions

Action approval flow: pending/ → operator approves → approved/ → executor runs.

Repository Structure

Getting Started

  1. Standardization: Follow the Infrastructure Standards.
  2. Deployment: See Deployment Conventions for how to roll out changes.
  3. SATURN: Remember that SATURN is the only node where commits should be made.

Documentation Index


Note: This repository documents the state of the homelab. Runtime state lives outside the repository in /opt/homelab.