homelab-codex-ws/docs/architecture/RECON-multiagent-2026-07-27.md
oskar 615b6ab2a8 docs(architecture): Etap 0 changes appendix in recon; telegram-frontend pattern in legacy doc
RECON appendix summarizes the whole etap-0 cleanup (topology dormant status,
hosts reconciliation, dead-code removal, legacy decision), carries the
runbook line for the stale chelsty pending action
(rm /opt/homelab/actions/pending/redeploy-chelsty-infra-ha-diag-agent.json —
verified the only chelsty-targeted pending file 2026-07-30), and lists
discrepancies found but out of scope: lustro watchtower, stale solaria/chelsty
world-state keys, piha naming-drift duplicate keys, gokapi still down with a
stuck pending redeploy, outline/joplin/umami still lacking services/ dirs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 15:36:29 +02:00

36 KiB
Raw Blame History

Architecture Recon — multi-agent systems (2026-07-27)

Read-only recon of the two multi-agent systems (control-plane, ai-cluster) across vps, piha, solaria, chelsty-infra. Runtime data collected 2026-07-27 from solaria (local), piha and vps (ssh). chelsty-infra was unreachable — Tailscale reports it offline, last seen 55 days ago, and the control-plane world state agrees (last_seen 2026-06-01, liveness dead). Everything about chelsty below is repo-only, marked as such. ai-cluster migration read from worktree task/ai-cluster-solaria @ b124e54.


A. Authority map

A1. Containers mounting /var/run/docker.sock (observed via docker inspect)

Node Container Mode
vps control-plane-executor rw
vps node-agent rw
vps ai-cluster-service-ops-worker-1 rw
vps stability-agent ro
piha node-agent rw
piha portainer rw
piha homepage rw
piha stability-agent ro
solaria node-agent rw
solaria stability-agent ro
chelsty-infra unverifiable (node offline); per repo compose: node-agent rw, stability-agent ro

After the ai-cluster migration lands, solaria additionally gets service-ops-worker with rw sock (services/ai-cluster/docker-compose.yml:133 in the worktree — no :ro).

A2. What each sock-holder is actually allowed to do

  • stability-agent (all nodes, ro): monitoring only — reads container state, emits filesystem events (containers_not_running, mqtt_unreachable as a TCP probe, disk/health). No restart code path exists.
  • node-agent (all nodes, rw): reads Docker state and pushes per-event JSON files to the VPS over ssh/rsync (VPS_EVENTS_HOST=100.95.58.48). It is also the on-node hands of the executor: it pulls its dispatch inbox (actions/dispatch/<node>/, rsync over the same ssh channel) and executes actions with a type whitelist of exactly {container_restart} (node_agent.py:105) plus a self-restart guard (SELF_RESTART_GUARD_NAMES = {"node-agent"}, :111). There is no allowlist of which containers it may restart — any name in an approved-and-dispatched action goes to container.restart() via its own sock. On solaria the Docker half is broken: the container starts with Docker unavailable: Permission denied — the base compose sets group_add: ["999"], piha overrides it to the host's docker gid (123), lustro to 991, but solaria has no override and its docker gid is 996. Solaria's node-agent therefore reports node health only and could not execute a dispatched restart.
  • control-plane-executor (vps, rw): executes only actions an operator moved to approved (HITL). Per action type: container_restart → writes a dispatch file for the target node-agent (it never touches a remote sock itself); disk_cleanup → real ssh to the node with a command-safety gate; alert_only → no-op success; redeploy → runs scripts/deploy/deploy-node.sh <node> <svc> inside the executor container — which cannot succeed as wired: the script ignores both arguments and deploys hostname's full service set from ${HOME}/homelab-codex-ws, a path that does not exist in the container → exit 1 → failed/. Redeploy is effectively a dead action type today (executor.py:106-119, deploy-node.sh has no $1/$2).
  • ai-cluster service-ops-worker (vps today, solaria after migration, rw): keyword-driven, no LLM, no human approval. Modes: diagnose (read-only docker ps/inspect/logs), safe_fixdocker restart <container>, deploy and repair are preview-only no-ops. The SERVICE_NAMES list in service_ops_worker.py:31-40 is a dead constant — never referenced. The only real gate is a command-shape allowlist (is_allowed_command, :113-137): docker restart <any-token> is permitted against any container on its host, resolved by exact-then-substring match over live docker ps output. The repo itself flags this (services/ai-cluster/service.yaml:62-66).
  • portainer (piha, rw): full unrestricted Docker management UI — outside both agent systems, no allowlist, no audit trail in the event store.
  • homepage (piha, rw): dashboard using the sock for container discovery; the mount is rw, so restriction is by application behavior only, not by mount.

A3. Restart-authority overlap on solaria after the ai-cluster move

The two paths have very different shapes:

  • service-ops-worker can restart anything docker ps shows on solaria: ollama, node-agent, planner-agent, stability-agent, node_exporter, and the whole ai-cluster stack including its own broker and itself — substring matching makes mosquitto or redis easy accidental targets. No approval, no coordination.
  • supervisor→executor→node-agent generates container_restart only for services in hosts/solaria/services.yaml desired state. Post-migration that is node-agent (refused by the node-agent self-guard), ollama (exact container name — executable), and ai-cluster — which the branch keys as a single service, a name matching no actual container (ai-cluster-openclaw-1, …), so a dispatched restart for it would fail to resolve. Whether the observer would even raise per-container incidents for the ai-cluster stack under that single key cannot be determined without deploying. Execution additionally requires solaria's node-agent to have working sock access, which it currently does not (A2).

So the clean overlap set is exactly ollama — the GPU workload — restartable by an unapproved keyword-matched MQTT task on one side and an approved HITL action on the other, with no shared lock, cooldown, or mutual awareness (supervisor cooldowns dedup only its own action IDs). The asymmetric remainder is arguably worse: everything else on the node is restartable only by the autonomous path, invisible to HITL.


B. Service inventory — repo vs reality

B4. services/ directories and owner_node

23 dirs; 4 lack service.yaml entirely (agent-system, control-plane, home-assistant, node-agent); node_exporter uses a different (flat) schema.

service owner_node service owner_node
brain-watchdog piha node_exporter per-host
fleet-prometheus vps npm vps
forgejo piha ollama solaria
gokapi vps paperless piha
ha-diag-agent per-host paperless-worker solaria
kb-postgres piha planner-agent solaria
kb-query piha stability-agent chelsty (not a real node name)
llm-gateway piha vikunja piha
mosquitto vps (README says piha) zigbee2mqtt piha
nextcloud piha agent-system / control-plane / home-assistant / node-agent no service.yaml

B5. Running containers (2026-07-27)

  • vps (24): control-plane ×4 (executor, supervisor, observer, operator-ui), node-agent, stability-agent, node_exporter, fleet-prometheus, ai-cluster ×6 (openclaw, codex-worker, planner-worker, service-ops-worker, redis, mosquitto), outline ×3, joplin ×2, npm, umami ×2, humanai-mailer + humanai-landing (no compose labels at all — not compose-managed).
  • piha (44): desired set ha-diag-agent, node-agent, brain-watchdog, vikunja(+db), llm-gateway, node_exporter, kb-postgres, kb-query — all running — plus ~28 containers with no hosts-manifest entry: agent-system ×4 (webui, runtime-materializer, telegram-bot, redis), stability-agent, ollama-piha, paperless ×3, immich ×4, forgejo ×3, zigbee2mqtt, audiobookshelf, grafana, wikijs ×2, code-server, portainer, prometheus, homepage, actual-budget, mqtt-exporter, owntracks ×3, vaultwarden, pihole-exporter (unlabeled), fail2ban-exporter, nginxproxymanager.
  • solaria (5): ollama, node-agent, planner-agent, stability-agent, node_exporter.
  • chelsty-infra: unverifiable (offline).

B6. Two-way diff

Manifest-but-not-running:

  • gokapi (hosts/vps/services.yaml) — not running on vps; a redeploy gokapi action sits in pending, unapproved.
  • Entire hosts/chelsty-infra/services.yaml set (ha-diag-agent, node-agent, mosquitto, zigbee2mqtt, frigate) — node offline 8 weeks, unverifiable.

Running-but-no-manifest (in the authoritative hosts/*/services.yaml):

  • vps: stability-agent, ai-cluster ×6, outline ×3, joplin ×2, npm, umami ×2, humanai-mailer, humanai-landing. (topology.yaml lists some of these for vps, but topology itself declares hosts/vps/services.yaml authoritative — see F20.)
  • piha: the ~28 containers above, including services that do have repo dirs (paperless, forgejo, zigbee2mqtt, agent-system, stability-agent) but no hosts/piha entry.
  • solaria: planner-agent, stability-agent, node_exporter (running; absent from hosts/solaria/services.yaml, which lists only node-agent + ollama).

B7. The shadow set — where stability-agent and friends come from

The premise "no dir under services/" is wrong for stability-agent: it lives at services/stability-agent/ (source, Dockerfile, compose, service.yaml). It is invisible to the normal pipeline because it deploys through its own scripts/deploy/deploy-stability-agent.sh (hardcoded Tailscale IPs) → per-node deploy-local.shdocker compose up -d --build --force-recreate, and scripts/deploy/deploy-node.sh:93-94 explicitly skips any service that has its own deploy-local.sh. It appears in no hosts/*/services.yaml; its service.yaml says owner_node: chelsty (a node that doesn't exist). Image is built on each node from the repo checkout. Role: per-node read-only watchdog emitting events.

Other members of the shadow set (deployed outside the declarative pipeline):

Thing Deploy path
agent-system (piha: webui, materializer, telegram-bot, redis) services/agent-system/deploy.sh (own docker compose up) — no service.yaml, no hosts entry
control-plane (vps) services/control-plane/deploy-local.sh + scripts/bootstrap/vps-control-plane.sh
frigate (chelsty) scripts/deploy/deploy-frigate.sh composing from hosts/chelsty-infra/runtime/frigate/ — full compose inside hosts/, not services/
kb-ingest host systemd unit (jobs/documents-ingest/systemd/), non-container
Home Assistant config scripts/ha/deploy.sh API push, outside compose entirely
humanai-mailer / humanai-landing (vps), pihole-exporter (piha) no compose labels — hand-run, no repo trace found
piha host mosquitto systemd OS package, no repo definition at all (see C8)
scripts/deploy/deploy-role.sh references roles/ which does not exist — dead script

C. MQTT topology

C8. Broker instances

Node Instance Bind Auth
piha host systemd mosquitto (OS package, not a container, not in repo) 0.0.0.0:1883 + [::]:1883 allow_anonymous true — open to LAN and Tailscale
vps mosquitto container (ai-cluster stack, pre-GitOps copy) 100.95.58.48:1883 (Tailscale IP only) password + ACL, user codex, topics codex/tasks+codex/results only
chelsty-infra repo: hosts/chelsty-infra/runtime/mosquitto/ listener 1883, all interfaces repo says allow_anonymous false + password file, but scripts/bootstrap/chelsty-runtime.sh:75-78 creates the password file empty; runtime unverifiable (offline)
solaria (future) worktree ai-cluster stack 100.100.231.104:1883 (Tailscale IP only) password + ACL, user codex

The repo's services/mosquitto/ manifest (owner_node vps, README "deployed on piha") matches nothing that actually runs — vps's broker belongs to the ai-cluster stack, piha's is a host package.

C9. Real clients per broker (24 h of logs)

  • piha (systemd broker): persistent, currently-established clients: zigbee2mqtt (container 172.21.0.2), mqtt-exporter (172.19.0.2), owntracks-recorder (172.23.0.3), Home Assistant at 192.168.31.7, ~11 LAN IoT devices (192.168.31.100.115), piha itself (.5), localhost. Churn: the lustro / pimirror Pi at 192.168.31.31 (MAC b8:27:eb = Raspberry Pi) made ~600 connections in 24 h, each with a fresh random UUID client-id (plus one pimirror-pimirror2); a second Pi at .19 connected twice.
  • vps (ai-cluster broker): zero new connections in the last 30 days. Last log entries are 2026-06-09: openclaw, codex-worker, vps-service-ops-1 (and planner) connecting as user codex. The four workers hold long-lived connections; the bus has been otherwise silent for ~7 weeks.
  • chelsty-infra: unverifiable (offline). Repo-expected clients: zigbee2mqtt, frigate (via host network, 127.0.0.1:1883), HA on chelsty-ha, mosquitto healthcheck ($SYS/broker/version).

C10. Topic map (from code, both checkouts)

Component Publishes Subscribes Broker
openclaw (FastAPI gateway) codex/tasks codex/results ai-cluster mosquitto (mosquitto:1883 in-stack)
codex-worker codex/results codex/tasks same
planner-worker codex/results codex/tasks same
service-ops-worker codex/results codex/tasks same
zigbee2mqtt (chelsty, repo) zigbee2mqtt/# same base chelsty broker
frigate (chelsty, repo) frigate topics chelsty broker (127.0.0.1)
zigbee2mqtt, owntracks, HA, IoT (piha, observed) device topics piha host broker
stability-agent no MQTT client; MQTT_HOST is only a TCP reachability probe target

Worker routing is by payload field target (role:dev / role:planner / role:service-ops / exact AGENT_ID), not by topic. The ACL permits exactly the two codex/* topics.

C11. Shared bus?

No. Control-plane and ai-cluster are entirely separate buses — control-plane does not use MQTT at all. No control-plane or node-agent component has an MQTT client; their bus is the filesystem event/action path over ssh/rsync. The only "MQTT" in control-plane is the event type name mqtt_unreachable produced by stability-agent's TCP probe. There are three unconnected broker islands (piha home-automation, vps→solaria ai-cluster, chelsty home-automation) and no bridge config anywhere.


D. Event pipeline economics

D12. Volume on disk

Node Store Files Size Oldest → newest
vps /opt/homelab/events/ total 36,851 56 MB data / 181 MB on disk 2026-05-17 → now
vps events/piha/ (per-event JSON) 29,437 131 MB on disk rolling
vps events/vps/ 4,523 33 MB rolling
vps events/lustro/ 2,673 13 MB rolling
vps events/solaria/ 207 2.3 MB rolling
piha local events/ (date-keyed jsonl) 20 29 MB 2026-05-17 → today
solaria local events/ 14 1 MB 2026-05-17 → 07-23

Two distinct stores coexist: the date-keyed append-only events/YYYY-MM-DD/<node>/events.jsonl (the documented design) and a newer node-keyed store of individual evt-*.json files on vps (events/<node>/evt-<node>-<ts>-<type>-<svc>.json), fed by node-agents pushing over ssh/rsync. Node-agents spool locally then delete after push: the empty local events/<node>/ directories have bloated directory inodes (26 MB on piha, 1.3 MB on solaria) — evidence of tens of thousands of files created and deleted.

Pruning: three code paths exist, none covers the bulk. No cron or systemd timer touches the events tree. (1) node_agent._cleanup_control_plane_fs (node_agent.py:673-771) runs every 60 s only on vps and deletes an event file only if its type is in {service_healthy, node_health} AND it is behind the observer checkpoint AND older than 3 days; it also prunes completed/failed actions

7 d and deploy logs >30 d. (2) scripts/maintenance/cleanup_event_backlog.py — manual, dry-run by default, same type filter. (3) rsync --remove-source-files clears each node's local spool after shipping (source-side only). Everything else — healthcheck_failed, containers_not_running, all ha_*, liveness, pressure — is retained indefinitely; that is the direct cause of the 29k-file piha backlog.

D13. Does the observer consume-and-delete?

Neither. The observer re-globs the entire tree recursively every 5 s (observer.py:894,948) and tracks position via a per-node-directory timestamp checkpoint persisted to /opt/homelab/state/observer_checkpoint.json (observer.py:225,356-360): a file counts as new iff its filename timestamp exceeds the checkpoint for parts[0] of its relative path. No per-event markers, no deletion, no archiving; unparseable files are quarantined to state/observer_failed_events/ (observer.py:250-268). Cost scales with total file count regardless of how many are new. Quirk: stability-agent writes to date-keyed dirs (events/YYYY-MM-DD/<node>/), so its checkpoint bucket is the date string, not the node — a new bucket every day.

D14. Why ~25k noisy events but only 2 action_results

Verified counts in the vps per-event store for piha: 19,300 ha_entity_unavailable_long, 3,376 healthcheck_failed, 2,490 containers_not_running, and exactly 2 action_result events (both node-exporter, 2026-07-23). Action pool on vps: 18 pending, 17 cancelled, 0 approved, 0 completed, 0 failed, 0 running.

The funnel, in order:

  1. Only dispatched container_restarts can ever produce an action_result (node_agent.py:993-1010; the only dispatchable type, node_agent.py:105). redeploy, disk_cleanup, alert_only never emit one. So 2 action_results = 2 container restarts ever executed, fleet-wide.
  2. Desired-state scoping. Drift actions are generated only for services listed in hosts/<node>/services.yaml (supervisor.py:309). A failing container not in the manifest (npm, paperless, mosquitto, immich, … — the majority on piha, B6) creates an incident in world state but no action, ever.
  3. Dedup-by-ID with no expiry. The action ID is deterministic (container-restart-<node>-<svc>, redeploy-<node>-<svc>, alert-ha-<suffix>-<node>, alert-<type>-<node>, disk-cleanup-<node>) and generation is skipped while that ID exists in pending/approved/running (supervisor.py:377-380 et al.). An unapproved pending action suppresses its ID forever. 19,300 ha_entity_unavailable_long events collapse onto the single ID alert-ha-entity-unavailable-piha; cooldowns (1 h for HA/node alerts, 30 min for websocket-restart) only apply after an action reaches completed/rejected/cancelled.
  4. Nothing is being approved. Executor only reads approved/ (executor.py:75-76); pending→approved happens only via operator-ui/webui/ Telegram. 18 actions sit pending (including redeploy gokapi for a genuinely down service and a redeploy targeting a node dead 8 weeks); 17 were auto-cancelled by the supervisor when services recovered.
  5. Even approval often couldn't act: redeploy is broken as wired (A2), and alert_only approval is a no-op success by design (executor.py:142-144).

D15. Supervisor event_type → action_type mapping (from code — differs from CLAUDE.md)

Path 1 — world-state drift (supervisor reads world/{services,incidents}.json, not raw events; supervisor.py:275-340):

Incident trigger_type Action Notes
containers_not_running container_restart risk low, conf 0.95
healthcheck_failed redeploy not alert-only, contrary to intuition — but redeploy is a dead end (A2)
service_unhealthy redeploy
deployment_failed redeploy
service missing from world state redeploy (missing_service)
node disk_pressure == high disk_cleanup skipped for chelsty-* (supervisor.py:45)
mqtt_unreachable dead branch — listed in CONTAINER_RESTART_TRIGGERS (supervisor.py:39) but the observer never creates incidents with that trigger_type

Path 2 — direct event-file routing (_process_ha_events, supervisor.py:557-578; globs all nodes every cycle, then if not etype.startswith("ha_"): continue apart from the node-alert set):

Event type Action Cooldown
ha_websocket_dead nominally container_restart (homeassistant) — in practice alert_only: HA_DIAG_SHADOW_MODE defaults "true" and nothing overrides it (supervisor.py:101,592-598) 1800 s
ha_websocket_recovered cancels the pending restart
6 ha_* alert types alert_only 3600 s
node_offline / node_stale / node_online alert_only 3600 s

Ignored entirely: service_healthy, node_health, high_cpu, high_memory, container_restarting, container_state_unexpected, action_result, deployment_started/completed, medium disk_pressure, mqtt_unreachable, and all stability-agent-only types (disk_usage_high, docker_api_error, agent_error). Additionally, stability-agent's own containers_not_running events carry service=None, which the observer skips when building incidents (observer.py:729) — stability-agent's flagship signal never opens an incident; only node-agent's equivalent does.

Dedup key: the deterministic action-ID filename (see D14). monitor: false (single occurrence, homeassistant @ chelsty-ha) is read in exactly one place — supervisor.py:203-207 — and suppresses only supervisor action generation; the observer still ingests the events, incidents still form, UIs still display them. HA events are also dropped for 300 s after a containers_not_running incident on homeassistant (supervisor.py:611-627), and each event file is routed at most once per supervisor process lifetime (in-memory set, lost on restart).


E. Solaria flapping

E16. Who emits node_offline / node_stale / node_online

The observer (control-plane, vps). Event payloads carry "source": "observer" and messages like Node solaria liveness fresh -> stale (last_seen 184s ago). Node-agent only sends heartbeats (node_health).

E17. Thresholds vs report interval

From code (services/control-plane/src/liveness.py:36-50): default nodes fresh ≤ 180 s, stale ≤ 600 s, dead > 600 s; remote nodes (chelsty-*) get fresh ≤ 900 s / dead > 3600 s. Confirmed at runtime by event payloads: stale fired at age 184 s, offline at 604 s. Node-agent interval: CHECK_INTERVAL=60 in every host override; consecutive node_health events land 63 s apart (60 s sleep + ~3 s work/push) — the observed ~63 s. So stale = ~3 missed heartbeats, dead = ~10. last_seen is refreshed by any ingested event from the node, not just heartbeats (observer.py:692).

E18. The 11 "flaps" — explanation

The premise "11 cycles in ~2.7 h" does not match the store: solaria has exactly 11 stale + 11 offline + 11 online events spread over Jul 11 → Jul 23, one cycle per day, plus today's node_online. Pattern (UTC): node_stale between 19:11 and 22:36 each evening, node_offline 7 minutes later, node_online the next day between 09:12 and 17:11 (gap Jul 2426, back online Jul 27 14:12).

Correlation: solaria's current boot is 2026-07-27 16:11 local — one minute before today's node_online. node-agent has RestartCount=0 and started at boot; its logs contain no ssh/rsync errors (the only error line is the Docker permission issue). Nothing in tailscaled logs beyond normal boot-time link setup.

Conclusion: solaria is not flapping — it is powered off every evening and powered on around midday. Each power-off produces exactly one stale→offline pair; each boot produces one online. piha, which runs 24/7, produced 4 such events total in the same store (roughly one real incident). No network, node-agent, or observer defect is involved.

E19. Is solaria's reporting path different?

Transport, interval, event format and shipping target are identical to piha (ssh/rsync push to 100.95.58.48, 60 s, VPS addressed by IP so solaria's disabled MagicDNS doesn't matter; same non-remote liveness tier). Three real differences: (1) the node-agent cannot read the Docker socket on solaria (base group_add: ["999"] vs host docker gid 996; piha/lustro override, solaria doesn't), so solaria emits no container-level events — only node_health; (2) scripts/deploy/orchestrate-deploy.sh:22 excludes solaria (skip set {saturn, solaria}) from fleet deploys, so its node-agent code can lag the repo; (3) the host's ~16 h/day power-off duty cycle, which is the entire flapping story.


F. Topology truth

F20. topology.yaml vs hosts/*/services.yaml vs reality — all disagreements

  1. inventory/topology.yaml:67 declares hosts/vps/services.yaml authoritative, then itself lists 10 services for vps where the hosts file lists 5 (delta: stability-agent, npm, outline, joplin, ai-cluster).
  2. outline, joplin, ai-cluster appear in topology only — no services/ dir on master, no hosts entry — yet all three run on vps. (services/outline, joplin, npm cutover is documented in CLAUDE.md but the running containers predate GitOps management; ai-cluster's repo definition exists only in the unmerged worktree.)
  3. gokapi: desired in hosts/vps, not running; the remediation action is stuck in pending.
  4. hosts/saturn/ has no services.yaml at all.
  5. stability-agent: service.yaml says owner_node: chelsty — no such node; runs on piha, vps, solaria (and presumably chelsty-infra); appears in no hosts file; topology lists it under vps only.
  6. mosquitto: service.yaml owner vps, README says piha, hosts list it only under chelsty-infra; reality: piha runs a host systemd broker (no repo definition), vps runs the ai-cluster container broker.
  7. zigbee2mqtt: owner_node piha, hosts list it only under chelsty-infra; reality: a zigbee2mqtt container runs on piha (and repo config also targets chelsty) — two instances, one manifest.
  8. solaria: hosts file lists node-agent + ollama; running adds planner-agent, stability-agent, node_exporter; topology lists node-agent only.
  9. piha: topology omits kb-query (hosts has it); ~28 running containers appear in neither (immich, forgejo, paperless, grafana, portainer, agent-system, …).
  10. topology says deployment.mode: pull, orchestrator: saturn; every deploy script SSH-pushes from saturn.
  11. World state tracks 6 nodes (vps, piha, solaria, lustro, chelsty-infra — chelsty-ha absent) while the recon scope assumed 4; lustro (MagicMirror Pi) is a full monitored node with node-agent and 1,951 service_healthy events for piper-tts.
  12. chelsty-infra: repo treats it as active (services, pending redeploy ha-diag-agent action); control plane has considered it dead since 2026-06-01 (~8 weeks).
  13. frigate is a full compose stack inside hosts/chelsty-infra/runtime/ — outside the services/ convention; services/home-assistant/ has no compose at all (config-push model); agent-system, control-plane, node-agent lack service.yaml.
  14. Schema drift: 21 service.yaml files nest under service:, node_exporter's is flat.

F21. Every monitor: false in the repo

Exactly one: hosts/chelsty-ha/services.yaml:12 for homeassistant, with a comment-only rationale (no reason: field): chelsty-ha has no node-agent, so there are no container events to track; HA is monitored indirectly via the chelsty-infra MQTT broker; re-enable once node-agent is bootstrapped there. Sole code reader: supervisor.py:203-207 (action generation only); a side effect at supervisor.py:529-531 auto-cancels existing pending actions for the service.

(Related but distinct: hosts/chelsty-infra/services.yaml:23-25 comments that node-agent there "monitors and emits events but does NO Docker cleanup".)


OPEN ARCHITECTURAL QUESTIONS

Decisions a human must make; no recommendations attached.

  1. Dual restart authority on solaria. After the ai-cluster migration the autonomous, keyword-triggered service-ops-worker can restart every container on solaria (its service allowlist is dead code), while the HITL path can restart the subset in desired state — cleanly overlapping on ollama, the GPU workload — with no shared lock, cooldown, or mutual awareness. Decide: does service-ops-worker keep an rw docker.sock on solaria; is a real target allowlist enforced; which system owns restarts for which services; and should most of the node remain restartable only by the non-HITL path. At stake: uncoordinated double-remediation, HITL bypass, GPU workloads killable by one MQTT message holding the codex credentials.
  2. Solaria's duty cycle vs its new role. Solaria is powered off ~16 h/day, and the migration makes it home of the ai-cluster bus (mosquitto, redis, openclaw, all workers) plus the GPU stack. Decide: is solaria a 24/7 service node or an opportunistic workstation, and either way, whether liveness monitoring and the ai-cluster control bus should live on a host with a nightly power-off. At stake: ai-cluster availability drops to solaria's uptime window; daily stale/offline/online alert noise is baked in.
  3. Event store lifecycle. Nothing prunes either event store; 37k files / 181 MB on vps and growing, plus directory-inode bloat from the spool-and-delete pattern on every node. Decide: retention policy, which of the two store formats (date-keyed jsonl vs per-event JSON) is canonical, and who owns pruning. At stake: unbounded disk growth on a 4 GB VPS, observer scan cost, inode exhaustion.
  4. The approval pipeline is not being operated — and partially cannot work. 18 pending actions (0 approved, 0 executed beyond the 2 node-exporter restarts), including a real outage (gokapi down) whose remediation type (redeploy) is broken as wired regardless of approval; a pending action suppresses its ID forever, so each incident class alerts at most once until someone drains the queue. Decide: pending-action TTL/expiry, escalation policy, whether alert_only belongs in the same queue as executable remediations, and whether redeploy should be fixed or removed. At stake: real incidents buried in a queue nobody drains; HITL becomes "no-op in the loop".
  5. GitOps boundary. The declarative desired state covers a minority of what runs: ~28 unmanaged containers on piha, a host-level anonymous MQTT broker with no repo definition, unlabeled hand-run containers on vps (humanai-*), and a shadow-deploy family (stability-agent, agent-system, frigate, kb-ingest, HA config push). Decide: what is in scope for drift detection — and either bring the rest in or declare it out. At stake: supervisor drift detection is structurally incomplete; the repo's authority claim is false today.
  6. chelsty-infra's status. Offline to the control plane since 2026-06-01, yet still a deploy target, still in manifests, with a pending remediation action. Decide: reclassify (dormant/seasonal site?) or repair connectivity, and how the supervisor should treat nodes offline for weeks. At stake: dead-node actions accumulating; monitoring credibility.
  7. One bus, three islands, or four. Messaging today: piha's anonymous home-automation broker, ai-cluster's authed codex/* bus, chelsty's local broker, plus control-plane's ssh/rsync file bus. Decide: is this separation intentional architecture (offline-first sites, isolated blast radii) or accident — and specifically whether the piha broker stays anonymous on 0.0.0.0 while HA, IoT devices, owntracks and exporters depend on it. At stake: auth model, single-points-of-failure placement, and whether "mqtt_unreachable" means anything consistent across nodes.
  8. Monitoring blind spots as policy. Solaria's node-agent has been blind to Docker since deployment (group_add mismatch); healthcheck_failed routes only to the dead redeploy path; stability-agent's containers_not_running (service=None) never opens an incident, making its flagship signal inert; ha_websocket_dead remediation has silently stayed in shadow mode; chelsty-ha is monitor:false pending a bootstrap that hasn't happened; unmanifested services get incidents but can never get actions. Decide which blind spots are accepted policy and which are defects, and whether "event emitted but unroutable" should be visible anywhere. At stake: the difference between "monitored" and "logged".

Etap 0 changes (2026-07-28)

Repo-only truth cleanup driven by this recon (commits 2026-07-28 → 2026-07-30, branch task/porzadki-topologia plus earlier etap-0 commits already on master). No runtime state was touched.

Topology / nodes

  • inventory/topology.yaml: chelsty-infra and chelsty-ha marked status: dormant (site hardware down since ~2026-06-01, pending physical revival — F20.12 / open question 6); per-node service enumerations removed — hosts/<node>/services.yaml is authoritative (F20.1-2); deployment.mode corrected to push (F20.10); lustro added as a full monitored node with its duty cycle documented (nightly ~23:30 power-off, one liveness cycle per day — same pattern as solaria, F20.11).
  • Supervisor reads status: dormant from inventory/topology.yaml (_load_dormant_nodes) and skips desired-state loading, disk-cleanup and node-alert generation for dormant nodes — replacing the hardcoded chelsty-name checks (topology is readable from the supervisor's runtime context, so no env-var fallback was needed).

hosts/ reconciliation (B6/F20)

  • hosts/solaria/services.yaml: added stability-agent, node_exporter, and planner-agent (monitor: false, legacy ai-cluster family, retirement candidate).
  • hosts/vps/services.yaml: added stability-agent, npm, outline, joplin-server, umami (all verified healthy in world state 2026-07-30); humanai-mailer / humanai-landing added as unmanaged: true + monitor: false (hand-run, no compose labels); ai-cluster deliberately gets NO entry — comment block points at ai-cluster-LEGACY.md.
  • hosts/piha/services.yaml: single comment block enumerating the ~28 known unmanaged containers (B5) — bringing them in is a later stage (open question 5).
  • hosts/saturn/services.yaml: created with an explicit empty service list (dev workstation / orchestrator, nothing monitored — F20.4).
  • hosts/lustro/services.yaml: added node-exporter and piper-tts (verified running); watchtower noted as deliberately unmanaged.
  • services/stability-agent/service.yaml: owner_node: chelstyper-host (B7/F20.5).
  • services/mosquitto/service.yaml: marked NOT DEPLOYED / legacy manifest — matches nothing that runs (C8/F20.6); kept for reference pending open question 7.

Dead code

  • scripts/deploy/deploy-role.sh deleted (referenced nonexistent roles/, B7).
  • Supervisor: mqtt_unreachable removed from CONTAINER_RESTART_TRIGGERS (dead branch — the observer never creates incidents with that trigger_type, D15).

Legacy

  • docs/architecture/ai-cluster-LEGACY.md: ai-cluster is retired in place, not migrated (bus idle since 2026-06-09, C9); branch task/ai-cluster-solaria stays unmerged as documentation; surviving patterns listed; runtime retirement runbook (stop stack on vps, observe free -m, remove containers) to be executed in a separate supervised session.

Runbook — stale chelsty action (from 1c, do in a runtime session)

  • On vps: rm /opt/homelab/actions/pending/redeploy-chelsty-infra-ha-diag-agent.json — the only chelsty-targeted pending action as of 2026-07-30 (verified via grep -l chelsty /opt/homelab/actions/pending/*.json).

Discrepancies found during this pass, NOT covered by etap 0

  • lustro also runs watchtower (auto-updater) — undocumented anywhere; whether unattended container updates on a monitored node are policy needs a decision.
  • World state carries stale solaria keys (solaria/executor, solaria/operator-ui, solaria/kb-postgres, solaria/paperless-worker) for containers that do not run there (B5 counted 5) — world state never prunes departed services; same mechanism keeps all chelsty-infra/* keys alive.
  • piha world state holds duplicate keys from naming drift (piha/node-exporter vs piha/node_exporter, immich-server vs immich_server, generic app/db/database/broker keys) — event-source naming is not normalized.
  • gokapi remains desired-but-down with its redeploy action stuck in pending (F20.3); the wider undrained queue (18 pending actions) is open question 4 — untouched here.
  • outline, joplin, umami run on vps with no services/<name> dir on master — their hosts entries added in etap 0 document this; actual GitOps cutover (compose in repo) is still pending (F20.2).