The group_add fix itself already landed on master (ddae57c) but the running
container still has GroupAdd=[999] — deploy pending from main checkout.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
37 KiB
Architecture Recon — multi-agent systems (2026-07-27)
Read-only recon of the two multi-agent systems (control-plane, ai-cluster) across vps,
piha, solaria, chelsty-infra. Runtime data collected 2026-07-27 from solaria (local),
piha and vps (ssh). chelsty-infra was unreachable — Tailscale reports it offline,
last seen 55 days ago, and the control-plane world state agrees (last_seen
2026-06-01, liveness dead). Everything about chelsty below is repo-only, marked as
such. ai-cluster migration read from worktree task/ai-cluster-solaria @ b124e54.
A. Authority map
A1. Containers mounting /var/run/docker.sock (observed via docker inspect)
| Node | Container | Mode |
|---|---|---|
| vps | control-plane-executor | rw |
| vps | node-agent | rw |
| vps | ai-cluster-service-ops-worker-1 | rw |
| vps | stability-agent | ro |
| piha | node-agent | rw |
| piha | portainer | rw |
| piha | homepage | rw |
| piha | stability-agent | ro |
| solaria | node-agent | rw |
| solaria | stability-agent | ro |
| chelsty-infra | unverifiable (node offline); per repo compose: node-agent rw, stability-agent ro | — |
After the ai-cluster migration lands, solaria additionally gets
service-ops-worker with rw sock (services/ai-cluster/docker-compose.yml:133
in the worktree — no :ro).
A2. What each sock-holder is actually allowed to do
- stability-agent (all nodes, ro): monitoring only — reads container state, emits
filesystem events (
containers_not_running,mqtt_unreachableas a TCP probe, disk/health). No restart code path exists. - node-agent (all nodes, rw): reads Docker state and pushes per-event JSON files
to the VPS over ssh/rsync (
VPS_EVENTS_HOST=100.95.58.48). It is also the on-node hands of the executor: it pulls its dispatch inbox (actions/dispatch/<node>/, rsync over the same ssh channel) and executes actions with a type whitelist of exactly{container_restart}(node_agent.py:105) plus a self-restart guard (SELF_RESTART_GUARD_NAMES = {"node-agent"},:111). There is no allowlist of which containers it may restart — any name in an approved-and-dispatched action goes tocontainer.restart()via its own sock. On solaria the Docker half is broken: the container starts withDocker unavailable: Permission denied— the base compose setsgroup_add: ["999"], piha overrides it to the host's docker gid (123), lustro to991, but solaria has no override and its docker gid is 996. Solaria's node-agent therefore reports node health only and could not execute a dispatched restart. - control-plane-executor (vps, rw): executes only actions an operator moved to
approved(HITL). Per action type:container_restart→ writes a dispatch file for the target node-agent (it never touches a remote sock itself);disk_cleanup→ real ssh to the node with a command-safety gate;alert_only→ no-op success;redeploy→ runsscripts/deploy/deploy-node.sh <node> <svc>inside the executor container — which cannot succeed as wired: the script ignores both arguments and deployshostname's full service set from${HOME}/homelab-codex-ws, a path that does not exist in the container → exit 1 →failed/. Redeploy is effectively a dead action type today (executor.py:106-119,deploy-node.shhas no$1/$2). - ai-cluster service-ops-worker (vps today, solaria after migration, rw):
keyword-driven, no LLM, no human approval. Modes:
diagnose(read-only docker ps/inspect/logs),safe_fix→docker restart <container>,deployandrepairare preview-only no-ops. TheSERVICE_NAMESlist inservice_ops_worker.py:31-40is a dead constant — never referenced. The only real gate is a command-shape allowlist (is_allowed_command,:113-137):docker restart <any-token>is permitted against any container on its host, resolved by exact-then-substring match over livedocker psoutput. The repo itself flags this (services/ai-cluster/service.yaml:62-66). - portainer (piha, rw): full unrestricted Docker management UI — outside both agent systems, no allowlist, no audit trail in the event store.
- homepage (piha, rw): dashboard using the sock for container discovery; the mount is rw, so restriction is by application behavior only, not by mount.
A3. Restart-authority overlap on solaria after the ai-cluster move
The two paths have very different shapes:
- service-ops-worker can restart anything
docker psshows on solaria: ollama, node-agent, planner-agent, stability-agent, node_exporter, and the whole ai-cluster stack including its own broker and itself — substring matching makesmosquittoorrediseasy accidental targets. No approval, no coordination. - supervisor→executor→node-agent generates
container_restartonly for services inhosts/solaria/services.yamldesired state. Post-migration that isnode-agent(refused by the node-agent self-guard),ollama(exact container name — executable), andai-cluster— which the branch keys as a single service, a name matching no actual container (ai-cluster-openclaw-1, …), so a dispatched restart for it would fail to resolve. Whether the observer would even raise per-container incidents for the ai-cluster stack under that single key cannot be determined without deploying. Execution additionally requires solaria's node-agent to have working sock access, which it currently does not (A2).
So the clean overlap set is exactly ollama — the GPU workload — restartable
by an unapproved keyword-matched MQTT task on one side and an approved HITL action
on the other, with no shared lock, cooldown, or mutual awareness (supervisor
cooldowns dedup only its own action IDs). The asymmetric remainder is arguably
worse: everything else on the node is restartable only by the autonomous path,
invisible to HITL.
B. Service inventory — repo vs reality
B4. services/ directories and owner_node
23 dirs; 4 lack service.yaml entirely (agent-system, control-plane,
home-assistant, node-agent); node_exporter uses a different (flat) schema.
| service | owner_node | service | owner_node | |
|---|---|---|---|---|
| brain-watchdog | piha | node_exporter | per-host | |
| fleet-prometheus | vps | npm | vps | |
| forgejo | piha | ollama | solaria | |
| gokapi | vps | paperless | piha | |
| ha-diag-agent | per-host | paperless-worker | solaria | |
| kb-postgres | piha | planner-agent | solaria | |
| kb-query | piha | stability-agent | chelsty (not a real node name) | |
| llm-gateway | piha | vikunja | piha | |
| mosquitto | vps (README says piha) | zigbee2mqtt | piha | |
| nextcloud | piha | agent-system / control-plane / home-assistant / node-agent | no service.yaml |
B5. Running containers (2026-07-27)
- vps (24): control-plane ×4 (executor, supervisor, observer, operator-ui), node-agent, stability-agent, node_exporter, fleet-prometheus, ai-cluster ×6 (openclaw, codex-worker, planner-worker, service-ops-worker, redis, mosquitto), outline ×3, joplin ×2, npm, umami ×2, humanai-mailer + humanai-landing (no compose labels at all — not compose-managed).
- piha (44): desired set ha-diag-agent, node-agent, brain-watchdog, vikunja(+db), llm-gateway, node_exporter, kb-postgres, kb-query — all running — plus ~28 containers with no hosts-manifest entry: agent-system ×4 (webui, runtime-materializer, telegram-bot, redis), stability-agent, ollama-piha, paperless ×3, immich ×4, forgejo ×3, zigbee2mqtt, audiobookshelf, grafana, wikijs ×2, code-server, portainer, prometheus, homepage, actual-budget, mqtt-exporter, owntracks ×3, vaultwarden, pihole-exporter (unlabeled), fail2ban-exporter, nginxproxymanager.
- solaria (5): ollama, node-agent, planner-agent, stability-agent, node_exporter.
- chelsty-infra: unverifiable (offline).
B6. Two-way diff
Manifest-but-not-running:
gokapi(hosts/vps/services.yaml) — not running on vps; aredeploy gokapiaction sits inpending, unapproved.- Entire
hosts/chelsty-infra/services.yamlset (ha-diag-agent, node-agent, mosquitto, zigbee2mqtt, frigate) — node offline 8 weeks, unverifiable.
Running-but-no-manifest (in the authoritative hosts/*/services.yaml):
- vps: stability-agent, ai-cluster ×6, outline ×3, joplin ×2, npm, umami ×2,
humanai-mailer, humanai-landing. (topology.yaml lists some of these for vps, but
topology itself declares
hosts/vps/services.yamlauthoritative — see F20.) - piha: the ~28 containers above, including services that do have repo dirs (paperless, forgejo, zigbee2mqtt, agent-system, stability-agent) but no hosts/piha entry.
- solaria: planner-agent, stability-agent, node_exporter (running; absent from hosts/solaria/services.yaml, which lists only node-agent + ollama).
B7. The shadow set — where stability-agent and friends come from
The premise "no dir under services/" is wrong for stability-agent: it lives at
services/stability-agent/ (source, Dockerfile, compose, service.yaml). It is
invisible to the normal pipeline because it deploys through its own
scripts/deploy/deploy-stability-agent.sh (hardcoded Tailscale IPs) → per-node
deploy-local.sh → docker compose up -d --build --force-recreate, and
scripts/deploy/deploy-node.sh:93-94 explicitly skips any service that has its
own deploy-local.sh. It appears in no hosts/*/services.yaml; its service.yaml
says owner_node: chelsty (a node that doesn't exist). Image is built on each node
from the repo checkout. Role: per-node read-only watchdog emitting events.
Other members of the shadow set (deployed outside the declarative pipeline):
| Thing | Deploy path |
|---|---|
| agent-system (piha: webui, materializer, telegram-bot, redis) | services/agent-system/deploy.sh (own docker compose up) — no service.yaml, no hosts entry |
| control-plane (vps) | services/control-plane/deploy-local.sh + scripts/bootstrap/vps-control-plane.sh |
| frigate (chelsty) | scripts/deploy/deploy-frigate.sh composing from hosts/chelsty-infra/runtime/frigate/ — full compose inside hosts/, not services/ |
| kb-ingest | host systemd unit (jobs/documents-ingest/systemd/), non-container |
| Home Assistant config | scripts/ha/deploy.sh API push, outside compose entirely |
| humanai-mailer / humanai-landing (vps), pihole-exporter (piha) | no compose labels — hand-run, no repo trace found |
| piha host mosquitto | systemd OS package, no repo definition at all (see C8) |
scripts/deploy/deploy-role.sh |
references roles/ which does not exist — dead script |
C. MQTT topology
C8. Broker instances
| Node | Instance | Bind | Auth |
|---|---|---|---|
| piha | host systemd mosquitto (OS package, not a container, not in repo) | 0.0.0.0:1883 + [::]:1883 |
allow_anonymous true — open to LAN and Tailscale |
| vps | mosquitto container (ai-cluster stack, pre-GitOps copy) |
100.95.58.48:1883 (Tailscale IP only) |
password + ACL, user codex, topics codex/tasks+codex/results only |
| chelsty-infra | repo: hosts/chelsty-infra/runtime/mosquitto/ |
listener 1883, all interfaces |
repo says allow_anonymous false + password file, but scripts/bootstrap/chelsty-runtime.sh:75-78 creates the password file empty; runtime unverifiable (offline) |
| solaria (future) | worktree ai-cluster stack | 100.100.231.104:1883 (Tailscale IP only) |
password + ACL, user codex |
The repo's services/mosquitto/ manifest (owner_node vps, README "deployed on
piha") matches nothing that actually runs — vps's broker belongs to the
ai-cluster stack, piha's is a host package.
C9. Real clients per broker (24 h of logs)
- piha (systemd broker): persistent, currently-established clients:
zigbee2mqtt (container 172.21.0.2), mqtt-exporter (172.19.0.2),
owntracks-recorder (172.23.0.3), Home Assistant at 192.168.31.7, ~11 LAN IoT
devices (192.168.31.100–.115), piha itself (.5), localhost. Churn: the lustro /
pimirror Pi at 192.168.31.31 (MAC b8:27:eb = Raspberry Pi) made ~600
connections in 24 h, each with a fresh random UUID client-id (plus one
pimirror-pimirror2); a second Pi at .19 connected twice. - vps (ai-cluster broker): zero new connections in the last 30 days. Last log
entries are 2026-06-09:
openclaw,codex-worker,vps-service-ops-1(and planner) connecting as usercodex. The four workers hold long-lived connections; the bus has been otherwise silent for ~7 weeks. - chelsty-infra: unverifiable (offline). Repo-expected clients: zigbee2mqtt,
frigate (via host network,
127.0.0.1:1883), HA on chelsty-ha, mosquitto healthcheck ($SYS/broker/version).
C10. Topic map (from code, both checkouts)
| Component | Publishes | Subscribes | Broker |
|---|---|---|---|
| openclaw (FastAPI gateway) | codex/tasks |
codex/results |
ai-cluster mosquitto (mosquitto:1883 in-stack) |
| codex-worker | codex/results |
codex/tasks |
same |
| planner-worker | codex/results |
codex/tasks |
same |
| service-ops-worker | codex/results |
codex/tasks |
same |
| zigbee2mqtt (chelsty, repo) | zigbee2mqtt/# |
same base | chelsty broker |
| frigate (chelsty, repo) | frigate topics | — | chelsty broker (127.0.0.1) |
| zigbee2mqtt, owntracks, HA, IoT (piha, observed) | device topics | — | piha host broker |
| stability-agent | — | — | no MQTT client; MQTT_HOST is only a TCP reachability probe target |
Worker routing is by payload field target (role:dev / role:planner /
role:service-ops / exact AGENT_ID), not by topic. The ACL permits exactly the
two codex/* topics.
C11. Shared bus?
No. Control-plane and ai-cluster are entirely separate buses — control-plane does
not use MQTT at all. No control-plane or node-agent component has an MQTT client;
their bus is the filesystem event/action path over ssh/rsync. The only "MQTT" in
control-plane is the event type name mqtt_unreachable produced by
stability-agent's TCP probe. There are three unconnected broker islands (piha
home-automation, vps→solaria ai-cluster, chelsty home-automation) and no bridge
config anywhere.
D. Event pipeline economics
D12. Volume on disk
| Node | Store | Files | Size | Oldest → newest |
|---|---|---|---|---|
| vps | /opt/homelab/events/ total |
36,851 | 56 MB data / 181 MB on disk | 2026-05-17 → now |
| vps | events/piha/ (per-event JSON) |
29,437 | 131 MB on disk | rolling |
| vps | events/vps/ |
4,523 | 33 MB | rolling |
| vps | events/lustro/ |
2,673 | 13 MB | rolling |
| vps | events/solaria/ |
207 | 2.3 MB | rolling |
| piha | local events/ (date-keyed jsonl) |
20 | 29 MB | 2026-05-17 → today |
| solaria | local events/ |
14 | 1 MB | 2026-05-17 → 07-23 |
Two distinct stores coexist: the date-keyed append-only
events/YYYY-MM-DD/<node>/events.jsonl (the documented design) and a newer
node-keyed store of individual evt-*.json files on vps
(events/<node>/evt-<node>-<ts>-<type>-<svc>.json), fed by node-agents pushing over
ssh/rsync. Node-agents spool locally then delete after push: the empty local
events/<node>/ directories have bloated directory inodes (26 MB on piha, 1.3 MB on
solaria) — evidence of tens of thousands of files created and deleted.
Pruning: three code paths exist, none covers the bulk. No cron or systemd timer
touches the events tree. (1) node_agent._cleanup_control_plane_fs
(node_agent.py:673-771) runs every 60 s only on vps and deletes an event file
only if its type is in {service_healthy, node_health} AND it is behind the
observer checkpoint AND older than 3 days; it also prunes completed/failed actions
7 d and deploy logs >30 d. (2)
scripts/maintenance/cleanup_event_backlog.py— manual, dry-run by default, same type filter. (3) rsync--remove-source-filesclears each node's local spool after shipping (source-side only). Everything else —healthcheck_failed,containers_not_running, allha_*, liveness, pressure — is retained indefinitely; that is the direct cause of the 29k-file piha backlog.
D13. Does the observer consume-and-delete?
Neither. The observer re-globs the entire tree recursively every 5 s
(observer.py:894,948) and tracks position via a per-node-directory timestamp
checkpoint persisted to /opt/homelab/state/observer_checkpoint.json
(observer.py:225,356-360): a file counts as new iff its filename timestamp exceeds
the checkpoint for parts[0] of its relative path. No per-event markers, no
deletion, no archiving; unparseable files are quarantined to
state/observer_failed_events/ (observer.py:250-268). Cost scales with total
file count regardless of how many are new. Quirk: stability-agent writes to
date-keyed dirs (events/YYYY-MM-DD/<node>/), so its checkpoint bucket is the
date string, not the node — a new bucket every day.
D14. Why ~25k noisy events but only 2 action_results
Verified counts in the vps per-event store for piha: 19,300
ha_entity_unavailable_long, 3,376 healthcheck_failed, 2,490
containers_not_running, and exactly 2 action_result events (both
node-exporter, 2026-07-23). Action pool on vps: 18 pending, 17 cancelled, 0
approved, 0 completed, 0 failed, 0 running.
The funnel, in order:
- Only dispatched
container_restarts can ever produce anaction_result(node_agent.py:993-1010; the only dispatchable type,node_agent.py:105).redeploy,disk_cleanup,alert_onlynever emit one. So 2action_results = 2 container restarts ever executed, fleet-wide. - Desired-state scoping. Drift actions are generated only for services listed
in
hosts/<node>/services.yaml(supervisor.py:309). A failing container not in the manifest (npm, paperless, mosquitto, immich, … — the majority on piha, B6) creates an incident in world state but no action, ever. - Dedup-by-ID with no expiry. The action ID is deterministic
(
container-restart-<node>-<svc>,redeploy-<node>-<svc>,alert-ha-<suffix>-<node>,alert-<type>-<node>,disk-cleanup-<node>) and generation is skipped while that ID exists inpending/approved/running(supervisor.py:377-380et al.). An unapproved pending action suppresses its ID forever. 19,300ha_entity_unavailable_longevents collapse onto the single IDalert-ha-entity-unavailable-piha; cooldowns (1 h for HA/node alerts, 30 min for websocket-restart) only apply after an action reachescompleted/rejected/cancelled. - Nothing is being approved. Executor only reads
approved/(executor.py:75-76); pending→approved happens only via operator-ui/webui/ Telegram. 18 actions sit pending (includingredeploy gokapifor a genuinely down service and a redeploy targeting a node dead 8 weeks); 17 were auto-cancelled by the supervisor when services recovered. - Even approval often couldn't act:
redeployis broken as wired (A2), andalert_onlyapproval is a no-op success by design (executor.py:142-144).
D15. Supervisor event_type → action_type mapping (from code — differs from CLAUDE.md)
Path 1 — world-state drift (supervisor reads world/{services,incidents}.json,
not raw events; supervisor.py:275-340):
| Incident trigger_type | Action | Notes |
|---|---|---|
containers_not_running |
container_restart |
risk low, conf 0.95 |
healthcheck_failed |
redeploy |
not alert-only, contrary to intuition — but redeploy is a dead end (A2) |
service_unhealthy |
redeploy |
|
deployment_failed |
redeploy |
|
| service missing from world state | redeploy (missing_service) |
|
node disk_pressure == high |
disk_cleanup |
skipped for chelsty-* (supervisor.py:45) |
mqtt_unreachable |
dead branch — listed in CONTAINER_RESTART_TRIGGERS (supervisor.py:39) but the observer never creates incidents with that trigger_type |
Path 2 — direct event-file routing (_process_ha_events,
supervisor.py:557-578; globs all nodes every cycle, then if not etype.startswith("ha_"): continue apart from the node-alert set):
| Event type | Action | Cooldown |
|---|---|---|
ha_websocket_dead |
nominally container_restart (homeassistant) — in practice alert_only: HA_DIAG_SHADOW_MODE defaults "true" and nothing overrides it (supervisor.py:101,592-598) |
1800 s |
ha_websocket_recovered |
cancels the pending restart | — |
6 ha_* alert types |
alert_only |
3600 s |
node_offline / node_stale / node_online |
alert_only |
3600 s |
Ignored entirely: service_healthy, node_health, high_cpu, high_memory,
container_restarting, container_state_unexpected, action_result,
deployment_started/completed, medium disk_pressure, mqtt_unreachable, and all
stability-agent-only types (disk_usage_high, docker_api_error, agent_error).
Additionally, stability-agent's own containers_not_running events carry
service=None, which the observer skips when building incidents
(observer.py:729) — stability-agent's flagship signal never opens an
incident; only node-agent's equivalent does.
Dedup key: the deterministic action-ID filename (see D14). monitor: false
(single occurrence, homeassistant @ chelsty-ha) is read in exactly one place —
supervisor.py:203-207 — and suppresses only supervisor action generation; the
observer still ingests the events, incidents still form, UIs still display them.
HA events are also dropped for 300 s after a containers_not_running incident on
homeassistant (supervisor.py:611-627), and each event file is routed at most once
per supervisor process lifetime (in-memory set, lost on restart).
E. Solaria flapping
E16. Who emits node_offline / node_stale / node_online
The observer (control-plane, vps). Event payloads carry "source": "observer"
and messages like Node solaria liveness fresh -> stale (last_seen 184s ago).
Node-agent only sends heartbeats (node_health).
E17. Thresholds vs report interval
From code (services/control-plane/src/liveness.py:36-50): default nodes
fresh ≤ 180 s, stale ≤ 600 s, dead > 600 s; remote nodes (chelsty-*) get
fresh ≤ 900 s / dead > 3600 s. Confirmed at runtime by event payloads: stale fired
at age 184 s, offline at 604 s. Node-agent interval: CHECK_INTERVAL=60 in every
host override; consecutive node_health events land 63 s apart (60 s sleep + ~3 s
work/push) — the observed ~63 s. So stale = ~3 missed heartbeats, dead = ~10.
last_seen is refreshed by any ingested event from the node, not just
heartbeats (observer.py:692).
E18. The 11 "flaps" — explanation
The premise "11 cycles in ~2.7 h" does not match the store: solaria has exactly 11
stale + 11 offline + 11 online events spread over Jul 11 → Jul 23, one cycle per
day, plus today's node_online. Pattern (UTC): node_stale between 19:11 and
22:36 each evening, node_offline 7 minutes later, node_online the next day
between 09:12 and 17:11 (gap Jul 24–26, back online Jul 27 14:12).
Correlation: solaria's current boot is 2026-07-27 16:11 local — one minute before
today's node_online. node-agent has RestartCount=0 and started at boot; its
logs contain no ssh/rsync errors (the only error line is the Docker permission
issue). Nothing in tailscaled logs beyond normal boot-time link setup.
Conclusion: solaria is not flapping — it is powered off every evening and powered on around midday. Each power-off produces exactly one stale→offline pair; each boot produces one online. piha, which runs 24/7, produced 4 such events total in the same store (roughly one real incident). No network, node-agent, or observer defect is involved.
E19. Is solaria's reporting path different?
Transport, interval, event format and shipping target are identical to piha
(ssh/rsync push to 100.95.58.48, 60 s, VPS addressed by IP so solaria's disabled
MagicDNS doesn't matter; same non-remote liveness tier). Three real differences:
(1) the node-agent cannot read the Docker socket on solaria (base
group_add: ["999"] vs host docker gid 996; piha/lustro override, solaria
doesn't), so solaria emits no container-level events — only node_health; (2)
scripts/deploy/orchestrate-deploy.sh:22 excludes solaria (skip set
{saturn, solaria}) from fleet deploys, so its node-agent code can lag the repo;
(3) the host's ~16 h/day power-off duty cycle, which is the entire flapping story.
F. Topology truth
F20. topology.yaml vs hosts/*/services.yaml vs reality — all disagreements
inventory/topology.yaml:67declareshosts/vps/services.yamlauthoritative, then itself lists 10 services for vps where the hosts file lists 5 (delta: stability-agent, npm, outline, joplin, ai-cluster).outline,joplin,ai-clusterappear in topology only — no services/ dir on master, no hosts entry — yet all three run on vps. (services/outline,joplin,npmcutover is documented in CLAUDE.md but the running containers predate GitOps management; ai-cluster's repo definition exists only in the unmerged worktree.)gokapi: desired in hosts/vps, not running; the remediation action is stuck inpending.hosts/saturn/has no services.yaml at all.stability-agent: service.yaml saysowner_node: chelsty— no such node; runs on piha, vps, solaria (and presumably chelsty-infra); appears in no hosts file; topology lists it under vps only.mosquitto: service.yaml ownervps, README says piha, hosts list it only under chelsty-infra; reality: piha runs a host systemd broker (no repo definition), vps runs the ai-cluster container broker.zigbee2mqtt: owner_nodepiha, hosts list it only under chelsty-infra; reality: a zigbee2mqtt container runs on piha (and repo config also targets chelsty) — two instances, one manifest.- solaria: hosts file lists node-agent + ollama; running adds planner-agent, stability-agent, node_exporter; topology lists node-agent only.
- piha: topology omits kb-query (hosts has it); ~28 running containers appear in neither (immich, forgejo, paperless, grafana, portainer, agent-system, …).
- topology says
deployment.mode: pull,orchestrator: saturn; every deploy script SSH-pushes from saturn. - World state tracks 6 nodes (vps, piha, solaria, lustro, chelsty-infra —
chelsty-ha absent) while the recon scope assumed 4;
lustro(MagicMirror Pi) is a full monitored node with node-agent and 1,951service_healthyevents for piper-tts. - chelsty-infra: repo treats it as active (services, pending
redeploy ha-diag-agentaction); control plane has considered it dead since 2026-06-01 (~8 weeks). frigateis a full compose stack insidehosts/chelsty-infra/runtime/— outside the services/ convention;services/home-assistant/has no compose at all (config-push model);agent-system,control-plane,node-agentlack service.yaml.- Schema drift: 21 service.yaml files nest under
service:, node_exporter's is flat.
F21. Every monitor: false in the repo
Exactly one: hosts/chelsty-ha/services.yaml:12 for homeassistant, with a
comment-only rationale (no reason: field): chelsty-ha has no node-agent, so there
are no container events to track; HA is monitored indirectly via the chelsty-infra
MQTT broker; re-enable once node-agent is bootstrapped there. Sole code reader:
supervisor.py:203-207 (action generation only); a side effect at
supervisor.py:529-531 auto-cancels existing pending actions for the service.
(Related but distinct: hosts/chelsty-infra/services.yaml:23-25 comments that
node-agent there "monitors and emits events but does NO Docker cleanup".)
OPEN ARCHITECTURAL QUESTIONS
Decisions a human must make; no recommendations attached.
- Dual restart authority on solaria. After the ai-cluster migration the
autonomous, keyword-triggered service-ops-worker can restart every container on
solaria (its service allowlist is dead code), while the HITL path can restart
the subset in desired state — cleanly overlapping on
ollama, the GPU workload — with no shared lock, cooldown, or mutual awareness. Decide: does service-ops-worker keep an rw docker.sock on solaria; is a real target allowlist enforced; which system owns restarts for which services; and should most of the node remain restartable only by the non-HITL path. At stake: uncoordinated double-remediation, HITL bypass, GPU workloads killable by one MQTT message holding thecodexcredentials. - Solaria's duty cycle vs its new role. Solaria is powered off ~16 h/day, and the migration makes it home of the ai-cluster bus (mosquitto, redis, openclaw, all workers) plus the GPU stack. Decide: is solaria a 24/7 service node or an opportunistic workstation, and either way, whether liveness monitoring and the ai-cluster control bus should live on a host with a nightly power-off. At stake: ai-cluster availability drops to solaria's uptime window; daily stale/offline/online alert noise is baked in.
- Event store lifecycle. Nothing prunes either event store; 37k files / 181 MB on vps and growing, plus directory-inode bloat from the spool-and-delete pattern on every node. Decide: retention policy, which of the two store formats (date-keyed jsonl vs per-event JSON) is canonical, and who owns pruning. At stake: unbounded disk growth on a 4 GB VPS, observer scan cost, inode exhaustion.
- The approval pipeline is not being operated — and partially cannot work.
18 pending actions (0 approved, 0 executed beyond the 2 node-exporter
restarts), including a real outage (
gokapidown) whose remediation type (redeploy) is broken as wired regardless of approval; a pending action suppresses its ID forever, so each incident class alerts at most once until someone drains the queue. Decide: pending-action TTL/expiry, escalation policy, whetheralert_onlybelongs in the same queue as executable remediations, and whether redeploy should be fixed or removed. At stake: real incidents buried in a queue nobody drains; HITL becomes "no-op in the loop". - GitOps boundary. The declarative desired state covers a minority of what runs: ~28 unmanaged containers on piha, a host-level anonymous MQTT broker with no repo definition, unlabeled hand-run containers on vps (humanai-*), and a shadow-deploy family (stability-agent, agent-system, frigate, kb-ingest, HA config push). Decide: what is in scope for drift detection — and either bring the rest in or declare it out. At stake: supervisor drift detection is structurally incomplete; the repo's authority claim is false today.
- chelsty-infra's status. Offline to the control plane since 2026-06-01, yet still a deploy target, still in manifests, with a pending remediation action. Decide: reclassify (dormant/seasonal site?) or repair connectivity, and how the supervisor should treat nodes offline for weeks. At stake: dead-node actions accumulating; monitoring credibility.
- One bus, three islands, or four. Messaging today: piha's anonymous
home-automation broker, ai-cluster's authed
codex/*bus, chelsty's local broker, plus control-plane's ssh/rsync file bus. Decide: is this separation intentional architecture (offline-first sites, isolated blast radii) or accident — and specifically whether the piha broker stays anonymous on 0.0.0.0 while HA, IoT devices, owntracks and exporters depend on it. At stake: auth model, single-points-of-failure placement, and whether "mqtt_unreachable" means anything consistent across nodes. - Monitoring blind spots as policy. Solaria's node-agent has been blind to
Docker since deployment (group_add mismatch);
healthcheck_failedroutes only to the dead redeploy path; stability-agent'scontainers_not_running(service=None) never opens an incident, making its flagship signal inert;ha_websocket_deadremediation has silently stayed in shadow mode; chelsty-ha is monitor:false pending a bootstrap that hasn't happened; unmanifested services get incidents but can never get actions. Decide which blind spots are accepted policy and which are defects, and whether "event emitted but unroutable" should be visible anywhere. At stake: the difference between "monitored" and "logged".
Etap 0 changes (2026-07-28)
Repo-only truth cleanup driven by this recon (commits 2026-07-28 → 2026-07-30,
branch task/porzadki-topologia plus earlier etap-0 commits already on master).
No runtime state was touched.
Topology / nodes
inventory/topology.yaml: chelsty-infra and chelsty-ha markedstatus: dormant(site hardware down since ~2026-06-01, pending physical revival — F20.12 / open question 6); per-node service enumerations removed —hosts/<node>/services.yamlis authoritative (F20.1-2);deployment.modecorrected topush(F20.10); lustro added as a full monitored node with its duty cycle documented (nightly ~23:30 power-off, one liveness cycle per day — same pattern as solaria, F20.11).- Supervisor reads
status: dormantfrominventory/topology.yaml(_load_dormant_nodes) and skips desired-state loading, disk-cleanup and node-alert generation for dormant nodes — replacing the hardcoded chelsty-name checks (topology is readable from the supervisor's runtime context, so no env-var fallback was needed).
hosts/ reconciliation (B6/F20)
hosts/solaria/services.yaml: added stability-agent, node_exporter, and planner-agent (monitor: false, legacy ai-cluster family, retirement candidate).hosts/vps/services.yaml: added stability-agent, npm, outline, joplin-server, umami (all verified healthy in world state 2026-07-30); humanai-mailer / humanai-landing added asunmanaged: true+monitor: false(hand-run, no compose labels); ai-cluster deliberately gets NO entry — comment block points atai-cluster-LEGACY.md.hosts/piha/services.yaml: single comment block enumerating the ~28 known unmanaged containers (B5) — bringing them in is a later stage (open question 5).hosts/saturn/services.yaml: created with an explicit empty service list (dev workstation / orchestrator, nothing monitored — F20.4).hosts/lustro/services.yaml: added node-exporter and piper-tts (verified running); watchtower noted as deliberately unmanaged.services/stability-agent/service.yaml:owner_node: chelsty→per-host(B7/F20.5).services/mosquitto/service.yaml: marked NOT DEPLOYED / legacy manifest — matches nothing that runs (C8/F20.6); kept for reference pending open question 7.
Dead code
scripts/deploy/deploy-role.shdeleted (referenced nonexistentroles/, B7).- Supervisor:
mqtt_unreachableremoved fromCONTAINER_RESTART_TRIGGERS(dead branch — the observer never creates incidents with that trigger_type, D15).
Legacy
docs/architecture/ai-cluster-LEGACY.md: ai-cluster is retired in place, not migrated (bus idle since 2026-06-09, C9); branchtask/ai-cluster-solariastays unmerged as documentation; surviving patterns listed; runtime retirement runbook (stop stack on vps, observefree -m, remove containers) to be executed in a separate supervised session.
Runbook — stale chelsty action (from 1c, do in a runtime session)
- On vps:
rm /opt/homelab/actions/pending/redeploy-chelsty-infra-ha-diag-agent.json— the only chelsty-targeted pending action as of 2026-07-30 (verified viagrep -l chelsty /opt/homelab/actions/pending/*.json).
Discrepancies found during this pass, NOT covered by etap 0
- lustro also runs
watchtower(auto-updater) — undocumented anywhere; whether unattended container updates on a monitored node are policy needs a decision. - World state carries stale solaria keys (
solaria/executor,solaria/operator-ui,solaria/kb-postgres,solaria/paperless-worker) for containers that do not run there (B5 counted 5) — world state never prunes departed services; same mechanism keeps allchelsty-infra/*keys alive. - piha world state holds duplicate keys from naming drift
(
piha/node-exportervspiha/node_exporter,immich-servervsimmich_server, genericapp/db/database/brokerkeys) — event-source naming is not normalized. gokapiremains desired-but-down with itsredeployaction stuck inpending(F20.3); the wider undrained queue (18 pending actions) is open question 4 — untouched here.outline,joplin,umamirun on vps with noservices/<name>dir on master — their hosts entries added in etap 0 document this; actual GitOps cutover (compose in repo) is still pending (F20.2).
Solaria node-agent docker gid (A2/E19) — status 2026-07-30
- Repo fix already on master (
ddae57c,group_add: "996"inhosts/solaria/runtime/node-agent/docker-compose.override.yml) but never deployed: running container inspected 2026-07-30 showsGroupAdd=[999]and still logsDocker unavailable: Permission denied. Deploy from the task worktree is blocked by design (deploy.shpreflight enforces master); pending operator run ofscripts/deploy/deploy.sh solariafrom the main checkout, then re-verify container events reach/opt/homelab/events/on vps.