126 plikow (md, yaml, sh, py) odwolywalo sie do sciezek sprzed migracji.
15 markdown-linkow [..](..) -> policzona sciezka WZGLEDNA wobec pliku
odsylajacego (wczesniej czesc z nich byla repo-root-relative i nie
rozwiazywala sie z katalogu, w ktorym lezala)
200 odwolan tekstowych (backticki, proza, yaml, importy w kodzie)
-> nowa sciezka repo-root-relative, zgodnie z konwencja repo
5 linkow rodzenstwa (gole nazwy plikow, np. "](DEPLOY.md)") — dzialaly
tylko w starym katalogu; przeliczone recznie
Objete m.in.: CLAUDE.md (scripts/onboard/README.md -> kb/runbooks/
node-onboarding-tool.md, docs/backlog.md -> kb/phases/backlog.md),
README.md, .claude/skills/, 20 session logow, kod jobow.
Ostatnie 5 odwolan pochodzi z tresci wciagnietej rebasem z origin/master
(session log 2026-07-31, override node-agenta na SOLARII, dwie pozycje
backlogu) — wskazywaly na docs/incidents/, docs/kb/modules/ i
services/narty27/README.md sprzed migracji.
Dodany wzajemny link miedzy kb/services/control-plane.md (stub kodu)
a kb/subsystems/control-plane.md (opis, deprecated) — dwa dokumenty o tym
samym systemie, latwe do pomylenia.
Weryfikacja na 790 plikach: 0 odwolan do starych sciezek,
0 martwych linkow markdown. Lint OKF: 190/190 plikow ZGODNE.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
659 lines
37 KiB
Markdown
659 lines
37 KiB
Markdown
---
|
||
okf: "0.1"
|
||
type: subsystem
|
||
visibility: private
|
||
status: active
|
||
updated: 2026-07-30
|
||
links: []
|
||
---
|
||
|
||
# Architecture Recon — multi-agent systems (2026-07-27)
|
||
|
||
Read-only recon of the two multi-agent systems (control-plane, ai-cluster) across vps,
|
||
piha, solaria, chelsty-infra. Runtime data collected 2026-07-27 from solaria (local),
|
||
piha and vps (ssh). **chelsty-infra was unreachable** — Tailscale reports it offline,
|
||
last seen 55 days ago, and the control-plane world state agrees (`last_seen`
|
||
2026-06-01, liveness `dead`). Everything about chelsty below is repo-only, marked as
|
||
such. ai-cluster migration read from worktree `task/ai-cluster-solaria` @ `b124e54`.
|
||
|
||
---
|
||
|
||
## A. Authority map
|
||
|
||
### A1. Containers mounting /var/run/docker.sock (observed via `docker inspect`)
|
||
|
||
| Node | Container | Mode |
|
||
|---|---|---|
|
||
| vps | control-plane-executor | **rw** |
|
||
| vps | node-agent | **rw** |
|
||
| vps | ai-cluster-service-ops-worker-1 | **rw** |
|
||
| vps | stability-agent | ro |
|
||
| piha | node-agent | **rw** |
|
||
| piha | portainer | **rw** |
|
||
| piha | homepage | **rw** |
|
||
| piha | stability-agent | ro |
|
||
| solaria | node-agent | **rw** |
|
||
| solaria | stability-agent | ro |
|
||
| chelsty-infra | *unverifiable (node offline)*; per repo compose: node-agent rw, stability-agent ro | — |
|
||
|
||
After the ai-cluster migration lands, solaria additionally gets
|
||
`service-ops-worker` with **rw** sock (`services/ai-cluster/docker-compose.yml:133`
|
||
in the worktree — no `:ro`).
|
||
|
||
### A2. What each sock-holder is actually allowed to do
|
||
|
||
- **stability-agent** (all nodes, ro): monitoring only — reads container state, emits
|
||
filesystem events (`containers_not_running`, `mqtt_unreachable` as a TCP probe,
|
||
disk/health). No restart code path exists.
|
||
- **node-agent** (all nodes, rw): reads Docker state and pushes per-event JSON files
|
||
to the VPS over ssh/rsync (`VPS_EVENTS_HOST=100.95.58.48`). It is also the on-node
|
||
hands of the executor: it pulls its dispatch inbox
|
||
(`actions/dispatch/<node>/`, rsync over the same ssh channel) and executes
|
||
actions with a **type whitelist of exactly `{container_restart}`**
|
||
(`node_agent.py:105`) plus a self-restart guard (`SELF_RESTART_GUARD_NAMES =
|
||
{"node-agent"}`, `:111`). There is **no allowlist of which containers it may
|
||
restart** — any name in an approved-and-dispatched action goes to
|
||
`container.restart()` via its own sock. On **solaria the Docker half is broken**:
|
||
the container starts with `Docker unavailable: Permission denied` — the base
|
||
compose sets `group_add: ["999"]`, piha overrides it to the host's docker gid
|
||
(`123`), lustro to `991`, but solaria has no override and its docker gid is 996.
|
||
Solaria's node-agent therefore reports node health only and could not execute a
|
||
dispatched restart.
|
||
- **control-plane-executor** (vps, rw): executes only actions an operator moved to
|
||
`approved` (HITL). Per action type: `container_restart` → writes a dispatch file
|
||
for the target node-agent (it never touches a remote sock itself);
|
||
`disk_cleanup` → real ssh to the node with a command-safety gate; `alert_only` →
|
||
no-op success; `redeploy` → runs `scripts/deploy/deploy-node.sh <node> <svc>`
|
||
inside the executor container — **which cannot succeed as wired**: the script
|
||
ignores both arguments and deploys `hostname`'s full service set from
|
||
`${HOME}/homelab-codex-ws`, a path that does not exist in the container → exit 1
|
||
→ `failed/`. Redeploy is effectively a dead action type today
|
||
(`executor.py:106-119`, `deploy-node.sh` has no `$1`/`$2`).
|
||
- **ai-cluster service-ops-worker** (vps today, solaria after migration, rw):
|
||
keyword-driven, **no LLM, no human approval**. Modes: `diagnose` (read-only
|
||
docker ps/inspect/logs), `safe_fix` → **`docker restart <container>`**, `deploy` and
|
||
`repair` are preview-only no-ops. The `SERVICE_NAMES` list in
|
||
`service_ops_worker.py:31-40` is a **dead constant — never referenced**. The only
|
||
real gate is a command-*shape* allowlist (`is_allowed_command`, `:113-137`):
|
||
`docker restart <any-token>` is permitted against **any container on its host**,
|
||
resolved by exact-then-**substring** match over live `docker ps` output. The repo
|
||
itself flags this (`services/ai-cluster/service.yaml:62-66`).
|
||
- **portainer** (piha, rw): full unrestricted Docker management UI — outside both
|
||
agent systems, no allowlist, no audit trail in the event store.
|
||
- **homepage** (piha, rw): dashboard using the sock for container discovery; the
|
||
mount is rw, so restriction is by application behavior only, not by mount.
|
||
|
||
### A3. Restart-authority overlap on solaria after the ai-cluster move
|
||
|
||
The two paths have very different shapes:
|
||
|
||
- **service-ops-worker** can restart **anything `docker ps` shows on solaria**:
|
||
ollama, node-agent, planner-agent, stability-agent, node_exporter, and the whole
|
||
ai-cluster stack including **its own broker and itself** — substring matching
|
||
makes `mosquitto` or `redis` easy accidental targets. No approval, no
|
||
coordination.
|
||
- **supervisor→executor→node-agent** generates `container_restart` only for
|
||
services in `hosts/solaria/services.yaml` desired state. Post-migration that is
|
||
`node-agent` (refused by the node-agent self-guard), `ollama` (exact container
|
||
name — executable), and `ai-cluster` — which the branch keys as a **single
|
||
service**, a name matching no actual container (`ai-cluster-openclaw-1`, …), so a
|
||
dispatched restart for it would fail to resolve. Whether the observer would even
|
||
raise per-container incidents for the ai-cluster stack under that single key
|
||
cannot be determined without deploying. Execution additionally requires solaria's
|
||
node-agent to have working sock access, which it currently does not (A2).
|
||
|
||
So the *clean* overlap set is exactly **`ollama`** — the GPU workload — restartable
|
||
by an unapproved keyword-matched MQTT task on one side and an approved HITL action
|
||
on the other, with no shared lock, cooldown, or mutual awareness (supervisor
|
||
cooldowns dedup only its own action IDs). The asymmetric remainder is arguably
|
||
worse: everything else on the node is restartable *only* by the autonomous path,
|
||
invisible to HITL.
|
||
|
||
---
|
||
|
||
## B. Service inventory — repo vs reality
|
||
|
||
### B4. services/ directories and owner_node
|
||
|
||
23 dirs; 4 lack `service.yaml` entirely (agent-system, control-plane,
|
||
home-assistant, node-agent); node_exporter uses a different (flat) schema.
|
||
|
||
| service | owner_node | | service | owner_node |
|
||
|---|---|---|---|---|
|
||
| brain-watchdog | piha | | node_exporter | per-host |
|
||
| fleet-prometheus | vps | | npm | vps |
|
||
| forgejo | piha | | ollama | solaria |
|
||
| gokapi | vps | | paperless | piha |
|
||
| ha-diag-agent | per-host | | paperless-worker | solaria |
|
||
| kb-postgres | piha | | planner-agent | solaria |
|
||
| kb-query | piha | | stability-agent | **chelsty** (not a real node name) |
|
||
| llm-gateway | piha | | vikunja | piha |
|
||
| mosquitto | **vps** (README says piha) | | zigbee2mqtt | piha |
|
||
| nextcloud | piha | | agent-system / control-plane / home-assistant / node-agent | *no service.yaml* |
|
||
|
||
### B5. Running containers (2026-07-27)
|
||
|
||
- **vps** (24): control-plane ×4 (executor, supervisor, observer, operator-ui),
|
||
node-agent, stability-agent, node_exporter, fleet-prometheus, ai-cluster ×6
|
||
(openclaw, codex-worker, planner-worker, service-ops-worker, redis, mosquitto),
|
||
outline ×3, joplin ×2, npm, umami ×2, humanai-mailer + humanai-landing (**no
|
||
compose labels at all** — not compose-managed).
|
||
- **piha** (44): desired set ha-diag-agent, node-agent, brain-watchdog, vikunja(+db),
|
||
llm-gateway, node_exporter, kb-postgres, kb-query — all running — plus ~28
|
||
containers with no hosts-manifest entry: agent-system ×4 (webui,
|
||
runtime-materializer, telegram-bot, redis), stability-agent, ollama-piha,
|
||
paperless ×3, immich ×4, forgejo ×3, zigbee2mqtt, audiobookshelf, grafana,
|
||
wikijs ×2, code-server, portainer, prometheus, homepage, actual-budget,
|
||
mqtt-exporter, owntracks ×3, vaultwarden, pihole-exporter (unlabeled),
|
||
fail2ban-exporter, nginxproxymanager.
|
||
- **solaria** (5): ollama, node-agent, planner-agent, stability-agent, node_exporter.
|
||
- **chelsty-infra**: unverifiable (offline).
|
||
|
||
### B6. Two-way diff
|
||
|
||
**Manifest-but-not-running:**
|
||
- `gokapi` (hosts/vps/services.yaml) — not running on vps; a `redeploy gokapi` action
|
||
sits in `pending`, unapproved.
|
||
- Entire `hosts/chelsty-infra/services.yaml` set (ha-diag-agent, node-agent,
|
||
mosquitto, zigbee2mqtt, frigate) — node offline 8 weeks, unverifiable.
|
||
|
||
**Running-but-no-manifest (in the authoritative hosts/*/services.yaml):**
|
||
- vps: stability-agent, ai-cluster ×6, outline ×3, joplin ×2, npm, umami ×2,
|
||
humanai-mailer, humanai-landing. (topology.yaml lists some of these for vps, but
|
||
topology itself declares `hosts/vps/services.yaml` authoritative — see F20.)
|
||
- piha: the ~28 containers above, including services that *do* have repo dirs
|
||
(paperless, forgejo, zigbee2mqtt, agent-system, stability-agent) but no
|
||
hosts/piha entry.
|
||
- solaria: planner-agent, stability-agent, node_exporter (running; absent from
|
||
hosts/solaria/services.yaml, which lists only node-agent + ollama).
|
||
|
||
### B7. The shadow set — where stability-agent and friends come from
|
||
|
||
The premise "no dir under services/" is wrong for stability-agent: it lives at
|
||
`services/stability-agent/` (source, Dockerfile, compose, service.yaml). It is
|
||
invisible to the normal pipeline because it deploys through its own
|
||
`scripts/deploy/deploy-stability-agent.sh` (hardcoded Tailscale IPs) → per-node
|
||
`deploy-local.sh` → `docker compose up -d --build --force-recreate`, and
|
||
`scripts/deploy/deploy-node.sh:93-94` explicitly **skips any service that has its
|
||
own `deploy-local.sh`**. It appears in no hosts/*/services.yaml; its service.yaml
|
||
says `owner_node: chelsty` (a node that doesn't exist). Image is built on each node
|
||
from the repo checkout. Role: per-node read-only watchdog emitting events.
|
||
|
||
Other members of the shadow set (deployed outside the declarative pipeline):
|
||
|
||
| Thing | Deploy path |
|
||
|---|---|
|
||
| agent-system (piha: webui, materializer, telegram-bot, redis) | `services/agent-system/deploy.sh` (own `docker compose up`) — no service.yaml, no hosts entry |
|
||
| control-plane (vps) | `services/control-plane/deploy-local.sh` + `scripts/bootstrap/vps-control-plane.sh` |
|
||
| frigate (chelsty) | `scripts/deploy/deploy-frigate.sh` composing from `hosts/chelsty-infra/runtime/frigate/` — full compose inside hosts/, not services/ |
|
||
| kb-ingest | host systemd unit (`jobs/documents-ingest/systemd/`), non-container |
|
||
| Home Assistant config | `scripts/ha/deploy.sh` API push, outside compose entirely |
|
||
| humanai-mailer / humanai-landing (vps), pihole-exporter (piha) | no compose labels — hand-run, no repo trace found |
|
||
| piha host mosquitto | systemd OS package, no repo definition at all (see C8) |
|
||
| `scripts/deploy/deploy-role.sh` | references `roles/` which **does not exist** — dead script |
|
||
|
||
---
|
||
|
||
## C. MQTT topology
|
||
|
||
### C8. Broker instances
|
||
|
||
| Node | Instance | Bind | Auth |
|
||
|---|---|---|---|
|
||
| piha | **host systemd mosquitto** (OS package, not a container, not in repo) | `0.0.0.0:1883` + `[::]:1883` | **`allow_anonymous true`** — open to LAN and Tailscale |
|
||
| vps | `mosquitto` container (ai-cluster stack, pre-GitOps copy) | `100.95.58.48:1883` (Tailscale IP only) | password + ACL, user `codex`, topics `codex/tasks`+`codex/results` only |
|
||
| chelsty-infra | repo: `hosts/chelsty-infra/runtime/mosquitto/` | `listener 1883`, all interfaces | repo says `allow_anonymous false` + password file, but `scripts/bootstrap/chelsty-runtime.sh:75-78` creates the password file **empty**; runtime unverifiable (offline) |
|
||
| solaria (future) | worktree ai-cluster stack | `100.100.231.104:1883` (Tailscale IP only) | password + ACL, user `codex` |
|
||
|
||
The repo's `services/mosquitto/` manifest (owner_node `vps`, README "deployed on
|
||
piha") matches **nothing that actually runs** — vps's broker belongs to the
|
||
ai-cluster stack, piha's is a host package.
|
||
|
||
### C9. Real clients per broker (24 h of logs)
|
||
|
||
- **piha (systemd broker):** persistent, currently-established clients:
|
||
zigbee2mqtt (container 172.21.0.2), mqtt-exporter (172.19.0.2),
|
||
owntracks-recorder (172.23.0.3), Home Assistant at 192.168.31.7, ~11 LAN IoT
|
||
devices (192.168.31.100–.115), piha itself (.5), localhost. Churn: the lustro /
|
||
pimirror Pi at 192.168.31.31 (MAC b8:27:eb = Raspberry Pi) made **~600
|
||
connections in 24 h, each with a fresh random UUID client-id** (plus one
|
||
`pimirror-pimirror2`); a second Pi at .19 connected twice.
|
||
- **vps (ai-cluster broker):** **zero new connections in the last 30 days.** Last log
|
||
entries are 2026-06-09: `openclaw`, `codex-worker`, `vps-service-ops-1` (and
|
||
planner) connecting as user `codex`. The four workers hold long-lived connections;
|
||
the bus has been otherwise silent for ~7 weeks.
|
||
- **chelsty-infra:** unverifiable (offline). Repo-expected clients: zigbee2mqtt,
|
||
frigate (via host network, `127.0.0.1:1883`), HA on chelsty-ha, mosquitto
|
||
healthcheck (`$SYS/broker/version`).
|
||
|
||
### C10. Topic map (from code, both checkouts)
|
||
|
||
| Component | Publishes | Subscribes | Broker |
|
||
|---|---|---|---|
|
||
| openclaw (FastAPI gateway) | `codex/tasks` | `codex/results` | ai-cluster mosquitto (`mosquitto:1883` in-stack) |
|
||
| codex-worker | `codex/results` | `codex/tasks` | same |
|
||
| planner-worker | `codex/results` | `codex/tasks` | same |
|
||
| service-ops-worker | `codex/results` | `codex/tasks` | same |
|
||
| zigbee2mqtt (chelsty, repo) | `zigbee2mqtt/#` | same base | chelsty broker |
|
||
| frigate (chelsty, repo) | frigate topics | — | chelsty broker (127.0.0.1) |
|
||
| zigbee2mqtt, owntracks, HA, IoT (piha, observed) | device topics | — | piha host broker |
|
||
| stability-agent | — | — | no MQTT client; `MQTT_HOST` is only a TCP reachability probe target |
|
||
|
||
Worker routing is by payload field `target` (`role:dev` / `role:planner` /
|
||
`role:service-ops` / exact `AGENT_ID`), not by topic. The ACL permits exactly the
|
||
two `codex/*` topics.
|
||
|
||
### C11. Shared bus?
|
||
|
||
**No. Control-plane and ai-cluster are entirely separate buses — control-plane does
|
||
not use MQTT at all.** No control-plane or node-agent component has an MQTT client;
|
||
their bus is the filesystem event/action path over ssh/rsync. The only "MQTT" in
|
||
control-plane is the event *type name* `mqtt_unreachable` produced by
|
||
stability-agent's TCP probe. There are three unconnected broker islands (piha
|
||
home-automation, vps→solaria ai-cluster, chelsty home-automation) and no bridge
|
||
config anywhere.
|
||
|
||
---
|
||
|
||
## D. Event pipeline economics
|
||
|
||
### D12. Volume on disk
|
||
|
||
| Node | Store | Files | Size | Oldest → newest |
|
||
|---|---|---|---|---|
|
||
| vps | `/opt/homelab/events/` total | 36,851 | 56 MB data / **181 MB on disk** | 2026-05-17 → now |
|
||
| vps | `events/piha/` (per-event JSON) | 29,437 | 131 MB on disk | rolling |
|
||
| vps | `events/vps/` | 4,523 | 33 MB | rolling |
|
||
| vps | `events/lustro/` | 2,673 | 13 MB | rolling |
|
||
| vps | `events/solaria/` | 207 | 2.3 MB | rolling |
|
||
| piha | local `events/` (date-keyed jsonl) | 20 | 29 MB | 2026-05-17 → today |
|
||
| solaria | local `events/` | 14 | 1 MB | 2026-05-17 → 07-23 |
|
||
|
||
Two distinct stores coexist: the date-keyed append-only
|
||
`events/YYYY-MM-DD/<node>/events.jsonl` (the documented design) and a newer
|
||
node-keyed store of **individual `evt-*.json` files** on vps
|
||
(`events/<node>/evt-<node>-<ts>-<type>-<svc>.json`), fed by node-agents pushing over
|
||
ssh/rsync. Node-agents spool locally then delete after push: the *empty* local
|
||
`events/<node>/` directories have bloated directory inodes (26 MB on piha, 1.3 MB on
|
||
solaria) — evidence of tens of thousands of files created and deleted.
|
||
|
||
**Pruning: three code paths exist, none covers the bulk.** No cron or systemd timer
|
||
touches the events tree. (1) `node_agent._cleanup_control_plane_fs`
|
||
(`node_agent.py:673-771`) runs every 60 s **only on vps** and deletes an event file
|
||
only if its type is in `{service_healthy, node_health}` AND it is behind the
|
||
observer checkpoint AND older than 3 days; it also prunes completed/failed actions
|
||
>7 d and deploy logs >30 d. (2) `scripts/maintenance/cleanup_event_backlog.py` —
|
||
manual, dry-run by default, same type filter. (3) rsync `--remove-source-files`
|
||
clears each node's local spool after shipping (source-side only). Everything else —
|
||
`healthcheck_failed`, `containers_not_running`, all `ha_*`, liveness, pressure —
|
||
is retained **indefinitely**; that is the direct cause of the 29k-file piha
|
||
backlog.
|
||
|
||
### D13. Does the observer consume-and-delete?
|
||
|
||
Neither. The observer re-globs the **entire** tree recursively every 5 s
|
||
(`observer.py:894,948`) and tracks position via a per-node-directory timestamp
|
||
checkpoint persisted to `/opt/homelab/state/observer_checkpoint.json`
|
||
(`observer.py:225,356-360`): a file counts as new iff its filename timestamp exceeds
|
||
the checkpoint for `parts[0]` of its relative path. No per-event markers, no
|
||
deletion, no archiving; unparseable files are quarantined to
|
||
`state/observer_failed_events/` (`observer.py:250-268`). Cost scales with total
|
||
file count regardless of how many are new. Quirk: stability-agent writes to
|
||
date-keyed dirs (`events/YYYY-MM-DD/<node>/`), so its checkpoint bucket is the
|
||
*date string*, not the node — a new bucket every day.
|
||
|
||
### D14. Why ~25k noisy events but only 2 `action_result`s
|
||
|
||
Verified counts in the vps per-event store for piha: 19,300
|
||
`ha_entity_unavailable_long`, 3,376 `healthcheck_failed`, 2,490
|
||
`containers_not_running`, and exactly 2 `action_result` events (both
|
||
`node-exporter`, 2026-07-23). Action pool on vps: **18 pending, 17 cancelled, 0
|
||
approved, 0 completed, 0 failed, 0 running.**
|
||
|
||
The funnel, in order:
|
||
|
||
1. **Only dispatched `container_restart`s can ever produce an `action_result`**
|
||
(`node_agent.py:993-1010`; the only dispatchable type, `node_agent.py:105`).
|
||
`redeploy`, `disk_cleanup`, `alert_only` never emit one. So 2 `action_result`s
|
||
= 2 container restarts ever executed, fleet-wide.
|
||
2. **Desired-state scoping.** Drift actions are generated only for services listed
|
||
in `hosts/<node>/services.yaml` (`supervisor.py:309`). A failing container not
|
||
in the manifest (npm, paperless, mosquitto, immich, … — the majority on piha,
|
||
B6) creates an incident in world state but **no action, ever**.
|
||
3. **Dedup-by-ID with no expiry.** The action ID is deterministic
|
||
(`container-restart-<node>-<svc>`, `redeploy-<node>-<svc>`,
|
||
`alert-ha-<suffix>-<node>`, `alert-<type>-<node>`, `disk-cleanup-<node>`) and
|
||
generation is skipped while that ID exists in `pending/approved/running`
|
||
(`supervisor.py:377-380` et al.). **An unapproved pending action suppresses its
|
||
ID forever.** 19,300 `ha_entity_unavailable_long` events collapse onto the
|
||
single ID `alert-ha-entity-unavailable-piha`; cooldowns (1 h for HA/node
|
||
alerts, 30 min for websocket-restart) only apply after an action reaches
|
||
`completed/rejected/cancelled`.
|
||
4. **Nothing is being approved.** Executor only reads `approved/`
|
||
(`executor.py:75-76`); pending→approved happens only via operator-ui/webui/
|
||
Telegram. 18 actions sit pending (including `redeploy gokapi` for a genuinely
|
||
down service and a redeploy targeting a node dead 8 weeks); 17 were
|
||
auto-cancelled by the supervisor when services recovered.
|
||
5. **Even approval often couldn't act:** `redeploy` is broken as wired (A2), and
|
||
`alert_only` approval is a no-op success by design (`executor.py:142-144`).
|
||
|
||
### D15. Supervisor event_type → action_type mapping (from code — differs from CLAUDE.md)
|
||
|
||
**Path 1 — world-state drift** (supervisor reads `world/{services,incidents}.json`,
|
||
not raw events; `supervisor.py:275-340`):
|
||
|
||
| Incident trigger_type | Action | Notes |
|
||
|---|---|---|
|
||
| `containers_not_running` | `container_restart` | risk low, conf 0.95 |
|
||
| `healthcheck_failed` | `redeploy` | **not** alert-only, contrary to intuition — but redeploy is a dead end (A2) |
|
||
| `service_unhealthy` | `redeploy` | |
|
||
| `deployment_failed` | `redeploy` | |
|
||
| service missing from world state | `redeploy` (`missing_service`) | |
|
||
| node `disk_pressure == high` | `disk_cleanup` | skipped for chelsty-* (`supervisor.py:45`) |
|
||
| `mqtt_unreachable` | **dead branch** — listed in `CONTAINER_RESTART_TRIGGERS` (`supervisor.py:39`) but the observer never creates incidents with that trigger_type | |
|
||
|
||
**Path 2 — direct event-file routing** (`_process_ha_events`,
|
||
`supervisor.py:557-578`; globs all nodes every cycle, then `if not
|
||
etype.startswith("ha_"): continue` apart from the node-alert set):
|
||
|
||
| Event type | Action | Cooldown |
|
||
|---|---|---|
|
||
| `ha_websocket_dead` | nominally `container_restart` (homeassistant) — **in practice `alert_only`**: `HA_DIAG_SHADOW_MODE` defaults `"true"` and nothing overrides it (`supervisor.py:101,592-598`) | 1800 s |
|
||
| `ha_websocket_recovered` | cancels the pending restart | — |
|
||
| 6 `ha_*` alert types | `alert_only` | 3600 s |
|
||
| `node_offline` / `node_stale` / `node_online` | `alert_only` | 3600 s |
|
||
|
||
**Ignored entirely:** `service_healthy`, `node_health`, `high_cpu`, `high_memory`,
|
||
`container_restarting`, `container_state_unexpected`, `action_result`,
|
||
`deployment_started/completed`, medium `disk_pressure`, `mqtt_unreachable`, and all
|
||
stability-agent-only types (`disk_usage_high`, `docker_api_error`, `agent_error`).
|
||
Additionally, stability-agent's own `containers_not_running` events carry
|
||
`service=None`, which the observer skips when building incidents
|
||
(`observer.py:729`) — **stability-agent's flagship signal never opens an
|
||
incident**; only node-agent's equivalent does.
|
||
|
||
Dedup key: the deterministic action-ID filename (see D14). `monitor: false`
|
||
(single occurrence, `homeassistant` @ chelsty-ha) is read in exactly one place —
|
||
`supervisor.py:203-207` — and suppresses only supervisor action generation; the
|
||
observer still ingests the events, incidents still form, UIs still display them.
|
||
HA events are also dropped for 300 s after a `containers_not_running` incident on
|
||
homeassistant (`supervisor.py:611-627`), and each event file is routed at most once
|
||
per supervisor process lifetime (in-memory set, lost on restart).
|
||
|
||
---
|
||
|
||
## E. Solaria flapping
|
||
|
||
### E16. Who emits node_offline / node_stale / node_online
|
||
|
||
The **observer** (control-plane, vps). Event payloads carry `"source": "observer"`
|
||
and messages like `Node solaria liveness fresh -> stale (last_seen 184s ago)`.
|
||
Node-agent only sends heartbeats (`node_health`).
|
||
|
||
### E17. Thresholds vs report interval
|
||
|
||
From code (`services/control-plane/src/liveness.py:36-50`): default nodes
|
||
fresh ≤ 180 s, stale ≤ 600 s, dead > 600 s; remote nodes (chelsty-*) get
|
||
fresh ≤ 900 s / dead > 3600 s. Confirmed at runtime by event payloads: stale fired
|
||
at age 184 s, offline at 604 s. Node-agent interval: `CHECK_INTERVAL=60` in every
|
||
host override; consecutive `node_health` events land 63 s apart (60 s sleep + ~3 s
|
||
work/push) — the observed ~63 s. So stale = ~3 missed heartbeats, dead = ~10.
|
||
`last_seen` is refreshed by *any* ingested event from the node, not just
|
||
heartbeats (`observer.py:692`).
|
||
|
||
### E18. The 11 "flaps" — explanation
|
||
|
||
The premise "11 cycles in ~2.7 h" does not match the store: solaria has exactly 11
|
||
stale + 11 offline + 11 online events **spread over Jul 11 → Jul 23, one cycle per
|
||
day**, plus today's `node_online`. Pattern (UTC): `node_stale` between 19:11 and
|
||
22:36 each evening, `node_offline` 7 minutes later, `node_online` the next day
|
||
between 09:12 and 17:11 (gap Jul 24–26, back online Jul 27 14:12).
|
||
|
||
Correlation: solaria's current boot is 2026-07-27 16:11 local — **one minute before
|
||
today's `node_online`**. node-agent has `RestartCount=0` and started at boot; its
|
||
logs contain no ssh/rsync errors (the only error line is the Docker permission
|
||
issue). Nothing in tailscaled logs beyond normal boot-time link setup.
|
||
|
||
**Conclusion: solaria is not flapping — it is powered off every evening and powered
|
||
on around midday.** Each power-off produces exactly one stale→offline pair; each
|
||
boot produces one online. piha, which runs 24/7, produced 4 such events total in the
|
||
same store (roughly one real incident). No network, node-agent, or observer defect
|
||
is involved.
|
||
|
||
### E19. Is solaria's reporting path different?
|
||
|
||
Transport, interval, event format and shipping target are identical to piha
|
||
(ssh/rsync push to `100.95.58.48`, 60 s, VPS addressed by IP so solaria's disabled
|
||
MagicDNS doesn't matter; same non-remote liveness tier). Three real differences:
|
||
(1) the node-agent **cannot read the Docker socket on solaria** (base
|
||
`group_add: ["999"]` vs host docker gid 996; piha/lustro override, solaria
|
||
doesn't), so solaria emits no container-level events — only node_health; (2)
|
||
`scripts/deploy/orchestrate-deploy.sh:22` **excludes solaria** (skip set
|
||
`{saturn, solaria}`) from fleet deploys, so its node-agent code can lag the repo;
|
||
(3) the host's ~16 h/day power-off duty cycle, which is the entire flapping story.
|
||
|
||
---
|
||
|
||
## F. Topology truth
|
||
|
||
### F20. topology.yaml vs hosts/*/services.yaml vs reality — all disagreements
|
||
|
||
1. `inventory/topology.yaml:67` declares `hosts/vps/services.yaml` authoritative,
|
||
then itself lists 10 services for vps where the hosts file lists 5 (delta:
|
||
stability-agent, npm, outline, joplin, ai-cluster).
|
||
2. `outline`, `joplin`, `ai-cluster` appear in topology only — no services/ dir on
|
||
master, no hosts entry — yet all three **run** on vps. (`services/outline`,
|
||
`joplin`, `npm` cutover is documented in CLAUDE.md but the running containers
|
||
predate GitOps management; ai-cluster's repo definition exists only in the
|
||
unmerged worktree.)
|
||
3. `gokapi`: desired in hosts/vps, **not running**; the remediation action is stuck
|
||
in `pending`.
|
||
4. `hosts/saturn/` has **no services.yaml** at all.
|
||
5. `stability-agent`: service.yaml says `owner_node: chelsty` — no such node; runs
|
||
on piha, vps, solaria (and presumably chelsty-infra); appears in no hosts file;
|
||
topology lists it under vps only.
|
||
6. `mosquitto`: service.yaml owner `vps`, README says piha, hosts list it only under
|
||
chelsty-infra; reality: piha runs a *host systemd* broker (no repo definition),
|
||
vps runs the ai-cluster container broker.
|
||
7. `zigbee2mqtt`: owner_node `piha`, hosts list it only under chelsty-infra;
|
||
reality: a zigbee2mqtt container runs on **piha** (and repo config also targets
|
||
chelsty) — two instances, one manifest.
|
||
8. solaria: hosts file lists node-agent + ollama; running adds planner-agent,
|
||
stability-agent, node_exporter; topology lists node-agent only.
|
||
9. piha: topology omits kb-query (hosts has it); ~28 running containers appear in
|
||
neither (immich, forgejo, paperless, grafana, portainer, agent-system, …).
|
||
10. topology says `deployment.mode: pull`, `orchestrator: saturn`; every deploy
|
||
script SSH-*pushes* from saturn.
|
||
11. World state tracks **6 nodes** (vps, piha, solaria, lustro, chelsty-infra —
|
||
chelsty-ha absent) while the recon scope assumed 4; `lustro` (MagicMirror Pi)
|
||
is a full monitored node with node-agent and 1,951 `service_healthy` events for
|
||
piper-tts.
|
||
12. chelsty-infra: repo treats it as active (services, pending `redeploy
|
||
ha-diag-agent` action); control plane has considered it **dead since
|
||
2026-06-01** (~8 weeks).
|
||
13. `frigate` is a full compose stack inside `hosts/chelsty-infra/runtime/` —
|
||
outside the services/ convention; `services/home-assistant/` has no compose at
|
||
all (config-push model); `agent-system`, `control-plane`, `node-agent` lack
|
||
service.yaml.
|
||
14. Schema drift: 21 service.yaml files nest under `service:`, node_exporter's is
|
||
flat.
|
||
|
||
### F21. Every `monitor: false` in the repo
|
||
|
||
Exactly one: `hosts/chelsty-ha/services.yaml:12` for `homeassistant`, with a
|
||
comment-only rationale (no `reason:` field): chelsty-ha has no node-agent, so there
|
||
are no container events to track; HA is monitored indirectly via the chelsty-infra
|
||
MQTT broker; re-enable once node-agent is bootstrapped there. Sole code reader:
|
||
`supervisor.py:203-207` (action generation only); a side effect at
|
||
`supervisor.py:529-531` auto-cancels existing pending actions for the service.
|
||
|
||
(Related but distinct: `hosts/chelsty-infra/services.yaml:23-25` comments that
|
||
node-agent there "monitors and emits events but does NO Docker cleanup".)
|
||
|
||
---
|
||
|
||
## OPEN ARCHITECTURAL QUESTIONS
|
||
|
||
Decisions a human must make; no recommendations attached.
|
||
|
||
1. **Dual restart authority on solaria.** After the ai-cluster migration the
|
||
autonomous, keyword-triggered service-ops-worker can restart every container on
|
||
solaria (its service allowlist is dead code), while the HITL path can restart
|
||
the subset in desired state — cleanly overlapping on `ollama`, the GPU
|
||
workload — with no shared lock, cooldown, or mutual awareness. Decide: does
|
||
service-ops-worker keep an rw docker.sock on solaria; is a real target
|
||
allowlist enforced; which system owns restarts for which services; and should
|
||
most of the node remain restartable *only* by the non-HITL path. At stake:
|
||
uncoordinated double-remediation, HITL bypass, GPU workloads killable by one
|
||
MQTT message holding the `codex` credentials.
|
||
2. **Solaria's duty cycle vs its new role.** Solaria is powered off ~16 h/day, and
|
||
the migration makes it home of the ai-cluster bus (mosquitto, redis, openclaw,
|
||
all workers) plus the GPU stack. Decide: is solaria a 24/7 service node or an
|
||
opportunistic workstation, and either way, whether liveness monitoring and the
|
||
ai-cluster control bus should live on a host with a nightly power-off. At stake:
|
||
ai-cluster availability drops to solaria's uptime window; daily
|
||
stale/offline/online alert noise is baked in.
|
||
3. **Event store lifecycle.** Nothing prunes either event store; 37k files / 181 MB
|
||
on vps and growing, plus directory-inode bloat from the spool-and-delete pattern
|
||
on every node. Decide: retention policy, which of the two store formats
|
||
(date-keyed jsonl vs per-event JSON) is canonical, and who owns pruning. At
|
||
stake: unbounded disk growth on a 4 GB VPS, observer scan cost, inode
|
||
exhaustion.
|
||
4. **The approval pipeline is not being operated — and partially cannot work.**
|
||
18 pending actions (0 approved, 0 executed beyond the 2 node-exporter
|
||
restarts), including a real outage (`gokapi` down) whose remediation type
|
||
(`redeploy`) is broken as wired regardless of approval; a pending action
|
||
suppresses its ID forever, so each incident class alerts at most once until
|
||
someone drains the queue. Decide: pending-action TTL/expiry, escalation policy,
|
||
whether `alert_only` belongs in the same queue as executable remediations, and
|
||
whether redeploy should be fixed or removed. At stake: real incidents buried in
|
||
a queue nobody drains; HITL becomes "no-op in the loop".
|
||
5. **GitOps boundary.** The declarative desired state covers a minority of what
|
||
runs: ~28 unmanaged containers on piha, a host-level anonymous MQTT broker with
|
||
no repo definition, unlabeled hand-run containers on vps (humanai-*), and a
|
||
shadow-deploy family (stability-agent, agent-system, frigate, kb-ingest, HA
|
||
config push). Decide: what is in scope for drift detection — and either bring
|
||
the rest in or declare it out. At stake: supervisor drift detection is
|
||
structurally incomplete; the repo's authority claim is false today.
|
||
6. **chelsty-infra's status.** Offline to the control plane since 2026-06-01, yet
|
||
still a deploy target, still in manifests, with a pending remediation action.
|
||
Decide: reclassify (dormant/seasonal site?) or repair connectivity, and how the
|
||
supervisor should treat nodes offline for weeks. At stake: dead-node actions
|
||
accumulating; monitoring credibility.
|
||
7. **One bus, three islands, or four.** Messaging today: piha's anonymous
|
||
home-automation broker, ai-cluster's authed `codex/*` bus, chelsty's local
|
||
broker, plus control-plane's ssh/rsync file bus. Decide: is this separation
|
||
intentional architecture (offline-first sites, isolated blast radii) or accident
|
||
— and specifically whether the piha broker stays anonymous on 0.0.0.0 while HA,
|
||
IoT devices, owntracks and exporters depend on it. At stake: auth model,
|
||
single-points-of-failure placement, and whether "mqtt_unreachable" means
|
||
anything consistent across nodes.
|
||
8. **Monitoring blind spots as policy.** Solaria's node-agent has been blind to
|
||
Docker since deployment (group_add mismatch); `healthcheck_failed` routes only
|
||
to the dead redeploy path; stability-agent's `containers_not_running`
|
||
(`service=None`) never opens an incident, making its flagship signal inert;
|
||
`ha_websocket_dead` remediation has silently stayed in shadow mode; chelsty-ha
|
||
is monitor:false pending a bootstrap that hasn't happened; unmanifested
|
||
services get incidents but can never get actions. Decide which blind spots are
|
||
accepted policy and which are defects, and whether "event emitted but
|
||
unroutable" should be visible anywhere. At stake: the difference between
|
||
"monitored" and "logged".
|
||
|
||
---
|
||
|
||
## Etap 0 changes (2026-07-28)
|
||
|
||
Repo-only truth cleanup driven by this recon (commits 2026-07-28 → 2026-07-30,
|
||
branch `task/porzadki-topologia` plus earlier etap-0 commits already on master).
|
||
No runtime state was touched.
|
||
|
||
**Topology / nodes**
|
||
|
||
- `inventory/topology.yaml`: chelsty-infra and chelsty-ha marked
|
||
`status: dormant` (site hardware down since ~2026-06-01, pending physical
|
||
revival — F20.12 / open question 6); per-node service enumerations removed —
|
||
`hosts/<node>/services.yaml` is authoritative (F20.1-2); `deployment.mode`
|
||
corrected to `push` (F20.10); lustro added as a full monitored node with its
|
||
duty cycle documented (nightly ~23:30 power-off, one liveness cycle per day —
|
||
same pattern as solaria, F20.11).
|
||
- Supervisor reads `status: dormant` from `inventory/topology.yaml`
|
||
(`_load_dormant_nodes`) and skips desired-state loading, disk-cleanup and
|
||
node-alert generation for dormant nodes — replacing the hardcoded
|
||
chelsty-name checks (topology is readable from the supervisor's runtime
|
||
context, so no env-var fallback was needed).
|
||
|
||
**hosts/ reconciliation (B6/F20)**
|
||
|
||
- `hosts/solaria/services.yaml`: added stability-agent, node_exporter, and
|
||
planner-agent (`monitor: false`, legacy ai-cluster family, retirement
|
||
candidate).
|
||
- `hosts/vps/services.yaml`: added stability-agent, npm, outline,
|
||
joplin-server, umami (all verified healthy in world state 2026-07-30);
|
||
humanai-mailer / humanai-landing added as `unmanaged: true` +
|
||
`monitor: false` (hand-run, no compose labels); ai-cluster deliberately gets
|
||
NO entry — comment block points at `ai-cluster-LEGACY.md`.
|
||
- `hosts/piha/services.yaml`: single comment block enumerating the ~28 known
|
||
unmanaged containers (B5) — bringing them in is a later stage (open
|
||
question 5).
|
||
- `hosts/saturn/services.yaml`: created with an explicit empty service list
|
||
(dev workstation / orchestrator, nothing monitored — F20.4).
|
||
- `hosts/lustro/services.yaml`: added node-exporter and piper-tts (verified
|
||
running); watchtower noted as deliberately unmanaged.
|
||
- `services/stability-agent/service.yaml`: `owner_node: chelsty` →
|
||
`per-host` (B7/F20.5).
|
||
- `services/mosquitto/service.yaml`: marked NOT DEPLOYED / legacy manifest —
|
||
matches nothing that runs (C8/F20.6); kept for reference pending open
|
||
question 7.
|
||
|
||
**Dead code**
|
||
|
||
- `scripts/deploy/deploy-role.sh` deleted (referenced nonexistent `roles/`,
|
||
B7).
|
||
- Supervisor: `mqtt_unreachable` removed from `CONTAINER_RESTART_TRIGGERS`
|
||
(dead branch — the observer never creates incidents with that trigger_type,
|
||
D15).
|
||
|
||
**Legacy**
|
||
|
||
- `kb/decisions/ai-cluster-legacy.md`: ai-cluster is retired in place,
|
||
not migrated (bus idle since 2026-06-09, C9); branch `task/ai-cluster-solaria`
|
||
stays unmerged as documentation; surviving patterns listed; runtime
|
||
retirement runbook (stop stack on vps, observe `free -m`, remove containers)
|
||
to be executed in a separate supervised session.
|
||
|
||
**Runbook — stale chelsty action (from 1c, do in a runtime session)**
|
||
|
||
- On vps: `rm /opt/homelab/actions/pending/redeploy-chelsty-infra-ha-diag-agent.json`
|
||
— the only chelsty-targeted pending action as of 2026-07-30 (verified via
|
||
`grep -l chelsty /opt/homelab/actions/pending/*.json`).
|
||
|
||
**Discrepancies found during this pass, NOT covered by etap 0**
|
||
|
||
- lustro also runs `watchtower` (auto-updater) — undocumented anywhere;
|
||
whether unattended container updates on a monitored node are policy needs a
|
||
decision.
|
||
- World state carries stale solaria keys (`solaria/executor`,
|
||
`solaria/operator-ui`, `solaria/kb-postgres`, `solaria/paperless-worker`) for
|
||
containers that do not run there (B5 counted 5) — world state never prunes
|
||
departed services; same mechanism keeps all `chelsty-infra/*` keys alive.
|
||
- piha world state holds duplicate keys from naming drift
|
||
(`piha/node-exporter` vs `piha/node_exporter`, `immich-server` vs
|
||
`immich_server`, generic `app`/`db`/`database`/`broker` keys) — event-source
|
||
naming is not normalized.
|
||
- `gokapi` remains desired-but-down with its `redeploy` action stuck in
|
||
`pending` (F20.3); the wider undrained queue (18 pending actions) is open
|
||
question 4 — untouched here.
|
||
- `outline`, `joplin`, `umami` run on vps with no `services/<name>` dir on
|
||
master — their hosts entries added in etap 0 document this; actual GitOps
|
||
cutover (compose in repo) is still pending (F20.2).
|
||
|
||
**Solaria node-agent docker gid (A2/E19) — status 2026-07-30**
|
||
|
||
- Repo fix already on master (`ddae57c`, `group_add: "996"` in
|
||
`hosts/solaria/runtime/node-agent/docker-compose.override.yml`) but never
|
||
deployed: running container inspected 2026-07-30 shows `GroupAdd=[999]` and
|
||
still logs `Docker unavailable: Permission denied`. Deploy from the task
|
||
worktree is blocked by design (`deploy.sh` preflight enforces master);
|
||
pending operator run of `scripts/deploy/deploy.sh solaria` from the main
|
||
checkout, then re-verify container events reach `/opt/homelab/events/` on vps.
|