homelab-codex-ws/kb/services/ha-diag-agent.md
oskar 6b85c7ef68 feat(kb): SPLIT service+runbook — 10 serwisow -> 20 dokumentow
Wzorzec mechaniczny: sekcje deploy/verify/install/testy wycinane do
kb/runbooks/<serwis>-*.md, reszta zostaje dokumentem type: service.
Wzajemne `links` w obie strony. Tresc sekcji nietknieta — przenoszone
doslownie, dodany wylacznie naglowek H1 nowego runbooka.

kb-query, paperless-worker, planner-agent, ha-diag-agent, ollama-piha,
narty27, home-assistant, ha-mcp, job-gmail-header-backfill, job-mail-body-ingest.

Weryfikacja: dla kazdego pliku multizbior niepustych linii
(main + runbook) == oryginal z HEAD. Zero zgubionych, zero dodanych.

Recon szacowal 13 splitow service+runbook; faktycznie 2-typowych jest 10,
pozostale 5 (paperless, nextcloud, gokapi, fleet-prometheus, deploy-runner)
sa 3-typowe i ida osobno jako splity wielotypowe.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 16:58:46 +02:00

3.5 KiB

okf type visibility status updated links
0.1 service private active 2026-06-11
../runbooks/ha-diag-agent-runbook.md

ha-diag-agent

Per-host Home Assistant diagnostic agent. Polls HA REST API on a schedule, emits structured events to /opt/homelab/events/<node>/, and exposes an HTTP API for health checks and manual check triggers.

Follows the same event-pipeline pattern as node-agent: filesystem-first, no direct supervisor integration, events processed by the VPS observer.

Architecture

APScheduler (interval-based REST checks)
  ├─ HeartbeatCheck            → pings /api/, emits ha_websocket_dead on failure
  ├─ UnavailableEntitiesCheck  → entity unavailable > threshold
  ├─ SystemHealthCheck         → /api/system_health per-integration status
  ├─ AutomationFailuresCheck   → automation last-run error traces
  └─ UpdatesAvailableCheck     → pending HA/integration updates

WebSocketMonitor (persistent, long-running — Phase 4b)
  └─ Maintains a live WS subscription to state_changed events
     Any traffic = HA is alive. Watchdog fires ha_websocket_dead on
     silence > 5min or on disconnect. Emits ha_websocket_recovered
     when the connection is restored after a dead alert.

FastAPI (port 8087, internal only — no host port mapping)
  GET  /health             → liveness probe (includes ws_connected field)
  POST /trigger/<check>    → run a named check on demand

SQLite (/data/ha_diag.db)
  entity_baseline          → last-known entity states
  check_history            → per-check run log
  alerts_sent              → dedup gate for alert events

The WebSocketMonitor is the only persistent-connection component; all other checks are APScheduler intervals (stateless REST polls).

Event Types

Type Severity Trigger
ha_websocket_dead error WS disconnect, silence > 5min, or /api/ unreachable
ha_websocket_recovered info WS reconnected after a dead alert (clears incident)
ha_integration_failed error Integration in error state
ha_entity_unavailable_long warning Entity unavailable > threshold
ha_automation_failing warning Automation last run errored
ha_update_available info HA or integration update pending
ha_recorder_lag warning Recorder write lag > threshold
ha_system_health_degraded warning System health check failed

Event routing in supervisor (Phase 5) maps these to notify actions. ha_websocket_recovered should be routed to clear any active ha_websocket_dead incident.

Deployment model

The agent is deployed per-host but targets a potentially remote HA instance:

Node Agent runs on HA lives on HA URL
piha piha piha (localhost) http://localhost:8123
chelsty-infra chelsty-infra chelsty-ha (HAOS VM, separate machine) http://100.70.180.90:8123

chelsty-infra note: Home Assistant runs on chelsty-ha, a dedicated Home Assistant OS VM. chelsty-infra is the hypervisor but does not run HA itself. The agent on chelsty-infra reaches HA over the Tailscale network (100.70.180.90:8123). If chelsty-ha gets a new Tailscale IP, update HA_URL in /opt/homelab/config/ha-diag-agent/.env on chelsty-infra.

Optional YAML config

Place /opt/homelab/config/ha-diag-agent/ha-diag-agent.yaml on the node. Values there are defaults; env vars take priority.

ha_url: http://homeassistant.local:8123
location_tag: ken
check_interval: 60