- shared aiohttp ClientSession in HAClient (Phase 1 Flag #2 fixed): make_session() factory, session injected at startup, closed on shutdown - Check.run() → list[CheckResult]: clean multi-event interface - first real diagnostic check: entity unavailable > 24h (INSERT OR IGNORE baseline preserves first-seen timestamp) - root cause grouping: emit ha_integration_failed instead of N entity events when ≥50% of integration's entities are unavailable (≥3 min) - alert deduplication via SQLite cooldown window (default 6h) - recovery clears baseline + dedup for immediate re-alert - configurable thresholds: duration, integration %, cooldown - 38 unit tests + 7 integration tests (42 pass, 3 skip w/o live HA) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| src/ha_diag | ||
| tests | ||
| docker-compose.yml | ||
| Dockerfile | ||
| env.example | ||
| healthcheck.sh | ||
| pyproject.toml | ||
| README.md | ||
| service.yaml | ||
ha-diag-agent
Per-host Home Assistant diagnostic agent. Polls HA REST API on a schedule,
emits structured events to /opt/homelab/events/<node>/, and exposes an
HTTP API for health checks and manual check triggers.
Follows the same event-pipeline pattern as node-agent: filesystem-first,
no direct supervisor integration, events processed by the VPS observer.
Architecture
APScheduler (every CHECK_INTERVAL s)
└─ HeartbeatCheck → pings /api/, emits ha_websocket_dead on failure
[Phase 3: EntityUnavailableCheck, SystemHealthCheck, UpdateCheck, ...]
FastAPI (port 8087)
GET /health → liveness probe
POST /trigger/<check> → run a named check on demand
SQLite (/data/ha_diag.db)
entity_baseline → last-known entity states
check_history → per-check run log
alerts_sent → dedup gate for alert events
Event Types
| Type | Severity | Trigger |
|---|---|---|
ha_websocket_dead |
error | HA /api/ unreachable |
ha_integration_failed |
error | Integration in error state |
ha_entity_unavailable_long |
warning | Entity unavailable > threshold |
ha_automation_failing |
warning | Automation last run errored |
ha_update_available |
info | HA or integration update pending |
ha_recorder_lag |
warning | Recorder write lag > threshold |
ha_system_health_degraded |
warning | System health check failed |
Event routing in supervisor (Phase 5) maps these to notify actions.
Deployment model
The agent is deployed per-host but targets a potentially remote HA instance:
| Node | Agent runs on | HA lives on | HA URL |
|---|---|---|---|
| piha | piha | piha (localhost) | http://localhost:8123 |
| chelsty-infra | chelsty-infra | chelsty-ha (HAOS VM, separate machine) | http://100.70.180.90:8123 |
chelsty-infra note: Home Assistant runs on chelsty-ha, a dedicated Home Assistant
OS VM. chelsty-infra is the hypervisor but does not run HA itself. The agent on
chelsty-infra reaches HA over the Tailscale network (100.70.180.90:8123). If chelsty-ha
gets a new Tailscale IP, update HA_URL in /opt/homelab/config/ha-diag-agent/.env on
chelsty-infra.
Deployment
# 1. Create config on target node
ssh oskar@<node-ip>
mkdir -p /opt/homelab/config/ha-diag-agent /var/lib/ha-diag-agent
cat > /opt/homelab/config/ha-diag-agent/.env << 'EOF'
HA_URL=http://homeassistant.local:8123 # or http://100.70.180.90:8123 for chelsty-infra
HA_TOKEN=<long-lived-token>
NODE_NAME=piha # or chelsty-infra
LOCATION_TAG=ken # or chelsty
CHECK_INTERVAL=60
EOF
# 2. Deploy
scripts/deploy/deploy.sh --service ha-diag-agent
# 3. Verify
docker ps --filter name=ha-diag-agent
curl http://localhost:8087/health
chelsty-infra note
chelsty-infra runs docker-compose v1 (1.29.2). Use docker-compose (hyphenated):
docker-compose -f docker-compose.yml up -d --build
HA long-lived token
In HA UI: Profile → Long-Lived Access Tokens → Create token.
Running Tests
cd services/ha-diag-agent
pip install -e ".[dev]"
pytest tests/ -v
Optional YAML config
Place /opt/homelab/config/ha-diag-agent/ha-diag-agent.yaml on the node.
Values there are defaults; env vars take priority.
ha_url: http://homeassistant.local:8123
location_tag: ken
check_interval: 60