homelab-codex-ws/services/home-assistant/DESIGN.md
oskar 0650eb857a feat(ha-mcp): read-only MCP server (faza 2a)
Own minimal MCP server exposing the live state of the HA instances in
services/home-assistant/instances.yaml to Claude Code over stdio — the
phase-2 "MCP read-only" gate in services/home-assistant/DESIGN.md.
Operator decision 2026-07-30: build our own rather than adopt hass-mcp,
so the tools reuse scripts/ha/lib/{ha_api,ha_ws}.py (one token-handling
story for the whole HA toolchain) and can answer from the repo and from
instances.yaml, which a generic server cannot.

Seven tools, all read-only, default instance `ken`: list_entities,
get_state, get_areas, find_entities_by_description, read_automation,
list_automations, instance_status.

Read-only by construction, not by policy: REST goes through ha_api.Client
(get/get_raw_text only — no POST method exists), WebSocket commands are
checked against a three-entry *_list allowlist before being sent, and
read_automation reads services/home-assistant/config/<instance>/ rather
than /api/config. Tests assert all three, including a grep guard that
fails if requests.post/call_service ever appears in the package. The
write path stays repo + scripts/ha/deploy.sh.

Details that follow from how this instance actually behaves:

- unavailable is never silent — every entity view carries unavailable +
  unavailable_since, every list a count. The 2026-07-23 audit traced ~15
  silently dead automations to conditions sitting on dead sensors.
- chelsty-ha (status: offline in instances.yaml) is answered from the
  file, never dialed — no 5s timeout for a known-offline LTE site.
- areas come from the WS registries (entity area_id > device area_id) with
  a storage-export fallback; area_source/area_note say which was used and
  what the offline export cannot resolve.
- PL->EN fuzzy matching, since the house is Polish and the entity_ids are
  transliterated English: "czujnik temperatury salon" ->
  sensor.thsalon_temperature, each hit explaining why it matched.
- 5s timeouts and errors returned as {"error": ...} inside a normal tool
  result — a missing token or an unreachable instance never crashes the
  server or hangs the agent.

Registered for Claude Code in the repo-root .mcp.json (new file) as `ha`,
via services/ha-mcp/run.sh (prefers the venv, falls back to system
python3). The mcp SDK lives in services/ha-mcp/.venv — rationale for venv
over --break-system-packages is in the README.

Tests: 42 offline (no network, no HA, no token) + a live read-only smoke
against ken — HA 2026.7.2, 1647 entities, 377 unavailable, 115
automations, 13 areas.
2026-07-30 16:47:27 +02:00

14 KiB

Home Assistant configs-as-code — design decisions

Status: phase 1 (partial). scripts/ha/deploy.sh implements the write path for the api adapter's automations/scripts/scenes scope (see "Deploy path" and "Sync model" below); dashboards/helpers and the docker-exec adapter have no write path yet. See docs/backlog.md for the tracking entry.

Phasing

Phase Scope
0 — Snapshot Import-only tooling (this skeleton). Pull /config + .storage from each instance into the repo, read-only. No deploy, no write path back to HA.
1 — Repo + CC Repo is the reviewable source of truth. Changes are authored in the repo (by a human or Claude Code) and pushed manually via the docker-exec / api adapters described below. Deploy has a hard drift-abort (see Sync model). Partially built: scripts/ha/deploy.sh covers the api adapter's automations/scripts/scenes scope; docker-exec deploy and dashboards/helpers writes are still open.
2 — MCP read-only Expose HA state (entities, areas, config) to agents via an MCP server in read-only mode, to let agents reason about the live instance without touching import/deploy paths. 2a done (operator decision 2026-07-30: own minimal MCP, not hass-mcp)services/ha-mcp/, seven read-only tools over stdio, registered for Claude Code in the repo-root .mcp.json. Read-only by construction: GET + allowlisted WS *_list reads only; automations are read from the repo, not the API. Open: exposing it to anything other than CC on the dev station (2b).
3 — Agents Agents propose changes (automations, scripts, scenes) through the same reviewable repo path used by humans; the human-in-the-loop approval flow from services/control-plane/ (pending → approved → executed) governs anything destructive. Telegram becomes a first-class interface alongside CC.

Each phase is a hard gate: no phase-N tooling depends on phase-(N+1) existing.

Scope

Full /config tree per instance, plus a curated export of .storage/* JSON files that are meaningfully version-controllable (registries, dashboards, helpers) — not secrets, tokens, or runtime databases. See .gitignore below for the exact exclusion list.

Canonical format

All YAML committed to the repo passes through one normalization function (scripts/ha/lib/normalize.py), used by every import path with no per-caller variation:

  • block style (no flow/inline collections)
  • fixed indent width 2
  • keys sorted (stable diffs)
  • unlimited line width (no wrapping)
  • UTF-8, no BOM

.storage/* files are JSON on disk in HA; on import they are parsed as JSON and re-emitted through the same YAML normalizer so the whole repo — config and storage-export alike — has one diff format.

Deploy path: adapter per instance

There is no single deploy mechanism — each HA instance gets an adapter behind a common interface (import.sh/eventual deploy.sh <instance>):

Instance Adapter Why
ken (RPi4, HAOS, LAN 192.168.31.7:8123) api Canonical home instance since the 2026-07-22 cutover (see Incident log). HAOS has no SSH access, so there is no docker exec/filesystem path — only the HA REST/websocket API is reachable. Full /config import is deferred until an alternative access path exists; for now the api adapter's import scope is limited to what the API exposes: automations, scripts, scenes, dashboards.
ken-legacy (piha, container homeassistant5) docker-exec over SSH (archive-only) Pre-migration container instance, superseded by ken at 31.7 (see Incident log) — same container filesystem access as the old ken entry (ssh oskar@piha "docker exec homeassistant5 ..."). Import only, for historical reference; never a deploy target.
chelsty-ha api Reachable over Tailscale at 100.70.180.90:8123 (confirmed working path — services/ha-diag-agent/DEPLOY.md already curls this for health checks). Config-as-code deploy will reuse the same reachability, calling the HA REST/websocket API rather than shelling into the container.

Open: a file adapter (direct bind-mount / SSH rsync to the config directory, bypassing docker exec) is worth revisiting once SSH access to the chelsty-ha VM itself is verified — see Open questions.

Sync model: bidirectional, asymmetric safety

  • Deploy (repo → instance): hard drift-abort. Before writing, the adapter re-imports the instance's current state and diffs it against the last-known-imported snapshot committed in the repo. Any unexpected difference aborts the deploy with a non-zero exit and a diff printed — never silently overwrites live drift. This applies unconditionally on PIHA (ken); untested/undecided whether it should be relaxed for chelsty-ha given its intermittent LTE uplink (see Open questions).
  • Import (instance → repo): the reverse direction is intentionally permissive. Running import.sh <instance> on a dev workstation pulls current state and, if it differs from the last commit, the operator commits it under a conventional message: drift(<instance>): <summary>. This is how out-of-band UI changes (made directly in the HA web UI) get captured back into git history instead of being silently overwritten by the next deploy.

Validation gate

check_config (HA's built-in config validator, invoked via the running container/instance — docker exec homeassistant5 python -m homeassistant --script check_config -c /config or the equivalent over the api adapter) is a mandatory gate before any deploy. A failing check_config blocks the deploy entirely; it is not a warning.

Change classification: reload vs. restart

Deploys default to the least disruptive mechanism:

  • reload — default for anything HA exposes a reload service for (automations, scripts, scenes, input_* helpers, template entities, etc.) via homeassistant.reload_config_entry / domain-specific *.reload services.
  • restart — only when the changed file requires it (e.g. configuration.yaml core changes, new integrations, .storage registry edits) and only when the deploy is invoked with an explicit --restart flag. No implicit restarts, ever — an unattended restart on chelsty-ha during a period of LTE unavailability would leave the site without automation until someone is physically present.

Split + normalization

On import, list-of-object YAML files are split one-object-per-file so diffs stay scoped to what actually changed:

  • automations.yamlconfig/<instance>/automations/<id>.yaml
  • scripts.yamlconfig/<instance>/scripts/<key>.yaml
  • scenes.yamlconfig/<instance>/scenes/<id>.yaml

Every other *.yaml file under /config is copied through the normalizer 1:1 (same relative path, same filename). Non-YAML files under /config (binaries, databases, secrets.yaml, logs — see .gitignore) are never copied into the repo.

Tokens

  • A dedicated deploy_agent HA user account (admin rights, local-only — never exposed through the public API/ingress) is created per instance, mirroring the existing diag_agent account pattern documented in services/ha-diag-agent/DEPLOY.md. Reusing diag_agent is explicitly rejected — deploy tooling and the diagnostic agent must be revocable independently.
  • Long-lived access tokens for deploy_agent live at ~/.config/ha-deploy/<instance>.token on PIHA (the control node where import/deploy tooling runs), chmod 600.
  • Tokens are never committed to the repo, never templated into env.example-style files, and never logged. Import scripts that need a token to hit /api/states fail soft (skip the fixtures step with a message) if the token file is absent, rather than aborting the whole import — see scripts/ha/import.sh.

Interface

  • Phase 0/1: Claude Code is the interface, with the operator's mobile CC client acting as the bridge for approvals made away from a desk.
  • Phase 3 (agents): Telegram joins CC as a first-class interface, reusing the existing Telegram bot / approval-queue pattern from services/control-plane/.

Incident log

2026-07-22 — two HA instances controlling the house in parallel

Symptom: automations firing twice from a single physical trigger — e.g. turn_on_led_nad_blatem_1 firing the same day from the same button press, mirror_on at 04:30 and gniazdka_w_lazience_on at 05:00 all firing on both instances.

How detected: comparing last_triggered from restore_state across the two instances showed identical automation IDs firing at the same times on both — the container on piha (homeassistant5, HA 2026.4.3, location_name KEN, mounted at /home/pi/homeassistant/config) never actually stopped running after the migration to the RPi4/HAOS instance at 192.168.31.7; it stayed alive and MQTT-connected, so both were independently reacting to the same physical events.

Root cause: instances.yaml had ken pointed at the piha container — that was the pre-migration instance, not the real one. The actual home instance had already moved to Home Assistant OS on a dedicated RPi4 (192.168.31.7:8123, ingress ha.kapala.org via NPM, confirmed HAOS via observer :4357, HACS installed, 118 automations), but the repo never followed the move.

Decision: 192.168.31.7 (HAOS/RPi4) is canonical ken. The piha container is renamed ken-legacy in instances.yaml, status: archived. Plan: archival import for historical reference → docker stop (not rm) → one week of observation → decide on docker rm. See docs/backlog.md for the ha-diag-agent re-pointing and wind-down follow-ups this incident generated.

Decyzje operatora po audycie 2026-07-23

Zobacz docs/audyt-automatyzacji-2026-07-23.md (sekcja "Do decyzji operatora", 17 punktów) — poniżej wyłącznie decyzje, które doprowadziły do zmian w fix-pack 1 (task/ha-fix-pack-1) albo świadomie do braku zmian. Reszta checklisty (baterie/re-pairing czujników, kalibracje TRV, xiaomi_miot, konsolidacja aliasów, higiena 4.x) zostaje otwarta w docs/backlog.md.

  • Pkt 6 (klima salon: sunset ubija też ręczne chłodzenie?) — decyzja: NIE. „Klima salon: wyłącz…" (1784804668795) ma teraz respektować input_boolean.klima_salon_auto — gałąź sunset/balkon działa tylko gdy klima_salon_auto = on (automat wyłącza tylko chłodzenie, którym sam zarządza). Wyjątek świadomy: trigger „klima_salon_auto przechodzi w off" nadal odpala suszenie parownika bez tego warunku (choose z osobną gałęzią) — to jest graceful shutdown samego automatu, nie efekt trybu auto, więc warunek auto = on by go zawsze ucinał (przełącznik właśnie zgasł). Zweryfikowano na żywym ken: helpery klimy (klima_salon_auto, _temp_docelowa, _tolerancja) istnieją (pkt 5 checklisty — obawa o martwe helpery z 1.1 nieaktualna).
  • Pkt 7 (enforcer sleep mode gasi światła cyklicznie całą noc) i pkt 9 (konsolidacja czterech nocnych wyłączników) — bez zmian w tym fix-packu. Oba wchłania przyszły projekt „architektura night_mode" (patrz docs/backlog.md) — punktowa łatka tu tylko dodałaby kolejny wariant do już przegęszczonego zestawu nakładających się automatyzacji (audyt 2.2).
  • Pkt 11 (OwnTracks: przywrócić czy skasować) i pkt 12 (Leave auto on: batch 02 — włączyć z powrotem?) — świadomie bez zmian; obie wymagają decyzji operatora spoza zakresu tego fix-packu (integracja/urlop), nie jednoliniowej poprawki.
  • On-leave alerty ruchu (1.2a) — świadomie bez zmian: martwe przez sprzęt (czujniki unavailable, pkt 1 checklisty), nie przez logikę automatyzacji; poprawka logiki niczego by nie naprawiła bez wymiany baterii/re-pairingu.

Zaimplementowane w tym fix-packu (bez decyzji operatora — jednoliniowe poprawki oczywistych bugów, audyt 2.6 i 2.3):

  • 1752086407230 „dry in Kuchnia": trigger moistnot_moist.
  • 1752085965483 „dry in Lazienka": literówka mesaagemessage.
  • 1700832676138 „…at 3am": time_pattern hours: '3' → punktowy trigger time: 03:00:00 (było 60 odpaleń/noc, teraz 1).
  • Pkt 8 (1764190493305 „after 15 minutes anyway" ma respektować ręczne „Disable AUTO off"?) — TAK: dodany warunek state: automation. turn_off_lights_in_kuchania = on (double-click wyłącza tę automatyzację; „anyway" teraz cofa się przed tym stanem zamiast go unieważniać po 15 min).

Konwencje automatyzacji

Ustalone po audycie 2026-07-23 (docs/audyt-automatyzacji-2026-07-23.md, sekcja 6 "Spójność stylistyczna", pkt 17 checklisty operatora — decyzja: TAK). Obowiązuje dla nowych automatyzacji od teraz; istniejące nie są migrowane hurtowo (patrz sekcja 4.1 audytu — konsolidacja/rename entity_id wymaga osobnej mapy referencji krzyżowych, to osobny task, nie efekt uboczny porządków).

  • entity_id, nie device automations. device_id/encja-UID (32-znakowy hex) są nieczytelne w YAML-u i kruche przy wymianie sprzętu — nowe urządzenie generuje nowy device_id, a automatyzacja umiera po cichu (dokładnie ten mechanizm ubił parę Tymka i ukrył rename occusalon, patrz audyt 4.1). Nowe automatyzacje używają action:/trigger:/condition: z entity_id: jawnym.
  • Alias po polsku, z prefiksem funkcjonalnym. Format "<Funkcja>: <opis>", np. "Klima salon: włącz chłodzenie i synchronizuj cel", "Pimirror: graceful shutdown przed odcięciem zasilania". Jeden język w aliasie (nie mieszanka PL/EN jak w automatyzacjach z audytu 6).
  • description: zawiera managed-by: repo. Odróżnia automatyzacje autorskie repo od tych z UI/migracji; przy okazji miejsce na kontekst (skąd odtworzone, jakie encje zaadaptowano — patrz przykład 1785164185794.yaml).

Open questions

  • What actually drives the phase-3 operational agent (a new agent process vs. extending an existing one in services/)?
  • Own minimal MCP server vs. adopting hass-mcp for phase 2 read-only access — decided 2026-07-30: own minimal server (services/ha-mcp/). Reusing scripts/ha/lib/{ha_api,ha_ws}.py keeps one token-handling and one read-only-by-construction story for the whole HA toolchain, and lets the tools answer from the repo (read_automation) and from instances.yaml (offline chelsty-ha) — neither of which a generic server knows about.
  • Whether SSH access to chelsty-ha itself (not just its HA API) is available/reliable enough to justify a file adapter there, which would let phase-1 tooling treat chelsty-ha more like ken for drift-checking purposes.