healthcheck_failed incidents fell through to redeploy, which is broken as wired (executor calls deploy-node.sh with arguments it ignores, at a path that does not exist in the container) — so 3376 healthcheck_failed events dead-ended with no working remediation (recon D14/D15). A container restart plausibly heals a failing healthcheck and rides the executor path that actually works; redeploy returns to the map once etap 2 fixes the executor. service_unhealthy / deployment_failed / missing_service stay on redeploy — theoretical until etap 2, kept so drift remains visible in pending actions (noted in comments). CLAUDE.md routing table updated to match; stale mqtt_unreachable example in the observer's trigger_type comment refreshed. Tests: trigger-type recognition and the end-to-end observer→supervisor reconcile test parametrized over both container_restart triggers, with an assertion that no redeploy action is also generated. Full control-plane suite: 147 passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| executor.py | ||
| index.html | ||
| liveness.py | ||
| operator_ui.py | ||
| supervisor.py | ||