homelab-codex-ws/kb/audits/redeploy-2026-08-05.md
oskar 7282a5e1d5 docs(recon): sciezka redeploy — fix jest w repo od 2026-08-03, nie jest wdrozony
Recon read-only zlecony pod teze "executor odpala deploy-node.sh z argumentami,
ktore skrypt ignoruje, na sciezce nieistniejacej w kontenerze". Teza byla
prawdziwa dla mastera do 2026-08-03; dzis nie jest. 79bfe8c + da151fc zastapily
to dispatchem do jobs/deploy-runner/ (systemd na hoscie wezla).

Realny stan: luka deployowa w trzech miejscach naraz — kontener executora
zbudowany 2026-07-22 (wciaz stary kod), /opt/homelab/actions/deploy/ nie istnieje
na VPS, deploy-runner nie jest zainstalowany na zadnym wezle. Zero akcji
kiedykolwiek osiagnelo stan terminalny (completed 0 / failed 0); jedyne dwa
action_result to reczne testy z 2026-07-23, nie remediacje z incydentu.

Zweryfikowane wzgledem fbf165f: healthcheck_failed idzie do container_restart,
nie do redeploy. Zywe incydenty w world state maja wylacznie trigger_type
containers_not_running / healthcheck_failed — czyli zaden nie generuje redeployu;
dzialaja tylko dryfy missing_service (2 pending).

Najostrzejszy problem projektowy: jedyny zywy emiter service_unhealthy
(node_agent.py:1104-1115) ma zahardkodowane service="control-plane", a
control-plane ma wlasne deploy-local.sh — deploy-service.sh:93-96 konczy sie
exit 3, wiec runner odmowi. Po wdrozeniu fixu redeploy sterowany incydentem
nadal nie wykona sie ani razu.

Ubocznie: w obrazie executora nie ma binarki ssh, wiec disk_cleanup
(executor.py:392-400) jest martwy tym samym defektem. Osobny task.

Dokument konczy sie sekcja FIX SHAPE — pieciopunktowa lista decyzji do podjecia,
bez rekomendacji.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:40:50 +02:00

668 lines
34 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
okf: "0.1"
type: audit
visibility: private
status: active
updated: 2026-08-05
as_of: 2026-08-05
links: []
---
# Recon — ścieżka remediacji `redeploy` (2026-08-05)
Read-only recon. Ground truth: `kb/subsystems/recon-multiagent.md` (A, D14, D15)
and `kb/audits/czujniki-2026-07-30.md` — both formerly at
`docs/architecture/RECON-*`, migrated into `kb/` by `00de810` / `9f77a72`.
Source read at `task/recon-redeploy` cut from master @ `4ecbdbb`. Runtime
evidence collected 2026-08-05 ~09:40 UTC from vps and piha over ssh; solaria is
the host this recon ran on (local reads).
---
## FINDING FIRST
**The break this recon was commissioned to characterize was already repaired in
the repo — two days after the task premise was written — and the repair has
never been deployed. `redeploy` is still dead in production, but for a
completely different reason than stated.**
The task brief describes master as: *"the executor calls deploy-node.sh with
args it ignores, on a path that doesn't exist inside the container."* That was
true of master until 2026-08-03. It is **not** true of master today:
- `79bfe8c` (2026-08-03) `feat(control-plane): host-side deploy runner — fix
broken redeploy path` and `da151fc` (2026-08-03) `fix(control-plane): redeploy
wykonywalny — dispatch do host-side deploy-runnera` replaced the
`deploy-node.sh` call with a dispatch-file handoff to a new host-level systemd
job, `jobs/deploy-runner/`.
- Master's `executor.py:125-144` no longer references `deploy-node.sh` at all;
`REPO_ROOT` survives only as a decorative env var (`executor.py:30-33`).
- The design decision, install runbook and open runtime steps are already
written up: `kb/decisions/deploy-runner-uzasadnienie.md`,
`kb/runbooks/deploy-runner-install.md`,
`kb/decisions/backlog-deploy-runner-instalacja.md`.
What is actually broken today is a **deploy gap in three places at once**:
| Layer | Repo (master @ `4ecbdbb`) | Runtime (2026-08-05) |
|---|---|---|
| executor code | dispatches to `actions/deploy/<node>/` | container built **2026-07-22**, still runs the `deploy-node.sh` code |
| `actions/deploy/` inbox | created by `_ensure_dirs` (`executor.py:79`) | **does not exist** on vps |
| deploy-runner on nodes | `jobs/deploy-runner/` + systemd units | **not installed** on vps, piha or solaria |
So the correct framing for the repair session is not "design a fix" — the fix is
designed, written, and 248-tests green. It is **"deploy the fix and prove it
end-to-end"**, plus resolving the four genuine design questions the repair left
open (see FIX SHAPE at the end), the sharpest of which is that **the only event
type that actually reaches a `redeploy` incident in live world state targets the
one service the new runner deliberately refuses to deploy.**
Secondary finding, unrelated to redeploy but discovered en route: the executor
container has **no `ssh` binary** (verified below), so `disk_cleanup`
the one action type that still shells out to ssh (`executor.py:392-400`) — is
dead by the same class of defect. It has never been exercised.
---
## A. GENERATION
### A1. `event_type → action_type` map as implemented on master @ `4ecbdbb`
Unchanged from recon D15 except for the `healthcheck_failed` move in `fbf165f`,
which the brief asked me to verify. **Verified: `healthcheck_failed` routes to
`container_restart`, not `redeploy`.**
`supervisor.py:52`:
```python
CONTAINER_RESTART_TRIGGERS = {"containers_not_running", "healthcheck_failed"}
```
**Path 1 — world-state drift** (`supervisor.py:334-395`, dispatch at
`:401-477`). The supervisor reads `world/{services,incidents}.json`, never raw
events. `trigger_type` comes from the incident the observer opened
(`supervisor.py:317-327` ← `observer.py:868`):
| Drift / incident `trigger_type` | Action | Where |
|---|---|---|
| `containers_not_running` | `container_restart` | `supervisor.py:52,410,423` |
| `healthcheck_failed` | `container_restart` | `supervisor.py:52` (added `fbf165f`, 2026-07-29) |
| `service_unhealthy` | `redeploy` | falls through `:447` |
| `deployment_failed` | `redeploy` | falls through `:447` |
| service absent from world state (`missing_service`, `trigger_type: None`) | `redeploy` | `supervisor.py:353-360``:447` |
| node `disk_pressure == high` | `disk_cleanup` | `supervisor.py:377-381,479+` |
Only services listed in `hosts/<node>/services.yaml` are considered
(`supervisor.py:350`); dormant nodes are skipped (`supervisor.py:342,378`).
**Path 2 — direct event-file routing** (`_process_ha_events`, `supervisor.py:395,
~630-660`) is unchanged and generates only `container_restart` (homeassistant,
shadow-downgraded to `alert_only`) and `alert_only`. **Path 2 never generates
`redeploy`.**
So: **three producers of `redeploy``service_unhealthy`, `deployment_failed`,
`missing_service` — all on path 1.** Section D9 shows only one of them fires in
practice, and it fires at the wrong target.
### A2. Fields on a `redeploy` action
Built at `supervisor.py:453-467`, verbatim:
```python
action = {
"action_id": action_id, # f"redeploy-{node}-{service}" (:413)
"timestamp": time.time(),
"type": "redeploy",
"node": node,
"service": service,
"risk_level": "guarded",
"confidence": 0.9,
"description": f"Redeploy {service} on {node} due to {drift['type']}",
"status": "pending",
"payload": {
"reason": drift["type"],
"svc_key": drift["svc_key"],
},
}
```
Written to `actions/pending/<action_id>.json` (`supervisor.py:469-471`).
**There is no `container_name`, no path, no compose reference, and no args
field.** Contrast `container_restart` (`supervisor.py:428-446`), which carries
`container_name` resolved via `_get_container_name(service)`. A redeploy action
carries exactly two addressing facts: `node` and `service`. Everything else —
which compose file, which override, which repo — is resolved on the node at
execution time from `service` alone.
Live example, the one non-chelsty pending redeploy on vps:
```json
{ "action_id": "redeploy-vps-gokapi", "timestamp": 1783615889.1762547,
"type": "redeploy", "node": "vps", "service": "gokapi",
"risk_level": "guarded", "confidence": 0.9,
"description": "Redeploy gokapi on vps due to missing_service",
"status": "pending",
"payload": { "reason": "missing_service", "svc_key": "vps/gokapi" } }
```
---
## B. EXECUTION — why it's dead
### B3. The redeploy handler
**Two answers, because repo and runtime disagree.**
**(a) Master @ `4ecbdbb``executor.py:125-144`.** No command is built at all.
The executor writes a dispatch file and returns:
```python
if action_type == "redeploy":
if not node or not service:
success = False
error_msg = (f"redeploy requires both node and service "
f"(node={node!r}, service={service!r})")
else:
self._dispatch_redeploy(action_id, node, service)
return # stays in running/ until action_result
```
`_dispatch_redeploy` (`executor.py:238-267`) writes
`actions/deploy/<node>/<action_id>.json` with `{action_id, type, node, service,
dispatched_at}`. Deliberately a **different inbox** from `container_restart`'s
`actions/dispatch/<node>/`, because node-agent deletes and failure-reports
anything in its own inbox that is not `container_restart`
(`executor.py:24-29` ← `node_agent.py:108,1011-1018`).
**(b) The executor actually running on vps** — image created
`2026-07-22T16:13:29Z`, i.e. 12 days before the fix. Its `/app/src/executor.py`
lines 106-119, read out of the live container:
```python
if action_type == "redeploy":
# Full service redeploy via the repo deploy script
cmd = [
str(REPO_ROOT / "scripts" / "deploy" / "deploy-node.sh"),
node,
service
]
logger.info(f"Running command: {' '.join(cmd)}")
result = subprocess.run(cmd, capture_output=True, text=True, cwd=str(REPO_ROOT))
```
`REPO_ROOT` = `/repo` (`os.getenv("REPO_ROOT", "/repo")`, running line 24;
master `executor.py:33`). Synchronous: `returncode == 0` → completed, else
failed with `stderr or stdout`.
### B4. Working directory / repo path, and does it exist in the container
The repo **is** reachable inside the executor container — the claim "a path that
doesn't exist inside the container" needs one level of precision. Two different
paths are involved and only the second is missing.
`docker inspect control-plane-executor` (vps):
```
Image=sha256:7603f09d… Created=2026-07-22T16:13:29.905995261Z
/opt/homelab -> /opt/homelab (rw)
/home/oskar/homelab-codex-ws -> /repo (rw)
/var/run/docker.sock -> /var/run/docker.sock (rw)
REPO_ROOT=/repo RUNTIME_PATH=/opt/homelab
```
Inside the container:
```
HOME=/home/homelab
/repo/scripts/deploy/deploy-node.sh ← EXISTS
ls: cannot access '/home/homelab/homelab-codex-ws': No such file or directory
git: MISSING docker: MISSING rsync: MISSING ssh: MISSING
bash: /usr/bin/bash python3: /usr/local/bin/python3 flock: /usr/bin/flock
```
So `cwd=/repo` is valid and the script is found and executed. It then fails on
**its own** internal path: `deploy-node.sh:8` hardcodes
`REPO_PATH="${HOME}/homelab-codex-ws"` = `/home/homelab/homelab-codex-ws`, which
does not exist, and `deploy-node.sh:18-21` exits 1.
Reproduced read-only inside the live container with the exact argv the running
executor would use:
```
rc= 1
stdout= --- Starting Deployment on 1CF38E89F150 ---
Error: Repository not found at /home/homelab/homelab-codex-ws
```
Note the hostname: `1CF38E89F150`, the container ID. Even with the repo present,
`deploy-node.sh:27-57` resolves the host directory by matching `os_hostname`
against `hosts/*/host.yaml` and would fail with *"No host directory found for
1cf38e89f150"* — the executor would have deployed **its own container's**
service set, never the action's `node`. And behind that, two more hard stops:
`git pull` (`deploy-node.sh:25`) and the compose invocation both need binaries
the image does not have.
**Four independent blockers, in the order they would be hit:** missing
`${HOME}/homelab-codex-ws` → hostname resolves to a container ID → no `git` → no
`docker`. The image mount is *not* one of them.
Note also `../..:/repo:ro` in `services/control-plane/docker-compose.yml:36,55,86`
— master mounts the repo read-only. The running container has it **rw**, another
marker that the container predates current master.
### B5. `deploy-node.sh` arg parsing, and the call chain
**`deploy-node.sh` has no argument parsing whatsoever.** There is no `getopts`,
no `case "$1"`, no `$1`/`$2` reference anywhere in its 115 lines. It is a
zero-argument script: it derives everything from `${HOME}` and `hostname`
(`:8,12,13`), and deploys the **whole service set** of whatever host it runs on
(`:60-78` reads `services.txt` / `services.yaml`, `:95-113` loops over all of
them). The `node` and `service` the old executor passed were silently discarded
by the shell.
Per-service deploy exists, but in a **different script**:
`scripts/deploy/deploy-service.sh`, extracted from `deploy-node.sh` by the
2026-08-03 fix. It takes `--repo`, `--host-dir`, `--service` plus optional
`--build-if-needed` / `--force-recreate` / `--remove-orphans`
(`deploy-service.sh:51-62`), validates the service name against
`^[a-z0-9][a-z0-9._-]{0,63}$` (`:72`), and is the single shared compose
invocation for both callers — deliberately, because a project-name mismatch plus
`--remove-orphans` is what wiped the control-plane on vps on 2026-06-25
(`deploy-service.sh:6-13`, `deploy-node.sh:86-92`).
**Call chains, with the host and container each link runs on:**
*Human deploy (SATURN-side dispatcher — confirms the solaria-gid-fix session's
model):*
```
operator @ SATURN: scripts/deploy/deploy.sh <target> [host shell, saturn]
preflight: must be on master, clean tree (deploy.sh:69-82)
→ ssh oskar@<target> 'cd ~/homelab-codex-ws && git pull
&& ./scripts/deploy/deploy-node.sh' (deploy.sh:199-202)
→ deploy-node.sh [host shell, target node]
→ deploy-service.sh --build-if-needed --remove-orphans × every service
(deploy-node.sh:100-105)
(target == control-plane takes a different branch:
deploy-control-plane.sh --ssh, deploy.sh:194-196)
```
*Agent redeploy, master @ `4ecbdbb`:*
```
supervisor [container control-plane-supervisor, vps]
→ /opt/homelab/actions/pending/redeploy-<node>-<svc>.json
operator (Telegram / operator-ui) → approved/
executor [container control-plane-executor, vps]
→ /opt/homelab/actions/deploy/<node>/<action_id>.json (executor.py:238-267)
deploy-runner [HOST systemd oneshot on <node>, NOT a container]
→ rsync-pull inbox from vps (deploy-runner.sh pull_actions)
→ python3 action.py validate (5 rules, action.py:78-125)
→ deploy-service.sh --repo … --host-dir … --service … --force-recreate
→ action.py emit-result → evt-<node>-<ts>-action_result-<svc>.json
→ rsync-push to vps events/
executor._reconcile_running_actions() (executor.py:269-318)
→ completed/ or failed/ (timeout REDEPLOY_TIMEOUT_SECS=900, executor.py:55)
```
The runner runs on the **host**, not in a container, on purpose: compose then
resolves relative bind mounts and the project name exactly as a human deploy
does (`deploy-runner.sh:26-31`). It also **never runs `git pull`** — a redeploy
reconciles the node to the checkout it already has; shipping code stays a human
`deploy.sh` (`deploy-runner.sh:32-33`).
### B6. Concrete failure mode today
**Nothing fires, and nothing ever has. No redeploy has ever been attempted,
approved, or executed — fleet-wide, ever.**
Action pool on vps, 2026-08-05:
```
pending 18 approved 0 running 0 completed 0 failed 0 rejected 0 cancelled 18
dispatch 1 deploy: ls: cannot access '/opt/homelab/actions/deploy': No such file or directory
```
`completed 0 / failed 0` is the whole story: **not one action has ever reached a
terminal state.** 18 redeploy actions exist on disk (16 cancelled, 2 pending);
none was ever approved, so the broken handler was never entered.
The two live ones:
```
/opt/homelab/actions/pending/redeploy-vps-gokapi.json reason: missing_service
/opt/homelab/actions/pending/redeploy-chelsty-infra-ha-diag-agent.json reason: missing_service
```
The chelsty one targets a node dead since 2026-06-01 and is already flagged for
manual removal (`kb/subsystems/recon-multiagent.md`, "Runbook — stale chelsty
action").
The executor's **entire log since 2026-07-22** is 144 lines and contains three
action executions, all hand-injected tests, zero supervisor-generated actions,
zero redeploys, and nothing at all after 2026-07-23:
```
2026-07-22 16:13:41 - Starting executor loop
2026-07-22 16:32:41 - Executing action: test-restart-piha-node_exporter-2
2026-07-22 16:32:41 - Dispatched container_restart … to node-agent on piha
2026-07-22 16:37:42 - ERROR - Action … timed out: Timed out after 300s waiting for node-agent on 'piha'
2026-07-23 11:54:05 - Executing action: test-e2e-1784807638
2026-07-23 11:54:05 - ERROR - Failed to move test-e2e-1784807638 to running: Expecting value: line 1 column 38
…(same pair repeating every 10 s for 11 minutes — malformed JSON never leaves approved/)…
2026-07-23 12:05:25 - Executing action: test-e2e-b
2026-07-23 12:05:25 - Dispatched container_restart test-e2e-b (container=node_exporter) to node-agent on piha
2026-07-23 12:05:56 - Action test-e2e-b completed
```
**Why nothing fires**, in the order the funnel closes:
1. `redeploy` needs `service_unhealthy`, `deployment_failed`, or
`missing_service`. Post-`fbf165f`, the two high-volume signals
(`containers_not_running` 2,490; `healthcheck_failed` 3,376 on piha) both go
to `container_restart`. Live incident census on vps confirms it —
`world/incidents.json` holds 3 incidents total, `trigger_type:
{containers_not_running: 2, healthcheck_failed: 1}`. **Zero
redeploy-producing incidents are open.**
2. `missing_service` still fires (it needs no incident at all), which is exactly
what the 2 pending redeploys are. But it is dedup-suppressed forever by its
own pending file (`supervisor.py:418-421`) — `redeploy-vps-gokapi` has held
its ID since 2026-07-19.
3. Nothing gets approved. `approved 0`, and the executor reads only `approved/`
(`executor.py:94-95`).
4. Even on approval the running executor would fail as in B4 — but this has
never actually happened, so there is **no failed/ artefact, no stack trace,
and no "Running command: …deploy-node.sh" log line to point at.** The
evidence for the break is the reproduction in B4, not a production incident.
---
## C. THE HEALTHY PATH FOR COMPARISON
### C7. `container_restart` traced end to end
Correction to the premise: the two `action_result` events of 2026-07-23 are
**manual e2e tests, not real remediations**. Payloads:
```json
{"action_id": "test-restart-piha-node_exporter-2", "success": true, "error": "", "node": "piha"}
{"action_id": "test-e2e-b", "success": true, "error": "", "node": "piha"}
```
Both hand-crafted IDs. A supervisor-generated ID would be
`container-restart-piha-node_exporter` (`supervisor.py:411`). And only **one**
of the two round-tripped cleanly: `test-restart-piha-node_exporter-2` was
dispatched 2026-07-22 16:32 UTC, timed out at 300 s, and its `action_result`
only landed 2026-07-23 11:22 UTC — ~19 h later, after the executor had already
filed it failed. `test-e2e-b` completed in 31 s. **So the reference path has
exactly one clean end-to-end demonstration, and one latency failure, and has
never run from a real incident.**
Path, verbatim:
1. `executor._execute_action` moves `approved/<id>.json``running/`
(`executor.py:104-112`), then `executor.py:146-157`:
`container_name = data.get("container_name") or service`
`_dispatch_container_restart(...)``return` (stays in `running/`).
2. `_dispatch_container_restart` (`executor.py:209-236`) writes
`actions/dispatch/<node>/<action_id>.json` = `{action_id, type, node,
service, container_name, dispatched_at}`.
3. node-agent on the node rsync-pulls its own subdir with
`--remove-source-files` (`node_agent.py:919-960`); on vps it reads the path
directly (`:933`).
4. `_execute_dispatched_action` (`node_agent.py:981-1051`) applies four gates —
idempotency (`:999`), node scoping (`:1003`), type whitelist
`ALLOWED_DISPATCH_ACTION_TYPES = {"container_restart"}` (`:108, :1011`),
self-restart guard `{"node-agent"}` (`:114, :1027`) — then
`self.docker_client.containers.get(container_name).restart()` (`:1043-1044`).
5. `_report_action_result` (`node_agent.py:1053+`) emits an `action_result`
event; the existing rsync ships it to vps.
6. `executor._reconcile_running_actions` (`:269-318`) matches it via
`_find_action_result` (`:320-356`) and moves the action to
`completed/`/`failed/`, or fails it after `ACTION_TIMEOUT_SECS=300`.
### C7b. Structural differences between the two handlers
On master they are now **deliberately near-identical** — steps 1, 2, 5 and 6 are
the same code, `DISPATCHED_ACTION_TYPES = {"container_restart", "redeploy"}`
(`executor.py:60`), and `_reconcile_running_actions` handles both with only the
timeout differing (`executor.py:299-302`). What remains different:
| | `container_restart` | `redeploy` |
|---|---|---|
| Inbox | `actions/dispatch/<node>/` | `actions/deploy/<node>/` (must not share — `executor.py:24-29`) |
| Executor on the node | **node-agent**, a container, own docker socket | **deploy-runner**, host systemd oneshot + timer, no container |
| Deployed how | `deploy.sh` / compose, already everywhere | manual per-node systemd install, **nowhere yet** |
| Poll cadence | `CHECK_INTERVAL=60` | `OnUnitActiveSec=60s` (matched on purpose) |
| Timeout | `ACTION_TIMEOUT_SECS=300` | `REDEPLOY_TIMEOUT_SECS=900` |
| Payload | carries `container_name` | carries only `service` |
| Validation | 4 gates in `node_agent.py` | 5 gates in `action.py:78-125`, incl. *service must be in `hosts/<node>/services.yaml`* and *compose file must exist* |
| Idempotency | `_already_processed` marker | `state/processed-deploy-actions/<id>.done` (`deploy-runner.sh:57,151-155`) |
| Concurrency | none needed | `flock` on `state/deploy-runner.lock` |
| Can act on node-agent itself | **no** (self-guard) | **yes** — deliberate; solaria's node-agent has been docker-blind for weeks (`deploy-runner.sh:26-31`) |
| Can act on control-plane / stability-agent | yes | **no**`deploy-service.sh:93-96` exits 3, runner reports failure with an explanation |
The last row is the important asymmetry: see D9.
### C8. How the executor reaches nodes
**It doesn't, and it cannot.** Verified inside the live container: `ssh: MISSING`,
`rsync: MISSING`, `git: MISSING`, `docker: MISSING`. There is no ssh client, no
key, no known_hosts. This is by design — `executor.py:200-207` and
`kb/phases/backlog.md` "Remediacja floty bez SSH": the VPS never initiates a
connection to a node. Every remote action is a **file the node comes and
collects**, over the ssh/rsync channel node-agent (and now deploy-runner)
already owns in the other direction. The executor's only outputs are writes into
the shared `/opt/homelab` mount.
Two consequences:
- **`disk_cleanup` is dead by the same defect.** `_execute_disk_cleanup`
(`executor.py:358-407`) builds `["ssh", *SSH_OPTIONS, f"{SSH_USER}@{node}",
…]` and calls `subprocess.run`. With no `ssh` binary this raises
`FileNotFoundError`, caught by `executor.py:175-177`, and the action fails with
`[Errno 2] No such file or directory: 'ssh'`. Never observed because
`failed/` is empty — no `disk_cleanup` was ever approved either. Out of scope
here; flagged as a follow-up.
- **Solaria's self-ssh problem is irrelevant to the executor.** `deploy.sh` runs
on SATURN and ssh's *to* the target, which is why it cannot deploy solaria
from solaria. The executor never ssh's anywhere; the connection is always
node → vps, initiated by the node. Solaria's node-agent already ships events
to vps over that channel daily, so the transport is proven. **Solaria's real
blocker is different and unrelated:** its node-agent has no docker socket
access (`group_add: "996"` fixed in repo @ `ddae57c`, still not deployed —
running container shows `GroupAdd=[999]`), which breaks `container_restart`
there. `redeploy` is unaffected, because deploy-runner runs on the host and
never touches node-agent.
---
## D. WHAT REDEPLOY IS FOR
### D9. When redeploy *should* fire vs `container_restart`
The semantic split the brief proposes is the one the code already encodes
(`supervisor.py:36-52, 423-426, 447-452`): **`container_restart` = the container
exists and a restart plausibly heals it; `redeploy` = the container is absent,
or its image/compose/env no longer matches the repo, so it must be re-created
from the manifest.** `deploy-service.sh --force-recreate` is exactly that and
nothing more (`deploy-service.sh:111-130`: `docker compose -f … [-f override]
[--env-file] up -d --force-recreate`, **no `--build`**, `deploy-service.sh:113-118`
— a redeploy is a reconcile, not a code ship, and building on the 4 GiB vps
mid-incident is an OOM risk).
Candidate event types for the redeploy side, and whether each actually reaches
an incident today:
| Candidate | Opens an incident? | Live volume | Verdict |
|---|---|---|---|
| `missing_service` (drift, not an event) | **n/a — needs no incident**; `supervisor.py:353-360` raises it purely from desired-vs-actual | 2 pending actions right now | **The only redeploy producer that actually fires.** Semantically the perfect fit: the container is gone, a restart is impossible. |
| `service_unhealthy` | **yes**`observer.py:771,792` | 34 events, vps only | **Fires, but is a trap — see below.** |
| `deployment_failed` | **yes**`observer.py:838-842` `_handle_deployment_failure` | **0 events fleet-wide** | Only emitter is `scripts/lib/log.sh:35,43` (deploy-time shell logging). Never reaches the vps event store. Dead in practice. |
| `containers_not_running` | yes, `observer.py:771` | 2,490 piha / 3 solaria / 1 vps | Correctly on `container_restart` — container exists. |
| `healthcheck_failed` | yes, `observer.py:771` | 3,376 piha | Correctly on `container_restart` since `fbf165f`. Escalating repeat failures to redeploy is a candidate, but needs a retry policy that does not exist. |
| `container_restarting` / `container_state_unexpected` | **no** — observational only, `observer.py:793-814` | 2 solaria | Deliberately not remediated. |
| stability-agent's `containers_not_running` | **no**`service=None` fails the `if service and service != "all"` guard at `observer.py:747`, so the incident branch is never reached | — | Known-broken producer (recon D15, `kb/audits/czujniki-2026-07-30.md`). |
**The trap.** `service_unhealthy` has exactly one live emitter in the whole
fleet: `node_agent.py:1104-1115`, which probes
`http://localhost:18180/summary` and hardcodes `service="control-plane"`. So
every `service_unhealthy` incident that could ever produce a `redeploy` targets
`vps/control-plane` — and `control-plane` is one of only two services with its
own `deploy-local.sh` (`services/control-plane/`, `services/stability-agent/`),
which `deploy-service.sh:93-96` refuses with exit 3, which the runner turns into
a **failed** action: *"control-plane owns its deploy path
(services/control-plane/deploy-local.sh) — needs an operator deploy, not an
automated redeploy"* (`deploy-runner.sh:179-181`). That refusal is correct — a
containerised control plane redeploying itself mid-incident is not something to
automate — but it means **the fixed redeploy path, once deployed, will still
produce zero successful incident-driven redeploys.** Only `missing_service`
will ever succeed.
Also worth stating plainly: **`redeploy` cannot repair a `service_unhealthy`
caused by bad code**, because the path deliberately does not `git pull`
(`deploy-runner.sh:32-33`) and deliberately does not `--build`. It repairs
*drift* — a container removed, a stale env file, a config change already in the
node's checkout — not *regression*.
### D10. Preconditions for a working redeploy
Per link in the chain, on master's architecture. Current status verified
2026-08-05.
**On the vps (control plane):**
| # | Precondition | Where | Status |
|---|---|---|---|
| 1 | executor running post-`da151fc` code | `executor.py:125-144` | ❌ container built 2026-07-22, runs `deploy-node.sh` code |
| 2 | `/opt/homelab/actions/deploy/<node>/` writable | `executor.py:79` (`_ensure_dirs`) | ❌ dir absent — created automatically by the new executor |
| 3 | repo checkout current | `/home/oskar/homelab-codex-ws` | ⚠️ at `efd8d2f`, one commit behind master `4ecbdbb`; **has** the fix (`jobs/deploy-runner/` dated Aug 3 17:28) |
| 4 | *not* needed: ssh, git, docker CLI in the executor image | — | ✅ by design |
**On each target node:**
| # | Precondition | Where | vps | piha | solaria |
|---|---|---|---|---|---|
| 5 | repo checkout at `REPO_PATH` | `deploy-runner.sh:73` | ✅ `efd8d2f` | ✅ `4ecbdbb` | (main checkout, not inspected) |
| 6 | `hosts/<node>/` exists in that checkout | `deploy-runner.sh:74` | ✅ | ✅ | ✅ |
| 7 | `scripts/deploy/deploy-service.sh` present + executable | `deploy-runner.sh:75-76` | ✅ Aug 3 | ✅ Aug 4 | — |
| 8 | `jobs/deploy-runner/` systemd unit + timer installed & enabled | `kb/runbooks/deploy-runner-install.md` | ❌ | ❌ | ❌ |
| 9 | `/opt/homelab/config/deploy-runner/env` with `NODE_NAME` (**empty `VPS_EVENTS_HOST` on vps**) | `env.example` | ❌ | ❌ | ❌ |
| 10 | service user in `docker` group | runbook | — | — | ✅ (`oskar` ∈ `docker` gid 996) |
| 11 | `python3` + PyYAML, `rsync`, `flock` | `action.py:60-64`, `deploy-runner.sh:70-76` | — | — | — |
| 12 | ssh key reaching vps (remote nodes only) | `deploy-runner.sh:62-66` | n/a | ✅ (node-agent already ships events) | ✅ |
**Per-action, at execution time** (`action.py:78-125` — all five must hold or the
runner reports `failed` without deploying):
| # | Rule | Note |
|---|---|---|
| 13 | `type == "redeploy"` | |
| 14 | `action["node"] == NODE_NAME` | defense in depth; inbox is already node-scoped |
| 15 | `action_id` / `service` match strict name patterns | no path traversal |
| 16 | **service listed in `hosts/<node>/services.yaml`** | repo desired state is the authority; a stale dispatch cannot pull an arbitrary stack |
| 17 | `services/<service>/docker-compose.yml` exists in the repo | |
| 18 | *(not a rule, an outcome)* service must **not** have `services/<svc>/deploy-local.sh` | else exit 3 → failed-with-explanation |
Checked against the one actionable pending action, `redeploy-vps-gokapi`: rule
16 ✅ (`hosts/vps/services.yaml:63`), rule 17 ✅ (`services/gokapi/docker-compose.yml`),
rule 18 ✅ (no `deploy-local.sh`). **`gokapi` is a valid end-to-end target the
moment the runner is installed on vps.**
Also required for the loop to *close*: the result event must land in
`/opt/homelab/events/<node>/` on the vps in `node_agent.emit_event()`'s exact
format, because `executor._find_action_result` (`:320-356`) globs
`evt-*-action_result-*.json` and parses the unix timestamp out of the
**filename**. `action.py:emit_result` reproduces that byte-for-byte and adds
`payload.source = "deploy-runner"`.
---
## FIX SHAPE — DECISIONS NEEDED
The transport and script layers are decided and written. What is left is one
sequencing decision and four genuine design choices the 2026-08-03 work left
open.
**1. Deploy vs re-derive.** The repair is `79bfe8c`/`da151fc` on master, 248
tests green, unexercised at runtime.
- *Deploy as written* — three runtime steps (`deploy.sh control-plane`; install
`jobs/deploy-runner/` on vps + piha + solaria per
`kb/runbooks/deploy-runner-install.md`; e2e on a benign piha service, then
`redeploy-vps-gokapi`). Fastest to a working loop; accepts a design nobody has
yet run.
- *Re-derive first* — treat the shipped design as a proposal, resolve 25 below,
then deploy once. Slower; avoids installing systemd units on three hosts twice.
Note that 25 are answerable either way, but their answers are cheaper to apply
**before** the units are on three nodes.
**2. Where the deploy actually runs: node-side runner vs saturn-side dispatcher.**
The shipped design is node-side (`deploy-runner` on each node, executor never
connects out).
- *Node-side (as shipped)*: preserves "remediacja floty bez SSH"; works on
solaria despite its broken node-agent socket and despite the
self-ssh problem; every node needs a host-level systemd install and its own
repo checkout, and the install is not covered by `deploy.sh` — it drifts
silently.
- *Executor → ssh SATURN → `deploy.sh <node>`*: reuses the human path exactly,
one install point instead of N, and inherits `deploy.sh`'s preflight gates.
But it inverts the no-SSH decision, puts a key to SATURN inside a
vps container, makes SATURN a runtime dependency of remediation (it is a
workstation, not a server), and `deploy.sh` is whole-node only — it cannot
express "redeploy just gokapi".
- *Executor → ssh direct to node → `deploy-service.sh`*: fewest moving parts, no
runner, no timer, no rsync. Requires an ssh client + fleet key in the executor
container — the exact thing the architecture forbids — and reintroduces the
self-ssh and reachability problems the pull model sidesteps.
**3. Per-service vs whole-node redeploy.** Shipped: per-service, and the action
payload carries no node-wide option.
- *Per-service*: minimal blast radius; matches the action's semantics
(`missing_service` names one service); a genuinely node-wide failure needs N
actions and N approvals.
- *Whole-node (i.e. `deploy-node.sh`, what the old executor accidentally
attempted)*: one action heals a node; but it touches every service including
healthy ones, runs `--build-if-needed`, and `--remove-orphans` is in play — the
2026-06-25 control-plane wipe is precedent for how that ends.
**4. How args flow — how much the action carries vs how much the node resolves.**
Shipped: the action carries `{node, service}` only; everything else (compose
path, override, env-file, force-recreate) is resolved on the node by
`deploy-service.sh` from repo state.
- *Thin action (as shipped)*: the repo is the single authority; an old or forged
dispatch cannot smuggle a path or a flag; but the operator approving the action
cannot see or influence what will actually run, and `--build`/`--pull` are not
expressible even when they are what is needed.
- *Fat action*: supervisor stamps compose path, override, flags into the payload
— visible at approval time, tunable per incident; every field becomes attack
surface the runner must re-validate, and payloads can go stale against the
node's checkout.
**5. What happens to `service_unhealthy` → `control-plane` (D9).** Once the
runner is installed, the only incident-driven redeploy producer will target the
one service the runner refuses.
- *Leave it*: the failed action with its explanatory message **is** the alert;
costs one failed action per control-plane outage and burns the dedup ID.
- *Downgrade to `alert_only` at the supervisor*: honest — nothing automated can
fix it — but drops the drift out of the remediation view.
- *Give control-plane a self-redeploy path* (analogous to
`deploy-control-plane.sh`): closes the loop, and lets a broken control plane
attempt to repair itself with its own hands. Highest risk on the list.
- *Widen `service_unhealthy` beyond the hardcoded control-plane probe*
(`node_agent.py:1104-1115`): would give redeploy a real, non-degenerate input
set for the first time — but that is a monitoring change, not a redeploy
change, and it interacts with the `healthcheck_failed` routing decided in
`fbf165f`.
**Out of scope here, found en route:** the executor's `disk_cleanup` handler
shells out to `ssh`, which does not exist in the image (C8) — same class of
defect as the original redeploy break, never observed because no `disk_cleanup`
was ever approved. Worth its own task.