Recon read-only zlecony pod teze "executor odpala deploy-node.sh z argumentami, ktore skrypt ignoruje, na sciezce nieistniejacej w kontenerze". Teza byla prawdziwa dla mastera do 2026-08-03; dzis nie jest.79bfe8c+da151fczastapily to dispatchem do jobs/deploy-runner/ (systemd na hoscie wezla). Realny stan: luka deployowa w trzech miejscach naraz — kontener executora zbudowany 2026-07-22 (wciaz stary kod), /opt/homelab/actions/deploy/ nie istnieje na VPS, deploy-runner nie jest zainstalowany na zadnym wezle. Zero akcji kiedykolwiek osiagnelo stan terminalny (completed 0 / failed 0); jedyne dwa action_result to reczne testy z 2026-07-23, nie remediacje z incydentu. Zweryfikowane wzgledemfbf165f: healthcheck_failed idzie do container_restart, nie do redeploy. Zywe incydenty w world state maja wylacznie trigger_type containers_not_running / healthcheck_failed — czyli zaden nie generuje redeployu; dzialaja tylko dryfy missing_service (2 pending). Najostrzejszy problem projektowy: jedyny zywy emiter service_unhealthy (node_agent.py:1104-1115) ma zahardkodowane service="control-plane", a control-plane ma wlasne deploy-local.sh — deploy-service.sh:93-96 konczy sie exit 3, wiec runner odmowi. Po wdrozeniu fixu redeploy sterowany incydentem nadal nie wykona sie ani razu. Ubocznie: w obrazie executora nie ma binarki ssh, wiec disk_cleanup (executor.py:392-400) jest martwy tym samym defektem. Osobny task. Dokument konczy sie sekcja FIX SHAPE — pieciopunktowa lista decyzji do podjecia, bez rekomendacji. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
34 KiB
| okf | type | visibility | status | updated | as_of | links |
|---|---|---|---|---|---|---|
| 0.1 | audit | private | active | 2026-08-05 | 2026-08-05 |
Recon — ścieżka remediacji redeploy (2026-08-05)
Read-only recon. Ground truth: kb/subsystems/recon-multiagent.md (A, D14, D15)
and kb/audits/czujniki-2026-07-30.md — both formerly at
docs/architecture/RECON-*, migrated into kb/ by 00de810 / 9f77a72.
Source read at task/recon-redeploy cut from master @ 4ecbdbb. Runtime
evidence collected 2026-08-05 ~09:40 UTC from vps and piha over ssh; solaria is
the host this recon ran on (local reads).
FINDING FIRST
The break this recon was commissioned to characterize was already repaired in
the repo — two days after the task premise was written — and the repair has
never been deployed. redeploy is still dead in production, but for a
completely different reason than stated.
The task brief describes master as: "the executor calls deploy-node.sh with args it ignores, on a path that doesn't exist inside the container." That was true of master until 2026-08-03. It is not true of master today:
79bfe8c(2026-08-03)feat(control-plane): host-side deploy runner — fix broken redeploy pathandda151fc(2026-08-03)fix(control-plane): redeploy wykonywalny — dispatch do host-side deploy-runnerareplaced thedeploy-node.shcall with a dispatch-file handoff to a new host-level systemd job,jobs/deploy-runner/.- Master's
executor.py:125-144no longer referencesdeploy-node.shat all;REPO_ROOTsurvives only as a decorative env var (executor.py:30-33). - The design decision, install runbook and open runtime steps are already
written up:
kb/decisions/deploy-runner-uzasadnienie.md,kb/runbooks/deploy-runner-install.md,kb/decisions/backlog-deploy-runner-instalacja.md.
What is actually broken today is a deploy gap in three places at once:
| Layer | Repo (master @ 4ecbdbb) |
Runtime (2026-08-05) |
|---|---|---|
| executor code | dispatches to actions/deploy/<node>/ |
container built 2026-07-22, still runs the deploy-node.sh code |
actions/deploy/ inbox |
created by _ensure_dirs (executor.py:79) |
does not exist on vps |
| deploy-runner on nodes | jobs/deploy-runner/ + systemd units |
not installed on vps, piha or solaria |
So the correct framing for the repair session is not "design a fix" — the fix is
designed, written, and 248-tests green. It is "deploy the fix and prove it
end-to-end", plus resolving the four genuine design questions the repair left
open (see FIX SHAPE at the end), the sharpest of which is that the only event
type that actually reaches a redeploy incident in live world state targets the
one service the new runner deliberately refuses to deploy.
Secondary finding, unrelated to redeploy but discovered en route: the executor
container has no ssh binary (verified below), so disk_cleanup —
the one action type that still shells out to ssh (executor.py:392-400) — is
dead by the same class of defect. It has never been exercised.
A. GENERATION
A1. event_type → action_type map as implemented on master @ 4ecbdbb
Unchanged from recon D15 except for the healthcheck_failed move in fbf165f,
which the brief asked me to verify. Verified: healthcheck_failed routes to
container_restart, not redeploy.
supervisor.py:52:
CONTAINER_RESTART_TRIGGERS = {"containers_not_running", "healthcheck_failed"}
Path 1 — world-state drift (supervisor.py:334-395, dispatch at
:401-477). The supervisor reads world/{services,incidents}.json, never raw
events. trigger_type comes from the incident the observer opened
(supervisor.py:317-327 ← observer.py:868):
Drift / incident trigger_type |
Action | Where |
|---|---|---|
containers_not_running |
container_restart |
supervisor.py:52,410,423 |
healthcheck_failed |
container_restart |
supervisor.py:52 (added fbf165f, 2026-07-29) |
service_unhealthy |
redeploy |
falls through :447 |
deployment_failed |
redeploy |
falls through :447 |
service absent from world state (missing_service, trigger_type: None) |
redeploy |
supervisor.py:353-360 → :447 |
node disk_pressure == high |
disk_cleanup |
supervisor.py:377-381,479+ |
Only services listed in hosts/<node>/services.yaml are considered
(supervisor.py:350); dormant nodes are skipped (supervisor.py:342,378).
Path 2 — direct event-file routing (_process_ha_events, supervisor.py:395, ~630-660) is unchanged and generates only container_restart (homeassistant,
shadow-downgraded to alert_only) and alert_only. Path 2 never generates
redeploy.
So: three producers of redeploy — service_unhealthy, deployment_failed,
missing_service — all on path 1. Section D9 shows only one of them fires in
practice, and it fires at the wrong target.
A2. Fields on a redeploy action
Built at supervisor.py:453-467, verbatim:
action = {
"action_id": action_id, # f"redeploy-{node}-{service}" (:413)
"timestamp": time.time(),
"type": "redeploy",
"node": node,
"service": service,
"risk_level": "guarded",
"confidence": 0.9,
"description": f"Redeploy {service} on {node} due to {drift['type']}",
"status": "pending",
"payload": {
"reason": drift["type"],
"svc_key": drift["svc_key"],
},
}
Written to actions/pending/<action_id>.json (supervisor.py:469-471).
There is no container_name, no path, no compose reference, and no args
field. Contrast container_restart (supervisor.py:428-446), which carries
container_name resolved via _get_container_name(service). A redeploy action
carries exactly two addressing facts: node and service. Everything else —
which compose file, which override, which repo — is resolved on the node at
execution time from service alone.
Live example, the one non-chelsty pending redeploy on vps:
{ "action_id": "redeploy-vps-gokapi", "timestamp": 1783615889.1762547,
"type": "redeploy", "node": "vps", "service": "gokapi",
"risk_level": "guarded", "confidence": 0.9,
"description": "Redeploy gokapi on vps due to missing_service",
"status": "pending",
"payload": { "reason": "missing_service", "svc_key": "vps/gokapi" } }
B. EXECUTION — why it's dead
B3. The redeploy handler
Two answers, because repo and runtime disagree.
(a) Master @ 4ecbdbb — executor.py:125-144. No command is built at all.
The executor writes a dispatch file and returns:
if action_type == "redeploy":
if not node or not service:
success = False
error_msg = (f"redeploy requires both node and service "
f"(node={node!r}, service={service!r})")
else:
self._dispatch_redeploy(action_id, node, service)
return # stays in running/ until action_result
_dispatch_redeploy (executor.py:238-267) writes
actions/deploy/<node>/<action_id>.json with {action_id, type, node, service, dispatched_at}. Deliberately a different inbox from container_restart's
actions/dispatch/<node>/, because node-agent deletes and failure-reports
anything in its own inbox that is not container_restart
(executor.py:24-29 ← node_agent.py:108,1011-1018).
(b) The executor actually running on vps — image created
2026-07-22T16:13:29Z, i.e. 12 days before the fix. Its /app/src/executor.py
lines 106-119, read out of the live container:
if action_type == "redeploy":
# Full service redeploy via the repo deploy script
cmd = [
str(REPO_ROOT / "scripts" / "deploy" / "deploy-node.sh"),
node,
service
]
logger.info(f"Running command: {' '.join(cmd)}")
result = subprocess.run(cmd, capture_output=True, text=True, cwd=str(REPO_ROOT))
REPO_ROOT = /repo (os.getenv("REPO_ROOT", "/repo"), running line 24;
master executor.py:33). Synchronous: returncode == 0 → completed, else
failed with stderr or stdout.
B4. Working directory / repo path, and does it exist in the container
The repo is reachable inside the executor container — the claim "a path that doesn't exist inside the container" needs one level of precision. Two different paths are involved and only the second is missing.
docker inspect control-plane-executor (vps):
Image=sha256:7603f09d… Created=2026-07-22T16:13:29.905995261Z
/opt/homelab -> /opt/homelab (rw)
/home/oskar/homelab-codex-ws -> /repo (rw)
/var/run/docker.sock -> /var/run/docker.sock (rw)
REPO_ROOT=/repo RUNTIME_PATH=/opt/homelab
Inside the container:
HOME=/home/homelab
/repo/scripts/deploy/deploy-node.sh ← EXISTS
ls: cannot access '/home/homelab/homelab-codex-ws': No such file or directory
git: MISSING docker: MISSING rsync: MISSING ssh: MISSING
bash: /usr/bin/bash python3: /usr/local/bin/python3 flock: /usr/bin/flock
So cwd=/repo is valid and the script is found and executed. It then fails on
its own internal path: deploy-node.sh:8 hardcodes
REPO_PATH="${HOME}/homelab-codex-ws" = /home/homelab/homelab-codex-ws, which
does not exist, and deploy-node.sh:18-21 exits 1.
Reproduced read-only inside the live container with the exact argv the running executor would use:
rc= 1
stdout= --- Starting Deployment on 1CF38E89F150 ---
Error: Repository not found at /home/homelab/homelab-codex-ws
Note the hostname: 1CF38E89F150, the container ID. Even with the repo present,
deploy-node.sh:27-57 resolves the host directory by matching os_hostname
against hosts/*/host.yaml and would fail with "No host directory found for
1cf38e89f150" — the executor would have deployed its own container's
service set, never the action's node. And behind that, two more hard stops:
git pull (deploy-node.sh:25) and the compose invocation both need binaries
the image does not have.
Four independent blockers, in the order they would be hit: missing
${HOME}/homelab-codex-ws → hostname resolves to a container ID → no git → no
docker. The image mount is not one of them.
Note also ../..:/repo:ro in services/control-plane/docker-compose.yml:36,55,86
— master mounts the repo read-only. The running container has it rw, another
marker that the container predates current master.
B5. deploy-node.sh arg parsing, and the call chain
deploy-node.sh has no argument parsing whatsoever. There is no getopts,
no case "$1", no $1/$2 reference anywhere in its 115 lines. It is a
zero-argument script: it derives everything from ${HOME} and hostname
(:8,12,13), and deploys the whole service set of whatever host it runs on
(:60-78 reads services.txt / services.yaml, :95-113 loops over all of
them). The node and service the old executor passed were silently discarded
by the shell.
Per-service deploy exists, but in a different script:
scripts/deploy/deploy-service.sh, extracted from deploy-node.sh by the
2026-08-03 fix. It takes --repo, --host-dir, --service plus optional
--build-if-needed / --force-recreate / --remove-orphans
(deploy-service.sh:51-62), validates the service name against
^[a-z0-9][a-z0-9._-]{0,63}$ (:72), and is the single shared compose
invocation for both callers — deliberately, because a project-name mismatch plus
--remove-orphans is what wiped the control-plane on vps on 2026-06-25
(deploy-service.sh:6-13, deploy-node.sh:86-92).
Call chains, with the host and container each link runs on:
Human deploy (SATURN-side dispatcher — confirms the solaria-gid-fix session's model):
operator @ SATURN: scripts/deploy/deploy.sh <target> [host shell, saturn]
preflight: must be on master, clean tree (deploy.sh:69-82)
→ ssh oskar@<target> 'cd ~/homelab-codex-ws && git pull
&& ./scripts/deploy/deploy-node.sh' (deploy.sh:199-202)
→ deploy-node.sh [host shell, target node]
→ deploy-service.sh --build-if-needed --remove-orphans × every service
(deploy-node.sh:100-105)
(target == control-plane takes a different branch:
deploy-control-plane.sh --ssh, deploy.sh:194-196)
Agent redeploy, master @ 4ecbdbb:
supervisor [container control-plane-supervisor, vps]
→ /opt/homelab/actions/pending/redeploy-<node>-<svc>.json
operator (Telegram / operator-ui) → approved/
executor [container control-plane-executor, vps]
→ /opt/homelab/actions/deploy/<node>/<action_id>.json (executor.py:238-267)
deploy-runner [HOST systemd oneshot on <node>, NOT a container]
→ rsync-pull inbox from vps (deploy-runner.sh pull_actions)
→ python3 action.py validate (5 rules, action.py:78-125)
→ deploy-service.sh --repo … --host-dir … --service … --force-recreate
→ action.py emit-result → evt-<node>-<ts>-action_result-<svc>.json
→ rsync-push to vps events/
executor._reconcile_running_actions() (executor.py:269-318)
→ completed/ or failed/ (timeout REDEPLOY_TIMEOUT_SECS=900, executor.py:55)
The runner runs on the host, not in a container, on purpose: compose then
resolves relative bind mounts and the project name exactly as a human deploy
does (deploy-runner.sh:26-31). It also never runs git pull — a redeploy
reconciles the node to the checkout it already has; shipping code stays a human
deploy.sh (deploy-runner.sh:32-33).
B6. Concrete failure mode today
Nothing fires, and nothing ever has. No redeploy has ever been attempted, approved, or executed — fleet-wide, ever.
Action pool on vps, 2026-08-05:
pending 18 approved 0 running 0 completed 0 failed 0 rejected 0 cancelled 18
dispatch 1 deploy: ls: cannot access '/opt/homelab/actions/deploy': No such file or directory
completed 0 / failed 0 is the whole story: not one action has ever reached a
terminal state. 18 redeploy actions exist on disk (16 cancelled, 2 pending);
none was ever approved, so the broken handler was never entered.
The two live ones:
/opt/homelab/actions/pending/redeploy-vps-gokapi.json reason: missing_service
/opt/homelab/actions/pending/redeploy-chelsty-infra-ha-diag-agent.json reason: missing_service
The chelsty one targets a node dead since 2026-06-01 and is already flagged for
manual removal (kb/subsystems/recon-multiagent.md, "Runbook — stale chelsty
action").
The executor's entire log since 2026-07-22 is 144 lines and contains three action executions, all hand-injected tests, zero supervisor-generated actions, zero redeploys, and nothing at all after 2026-07-23:
2026-07-22 16:13:41 - Starting executor loop
2026-07-22 16:32:41 - Executing action: test-restart-piha-node_exporter-2
2026-07-22 16:32:41 - Dispatched container_restart … to node-agent on piha
2026-07-22 16:37:42 - ERROR - Action … timed out: Timed out after 300s waiting for node-agent on 'piha'
2026-07-23 11:54:05 - Executing action: test-e2e-1784807638
2026-07-23 11:54:05 - ERROR - Failed to move test-e2e-1784807638 to running: Expecting value: line 1 column 38
…(same pair repeating every 10 s for 11 minutes — malformed JSON never leaves approved/)…
2026-07-23 12:05:25 - Executing action: test-e2e-b
2026-07-23 12:05:25 - Dispatched container_restart test-e2e-b (container=node_exporter) to node-agent on piha
2026-07-23 12:05:56 - Action test-e2e-b completed
Why nothing fires, in the order the funnel closes:
redeployneedsservice_unhealthy,deployment_failed, ormissing_service. Post-fbf165f, the two high-volume signals (containers_not_running2,490;healthcheck_failed3,376 on piha) both go tocontainer_restart. Live incident census on vps confirms it —world/incidents.jsonholds 3 incidents total,trigger_type: {containers_not_running: 2, healthcheck_failed: 1}. Zero redeploy-producing incidents are open.missing_servicestill fires (it needs no incident at all), which is exactly what the 2 pending redeploys are. But it is dedup-suppressed forever by its own pending file (supervisor.py:418-421) —redeploy-vps-gokapihas held its ID since 2026-07-19.- Nothing gets approved.
approved 0, and the executor reads onlyapproved/(executor.py:94-95). - Even on approval the running executor would fail as in B4 — but this has never actually happened, so there is no failed/ artefact, no stack trace, and no "Running command: …deploy-node.sh" log line to point at. The evidence for the break is the reproduction in B4, not a production incident.
C. THE HEALTHY PATH FOR COMPARISON
C7. container_restart traced end to end
Correction to the premise: the two action_result events of 2026-07-23 are
manual e2e tests, not real remediations. Payloads:
{"action_id": "test-restart-piha-node_exporter-2", "success": true, "error": "", "node": "piha"}
{"action_id": "test-e2e-b", "success": true, "error": "", "node": "piha"}
Both hand-crafted IDs. A supervisor-generated ID would be
container-restart-piha-node_exporter (supervisor.py:411). And only one
of the two round-tripped cleanly: test-restart-piha-node_exporter-2 was
dispatched 2026-07-22 16:32 UTC, timed out at 300 s, and its action_result
only landed 2026-07-23 11:22 UTC — ~19 h later, after the executor had already
filed it failed. test-e2e-b completed in 31 s. So the reference path has
exactly one clean end-to-end demonstration, and one latency failure, and has
never run from a real incident.
Path, verbatim:
executor._execute_actionmovesapproved/<id>.json→running/(executor.py:104-112), thenexecutor.py:146-157:container_name = data.get("container_name") or service→_dispatch_container_restart(...)→return(stays inrunning/)._dispatch_container_restart(executor.py:209-236) writesactions/dispatch/<node>/<action_id>.json={action_id, type, node, service, container_name, dispatched_at}.- node-agent on the node rsync-pulls its own subdir with
--remove-source-files(node_agent.py:919-960); on vps it reads the path directly (:933). _execute_dispatched_action(node_agent.py:981-1051) applies four gates — idempotency (:999), node scoping (:1003), type whitelistALLOWED_DISPATCH_ACTION_TYPES = {"container_restart"}(:108, :1011), self-restart guard{"node-agent"}(:114, :1027) — thenself.docker_client.containers.get(container_name).restart()(:1043-1044)._report_action_result(node_agent.py:1053+) emits anaction_resultevent; the existing rsync ships it to vps.executor._reconcile_running_actions(:269-318) matches it via_find_action_result(:320-356) and moves the action tocompleted//failed/, or fails it afterACTION_TIMEOUT_SECS=300.
C7b. Structural differences between the two handlers
On master they are now deliberately near-identical — steps 1, 2, 5 and 6 are
the same code, DISPATCHED_ACTION_TYPES = {"container_restart", "redeploy"}
(executor.py:60), and _reconcile_running_actions handles both with only the
timeout differing (executor.py:299-302). What remains different:
container_restart |
redeploy |
|
|---|---|---|
| Inbox | actions/dispatch/<node>/ |
actions/deploy/<node>/ (must not share — executor.py:24-29) |
| Executor on the node | node-agent, a container, own docker socket | deploy-runner, host systemd oneshot + timer, no container |
| Deployed how | deploy.sh / compose, already everywhere |
manual per-node systemd install, nowhere yet |
| Poll cadence | CHECK_INTERVAL=60 |
OnUnitActiveSec=60s (matched on purpose) |
| Timeout | ACTION_TIMEOUT_SECS=300 |
REDEPLOY_TIMEOUT_SECS=900 |
| Payload | carries container_name |
carries only service |
| Validation | 4 gates in node_agent.py |
5 gates in action.py:78-125, incl. service must be in hosts/<node>/services.yaml and compose file must exist |
| Idempotency | _already_processed marker |
state/processed-deploy-actions/<id>.done (deploy-runner.sh:57,151-155) |
| Concurrency | none needed | flock on state/deploy-runner.lock |
| Can act on node-agent itself | no (self-guard) | yes — deliberate; solaria's node-agent has been docker-blind for weeks (deploy-runner.sh:26-31) |
| Can act on control-plane / stability-agent | yes | no — deploy-service.sh:93-96 exits 3, runner reports failure with an explanation |
The last row is the important asymmetry: see D9.
C8. How the executor reaches nodes
It doesn't, and it cannot. Verified inside the live container: ssh: MISSING,
rsync: MISSING, git: MISSING, docker: MISSING. There is no ssh client, no
key, no known_hosts. This is by design — executor.py:200-207 and
kb/phases/backlog.md "Remediacja floty bez SSH": the VPS never initiates a
connection to a node. Every remote action is a file the node comes and
collects, over the ssh/rsync channel node-agent (and now deploy-runner)
already owns in the other direction. The executor's only outputs are writes into
the shared /opt/homelab mount.
Two consequences:
disk_cleanupis dead by the same defect._execute_disk_cleanup(executor.py:358-407) builds["ssh", *SSH_OPTIONS, f"{SSH_USER}@{node}", …]and callssubprocess.run. With nosshbinary this raisesFileNotFoundError, caught byexecutor.py:175-177, and the action fails with[Errno 2] No such file or directory: 'ssh'. Never observed becausefailed/is empty — nodisk_cleanupwas ever approved either. Out of scope here; flagged as a follow-up.- Solaria's self-ssh problem is irrelevant to the executor.
deploy.shruns on SATURN and ssh's to the target, which is why it cannot deploy solaria from solaria. The executor never ssh's anywhere; the connection is always node → vps, initiated by the node. Solaria's node-agent already ships events to vps over that channel daily, so the transport is proven. Solaria's real blocker is different and unrelated: its node-agent has no docker socket access (group_add: "996"fixed in repo @ddae57c, still not deployed — running container showsGroupAdd=[999]), which breakscontainer_restartthere.redeployis unaffected, because deploy-runner runs on the host and never touches node-agent.
D. WHAT REDEPLOY IS FOR
D9. When redeploy should fire vs container_restart
The semantic split the brief proposes is the one the code already encodes
(supervisor.py:36-52, 423-426, 447-452): container_restart = the container
exists and a restart plausibly heals it; redeploy = the container is absent,
or its image/compose/env no longer matches the repo, so it must be re-created
from the manifest. deploy-service.sh --force-recreate is exactly that and
nothing more (deploy-service.sh:111-130: docker compose -f … [-f override] [--env-file] up -d --force-recreate, no --build, deploy-service.sh:113-118
— a redeploy is a reconcile, not a code ship, and building on the 4 GiB vps
mid-incident is an OOM risk).
Candidate event types for the redeploy side, and whether each actually reaches an incident today:
| Candidate | Opens an incident? | Live volume | Verdict |
|---|---|---|---|
missing_service (drift, not an event) |
n/a — needs no incident; supervisor.py:353-360 raises it purely from desired-vs-actual |
2 pending actions right now | The only redeploy producer that actually fires. Semantically the perfect fit: the container is gone, a restart is impossible. |
service_unhealthy |
yes — observer.py:771,792 |
34 events, vps only | Fires, but is a trap — see below. |
deployment_failed |
yes — observer.py:838-842 _handle_deployment_failure |
0 events fleet-wide | Only emitter is scripts/lib/log.sh:35,43 (deploy-time shell logging). Never reaches the vps event store. Dead in practice. |
containers_not_running |
yes, observer.py:771 |
2,490 piha / 3 solaria / 1 vps | Correctly on container_restart — container exists. |
healthcheck_failed |
yes, observer.py:771 |
3,376 piha | Correctly on container_restart since fbf165f. Escalating repeat failures to redeploy is a candidate, but needs a retry policy that does not exist. |
container_restarting / container_state_unexpected |
no — observational only, observer.py:793-814 |
2 solaria | Deliberately not remediated. |
stability-agent's containers_not_running |
no — service=None fails the if service and service != "all" guard at observer.py:747, so the incident branch is never reached |
— | Known-broken producer (recon D15, kb/audits/czujniki-2026-07-30.md). |
The trap. service_unhealthy has exactly one live emitter in the whole
fleet: node_agent.py:1104-1115, which probes
http://localhost:18180/summary and hardcodes service="control-plane". So
every service_unhealthy incident that could ever produce a redeploy targets
vps/control-plane — and control-plane is one of only two services with its
own deploy-local.sh (services/control-plane/, services/stability-agent/),
which deploy-service.sh:93-96 refuses with exit 3, which the runner turns into
a failed action: "control-plane owns its deploy path
(services/control-plane/deploy-local.sh) — needs an operator deploy, not an
automated redeploy" (deploy-runner.sh:179-181). That refusal is correct — a
containerised control plane redeploying itself mid-incident is not something to
automate — but it means the fixed redeploy path, once deployed, will still
produce zero successful incident-driven redeploys. Only missing_service
will ever succeed.
Also worth stating plainly: redeploy cannot repair a service_unhealthy
caused by bad code, because the path deliberately does not git pull
(deploy-runner.sh:32-33) and deliberately does not --build. It repairs
drift — a container removed, a stale env file, a config change already in the
node's checkout — not regression.
D10. Preconditions for a working redeploy
Per link in the chain, on master's architecture. Current status verified 2026-08-05.
On the vps (control plane):
| # | Precondition | Where | Status |
|---|---|---|---|
| 1 | executor running post-da151fc code |
executor.py:125-144 |
❌ container built 2026-07-22, runs deploy-node.sh code |
| 2 | /opt/homelab/actions/deploy/<node>/ writable |
executor.py:79 (_ensure_dirs) |
❌ dir absent — created automatically by the new executor |
| 3 | repo checkout current | /home/oskar/homelab-codex-ws |
⚠️ at efd8d2f, one commit behind master 4ecbdbb; has the fix (jobs/deploy-runner/ dated Aug 3 17:28) |
| 4 | not needed: ssh, git, docker CLI in the executor image | — | ✅ by design |
On each target node:
| # | Precondition | Where | vps | piha | solaria |
|---|---|---|---|---|---|
| 5 | repo checkout at REPO_PATH |
deploy-runner.sh:73 |
✅ efd8d2f |
✅ 4ecbdbb |
(main checkout, not inspected) |
| 6 | hosts/<node>/ exists in that checkout |
deploy-runner.sh:74 |
✅ | ✅ | ✅ |
| 7 | scripts/deploy/deploy-service.sh present + executable |
deploy-runner.sh:75-76 |
✅ Aug 3 | ✅ Aug 4 | — |
| 8 | jobs/deploy-runner/ systemd unit + timer installed & enabled |
kb/runbooks/deploy-runner-install.md |
❌ | ❌ | ❌ |
| 9 | /opt/homelab/config/deploy-runner/env with NODE_NAME (empty VPS_EVENTS_HOST on vps) |
env.example |
❌ | ❌ | ❌ |
| 10 | service user in docker group |
runbook | — | — | ✅ (oskar ∈ docker gid 996) |
| 11 | python3 + PyYAML, rsync, flock |
action.py:60-64, deploy-runner.sh:70-76 |
— | — | — |
| 12 | ssh key reaching vps (remote nodes only) | deploy-runner.sh:62-66 |
n/a | ✅ (node-agent already ships events) | ✅ |
Per-action, at execution time (action.py:78-125 — all five must hold or the
runner reports failed without deploying):
| # | Rule | Note |
|---|---|---|
| 13 | type == "redeploy" |
|
| 14 | action["node"] == NODE_NAME |
defense in depth; inbox is already node-scoped |
| 15 | action_id / service match strict name patterns |
no path traversal |
| 16 | service listed in hosts/<node>/services.yaml |
repo desired state is the authority; a stale dispatch cannot pull an arbitrary stack |
| 17 | services/<service>/docker-compose.yml exists in the repo |
|
| 18 | (not a rule, an outcome) service must not have services/<svc>/deploy-local.sh |
else exit 3 → failed-with-explanation |
Checked against the one actionable pending action, redeploy-vps-gokapi: rule
16 ✅ (hosts/vps/services.yaml:63), rule 17 ✅ (services/gokapi/docker-compose.yml),
rule 18 ✅ (no deploy-local.sh). gokapi is a valid end-to-end target the
moment the runner is installed on vps.
Also required for the loop to close: the result event must land in
/opt/homelab/events/<node>/ on the vps in node_agent.emit_event()'s exact
format, because executor._find_action_result (:320-356) globs
evt-*-action_result-*.json and parses the unix timestamp out of the
filename. action.py:emit_result reproduces that byte-for-byte and adds
payload.source = "deploy-runner".
FIX SHAPE — DECISIONS NEEDED
The transport and script layers are decided and written. What is left is one sequencing decision and four genuine design choices the 2026-08-03 work left open.
1. Deploy vs re-derive. The repair is 79bfe8c/da151fc on master, 248
tests green, unexercised at runtime.
- Deploy as written — three runtime steps (
deploy.sh control-plane; installjobs/deploy-runner/on vps + piha + solaria perkb/runbooks/deploy-runner-install.md; e2e on a benign piha service, thenredeploy-vps-gokapi). Fastest to a working loop; accepts a design nobody has yet run. - Re-derive first — treat the shipped design as a proposal, resolve 2–5 below, then deploy once. Slower; avoids installing systemd units on three hosts twice.
Note that 2–5 are answerable either way, but their answers are cheaper to apply before the units are on three nodes.
2. Where the deploy actually runs: node-side runner vs saturn-side dispatcher.
The shipped design is node-side (deploy-runner on each node, executor never
connects out).
- Node-side (as shipped): preserves "remediacja floty bez SSH"; works on
solaria despite its broken node-agent socket and despite the
self-ssh problem; every node needs a host-level systemd install and its own
repo checkout, and the install is not covered by
deploy.sh— it drifts silently. - Executor → ssh SATURN →
deploy.sh <node>: reuses the human path exactly, one install point instead of N, and inheritsdeploy.sh's preflight gates. But it inverts the no-SSH decision, puts a key to SATURN inside a vps container, makes SATURN a runtime dependency of remediation (it is a workstation, not a server), anddeploy.shis whole-node only — it cannot express "redeploy just gokapi". - Executor → ssh direct to node →
deploy-service.sh: fewest moving parts, no runner, no timer, no rsync. Requires an ssh client + fleet key in the executor container — the exact thing the architecture forbids — and reintroduces the self-ssh and reachability problems the pull model sidesteps.
3. Per-service vs whole-node redeploy. Shipped: per-service, and the action payload carries no node-wide option.
- Per-service: minimal blast radius; matches the action's semantics
(
missing_servicenames one service); a genuinely node-wide failure needs N actions and N approvals. - Whole-node (i.e.
deploy-node.sh, what the old executor accidentally attempted): one action heals a node; but it touches every service including healthy ones, runs--build-if-needed, and--remove-orphansis in play — the 2026-06-25 control-plane wipe is precedent for how that ends.
4. How args flow — how much the action carries vs how much the node resolves.
Shipped: the action carries {node, service} only; everything else (compose
path, override, env-file, force-recreate) is resolved on the node by
deploy-service.sh from repo state.
- Thin action (as shipped): the repo is the single authority; an old or forged
dispatch cannot smuggle a path or a flag; but the operator approving the action
cannot see or influence what will actually run, and
--build/--pullare not expressible even when they are what is needed. - Fat action: supervisor stamps compose path, override, flags into the payload — visible at approval time, tunable per incident; every field becomes attack surface the runner must re-validate, and payloads can go stale against the node's checkout.
5. What happens to service_unhealthy → control-plane (D9). Once the
runner is installed, the only incident-driven redeploy producer will target the
one service the runner refuses.
- Leave it: the failed action with its explanatory message is the alert; costs one failed action per control-plane outage and burns the dedup ID.
- Downgrade to
alert_onlyat the supervisor: honest — nothing automated can fix it — but drops the drift out of the remediation view. - Give control-plane a self-redeploy path (analogous to
deploy-control-plane.sh): closes the loop, and lets a broken control plane attempt to repair itself with its own hands. Highest risk on the list. - Widen
service_unhealthybeyond the hardcoded control-plane probe (node_agent.py:1104-1115): would give redeploy a real, non-degenerate input set for the first time — but that is a monitoring change, not a redeploy change, and it interacts with thehealthcheck_failedrouting decided infbf165f.
Out of scope here, found en route: the executor's disk_cleanup handler
shells out to ssh, which does not exist in the image (C8) — same class of
defect as the original redeploy break, never observed because no disk_cleanup
was ever approved. Worth its own task.