homelab-codex-ws/services/node-agent
Oskar Kapala d483274037 fix(node-agent): batch rsync, backlog trim, timeout 120s, backlog warn
Root cause of fleet staleness since ~OOM 2026-06-01: events/<node>/
grew to 291k files (1.2G on piha); rsync of the whole dir exceeded the
30s subprocess timeout every cycle; --remove-source-files never ran;
backlog compounded silently.

Four fixes:

1. Batch shipping (SHIP_BATCH_SIZE=1000): rsync sends only the oldest
   1000 files per cycle via --files-from=- on stdin instead of the whole
   directory.  Backlog drains across cycles; each push fits within timeout.

2. Timeout 30s → 120s: wider margin for large batches and slow links
   (piha → vps over Tailscale).

3. Backlog trim safety-net (_trim_events_backlog): if events/<node>/ exceeds
   BACKLOG_MAX (5000) files, oldest files are deleted to bring count back to
   5000.  Called each cycle before shipping.  Breaks the death-spiral
   independently of rsync success.  VPS excluded (uses _cleanup_control_plane_fs).

4. Backlog visibility: WARNING log with file count when unsent events
   exceed BACKLOG_WARN (2000).  "events backlog: N unsent files" — no more
   silent accumulation.

SHIP_BATCH_SIZE is env-configurable for tuning per-node.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-09 18:16:08 +02:00
..
src fix(node-agent): batch rsync, backlog trim, timeout 120s, backlog warn 2026-06-09 18:16:08 +02:00
docker-compose.yml fix(node-agent): run as uid 1000 with docker group access 2026-06-03 18:20:31 +02:00
Dockerfile fix(node-agent): run as uid 1000 with docker group access 2026-06-03 18:20:31 +02:00