Root cause of fleet staleness since ~OOM 2026-06-01: events/<node>/ grew to 291k files (1.2G on piha); rsync of the whole dir exceeded the 30s subprocess timeout every cycle; --remove-source-files never ran; backlog compounded silently. Four fixes: 1. Batch shipping (SHIP_BATCH_SIZE=1000): rsync sends only the oldest 1000 files per cycle via --files-from=- on stdin instead of the whole directory. Backlog drains across cycles; each push fits within timeout. 2. Timeout 30s → 120s: wider margin for large batches and slow links (piha → vps over Tailscale). 3. Backlog trim safety-net (_trim_events_backlog): if events/<node>/ exceeds BACKLOG_MAX (5000) files, oldest files are deleted to bring count back to 5000. Called each cycle before shipping. Breaks the death-spiral independently of rsync success. VPS excluded (uses _cleanup_control_plane_fs). 4. Backlog visibility: WARNING log with file count when unsent events exceed BACKLOG_WARN (2000). "events backlog: N unsent files" — no more silent accumulation. SHIP_BATCH_SIZE is env-configurable for tuning per-node. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| agent-system | ||
| brain-watchdog | ||
| control-plane | ||
| forgejo | ||
| ha-diag-agent | ||
| mosquitto | ||
| node-agent | ||
| node_exporter | ||
| npm | ||
| ollama | ||
| planner-agent | ||
| stability-agent | ||
| zigbee2mqtt | ||
| .gitkeep | ||