Root cause of fleet staleness since ~OOM 2026-06-01: events/<node>/ grew to 291k files (1.2G on piha); rsync of the whole dir exceeded the 30s subprocess timeout every cycle; --remove-source-files never ran; backlog compounded silently. Four fixes: 1. Batch shipping (SHIP_BATCH_SIZE=1000): rsync sends only the oldest 1000 files per cycle via --files-from=- on stdin instead of the whole directory. Backlog drains across cycles; each push fits within timeout. 2. Timeout 30s → 120s: wider margin for large batches and slow links (piha → vps over Tailscale). 3. Backlog trim safety-net (_trim_events_backlog): if events/<node>/ exceeds BACKLOG_MAX (5000) files, oldest files are deleted to bring count back to 5000. Called each cycle before shipping. Breaks the death-spiral independently of rsync success. VPS excluded (uses _cleanup_control_plane_fs). 4. Backlog visibility: WARNING log with file count when unsent events exceed BACKLOG_WARN (2000). "events backlog: N unsent files" — no more silent accumulation. SHIP_BATCH_SIZE is env-configurable for tuning per-node. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| src | ||
| docker-compose.yml | ||
| Dockerfile | ||