Oskar Kapala
c5dd1ecb56
docs(kb): kb-02 architektura dokumentow — master + moduly 0-3 (split Paperless PIHA/OCR SOLARIA)
2026-07-01 21:56:07 +02:00
Oskar Kapala
727cb999ef
docs(sesja): inwentaryzacja floty 2026-06-30 + dysk SATURN + safeclean; backlog 23 rozjazdy
2026-06-30 19:56:50 +02:00
Oskar Kapala
229f85bd9b
docs(infra): inwentaryzacja floty 2026-06-30 — 23 rozjazdy repo↔rzeczywistość
2026-06-30 19:40:50 +02:00
Oskar Kapala
4055a8ffab
docs(backlog): oznacz kroki 4+5 Prometheus jako ZROBIONE, dopisz tech-debt log PROMETHEUS_URL
...
Kroki 4 (reguły liveness d417000 ) i 5 (watchdog poll 62d6fc0 ) z planu monitoringu zamknięte.
Nowy wpis aktywny: brain-watchdog nie loguje PROMETHEUS_URL przy starcie — utrudnia weryfikację.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 19:29:24 +02:00
Oskar Kapala
0656635793
docs(sesja): fleet-prometheus liveness — reguły NodeDown + watchdog→Prometheus poll
...
Sesja 2026-06-30: commits d417000 (rules/liveness.yml, NodeDown inactive=OK) i 62d6fc0
(brain-watchdog/check_prometheus_alerts, architektura A, 12 testów). Incydent PIHA
divergent branches — praca uratowana verify-before-reset. PENDING: end-to-end firing.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 19:29:20 +02:00
Oskar Kapala
831bbedffa
docs(sesja): kapala.org → Cloudflare wildcard DNS-01, HA i Immich na mesh
...
Sesja 2026-06-30: migracja DNS kapala.org z 42.pl na Cloudflare, wildcard
*.kapala.org przez DNS-01, ha/immich.kapala.org przez NPM@PIHA mesh-only.
- nowy session log z root-cause buga "unrecognized name" (NPM duplikuje
proxy_http_version przy Websockets ON) + wzorzec migracji usługi na mesh
- backlog: gotchas (NPM WS config, CF auto-proxy DKIM) + TODO (migracja
okit.pl, foty renew, cleanup ha-ken add-on, stale ha.okit.pl)
- topology: nowa sekcja ingress (ha/immich → NPM@PIHA, wildcard cert)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 16:23:57 +02:00
Oskar Kapala
48107ef2c1
docs(backlog): zamknij deploy.sh vps bug + 2 nowe wpisy (prom-reload, chelsty node_exporter)
...
- 🔴 deploy.sh vps↔control-plane → Zamknięte (fix 3b71707 , guard deploy-local.sh)
- krok 2 planu floty (targety 100.x) → ZROBIONE (7d4014e )
- NOWE: deploy-node.sh nie reloaduje config-driven serwisów (cicha rozbieżność deploy↔config)
- NOWE: zbadać chelsty/chelsty-infra node_exporter DOWN z VPS
- Problem B (ghost kontenery) pozostaje otwarty osobno
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 17:30:26 +02:00
Oskar Kapala
39ab34643d
docs(kb): sesja 2026-06-26 — fleet-prometheus targety floty + zamknięcie buga deploy.sh vps
...
- targety node_exporter floty (piha/solaria/lustro) wdrożone (7d4014e )
- fix deploy.sh vps↔control-plane potwierdzony w boju (3b71707 )
- PENDING: health-verify targetów w /api/v1/targets nie potwierdzony
- lekcja: nie commitować na master równolegle gdy CC pracuje na wątku (potrójny rebase)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 17:30:19 +02:00
oskar
7a95ab2fff
docs(kb): sesja 2026-06-25 — bulk import Gmail uruchomiony (225k kopert na PIHA)
2026-06-26 16:33:17 +02:00
Oskar Kapala
1529911012
docs(backlog): add anomaly-detection liveness idea (mózg learns per-node daily pattern)
2026-06-26 16:20:54 +02:00
oskar
c0ffb6abf7
docs(kb): sesja 2026-06-22 — spine relokowany na PIHA + przygotowanie hosta
...
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 18:52:35 +02:00
oskar
c32e050d56
docs(kb): sesja 2026-06-24 — importer Gmail gotowy + konwencja jobs/
2026-06-25 17:41:01 +02:00
oskar
02d0fa391f
docs(deployment): add critical warning — deploy.sh vps destroys control-plane (project-name divergence, 2026-06-25 incident)
...
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-25 16:30:53 +02:00
oskar
a17a11bc9d
docs(backlog): add 2026-06-25 bugs (deploy.sh vps mina, ghost kontenery, supervisor bez akcji, phantom world-state); close flaky tests + env-file fixes
...
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-25 16:30:51 +02:00
oskar
2655b4c6f4
docs(sessions): add 2026-06-25 session log (fleet-prometheus etap 1 domknięty + incydent mózgu)
...
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-25 16:30:47 +02:00
oskar
afc0a52e3e
docs(backlog): add Compose state-drift and flaky-test tech-debts from 2026-06-24
...
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-24 23:37:52 +02:00
oskar
a55c0928e6
docs(sessions): add 2026-06-24 session log (fleet-prometheus etap 1)
...
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-24 23:37:49 +02:00
oskar
ed54c2777a
docs(backlog): swap VPS DONE; sekcja planu Prometheus-as-truth floty
...
Swap 2-4GB na VPS -> Zamkniete (2026-06-22). Nowa sekcja planu monitoringu
floty (Prometheus pull up{}, osobny instance, bez Alertmanagera) z krokami.
Zastepuje szkic blackbox+Alertmanager z 2026-06-17.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 22:15:24 +02:00
oskar
1136a48222
docs(session): 2026-06-22 — decyzja Prometheus-as-truth dla liveness floty
...
Osobny fleet-Prometheus (pull, up{}) zastępuje warstwę wykrywania
node-agent->rsync->observer. Bez Alertmanagera (brain-watchdog drugie wejście).
Placement VPS. Swap 4G done. Plan kroków OTWARTE.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 22:15:16 +02:00
oskar
2b3cb89144
refactor(kb-postgres): relokacja SOLARIA→PIHA — arm64, mem_limit 1g, tuning pod małą maszynę
...
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 19:59:58 +02:00
oskar
bbfbb698f8
docs(kb): etap 1 fundament done + session log 2026-06-17
...
- kb-00-overview.md: etap 1 oznaczony jako DONE (kb-postgres + koperta +
packages/kb-mail), następny krok = etap 2 bulk Gmail; dodana sekcja
konwencji packages/
- CLAUDE.md: sekcja "Shared Python Libraries (packages/)" — konwencja,
layout, instalacja w Dockerfile
- docs/sessions/2026-06-17-kb-foundations.md: pełny log sesji (architektura
KB, decyzje, etap 1 zbudowany, poprawki spójności, następny krok)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-19 20:02:25 +02:00
oskar
214a6c3877
docs(kb): foundacja — overview + projekt maili (2 żywe źródła)
2026-06-19 20:02:25 +02:00
oskar
643ca209f4
docs(session): 2026-06-17 — vikunja OIDC+gitops, observer heartbeat-TTL, panel-source finding + backlog
2026-06-17 22:10:52 +02:00
oskar
5f1528e4ab
feat(observer): 3-state node liveness (fresh/stale/dead) + transitions + read-time net
...
Fixes the "dead node shown NOMINAL" silent outage: node status was set only by
events and never expired, so a node that crashed/lost connectivity stayed
"online" forever (chelsty-infra was online for 16d, piha ~6d). The only thing
that flipped status to offline was a node_offline event, which an unreachable
node can never emit.
Now node status is derived from freshness (now - last_seen), recomputed every
observer cycle (incl. cycles with no new events):
- always-on: fresh <=180s, stale 180-600s, dead >600s (3x the 60s heartbeat)
- remote/LTE (chelsty-*): fresh <=900s, stale 900-3600s, dead >3600s
Thresholds + tier logic live in ONE shared helper, services/control-plane/src/
liveness.py, imported by the observer and both operator UIs (bind-mounted into
the agent-system webui image). No 3x copy.
Transitions are not silent: the observer emits node_stale / node_offline /
node_online (recovery) events tagged source=observer (skipped on re-ingest so
they never reset last_seen), routed by the supervisor to alert_only actions.
Read-time safety net: both UIs recompute liveness from last_seen at request
time, so a stalled observer still surfaces dead nodes. Services inherit their
node's liveness (cascade, variant B) without mutating services.json.
Replaces the earlier binary NODE_OFFLINE_TTL_SECS flip.
Tests: liveness unit tests, observer 3-state + transitions/recovery/baseline +
self-event skip, operator_ui read-time net + cascade, supervisor node-event
routing. 89 passed. docker compose config valid for both stacks.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 20:07:25 +02:00
Oskar Kapala
31b5981174
docs: session 2026-06-11 20:35
2026-06-11 20:35:23 +02:00
Oskar Kapala
c1acee7acf
docs: session 2026-06-11 20:19
2026-06-11 20:19:26 +02:00
Oskar Kapala
a0bfd96870
docs: session 2026-06-11 — lustro ssh shipping fix + ha-diag-agent piha + backlog/flota-bomba
...
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 14:18:00 +02:00
Oskar Kapala
5c2516d097
docs: session 2026-06-09 + skill/backlog update
...
- docs/sessions/2026-06-09-flota-recovery-lustro-register.md: flota
recovery (root cause aerbot group, 3 warstwy maskujące), lustro register
stan+plan, fix-event-bloat i OOM pending, worktree gotcha
- docs/backlog.md: nowy plik — tech-debt tracker; wpisy: --omit-dir-times,
oskar∈aerbot deklaratywnie, worktree per task, observer staleness
- .claude/skills/node-onboarding/SKILL.md: step table aktualizacja (PROVEN:
20-base, 30-node-agent; WRITTEN: 40-register, 50-verify), 3 nowe gotchas
(rsync perm, observer restart, worktree branch)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-09 20:38:35 +02:00
Oskar Kapala
9b2a1b4e9a
docs(backlog): observer staleness — dead node shows NOMINAL (heartbeat TTL)
2026-06-09 12:16:59 +02:00
Oskar Kapala
85e056046c
docs(session): worktree hygiene update + marker gap note
2026-06-09 11:38:41 +02:00
Oskar Kapala
e59eb12da3
docs: session log 2026-06-08 — LUSTRO onboarding
...
Records the onboarding session for LUSTRO (RPi4, KEN site):
node facts from preflight, key decisions (user pi/uid-1000, IP
over mDNS, zram target), 00-access status, tool bugs fixed
(dry-run propagation, yaml_get greedy-colon + inline comment,
ssh known-hosts in verify), open items for next session
(worktree hygiene first, bootstrap-runtime, node-agent, register,
verify, mm-watch).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-08 22:31:03 +02:00
Oskar Kapala
5ccdfa0ca6
docs: add planner-agent docs and session summary 2026-05-27
...
- services/planner-agent/README.md: full service doc (what it does,
LLM fallback chain, env vars, deploy steps, local run, redis-cli
end-to-end test, healthcheck)
- README.md: add Agent System section with all agents and their roles
- docs/sessions/2026-05-27-planner-agent.md: session summary (built
files, architectural decisions, problems + solutions, deployment
status, pending work)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-27 22:35:59 +02:00
Oskar Kapala
603e10a364
docs: session summary 2026-05-27 + update observer/control-plane/chelsty docs
...
docs/sessions/2026-05-27.md (new):
- Full session record: problems found, all commits shipped, end state
- Written in Polish per operator preference for session notes
- Known limitations: SLZB-06U offline, ezsp→ember migration pending
docs/observer-runtime.md:
- Document per-node checkpoint format (replaces old global checkpoint)
- Add service_healthy / service_recovered resolution behavior
- Document ghost key pruning (_prune_stale_world patterns)
- Add event type reference table (negative vs positive)
docs/vps-control-plane.md:
- Add container names and network_mode: host detail
- Document monitor:false, NODE_ALIAS_MAP, auto-cancel behavior
- Add piha agent-system materializer integration note
- Rewrite recovery section with actionable bootstrap-flood diagnosis
- Add action state machine (pending→approved→running→completed/cancelled)
docs/chelsty-runtime.md:
- Add chelsty-infra/chelsty-ha node table
- Document docker-compose v1 constraint (always use docker-compose, not docker compose)
- Add mosquitto network_mode:host + z2m extra_hosts:host-gateway explanation
- Add z2m config writable requirement (EROFS failure mode documented)
- Add chelsty-ha monitor:false rationale
- Add minimal configuration.yaml template for z2m
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-27 16:18:31 +02:00
oskar
dc483ae31a
docs(chelsty): update docs and topology for site/node split
...
- chelsty-runtime.md: references chelsty-infra and chelsty-ha nodes
- chelsty-stability-agent.md: scoped to chelsty-infra
- topology.yaml: chelsty monolith replaced with chelsty-infra + chelsty-ha
2026-05-20 14:23:57 +02:00
oskar
8a12b7ff17
docs: uzupelnij dokumentacje pod katem agentow AI
...
Co-authored-by: Junie <junie@jetbrains.com>
2026-05-20 12:06:23 +02:00
oskar
9f20dcae05
Add control plane deploy script and fix UI healthcheck
2026-05-18 21:34:57 +02:00
oskar
b129f03837
Fix stability agent fleet deploy scripts
2026-05-17 21:09:06 +02:00
oskar
b7faac00c5
Add executable stability agent fleet deploy scripts
2026-05-17 17:32:10 +02:00
oskar
8f305ba3df
Merge VPS control plane deployment and observer runtime
2026-05-17 17:30:04 +02:00
oskar
c9ddfa9ac1
Roll out stability agent to homelab nodes
2026-05-17 15:54:19 +02:00
oskar
8d0f2379ba
Add CHELSTY stability agent
2026-05-15 18:51:45 +02:00
Oskar Kapala
2029457f57
Implement VPS control-plane deployment profile
2026-05-12 20:19:05 +02:00
Oskar Kapala
8f5b905015
Implement observer runtime world synthesis engine
2026-05-12 14:07:03 +02:00
Oskar Kapala
431d777989
Implement filesystem-first runtime event system
2026-05-12 13:38:25 +02:00
Oskar Kapala
0eeb0ac600
Implement reproducible node onboarding
2026-05-12 13:18:00 +02:00
Oskar Kapala
81bce00bf3
Bootstrap CHELSTY runtime stack
2026-05-11 21:36:10 +02:00
Oskar Kapala
b524a3886a
Harden deployment runtime framework
2026-05-11 21:20:13 +02:00
Oskar Kapala
5947ddd03d
Implement staged deployment runtime
2026-05-11 21:04:24 +02:00
Oskar Kapala
31e84a139c
Add CHELSTY home automation inventory model
...
# Conflicts:
# hosts/chelsty/networking.yaml
2026-05-11 20:54:54 +02:00
Oskar Kapala
bbdbdb8321
Add node capability model
2026-05-11 20:46:50 +02:00
Oskar Kapala
9b85ec5e5f
Add topology inventory foundation
2026-05-10 22:05:16 +02:00
Oskar Kapala
d0540f7eb8
Add infrastructure standards and deployment conventions
2026-05-07 21:16:03 +02:00
Oskar Kapala
4064f38b28
Document VPS connectivity diagnostics
2026-04-15 20:38:09 +02:00
Oskar Kapala
03281b989a
Document Hetzner VPS handoff
2026-04-15 17:46:42 +02:00
Oskar Kapala
a1a74f30ba
Document current homelab state
2026-04-15 17:37:25 +02:00