Commit graph

371 commits

Author SHA1 Message Date
Oskar Kapala b51af01fa7 fix(capabilities): saturn RAM/storage + solaria CPU/storage — reconcile with real hardware (audit 2026-06-30)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-03 14:45:53 +02:00
oskar 8ccd58b8a1 docs(sesja): modul 0 wykonany (+1Gi PIHA) + migracja forgejo/vikunja na kapala.org (OIDC gotcha) 2026-07-02 17:37:07 +02:00
oskar bb9ddde91d docs(backlog): zamkniecia po reconie 2026-07-02 + followupy z rozbrajania min
Zamkniete: bug B (ghost kontenery — zniknely), pending poll-Prometheus-watchdog
(potwierdzony), miny #1/#2/#3 z inwentaryzacji. Dodane followupy: wpisy hostowe
forgejo/mosquitto, mem_limit mosquitto@VPS, mosquitto per-host (chelsty-infra),
broker :1883 w topology, pi-watchtower-1 restart-loop, alias lustro, mem_limit
fleet-prometheus. Ocena joplin-db postgres:18 zdezaktualizowana (PG18 GA).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 17:30:48 +02:00
oskar 0bcf5e505f docs(sessions): sesja 2026-07-02 — recon-weryfikacja inwentaryzacji, trzy miny rozbrojone
Recon Fable (57a6dff): bilans 23 rozjazdow (20 aktualnych / 2 zmienione /
1 wyjasniony), dwa pendingi domkniete (poll Prometheus w brain-watchdog
potwierdzony; ghost kontenery B zniknely). Miny: #1 PIHA checkout
(lekcja checkout-vs-reset), #2 slepy control-plane na SATURN (compose down),
#3 owner_node (886bc85).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 17:30:48 +02:00
oskar b6d55949ba fix(vikunja): PUBLICURL na vikunja.kapala.org 2026-07-02 17:28:56 +02:00
oskar f454ac7448 fix(vikunja): OIDC na forgejo.kapala.org + redirect vikunja.kapala.org (cert okit.pl wygasl) 2026-07-02 17:24:51 +02:00
oskar 886bc85e0b fix(inventory): correct owner_node — forgejo→piha, mosquitto→vps (per 2026-07-02 verify)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 16:48:12 +02:00
oskar 6b35afd2ea docs(infra): egzekucja odchudzania PIHA — ES+diskover ubite (+1.0Gi), llm-gateway/immich zostają
Faza 2 modułu 0 po review Oskara: elasticsearch+diskover usunięte (compose down,
dane esdata zostawione na dysku), available 2.8Gi -> 3.8Gi, kryterium >=1.5Gi
spełnione. llm-gateway udokumentowany (własny router LLM -> Ollama@SOLARIA,
źródło tylko w /opt/llm-gateway — archiwizacja w backlogu); immich zostaje na
PIHA na stałe (24/7, SOLARIA sesyjna).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 16:44:57 +02:00
oskar f1e302d19b docs(infra): audyt odchudzania PIHA — 917Mi bezpieczne (ES+diskover+llm-gateway), immich->SOLARIA kandydat 2026-07-02 16:11:41 +02:00
oskar 57a6dffc5e docs(infra): weryfikacja inwentaryzacji 2026-06-30 — stan na 2026-07-02
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 15:26:18 +02:00
oskar 22adfb1c8e docs(kb): kb-02 moduly 4-5 — Nextcloud + ingest->koperta (komplet szkieletow 0-5) 2026-07-02 15:09:40 +02:00
Oskar Kapala c5dd1ecb56 docs(kb): kb-02 architektura dokumentow — master + moduly 0-3 (split Paperless PIHA/OCR SOLARIA) 2026-07-01 21:56:07 +02:00
Oskar Kapala 727cb999ef docs(sesja): inwentaryzacja floty 2026-06-30 + dysk SATURN + safeclean; backlog 23 rozjazdy 2026-06-30 19:56:50 +02:00
Oskar Kapala 229f85bd9b docs(infra): inwentaryzacja floty 2026-06-30 — 23 rozjazdy repo↔rzeczywistość 2026-06-30 19:40:50 +02:00
Oskar Kapala 4055a8ffab docs(backlog): oznacz kroki 4+5 Prometheus jako ZROBIONE, dopisz tech-debt log PROMETHEUS_URL
Kroki 4 (reguły liveness d417000) i 5 (watchdog poll 62d6fc0) z planu monitoringu zamknięte.
Nowy wpis aktywny: brain-watchdog nie loguje PROMETHEUS_URL przy starcie — utrudnia weryfikację.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 19:29:24 +02:00
Oskar Kapala 0656635793 docs(sesja): fleet-prometheus liveness — reguły NodeDown + watchdog→Prometheus poll
Sesja 2026-06-30: commits d417000 (rules/liveness.yml, NodeDown inactive=OK) i 62d6fc0
(brain-watchdog/check_prometheus_alerts, architektura A, 12 testów). Incydent PIHA
divergent branches — praca uratowana verify-before-reset. PENDING: end-to-end firing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 19:29:20 +02:00
Oskar Kapala 62d6fc066b feat(brain-watchdog): poll Prometheus /api/v1/alerts as second alert source (Telegram)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 18:51:34 +02:00
Oskar Kapala d4170004e7 feat(fleet-prometheus): add liveness alert rules for always-on nodes (vps, piha)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 16:42:52 +02:00
Oskar Kapala 831bbedffa docs(sesja): kapala.org → Cloudflare wildcard DNS-01, HA i Immich na mesh
Sesja 2026-06-30: migracja DNS kapala.org z 42.pl na Cloudflare, wildcard
*.kapala.org przez DNS-01, ha/immich.kapala.org przez NPM@PIHA mesh-only.

- nowy session log z root-cause buga "unrecognized name" (NPM duplikuje
  proxy_http_version przy Websockets ON) + wzorzec migracji usługi na mesh
- backlog: gotchas (NPM WS config, CF auto-proxy DKIM) + TODO (migracja
  okit.pl, foty renew, cleanup ha-ken add-on, stale ha.okit.pl)
- topology: nowa sekcja ingress (ha/immich → NPM@PIHA, wildcard cert)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 16:23:57 +02:00
Oskar Kapala 48107ef2c1 docs(backlog): zamknij deploy.sh vps bug + 2 nowe wpisy (prom-reload, chelsty node_exporter)
- 🔴 deploy.sh vps↔control-plane → Zamknięte (fix 3b71707, guard deploy-local.sh)
- krok 2 planu floty (targety 100.x) → ZROBIONE (7d4014e)
- NOWE: deploy-node.sh nie reloaduje config-driven serwisów (cicha rozbieżność deploy↔config)
- NOWE: zbadać chelsty/chelsty-infra node_exporter DOWN z VPS
- Problem B (ghost kontenery) pozostaje otwarty osobno

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 17:30:26 +02:00
Oskar Kapala 39ab34643d docs(kb): sesja 2026-06-26 — fleet-prometheus targety floty + zamknięcie buga deploy.sh vps
- targety node_exporter floty (piha/solaria/lustro) wdrożone (7d4014e)
- fix deploy.sh vps↔control-plane potwierdzony w boju (3b71707)
- PENDING: health-verify targetów w /api/v1/targets nie potwierdzony
- lekcja: nie commitować na master równolegle gdy CC pracuje na wątku (potrójny rebase)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 17:30:19 +02:00
Oskar Kapala 7d4014e512 feat(fleet-prometheus): add piha/solaria/lustro node_exporter scrape targets
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:41:09 +02:00
oskar 7a95ab2fff docs(kb): sesja 2026-06-25 — bulk import Gmail uruchomiony (225k kopert na PIHA) 2026-06-26 16:33:17 +02:00
Oskar Kapala 1529911012 docs(backlog): add anomaly-detection liveness idea (mózg learns per-node daily pattern) 2026-06-26 16:20:54 +02:00
Oskar Kapala 3b71707660 fix(deploy): exclude control-plane from node deploy loop — prevents brain teardown via project-name divergence
deploy.sh vps iterates hosts/vps/services.yaml in deploy-node.sh, which
includes control-plane. But control-plane also has a dedicated deploy path
(deploy.sh control-plane → deploy-control-plane.sh → deploy-local.sh) that
runs compose from services/control-plane/. The node loop runs compose from
${REPO_PATH}, so on Compose versions that derive the project name from cwd
the two paths own the containers under different project names. The loop's
`up -d --remove-orphans` then Recreates and tears down the running brain
(observer/supervisor/executor/ui), aborting the loop under set -e. This
wiped the VPS control-plane on 2026-06-25.

Generic guard: skip any service that ships its own services/<svc>/deploy-local.sh.
control-plane stays in services.yaml so the gate (pytest+build) still covers it;
only the destructive loop deploy is skipped. Protects future services with a
dedicated deploy path too.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 15:16:42 +02:00
oskar c0ffb6abf7 docs(kb): sesja 2026-06-22 — spine relokowany na PIHA + przygotowanie hosta
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 18:52:35 +02:00
oskar c32e050d56 docs(kb): sesja 2026-06-24 — importer Gmail gotowy + konwencja jobs/ 2026-06-25 17:41:01 +02:00
oskar 254a6808bc fix(piha): capabilities — realny RAM/NVMe + gitignore build dirs
hosts/piha/capabilities.yaml: memory 4→8 GB, storage sd-card/32 GB → NVMe ~477 GB (/home 410 GB)
.gitignore: packages/*/build/ i jobs/*/build/ (wyciekały po pip install -e)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-25 17:07:11 +02:00
oskar f234280b3a refactor(kb-mail): importer Gmail — entities załączników, --limit, batch, bez Dockera, DSN→PIHA
- Usuwa Dockerfile i docker-compose.yml; job odpalany lokalnie na PIHA (pip install -e)
- _parse_attachments: manifest MIME → entities[]{type,filename,content_type,size,sha256};
  bajty zostają w .eml, wyciąganie/OCR = faza 2
- Batch inserty co 500 wpisów (executemany + ON CONFLICT DO NOTHING); idempotentny
  na skipped przez _eml_ref (mirrors archive._UNSAFE); pełna wznawialność
- --limit N: ucina pętlę po N wiadomościach do testów na próbce
- epoch_fallback: licznik + WARNING gdy Date nieparsowalne/brak
- Nowe stats: msgs_with_attachments, total_attachments, total_attachment_bytes
- DSN w docstringu: localhost:5433/kb i piha:5433/kb; usunięto solaria:5433
- 9 nowych testów (24 razem), wszystkie zielone

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-25 17:07:11 +02:00
oskar 6dc1d325f1 feat(kb-mail): etap 2 — jednorazowy bulk importer Gmail (mbox → archiwum)
One-shot job `jobs/gmail-bulk-import/` wczytuje plik .mbox z Google Takeout
i importuje każdą wiadomość do archiwum .eml + opcjonalnie do koperty w DB.
Idempotentny (FileExistsError → skip; ON CONFLICT DO NOTHING w DB).
15 testów jednostkowych (bez DB, bez zewnętrznych serwisów) — wszystkie zielone.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-25 17:07:11 +02:00
oskar 02d0fa391f docs(deployment): add critical warning — deploy.sh vps destroys control-plane (project-name divergence, 2026-06-25 incident)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-25 16:30:53 +02:00
oskar a17a11bc9d docs(backlog): add 2026-06-25 bugs (deploy.sh vps mina, ghost kontenery, supervisor bez akcji, phantom world-state); close flaky tests + env-file fixes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-25 16:30:51 +02:00
oskar 2655b4c6f4 docs(sessions): add 2026-06-25 session log (fleet-prometheus etap 1 domknięty + incydent mózgu)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-25 16:30:47 +02:00
oskar 992ff7ca7c test(control-plane): isolate _make_observer_simple module-state leak — fixes flaky test_incident_lifecycle
The test_run_once_* cases were flaky/order-dependent. Root cause: observer.observer
derives OBSERVER_STATE_FILE from STATE_DIR at import time. The helper patched
STATE_DIR but never OBSERVER_STATE_FILE, so run_once()/_save_checkpoint() wrote the
checkpoint to the real /opt/homelab/state/observer_checkpoint.json. Those node_checkpoints
(tmp paths tagged with a pytest run number) leaked across tests and across pytest runs;
run_once's `file_path > checkpoint` string compare then skipped/kept events based on
run-number ordering. The helper also never restored the module globals it overwrote.

Replace both ad-hoc helpers with an autouse monkeypatch fixture that redirects every
observer path — including OBSERVER_STATE_FILE — into the per-test tmp_path and reverts
them afterward. Tests no longer touch real disk and are deterministic.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 14:48:03 +02:00
oskar 686aca7060 fix(deploy-node): pass --env-file per-service so env-interpolated binds resolve (fleet-prometheus 0.0.0.0 leak)
Without --env-file, docker compose resolved variables from the repo root
(cwd), not from services/<service>/.env where the file actually lives.
This caused ${TAILSCALE_BIND_IP} to expand to empty string, binding
fleet-prometheus on 0.0.0.0:9090 instead of the Tailscale-only IP —
a security hole on the public VPS.

Guard mirrors the existing override-file pattern: only add --env-file
when the file exists, so services without .env continue to work as
before. Flag is injected into COMPOSE_CMD (before the `up` subcommand)
so docker compose sees it as a global option.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-25 13:41:55 +02:00
oskar afc0a52e3e docs(backlog): add Compose state-drift and flaky-test tech-debts from 2026-06-24
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-24 23:37:52 +02:00
oskar a55c0928e6 docs(sessions): add 2026-06-24 session log (fleet-prometheus etap 1)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-24 23:37:49 +02:00
oskar ae739b5037 fix(deploy-node): resolve HOST_DIR via os_hostname lookup, not bare hostname
Replaces the brittle `hosts/$(hostname|lower)` assumption with a scan
of hosts/*/host.yaml for a matching os_hostname field. This fixes VPS
where the OS hostname (ubuntu-4gb-hel1-1) never matched the repo
directory (hosts/vps/), causing a silent "No services found" false-green.

Fallback to lowercase-hostname dir preserved for nodes that haven't yet
received the os_hostname field; exits 1 with a clear message if neither
match nor fallback directory exists.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-24 22:24:40 +02:00
oskar 644f32d2f7 feat(hosts): add os_hostname field to all host.yaml files
Declares the real OS-level hostname for each node so deploy-node.sh
can resolve the correct hosts/ directory without assuming
OS-hostname == logical name. Only VPS differs: its OS hostname is
ubuntu-4gb-hel1-1 while the repo directory is hosts/vps/.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-24 22:24:34 +02:00
oskar 43c47a0a55 fix(fleet-prometheus): bind listen socket to Tailscale IP, not 0.0.0.0
Previously published "9090:9090" → Docker bound to 0.0.0.0 (all host interfaces,
including the public Hetzner IP), leaving tailscale-internal enforced only by the
VPS firewall. Now bind explicitly to the VPS Tailscale interface for
defense-in-depth: the port does not exist on the public IP at all.

- ports -> "${TAILSCALE_BIND_IP}:9090:9090"
- env.example: add TAILSCALE_BIND_IP (VPS Tailscale IP, verify via `tailscale ip -4`)
- README: deploy section — .env is mandatory; a missing .env makes Compose
  silently bind 0.0.0.0 (warns, does not fail), so use --env-file and verify host_ip

Smoke: config with --env-file and with co-located .env both resolve
host_ip=100.95.58.48; with .env absent Compose warns and falls back to 0.0.0.0
(documented). .env is gitignored (global *.env rule).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 18:07:18 +02:00
oskar 01a0314176 chore(vps): register fleet-prometheus in topology and services
- hosts/vps/services.yaml: fleet-prometheus entry (role: fleet-liveness-source,
  exposure: tailscale-internal, port 9090, runtime paths, depends_on node_exporter)
- inventory/topology.yaml: add fleet-prometheus to the vps services list

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 18:07:18 +02:00
oskar 4518b15f98 feat(fleet-prometheus): scaffold fleet liveness Prometheus (VPS)
Clean scaffold of a Prometheus instance dedicated to fleet liveness, deliberately
separate from the home `prom` on PIHA. Scrapes only itself + the VPS-local
node_exporter for now.

- image pinned to prom/prometheus:v3.5.0 (LTS) — no :latest anti-pattern
- TSDB retention 15d / 2GB on the tight VPS; --web.enable-lifecycle for reloads
- mem_limit 512m, oom_score_adj +200 (sacrificial OOM victim before control-plane)
- tailscale-internal exposure via plain published 9090, mirroring control-plane
- node_exporter scraped via host.docker.internal:9100 (it runs network_mode: host)
- secret-free prometheus.yml; alerting/rule_files left as commented placeholders
- in-container healthcheck via busybox wget (present in the image; curl is not)

Smoke: docker compose config OK; promtool check config SUCCESS; up -d ->
/-/healthy + /-/ready OK; both targets (prometheus, fleet-node) report up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 18:07:18 +02:00
oskar 76ae6e58ef docs(tech-debt): cleanup artefaktow paste-error w ~oskar (SOLARIA, niski prio) 2026-06-24 17:13:51 +02:00
oskar c1f4172340 docs(hosts/vps): notatka host-state — /swapfile 4G + swappiness=10
Celowy stan host-level (poza GitOps) udokumentowany w manifescie, zeby przy
odtwarzaniu VPS nie zniknal cicho. Tylko opisowy blok host_state, bez zmian
struktury runtime/deploy.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 22:15:24 +02:00
oskar ed54c2777a docs(backlog): swap VPS DONE; sekcja planu Prometheus-as-truth floty
Swap 2-4GB na VPS -> Zamkniete (2026-06-22). Nowa sekcja planu monitoringu
floty (Prometheus pull up{}, osobny instance, bez Alertmanagera) z krokami.
Zastepuje szkic blackbox+Alertmanager z 2026-06-17.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 22:15:24 +02:00
oskar 1136a48222 docs(session): 2026-06-22 — decyzja Prometheus-as-truth dla liveness floty
Osobny fleet-Prometheus (pull, up{}) zastępuje warstwę wykrywania
node-agent->rsync->observer. Bez Alertmanagera (brain-watchdog drugie wejście).
Placement VPS. Swap 4G done. Plan kroków OTWARTE.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 22:15:16 +02:00
oskar 2b3cb89144 refactor(kb-postgres): relokacja SOLARIA→PIHA — arm64, mem_limit 1g, tuning pod małą maszynę
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 19:59:58 +02:00
oskar bbfbb698f8 docs(kb): etap 1 fundament done + session log 2026-06-17
- kb-00-overview.md: etap 1 oznaczony jako DONE (kb-postgres + koperta +
  packages/kb-mail), następny krok = etap 2 bulk Gmail; dodana sekcja
  konwencji packages/
- CLAUDE.md: sekcja "Shared Python Libraries (packages/)" — konwencja,
  layout, instalacja w Dockerfile
- docs/sessions/2026-06-17-kb-foundations.md: pełny log sesji (architektura
  KB, decyzje, etap 1 zbudowany, poprawki spójności, następny krok)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-19 20:02:25 +02:00
oskar 1666511475 feat(kb-mail): fundament — pgvector spine, koperta, archiwum, pakiet domeny
- services/kb-postgres: pgvector/pgvector:pg16 na SOLARIA (:5433), named
  volume, init/001_envelope.sql (CREATE EXTENSION vector + zamrożona tabela
  envelope: id/source/ts/geo/raw_ref/entities), service.yaml, healthcheck,
  README z poprawnym mechanizmem deploy (deploy-node.sh składa dwa -f)
- hosts/solaria/runtime/kb-postgres/docker-compose.override.yml: mem_limit 4g
- inventory/topology.yaml + hosts/solaria/services.yaml: kb-postgres wpisany
- packages/kb-mail: nowa konwencja shared lib (pip install /repo/packages/<lib>/)
  envelope.py — @dataclass Envelope, walidacja tz-aware ts
  db.py       — insert_envelope / get_envelope (asyncpg, ON CONFLICT DO NOTHING)
  archive.py  — save_eml append-only (asyncio.to_thread, FileExistsError na dup)
  tests: 15 unit pass + 5 integration (@pytest.mark.integration, wymaga KB_TEST_DSN)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-19 20:02:25 +02:00
oskar 214a6c3877 docs(kb): foundacja — overview + projekt maili (2 żywe źródła) 2026-06-19 20:02:25 +02:00