Compare commits

...

5 commits

Author SHA1 Message Date
oskar 369ba31dde chore(kb): whitelist --check — /opt/homelab jako swiadomy wyjatek
Decyzja operatora 2026-08-04: /opt/homelab to standardowa sciezka deploy rootu
homelaba, ta sama na kazdym wezle, opisana wprost w publicznej czesci modelu
(standards, service-model, observer, event-system). Nie ujawnia sekretow ani
topologii, wiec zostaje na stronie publicznej.

scripts/kb/check_whitelist.txt: wpis `/opt/homelab` z uzasadnieniem i data.
Zakres wyjatku jest waski — sprawdzone, ze wycisza wylacznie warianty
`/opt/homelab/...`; `/home/oskar/...`, `/opt/other/...`, adresy RFC1918 i porty
dalej zapalaja czerwone.

gen_pages.py: load_whitelist() obcina komentarz `#` w dowolnym miejscu linii,
nie tylko na jej poczatku — wyjatek ma stac obok uzasadnienia, a nie osobno.
Zaden ze skanowanych wzorcow nie zawiera `#`, wiec obciecie jest bezpieczne.

kb/runbooks/kb-site-deploy.md §2: zapisany aktualny stan bramki (exit 0, 22
trafienia wyciszone) zamiast opisu decyzji do podjecia.

Po zmianie: python3 scripts/kb/gen_pages.py --check -> CZYSTO, exit 0.
2026-08-04 18:05:45 +02:00
oskar 41d3543082 docs(kb): kb-site — dokument serwisu (public) + runbook deployu (private)
kb/services/kb-site.md (type: service, visibility: public) — opis samej
wystawki: co jest publikowane (regula fail-closed na `visibility`), jak
renderowane sa odnosniki do dokumentow nieopublikowanych, struktura builda,
stopka z commitem, bramka --check. Napisany tak, zeby sam przechodzil --check:
zero adresow, portow i sciezek hosta — szczegoly operacyjne siedza w runbooku.

kb/runbooks/kb-site-deploy.md (type: runbook, visibility: private) — pelna
procedura: deploy na PIHA, generacja + kontrola wyciekow jako bramka
publikacji, podmiana tresci w wolumenie helperem alpine (wzorzec z
services/narty27/README.md, rozszerzony z pliku na drzewo), rekord A w
Cloudflare + Pi-hole local DNS (split-horizon), proxy host i cert przez
scripts/npm/npm_api.py, weryfikacja, rutynowa aktualizacja jednym lancuchem
&&, tabela rollback/typowe problemy.

Dwie rzeczy zapisane wprost, bo latwo je przeoczyc:
- kolejnosc DNS -> NPM host -> cert (HTTP-01 wymaga dzialajacego vhosta),
- `rm -rf /content/*` przed rozpakowaniem — bez tego dokument przelaczony z
  public na private zostaje w wolumenie i dalej jest serwowany.

Runbook notuje tez aktualny wynik --check (22 trafienia /opt/homelab w
dokumentach public) jako decyzje do podjecia przed pierwsza publikacja:
wyczyscic zrodla albo swiadomie wpisac wyjatek do whitelisty.
2026-08-04 18:01:56 +02:00
oskar 2c5894d70b feat(kb-site): serwis nginx na PIHA + manifesty (wzorzec narty27)
services/kb-site/ — nginx:alpine serwujacy wolumen kb-site_content
(:ro, /usr/share/nginx/html), pelny layout z CLAUDE.md: docker-compose.yml,
service.yaml, README-wskaznik, env.example (swiadomie pusty — brak sekretow),
healthcheck.sh. Wpis w hosts/piha/services.yaml.

Port hosta 8250, NIE 8240 z opisu zadania: 8240 jest juz zajete przez narty27
(services/narty27/docker-compose.yml). Blok statyczny PIHA to 8210 paperless,
8220 nextcloud, 8230 kb-query, 8240 narty27 -> 8250 to nastepny wolny.
Uzgodnione z operatorem.

exposure: public — w odroznieniu od narty27 ta wystawka ma byc dostepna z
internetu przez vhost npm@PIHA kb.okit.pl; bind :8250 jest upstreamem proxy,
nie punktem wejscia.

Tresc jest czystym artefaktem repo (wyjscie scripts/kb/gen_pages.py), zyje
wylacznie w wolumenie kb-site_kb-site_content — bez binda pod /opt/homelab/data
i bez zadania backupu: odtworzeniem jest regeneracja z kb/.

Sprawdzone lokalnie: docker compose config -q, bash -n healthcheck.sh,
yaml.safe_load na obu manifestach.
2026-08-04 17:59:05 +02:00
oskar 9591c8b688 feat(kb): gen_pages.py --check — skan wygenerowanego HTML pod katem wyciekow
Tryb --check nie generuje niczego: skanuje build/kb-site/**/*.html linia po
linii i przerywa z exit 1 na pierwszym zestawie trafien. Wzorce:

  ip-rfc1918    192.168./10./172.16-31.
  ip-tailscale  100.64-127.
  ip-public-v4  reszta poprawnych adresow v4 (loopback, link-local, multicast,
                broadcast i pule dokumentacyjne RFC 5737 sa neutralne)
  ip-v6         adresy z "::" albo >=4 grupami (3 grupy to zwykle godzina)
  port          :NNNN w zakresie 1024-65535
  path-host     /home/... i /opt/...
  token-hex     ciagi hex >=32 znakow
  token-b64     ciagi base64-podobne >=40 znakow mieszajace cyfry i litery
                (sciezki absolutne odsiane — raportuje je path-host)

Raport: plik:linia [wzorzec] trafienie, na koncu licznik per wzorzec.

scripts/kb/check_whitelist.txt — swiadome wyjatki, na start PUSTY (same
komentarze z opisem formatu). Wpis to `<fragment>` albo
`<sciezka strony>|<fragment>`; fragment dopasowuje sie jako podciag trafienia,
wiec jeden wpis `/opt/homelab` wycisza wszystkie warianty.

Uruchomione lokalnie na 11 wygenerowanych stronach: 22 trafienia, wszystkie
path-host (/opt/homelab/... w dokumentach public), zero IP, portow i tokenow.
Whitelist zostaje pusta — decyzja co z tymi sciezkami zrobic nalezy do
operatora (patrz kb/runbooks/kb-site-deploy.md).
2026-08-04 17:57:37 +02:00
oskar 65093815a8 feat(kb): generator publicznej warstwy KB (gen_pages.py)
scripts/kb/gen_pages.py — kb/**/*.md -> build/kb-site/ (index.html pogrupowany
per type + strona na dokument). Renderer markdown na samej bibliotece
standardowej, wzorowany na ~/narty-2027/saalbach-kb/gen_pages.py; parser
frontmattera wspoldzielony z check_okf.py, zeby lint i generator widzialy
frontmatter tak samo.

Kwalifikacja fail-closed: publikowany jest wylacznie dokument z jawnym
`visibility: public`. Brak frontmattera, niepoprawny YAML, brak pola albo inna
wartosc = private. Linki do dokumentow nieopublikowanych nie sa renderowane jako
linki — zostaje etykieta z dopiskiem [private]; render_inline ma druga bramke
(linkuje tylko http/mailto/kotwice/.html), wiec martwy odnosnik nie ma jak
przeciec na strone.

BASE_URL jest parametrem (--base-url, domyslnie https://kb.okit.pl) i trafia do
<link rel="canonical">. Stopka kazdej strony: data generacji + git rev-parse
--short HEAD.

check_okf.py: EXCLUDE_DIRS = ("build",) — wyjscie generatora nie jest zrodlem
i nie podlega lintowi. build/ dopisany do .gitignore.

Uruchomione lokalnie: 150 dokumentow kb/, 10 public, 140 pominietych.
2026-08-04 17:57:15 +02:00
12 changed files with 1330 additions and 1 deletions

2
.gitignore vendored
View file

@ -22,6 +22,8 @@ venv/
*.egg-info/
packages/*/build/
jobs/*/build/
# wyjscie generatorow (scripts/kb/gen_pages.py -> build/kb-site/) — artefakt, nie zrodlo
build/
# Tools
.aider*

View file

@ -165,6 +165,26 @@ services:
# No backup job, no /opt/homelab/data bind.
config_path: services/narty27
kb-site:
role: static-html-host # public KB slice generated by scripts/kb/gen_pages.py
deployment_model: docker-compose
exposure: public # via npm@PIHA vhost kb.okit.pl; container binds LAN :8250
offline_required: false
depends_on:
local: []
external: []
ports:
- name: http
container_port: 80
host_port: 8250
protocol: tcp
runtime:
# No config and no secrets. Content is a pure artifact of the repo: it
# lives only in the Docker named volume kb-site_kb-site_content and is
# refreshed by regenerating from kb/ (kb/runbooks/kb-site-deploy.md).
# No backup job, no /opt/homelab/data bind.
config_path: services/kb-site
# --- Known unmanaged containers on piha (recon B5/B6, 2026-07-27) ----------
# ~28 running containers have no entry above and are deliberately NOT being
# added piecemeal — bringing them under desired state is a later stage

View file

@ -0,0 +1,246 @@
---
okf: "0.1"
type: runbook
visibility: private
status: active
updated: 2026-08-04
links:
- ../services/kb-site.md
- ./npm-api.md
- ../phases/okit-cloudflare.md
---
# kb-site — deploy, publikacja treści, ingress (kb.okit.pl)
Serwis: `services/kb-site/` (nginx:alpine na PIHA, host port **8250**).
Generator: `scripts/kb/gen_pages.py`. Publiczny adres: `https://kb.okit.pl`.
Kolejność ma znaczenie: **DNS → NPM host → cert**. Certyfikat Let's Encrypt
leci challenge'em HTTP-01, więc rekord A musi już wskazywać na łącze domowe,
a vhost musi już odpowiadać na porcie 80, zanim zamówisz cert.
---
## 0. Wymagania wstępne
- PIHA ma checkout repo w `~/homelab-codex-ws` i działającego Dockera.
- npm@PIHA działa (`http://192.168.31.5:81`), port 80/443 z routera jest
przekierowany na PIHA — tak samo jak dla `vikunja.okit.pl`.
- `scripts/npm/.env` wypełniony poświadczeniami (patrz `kb/runbooks/npm-api.md`).
- Strefa `okit.pl` w Cloudflare (patrz `kb/phases/okit-cloudflare.md`).
---
## 1. Deploy serwisu na PIHA
```bash
# na PIHA
cd ~/homelab-codex-ws
git pull
docker compose -f services/kb-site/docker-compose.yml up -d
```
Bez `.env` i bez override'a — serwis nie ma konfiguracji ani sekretów.
Kontener wstanie, ale healthcheck będzie **unhealthy do czasu wgrania treści**
(pusty wolumen = brak `index.html`). To normalne między krokiem 1 a 3.
```bash
docker ps --filter name=kb-site
docker volume ls | grep kb-site # kb-site_kb-site_content
```
---
## 2. Generacja treści + kontrola wycieków (SATURN / SOLARIA)
Generujemy w checkoucie repo, na węźle z którego pracujesz — nie na PIHA.
```bash
python3 scripts/kb/gen_pages.py --base-url https://kb.okit.pl
python3 scripts/kb/gen_pages.py --check
echo $? # 0 = czysto, 1 = trafienia
```
`--check` skanuje **wygenerowany HTML**, nie źródła: adresy IP (RFC1918,
Tailscale 100.64/10, publiczne v4/v6), porty 1024-65535, ścieżki `/home/`
i `/opt/`, długie hexy i ciągi base64 wyglądające na tokeny.
**Kontrola jest bramką publikacji — nie kopiuj treści na PIHA przy exit 1.**
Stan na 2026-08-04 (11 dokumentów public): **exit 0 — czysto, 22 trafienia
wyciszone whitelistą.** Wszystkie wyciszone to `path-host` `/opt/homelab/...`
w dokumentach opisujących layout runtime (`subsystems/observer`,
`subsystems/standards`, `subsystems/service-model`, `subsystems/event-system`,
`subsystems/agent-operating-procedures`, `subsystems/action-approval-model`,
`runbooks/node-onboarding`). Zero IP, zero portów, zero tokenów.
Decyzja operatora 2026-08-04: `/opt/homelab` jest świadomym wyjątkiem —
standardowa ścieżka deploy rootu, ta sama na każdym węźle, nie ujawnia
sekretów ani topologii. Wpis siedzi w `scripts/kb/check_whitelist.txt` razem
z uzasadnieniem. Wycisza wyłącznie warianty `/opt/homelab/...`; `/home/...`,
adresy, porty i tokeny dalej zapalają czerwone.
Każdy kolejny wyjątek podlega tej samej regule: najpierw próba wyczyszczenia
źródła, wpis do whitelisty dopiero jako świadoma decyzja, zawsze z komentarzem
i datą.
---
## 3. Wgranie treści do wolumenu (helper alpine)
Wzorzec z `services/narty27/README.md`, rozszerzony z jednego pliku na drzewo:
`docker cp` nie zapisze do montowania `:ro`, więc treść wjeżdża jednorazowym
kontenerem pomocniczym, który montuje wolumen zapisywalnie.
```bash
# 1. spakuj build (na węźle generującym)
tar -czf /tmp/kb-site.tgz -C build/kb-site .
# 2. wyślij na PIHA
scp /tmp/kb-site.tgz piha:/tmp/kb-site.tgz
# 3. podmień zawartość wolumenu przez helper
ssh piha 'docker run --rm \
-v kb-site_kb-site_content:/content \
-v /tmp:/src:ro \
alpine sh -c "rm -rf /content/* /content/.[!.]* 2>/dev/null; \
tar -xzf /src/kb-site.tgz -C /content"'
# 4. posprzątaj staging
rm /tmp/kb-site.tgz
ssh piha 'rm /tmp/kb-site.tgz'
```
`rm -rf /content/*` przed rozpakowaniem jest **obowiązkowe**: bez tego strona
dokumentu przełączonego z `public` na `private` zostałaby w wolumenie i dalej
serwowała treść, której już nie publikujemy. Publikacja jest podmianą całości,
nie dogrywaniem plików.
Weryfikacja z PIHA:
```bash
services/kb-site/healthcheck.sh
curl -sI http://192.168.31.5:8250/index.html | head -1
```
---
## 4. Cloudflare — rekord A
```
Type: A
Name: kb
Content: <publiczny IP łącza domowego> # ten sam, co ma vikunja.okit.pl
Proxy: DNS only (szara chmurka)
TTL: Auto
```
Docelowy IP weź z istniejącego rekordu `vikunja` w tej samej strefie zamiast
wpisywać z pamięci — łącze domowe potrafi zmienić adres.
**DNS only, nie proxied**: challenge HTTP-01 w kroku 5 musi trafić na nginx,
a nie na krawędź Cloudflare. Włączenie proxy to osobna decyzja, po tym jak cert
już działa.
Split-horizon (patrz `kb/phases/okit-cloudflare.md`): w Pi-hole na PIHA dodaj
Local DNS Record `kb.okit.pl → 192.168.31.5`, żeby klienci w LAN szli prosto do
npm, a nie przez hairpin NAT na routerze.
Sprawdzenie propagacji (z hosta poza LAN albo przez publiczny resolver):
```bash
dig +short kb.okit.pl @1.1.1.1
```
---
## 5. NPM — proxy host + certyfikat
Wszystkie komendy `npm_api.py`**dry-run domyślnie**; realna zmiana dopiero
z `--apply`. Puść najpierw bez `--apply` i przeczytaj payload.
```bash
# 5a. host: kb.okit.pl -> PIHA:8250
python3 scripts/npm/npm_api.py --npm piha create-host \
--domain kb.okit.pl \
--forward-host 192.168.31.5 --forward-port 8250 \
--block-exploits --http2-support
# ...i to samo z --apply
# 5b. cert Let's Encrypt (HTTP-01 — wymaga działającego kroku 4 i 5a)
python3 scripts/npm/npm_api.py --npm piha create-cert --domain kb.okit.pl --apply
# 5c. podepnij cert pod host i wymuś HTTPS
python3 scripts/npm/npm_api.py --npm piha list-hosts # weź HOST_ID
python3 scripts/npm/npm_api.py --npm piha list-certs # weź CERT_ID
python3 scripts/npm/npm_api.py --npm piha set-cert HOST_ID CERT_ID --ssl-forced --apply
```
Jeśli w npm@PIHA istnieje już wildcard `*.okit.pl` (faza 1 z
`kb/phases/okit-cloudflare.md`), pomiń 5b i w 5c podepnij jego `CERT_ID`
jeden cert DNS-01 zamiast kolejnego per-host HTTP-01.
Diagnostyka nieudanego certu: `ssh piha 'docker logs npm --tail 100'` — NPM
zwraca wyjście certbota w treści błędu, a przy porażce challenge'u kasuje
wiersz certyfikatu (szczegóły w `kb/runbooks/npm-api.md`).
---
## 6. Weryfikacja końcowa
```bash
curl -sI https://kb.okit.pl/ | head -1 # 200
curl -s https://kb.okit.pl/ | grep -c 'class="cards"' # spis jest
curl -s https://kb.okit.pl/subsystems/observer.html | tail -5 # stopka: data + commit
```
W przeglądarce: `https://kb.okit.pl` — spis pogrupowany per type, kłódka bez
ostrzeżeń, wejście w dowolną kartę, powrót linkiem „All documents".
Kontrola treści (ręczna, jednorazowa po pierwszej publikacji): przejrzyj spis
i potwierdź, że nie ma tam nic, czego nie chcesz mieć w internecie. Generator
pilnuje pola `visibility`, ale to człowiek decyduje, co dostaje `public`.
---
## 7. Rutynowa aktualizacja treści
Po każdej zmianie w `kb/**/*.md`, która dotyczy dokumentów `public`:
```bash
python3 scripts/kb/gen_pages.py --base-url https://kb.okit.pl \
&& python3 scripts/kb/gen_pages.py --check \
&& tar -czf /tmp/kb-site.tgz -C build/kb-site . \
&& scp /tmp/kb-site.tgz piha:/tmp/kb-site.tgz \
&& ssh piha 'docker run --rm -v kb-site_kb-site_content:/content -v /tmp:/src:ro \
alpine sh -c "rm -rf /content/* /content/.[!.]* 2>/dev/null; tar -xzf /src/kb-site.tgz -C /content"' \
&& ssh piha 'rm /tmp/kb-site.tgz' \
&& rm /tmp/kb-site.tgz
```
Łańcuch na `&&` jest celowy: `--check` z exit 1 zatrzymuje publikację.
Kontener nie wymaga restartu — nginx czyta wolumen na bieżąco.
---
## 8. Rollback i typowe problemy
| Objaw | Przyczyna | Reakcja |
|---|---|---|
| healthcheck unhealthy, 404 na `/` | pusty wolumen — treść nigdy nie wjechała | krok 3 |
| stara strona po publikacji | pominięte `rm -rf /content/*` | powtórz krok 3 w całości |
| dokument nie pojawia się na stronie | brak `visibility: public` albo niepoprawny frontmatter (fail-closed) | `python3 scripts/kb/check_okf.py`, potem regeneracja |
| link renderuje się jako tekst z `[private]` | cel jest prywatny albo nie istnieje | tak ma być — to nie błąd |
| `--check` zapala nowe trafienie | do dokumentu `public` wjechał adres/ścieżka/token | popraw źródło; whitelist tylko świadomie |
| cert nie schodzi | rekord A nie propagował, proxy Cloudflare włączone, albo port 80 niedostępny | krok 4, potem 5b ponownie |
Awaryjne wygaszenie wystawki (bez ruszania NPM i DNS):
```bash
ssh piha 'cd ~/homelab-codex-ws && docker compose -f services/kb-site/docker-compose.yml down'
```
Wolumen zostaje; `up -d` przywraca stan sprzed. Pełne odtworzenie od zera to
kroki 1-3 — treść jest artefaktem repo, nie danymi.

74
kb/services/kb-site.md Normal file
View file

@ -0,0 +1,74 @@
---
okf: "0.1"
type: service
visibility: public
status: active
updated: 2026-08-04
links:
- ../runbooks/kb-site-deploy.md
---
# kb-site
Public slice of this knowledge base, served as static HTML at `kb.okit.pl`.
Plain `nginx:alpine` on the PIHA node reading one Docker named volume — no
build step at runtime, no database, no dependencies.
The site is a rendering, not a source. Every page is generated from the
Markdown documents of the knowledge base by `scripts/kb/gen_pages.py` and
copied into the volume; the HTML is never edited by hand and never committed.
## What gets published
Only documents that carry an explicit `visibility: public` field in their
frontmatter. The generator is **fail-closed**: a document with no frontmatter,
with unparseable frontmatter, with no `visibility` field, or with any other
value is treated as private and stays out of the build.
The same rule applies to cross-references. A link pointing at a document that
was not published is not rendered as a link — only the label survives, marked
`[private]`. A public page therefore never exposes the location of an internal
document and never produces a dead link.
Frontmatter shown on a page is deliberately partial: type, status and the last
update date. The `links` field is omitted, because a path to a private document
is already a leak of its name.
## Structure of the build
```
build/kb-site/
index.html list of all published documents, grouped by type
<directory>/<name>.html one page per document, mirroring the source tree
```
Each page carries a footer with the generation timestamp and the short commit
hash of the repository state it was rendered from, so any page can be traced
back to an exact revision.
## Leak gate
`gen_pages.py --check` re-reads the generated HTML — not the sources — and
fails on anything that looks like infrastructure detail leaking into a public
page: private, carrier-grade and public IP addresses (v4 and v6), high service
port numbers, absolute host paths, and long hex or base64 strings that look
like credentials. Deliberate exceptions live in a whitelist file that starts
out empty, so every exception is a recorded decision.
The check is a release gate: content is copied to the host only after it
passes.
## Operations
Deployment, content refresh, reverse-proxy and DNS setup are described in the
[kb-site deployment runbook](../runbooks/kb-site-deploy.md), which is internal —
on this site the reference above is plain text, exactly as described in the
previous section.
## Content lifetime
The volume holds an artifact, not data. There is no backup job — recovery is a
regeneration from the repository. Because the generator writes a fresh tree on
every run and the publish step replaces the volume contents wholesale, a
document that flips from public to private disappears from the site on the next
publish.

View file

@ -21,7 +21,8 @@ reguły tego repo:
Zakres domyślny: kb/ oraz docs/sessions/, z wyłączeniem README-wskaźników
(POINTER_GLOBS) te nawigacją do kb-doca, nie dokumentami KB, i celowo nie
mają frontmattera OKF. Reszta repo (CLAUDE.md, README.md, .claude/skills/ itd.)
mają frontmattera OKF oraz katalogów z artefaktami (EXCLUDE_DIRS: build/).
Reszta repo (CLAUDE.md, README.md, .claude/skills/ itd.)
leży poza bazą wiedzy i nie podlega walidacji.
Tylko biblioteka standardowa: minimalny parser YAML wystarczający dla
@ -64,6 +65,11 @@ POINTER_GLOBS = (
"hosts/*/README.md",
)
# Katalogi z artefaktami generatorow (build/kb-site/ z scripts/kb/gen_pages.py).
# Wyjscie generatora nie jest zrodlem — nie ma podlegac lintowi ani teraz, ani
# gdyby SCOPE kiedys sie poszerzyl.
EXCLUDE_DIRS = ("build",)
DATE_RE = re.compile(r"^\d{4}-\d{2}-\d{2}$")
@ -71,6 +77,10 @@ def is_pointer(rel: str) -> bool:
return any(PurePosixPath(rel).match(pat) for pat in POINTER_GLOBS)
def is_excluded(rel: str) -> bool:
return any(part in EXCLUDE_DIRS for part in PurePosixPath(rel).parts)
def split_frontmatter(text: str) -> tuple[str | None, str]:
"""Zwraca (blok frontmattera lub None, reszta dokumentu)."""
if not text.startswith("---\n"):
@ -140,6 +150,7 @@ def scope_files(root: Path) -> list[Path]:
p for p in base.rglob("*.md")
if ".git" not in p.parts
and not is_pointer(p.relative_to(root).as_posix())
and not is_excluded(p.relative_to(root).as_posix())
)
return sorted(set(files))

View file

@ -0,0 +1,26 @@
# Świadome wyjątki dla `python3 scripts/kb/gen_pages.py --check`.
#
# --check skanuje WYGENEROWANY HTML (build/kb-site/) i przerywa z exit 1, gdy
# znajdzie adres IP, port, ścieżkę hosta (/home/, /opt/) albo coś, co wygląda
# na token. Ten plik wycisza pojedyncze, świadomie zaakceptowane trafienia.
#
# Format — jeden wpis na linię, `#` zaczyna komentarz (także w środku linii):
#
# <fragment> wycisza KAŻDE trafienie zawierające <fragment>
# <ścieżka strony>|<fragment> to samo, ale tylko w tej jednej stronie
#
# Ścieżka strony jest względna wobec katalogu wyjściowego, np.
# `subsystems/observer.html`. Fragment dopasowuje się jako podciąg trafienia,
# więc wpis `/opt/homelab` wycisza wszystkie warianty `/opt/homelab/...`.
#
# Zasada: najpierw popraw źródło w kb/ (dokument `visibility: public` nie
# powinien zawierać adresów ani ścieżek konkretnego hosta). Whitelist jest
# ostatecznością dla rzeczy, które naprawdę mają zostać na stronie — każdy wpis
# ma mieć obok uzasadnienie i datę decyzji.
# Decyzja operatora 2026-08-04: standardowa ścieżka deploy rootu homelaba,
# opisana wprost w publicznej części modelu (kb/subsystems/standards.md,
# service-model, observer, event-system). Nie ujawnia sekretów ani topologii —
# ta sama ścieżka jest na każdym węźle i nie mówi nic o tym, co na nim stoi.
# Wycisza warianty `/opt/homelab/...`; `/home/...` NIE jest objęte.
/opt/homelab

853
scripts/kb/gen_pages.py Normal file
View file

@ -0,0 +1,853 @@
#!/usr/bin/env python3
"""Generator publicznej warstwy bazy wiedzy (kb.okit.pl) z dokumentów OKF v0.1.
Wzorzec: ~/narty-2027/saalbach-kb/gen_pages.py renderer markdown na samej
bibliotece standardowej plus prosty, czytelny szablon HTML. Tutaj dochodzi
jedna reguła nadrzędna nad wszystkim innym:
PUBLIKUJEMY WYŁĄCZNIE DOKUMENTY Z `visibility: public`.
Kwalifikacja jest **fail-closed** dokument trafia na stronę tylko wtedy, gdy
da się jednoznacznie stwierdzić, że jest publiczny. Brak frontmattera, niepoprawny
YAML, brak pola `visibility`, pusta albo nierozpoznana wartość = PRIVATE. Każdy
inny wariant (np. domyślnie public, gdy nic nie napisano") oznaczałby, że
przeoczony frontmatter wypycha wewnętrzny dokument do internetu.
Ta sama reguła obowiązuje linki: odnośnik do dokumentu, który nie został
opublikowany, NIE jest renderowany jako link zostaje sam tekst etykiety
z dopiskiem `[private]`. Dzięki temu strona publiczna nigdy nie wskazuje
ścieżek prywatnych dokumentów ani nie generuje martwych 404.
Wejście: kb/**/*.md
Wyjście: build/kb-site/
index.html spis stron pogrupowany per `type`
<katalog>/<nazwa>.html strona na dokument (lustro drzewa kb/)
Stopka każdej strony: data generacji + krótki hash commita (`git rev-parse --short HEAD`).
Tryb `--check` nie generuje niczego skanuje JUŻ WYGENEROWANY katalog wyjściowy
w poszukiwaniu wycieków (adresy IP, porty, ścieżki hosta, tokeny). Świadome
wyjątki trzymamy w scripts/kb/check_whitelist.txt. Trafienie = exit 1.
Tylko biblioteka standardowa (jak scripts/npm/npm_api.py) skrypt uruchamiany
doraźnie z SATURN/SOLARIA, bez własnego obrazu i bez `pip install`. Parser
frontmattera jest współdzielony z check_okf.py, żeby obie ścieżki widziały
frontmatter dokładnie tak samo.
Uruchomienie:
python3 scripts/kb/gen_pages.py [--base-url URL] [--out KATALOG]
python3 scripts/kb/gen_pages.py --check [--out KATALOG] [--whitelist PLIK]
"""
from __future__ import annotations
import argparse
import html
import posixpath
import re
import shutil
import subprocess
import sys
from dataclasses import dataclass
from datetime import datetime, timezone
from pathlib import Path
# Ten sam parser frontmattera co walidator OKF — check_okf.py leży w tym samym
# katalogu, więc import działa niezależnie od katalogu uruchomienia.
sys.path.insert(0, str(Path(__file__).resolve().parent))
from check_okf import parse_yaml, split_frontmatter # noqa: E402
REPO_ROOT = Path(__file__).resolve().parents[2]
KB_DIR = REPO_ROOT / "kb"
DEFAULT_OUT = REPO_ROOT / "build" / "kb-site"
DEFAULT_BASE_URL = "https://kb.okit.pl"
DEFAULT_WHITELIST = Path(__file__).resolve().parent / "check_whitelist.txt"
SITE_NAME = "homelab-codex — knowledge base"
SITE_LEAD = (
"Public slice of the homelab knowledge base. Only documents marked "
"<code>visibility: public</code> are published here; everything else stays "
"in the private repository."
)
PUBLIC = "public"
# Nagłówki grup na stronie spisu. Kolejność listy = kolejność sekcji; typy spoza
# listy trafiają na koniec, alfabetycznie, pod własną nazwą.
TYPE_ORDER = [
("subsystem", "Subsystems"),
("service", "Services"),
("node", "Nodes"),
("runbook", "Runbooks"),
("decision", "Decisions"),
("incident", "Incidents"),
("phase", "Phases"),
("audit", "Audits"),
("session-log", "Session logs"),
]
# --- model ------------------------------------------------------------
@dataclass
class Doc:
id: str # ścieżka względem kb/ bez rozszerzenia, np. "subsystems/observer"
path: Path
frontmatter: dict
body: str
public: bool
@property
def page(self) -> str:
"""Ścieżka strony względem katalogu wyjściowego."""
return f"{self.id}.html"
@property
def doc_type(self) -> str:
return str(self.frontmatter.get("type") or "").strip() or "?"
@property
def title(self) -> str:
heading = first_heading(self.body)
return heading or self.id.rsplit("/", 1)[-1].replace("-", " ")
# Nie każdy dokument zaczyna się od `#` — np. kb/subsystems/action-approval-model.md
# otwiera się `###`. Tytułem strony jest pierwszy nagłówek DOWOLNEGO poziomu.
_ANY_HEADING_RE = re.compile(r"^#{1,6}\s+(.+)$", re.M)
def first_heading(body: str) -> str | None:
match = _ANY_HEADING_RE.search(body)
return match.group(1).strip() if match else None
def is_public(frontmatter: dict | None) -> bool:
"""Fail-closed: publiczny jest wyłącznie dokument z jawnym `visibility: public`."""
if not isinstance(frontmatter, dict):
return False
return str(frontmatter.get("visibility", "")).strip() == PUBLIC
def load_docs() -> list[Doc]:
docs: list[Doc] = []
for path in sorted(KB_DIR.rglob("*.md")):
if ".git" in path.parts:
continue
text = path.read_text(encoding="utf-8")
frontmatter: dict = {}
body = text
try:
block, rest = split_frontmatter(text)
if block is not None:
# Niepoprawny YAML => frontmatter nieznany => dokument prywatny.
frontmatter = parse_yaml(block)
body = rest
except ValueError:
frontmatter = {}
body = text
docs.append(
Doc(
id=path.relative_to(KB_DIR).with_suffix("").as_posix(),
path=path,
frontmatter=frontmatter,
body=body.lstrip("\n"),
public=is_public(frontmatter),
)
)
return docs
# --- linki ------------------------------------------------------------
# Etykieta linku bywa złamana na dwie linie (np. kb/decisions/ai-cluster-legacy.md),
# więc klasa [^\]] musi łapać też znak nowej linii — domyślnie łapie.
_MD_LINK_RE = re.compile(r"\[([^\]]*)\]\(([^)\s]+)((?:#[^)\s]*)?)\)")
_EXTERNAL_PREFIXES = ("http://", "https://", "mailto:", "#")
def rewrite_links(doc: Doc, published: dict[str, str]) -> str:
"""Linki wewnątrz-repowe → strony; wszystko nieopublikowane → zwykły tekst.
`published` mapuje id dokumentu (ścieżka względem kb/ bez rozszerzenia) na
ścieżkę wygenerowanej strony. Cel spoza tej mapy jest z definicji niepubliczny
(private albo w ogóle nie jest dokumentem KB) i traci link.
"""
here = posixpath.dirname(doc.page)
def repl(match: re.Match) -> str:
label, target, fragment = match.group(1), match.group(2), match.group(3)
if target.startswith(_EXTERNAL_PREFIXES):
return match.group(0)
if not target.endswith(".md"):
# Odnośnik do pliku repo (compose, yaml, skrypt) — na stronie
# publicznej nie ma czego pokazać, więc zostaje sama etykieta.
return label
if target.startswith("/"):
rel = target.lstrip("/")
rel = rel[3:] if rel.startswith("kb/") else rel
else:
rel = posixpath.normpath(
posixpath.join(posixpath.dirname(doc.id), target)
)
doc_id = rel[:-3]
page = published.get(doc_id)
if page is None:
return f"{label} [private]"
return f"[{label}]({posixpath.relpath(page, here or '.')}{fragment})"
return _MD_LINK_RE.sub(repl, doc.body)
# --- renderer markdown (biblioteka standardowa) ------------------------
_CODE_RE = re.compile(r"`([^`]+)`")
_LINK_RE = re.compile(r"\[([^\]]*)\]\(([^)\s]+)\)")
_BOLD_RE = re.compile(r"\*\*(.+?)\*\*", re.S)
_ITALIC_STAR_RE = re.compile(r"(?<!\*)\*(?!\s)([^*]+?)(?<!\s)\*(?!\*)")
_ITALIC_UNDER_RE = re.compile(r"(?<![\w\\])_(?!\s)([^_]+?)(?<!\s)_(?![\w])")
_HEADING_RE = re.compile(r"^(#{1,6})\s+(.*)$")
_HR_RE = re.compile(r"^(-{3,}|\*{3,}|_{3,})$")
_LIST_RE = re.compile(r"^\s*([-*+]|\d+[.)])\s+(.*)$")
_TASK_RE = re.compile(r"^\[([ xX])\]\s+(.*)$")
_TABLE_SEP_RE = re.compile(r"^\s*\|?\s*:?-{2,}:?\s*(\|\s*:?-{2,}:?\s*)*\|?\s*$")
def _linkable(href: str) -> bool:
"""Druga bramka fail-closed: linkujemy tylko adresy zewnętrzne, kotwice
i wygenerowane strony. Cokolwiek innego przeciekło przez rewrite_links()
(ścieżka repo, .md bez odpowiednika) zostaje tekstem, nie martwym linkiem."""
return href.startswith(_EXTERNAL_PREFIXES) or href.split("#", 1)[0].endswith(".html")
def render_inline(text: str) -> str:
codes: list[str] = []
def stash(match: re.Match) -> str:
codes.append(match.group(1))
return f"\x00c{len(codes) - 1}\x00"
text = _CODE_RE.sub(stash, text)
text = html.escape(text, quote=False)
def link(match: re.Match) -> str:
label, href = match.group(1), match.group(2)
if not _linkable(href):
return label
attrs = ""
if href.startswith(("http://", "https://", "mailto:")):
attrs = ' target="_blank" rel="noopener" class="ext"'
return f'<a href="{html.escape(href, quote=True)}"{attrs}>{label}</a>'
text = _LINK_RE.sub(link, text)
text = _BOLD_RE.sub(r"<strong>\1</strong>", text)
text = _ITALIC_STAR_RE.sub(r"<em>\1</em>", text)
text = _ITALIC_UNDER_RE.sub(r"<em>\1</em>", text)
for i, code in enumerate(codes):
text = text.replace(f"\x00c{i}\x00", f"<code>{html.escape(code, quote=False)}</code>")
return text
def _alignments(sep: str) -> list[str]:
out = []
for cell in _cells(sep):
left, right = cell.startswith(":"), cell.endswith(":")
out.append("center" if left and right else "right" if right else "left")
return out
def _cells(row: str) -> list[str]:
return [c.strip() for c in row.strip().strip("|").split("|")]
def _starts_block(line: str) -> bool:
s = line.strip()
return (
not s
or bool(_HEADING_RE.match(s))
or bool(_HR_RE.match(s))
or s.startswith(("```", ">"))
or bool(_LIST_RE.match(line))
)
def render_blocks(text: str) -> str:
lines = text.split("\n")
out: list[str] = []
i = 0
while i < len(lines):
line = lines[i]
stripped = line.strip()
if not stripped:
i += 1
continue
if stripped.startswith("```"):
i += 1
code: list[str] = []
while i < len(lines) and not lines[i].strip().startswith("```"):
code.append(lines[i])
i += 1
i += 1
out.append(f"<pre><code>{html.escape(chr(10).join(code), quote=False)}</code></pre>")
continue
heading = _HEADING_RE.match(stripped)
if heading:
level = len(heading.group(1))
out.append(f"<h{level}>{render_inline(heading.group(2))}</h{level}>")
i += 1
continue
if _HR_RE.match(stripped):
out.append("<hr>")
i += 1
continue
if "|" in stripped and i + 1 < len(lines) and _TABLE_SEP_RE.match(lines[i + 1]):
header = _cells(stripped)
align = _alignments(lines[i + 1])
i += 2
rows = []
while i < len(lines) and "|" in lines[i] and lines[i].strip():
rows.append(_cells(lines[i]))
i += 1
head = "".join(
f'<th style="text-align:{align[n] if n < len(align) else "left"}">'
f"{render_inline(c)}</th>"
for n, c in enumerate(header)
)
body = "".join(
"<tr>"
+ "".join(
f'<td style="text-align:{align[n] if n < len(align) else "left"}">'
f"{render_inline(c)}</td>"
for n, c in enumerate(row)
)
+ "</tr>"
for row in rows
)
out.append(
f'<div class="table-wrap"><table><thead><tr>{head}</tr></thead>'
f"<tbody>{body}</tbody></table></div>"
)
continue
if stripped.startswith(">"):
quoted = []
while i < len(lines) and lines[i].strip().startswith(">"):
quoted.append(lines[i].strip()[1:].lstrip())
i += 1
out.append(f"<blockquote>{render_blocks(chr(10).join(quoted))}</blockquote>")
continue
item = _LIST_RE.match(line)
if item:
ordered = item.group(1)[0].isdigit()
items: list[str] = []
while i < len(lines):
current = lines[i]
match = _LIST_RE.match(current)
if match:
items.append(match.group(2))
i += 1
elif current.startswith((" ", "\t")) and current.strip() and items:
items[-1] += " " + current.strip()
i += 1
else:
break
rendered = []
for raw in items:
task = _TASK_RE.match(raw)
if task:
checked = " checked" if task.group(1).lower() == "x" else ""
rendered.append(
f'<li class="task"><input type="checkbox" disabled{checked}> '
f"{render_inline(task.group(2))}</li>"
)
else:
rendered.append(f"<li>{render_inline(raw)}</li>")
tag = "ol" if ordered else "ul"
out.append(f"<{tag}>{''.join(rendered)}</{tag}>")
continue
paragraph = [stripped]
i += 1
while i < len(lines) and not _starts_block(lines[i]):
if "|" in lines[i] and i + 1 < len(lines) and _TABLE_SEP_RE.match(lines[i + 1]):
break
paragraph.append(lines[i].strip())
i += 1
out.append(f"<p>{render_inline(' '.join(paragraph))}</p>")
return "\n".join(out)
# --- szablon ----------------------------------------------------------
CSS = """
:root { color-scheme: light dark; --fg:#0f172a; --muted:#64748b; --bg:#ffffff;
--line:#e2e8f0; --accent:#0d9488; --chip:#f1f5f9; --quote:#f8fafc; }
@media (prefers-color-scheme: dark) {
:root { --fg:#e2e8f0; --muted:#94a3b8; --bg:#0f172a; --line:#334155;
--accent:#2dd4bf; --chip:#1e293b; --quote:#1e293b; }
}
* { box-sizing: border-box; }
body { margin: 0 auto; padding: 1.25rem 1rem 4rem; max-width: 46rem;
font-family: system-ui, -apple-system, "Segoe UI", Roboto, sans-serif;
font-size: 17px; line-height: 1.55; color: var(--fg); background: var(--bg);
-webkit-text-size-adjust: 100%; }
nav { font-size: .85rem; margin-bottom: 1.25rem; }
nav a { color: var(--muted); }
h1 { font-size: 1.55rem; line-height: 1.25; margin: 0 0 .35rem; }
h2 { font-size: 1.2rem; margin: 2rem 0 .5rem; padding-top: .5rem;
border-top: 1px solid var(--line); }
h3 { font-size: 1.02rem; margin: 1.4rem 0 .4rem; }
h4, h5, h6 { font-size: .95rem; margin: 1.1rem 0 .3rem; }
p, ul, ol { margin: .6rem 0; }
ul, ol { padding-left: 1.3rem; }
li { margin: .25rem 0; }
li.task { list-style: none; margin-left: -1.3rem; }
li.task input { margin-right: .45rem; }
a { color: var(--accent); }
a.ext::after { content: " \\2197"; font-size: .8em; color: var(--muted); }
code { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: .88em;
background: var(--chip); padding: .1em .35em; border-radius: 4px; }
pre { background: var(--chip); padding: .75rem; border-radius: 6px; overflow-x: auto; }
pre code { background: none; padding: 0; }
blockquote { margin: .9rem 0; padding: .6rem .9rem; background: var(--quote);
border-left: 3px solid var(--accent); border-radius: 0 6px 6px 0; }
blockquote p { margin: .2rem 0; }
hr { border: 0; border-top: 1px solid var(--line); margin: 1.5rem 0; }
.table-wrap { overflow-x: auto; margin: .8rem 0; -webkit-overflow-scrolling: touch; }
table { border-collapse: collapse; width: 100%; font-size: .92rem; }
th, td { border: 1px solid var(--line); padding: .4rem .55rem; vertical-align: top; }
th { background: var(--chip); font-weight: 600; }
.chip { display: inline-block; background: var(--chip); color: var(--muted);
font-size: .72rem; text-transform: uppercase; letter-spacing: .04em;
padding: .15rem .5rem; border-radius: 999px; }
.lead { color: var(--muted); margin: .35rem 0 1.25rem; }
.meta { color: var(--muted); font-size: .82rem; margin: .2rem 0 1.5rem; }
.cards { list-style: none; padding: 0; margin: 0; }
.cards li { border: 1px solid var(--line); border-radius: 8px; padding: .8rem .9rem;
margin: .6rem 0; }
.cards a { font-weight: 600; text-decoration: none; }
.cards p { margin: .3rem 0 0; font-size: .9rem; color: var(--muted); }
.cards .chip { margin-left: .35rem; }
footer { margin-top: 3rem; padding-top: 1rem; border-top: 1px solid var(--line);
font-size: .78rem; color: var(--muted); }
"""
def page_shell(
title: str,
depth: int,
content: str,
*,
canonical: str,
stamp: str,
nav: str | None = None,
) -> str:
up = "../" * depth
if nav is None:
nav = f'<a href="{up}index.html">← All documents</a>'
nav_html = f"<nav>{nav}</nav>\n" if nav else ""
return f"""<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>{html.escape(title)} {html.escape(SITE_NAME)}</title>
<link rel="canonical" href="{html.escape(canonical, quote=True)}">
<style>{CSS}</style>
</head>
<body>
{nav_html}{content}
<footer>{stamp}</footer>
</body>
</html>
"""
def meta_line(doc: Doc) -> str:
"""Krótki pasek metadanych. Świadomie BEZ pola `links` — wpisy potrafią
wskazywać dokumenty prywatne, a sama ścieżka to już wyciek nazwy."""
parts = [f'<span class="chip">{html.escape(doc.doc_type)}</span>']
status = str(doc.frontmatter.get("status") or "").strip()
updated = str(doc.frontmatter.get("updated") or "").strip()
if status:
parts.append(html.escape(status))
if updated:
parts.append("updated " + html.escape(updated))
return '<p class="meta">' + " · ".join(parts) + "</p>"
def render_page(doc: Doc, published: dict[str, str], base_url: str, stamp: str) -> str:
rendered = render_blocks(rewrite_links(doc, published))
# Pierwszy nagłówek treści awansujemy na <h1> strony zamiast dokładać drugi
# tytuł nad nim (dokument może zaczynać się od `#` albo od `###`).
match = re.match(r"\s*<h[1-6]>(.*?)</h[1-6]>", rendered, re.S)
if match:
headline = match.group(1)
rendered = rendered[match.end() :]
else:
headline = html.escape(doc.title)
content = f"<h1>{headline}</h1>\n{meta_line(doc)}\n{rendered}"
canonical = f"{base_url.rstrip('/')}/{doc.page}"
return page_shell(doc.title, doc.id.count("/"), content, canonical=canonical, stamp=stamp)
_INLINE_MARKUP_RE = re.compile(r"[*`_]|\[private\]")
def summary(doc: Doc, limit: int = 190) -> str:
"""Pierwszy akapit treści (bez nagłówka) jako zajawka na spisie."""
body = _ANY_HEADING_RE.sub("", doc.body, count=1)
paragraph = ""
for chunk in body.split("\n\n"):
text = " ".join(line.strip() for line in chunk.strip().splitlines())
if not text or text.startswith(("#", "|", ">", "```", "-", "*")):
continue
paragraph = text
break
if not paragraph:
return ""
paragraph = _MD_LINK_RE.sub(r"\1", paragraph)
paragraph = _INLINE_MARKUP_RE.sub("", paragraph)
if len(paragraph) > limit:
paragraph = paragraph[:limit].rsplit(" ", 1)[0] + ""
return paragraph
def render_index(docs: list[Doc], base_url: str, stamp: str) -> str:
known = [t for t, _ in TYPE_ORDER]
labels = dict(TYPE_ORDER)
by_type: dict[str, list[Doc]] = {}
for doc in docs:
by_type.setdefault(doc.doc_type, []).append(doc)
order = [t for t in known if t in by_type]
order += sorted(t for t in by_type if t not in known)
sections = []
for doc_type in order:
items = sorted(by_type[doc_type], key=lambda d: d.title.lower())
cards = []
for doc in items:
lead = summary(doc)
cards.append(
f'<li><a href="{html.escape(doc.page, quote=True)}">'
f"{html.escape(doc.title)}</a>"
f'<span class="chip">{html.escape(doc.doc_type)}</span>'
+ (f"<p>{html.escape(lead)}</p>" if lead else "")
+ "</li>"
)
heading = labels.get(doc_type, doc_type)
sections.append(
f"<h2>{html.escape(heading)} <span class=\"chip\">{len(items)}</span></h2>\n"
f'<ul class="cards">{"".join(cards)}</ul>'
)
content = (
f"<h1>{html.escape(SITE_NAME)}</h1>\n"
f'<p class="lead">{SITE_LEAD}</p>\n' + "\n".join(sections)
)
return page_shell(
"Index",
0,
content,
canonical=base_url.rstrip("/") + "/",
stamp=stamp,
nav="",
)
# --- budowa -----------------------------------------------------------
def git_commit() -> str:
try:
out = subprocess.run(
["git", "rev-parse", "--short", "HEAD"],
cwd=REPO_ROOT,
capture_output=True,
text=True,
check=True,
)
return out.stdout.strip() or "unknown"
except (OSError, subprocess.CalledProcessError):
return "unknown"
def build(out_dir: Path, base_url: str) -> list[Doc]:
docs = load_docs()
public = [d for d in docs if d.public]
published = {d.id: d.page for d in public}
generated = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M UTC")
stamp = (
f"Generated {generated} · commit <code>{html.escape(git_commit())}</code> "
f"· <code>scripts/kb/gen_pages.py</code> — do not edit the HTML, "
f"edit the source document and regenerate."
)
if out_dir.exists():
shutil.rmtree(out_dir)
out_dir.mkdir(parents=True)
for doc in public:
target = out_dir / doc.page
target.parent.mkdir(parents=True, exist_ok=True)
target.write_text(render_page(doc, published, base_url, stamp), encoding="utf-8")
(out_dir / "index.html").write_text(
render_index(public, base_url, stamp), encoding="utf-8"
)
print(f"Repo: {REPO_ROOT}")
print(f"Źródło: {KB_DIR.relative_to(REPO_ROOT)}/**/*.md")
print(f"Wyjście: {out_dir}")
print(f"BASE_URL: {base_url}")
print(
f"Dokumenty: {len(docs)} razem, {len(public)} public, "
f"{len(docs) - len(public)} pominiętych (private / brak frontmattera)"
)
by_type: dict[str, list[Doc]] = {}
for doc in public:
by_type.setdefault(doc.doc_type, []).append(doc)
print()
for doc_type in sorted(by_type):
print(f" {doc_type} ({len(by_type[doc_type])}):")
for doc in sorted(by_type[doc_type], key=lambda d: d.id):
print(f" {doc.page} ← kb/{doc.id}.md")
print()
print(f" index.html ({len(public)} pozycji)")
print()
print("Kontrola wycieków: python3 scripts/kb/gen_pages.py --check")
return public
# --- tryb --check: skan wygenerowanego HTML ---------------------------
_OCTET = r"(?:25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)"
_IPV4_RE = re.compile(rf"(?<![\w.]){_OCTET}(?:\.{_OCTET}){{3}}(?![\w.])")
_PRIVATE_IPV4_RE = re.compile(
r"^(?:192\.168\.|10\.|172\.(?:1[6-9]|2\d|3[01])\.)"
)
_TAILSCALE_IPV4_RE = re.compile(
r"^100\.(?:6[4-9]|[7-9]\d|1[01]\d|12[0-7])\."
)
# Adresy, które nie są wyciekiem: loopback, link-local, multicast/broadcast,
# „to sieć", maski i pule dokumentacyjne z RFC 5737.
_NEUTRAL_IPV4_RE = re.compile(
r"^(?:127\.|169\.254\.|0\.|22[4-9]\.|2[3-5]\d\.|255\.|"
r"192\.0\.2\.|198\.51\.100\.|203\.0\.113\.)"
)
# IPv6: wymagamy albo "::", albo >=4 grup — trzy grupy to zwykle godzina 10:15:30.
_IPV6_RE = re.compile(
r"(?<![\w:])(?:"
r"(?:[0-9a-fA-F]{1,4}:){3,7}[0-9a-fA-F]{1,4}"
r"|(?:[0-9a-fA-F]{1,4}:){1,7}:(?:[0-9a-fA-F]{1,4}(?::[0-9a-fA-F]{1,4}){0,6})?"
r"|::(?:[0-9a-fA-F]{1,4}:){0,6}[0-9a-fA-F]{1,4}"
r")(?![\w:])"
)
_NEUTRAL_IPV6 = {"::", "::1"}
# Port leci zaraz po hoście albo adresie (`localhost:8123`, `192.168.31.5:8240`),
# więc przed dwukropkiem stoi znak słowa — żadnego lookbehindu na \w. Odsiewamy
# tylko `::` z adresów IPv6; te i tak raportuje wzorzec `ip-v6`.
_PORT_RE = re.compile(r"(?<!:):(\d{4,5})(?!\d)")
_PATH_RE = re.compile(r"/(?:home|opt)/[\w.\-/]*")
_HEX_TOKEN_RE = re.compile(r"(?<![\w])[0-9a-fA-F]{32,}(?![\w])")
_B64_TOKEN_RE = re.compile(r"(?<![\w+/=-])[A-Za-z0-9+/_-]{40,}={0,2}(?![\w+/=-])")
def _ipv4_hits(text: str) -> list[tuple[str, str]]:
hits = []
for match in _IPV4_RE.finditer(text):
value = match.group(0)
if _PRIVATE_IPV4_RE.match(value):
hits.append(("ip-rfc1918", value))
elif _TAILSCALE_IPV4_RE.match(value):
hits.append(("ip-tailscale", value))
elif _NEUTRAL_IPV4_RE.match(value):
continue
else:
hits.append(("ip-public-v4", value))
return hits
def _ipv6_hits(text: str) -> list[tuple[str, str]]:
hits = []
for match in _IPV6_RE.finditer(text):
value = match.group(0)
if value in _NEUTRAL_IPV6:
continue
hits.append(("ip-v6", value))
return hits
def _port_hits(text: str) -> list[tuple[str, str]]:
hits = []
for match in _PORT_RE.finditer(text):
port = int(match.group(1))
if 1024 <= port <= 65535:
hits.append(("port", match.group(0)))
return hits
def _token_hits(text: str) -> list[tuple[str, str]]:
hits = []
for match in _HEX_TOKEN_RE.finditer(text):
hits.append(("token-hex", match.group(0)))
for match in _B64_TOKEN_RE.finditer(text):
value = match.group(0)
# Ścieżka absolutna nie jest tokenem — alfabet base64 zawiera "/", więc
# długie `/opt/homelab/events/...` łapało się tu jako fałszywy alarm.
# Ścieżki hosta i tak raportuje osobny wzorzec `path-host`.
if value.startswith("/"):
continue
# Bez cyfry i bez litery to nie jest sekret, tylko długie słowo albo
# ciąg myślników — sekrety mieszają jedno z drugim.
if any(c.isdigit() for c in value) and any(c.isalpha() for c in value):
hits.append(("token-b64", value))
return hits
def scan_line(text: str) -> list[tuple[str, str]]:
hits = _ipv4_hits(text) + _ipv6_hits(text) + _port_hits(text)
hits += [("path-host", m.group(0)) for m in _PATH_RE.finditer(text)]
hits += _token_hits(text)
return hits
def load_whitelist(path: Path) -> list[tuple[str | None, str]]:
"""Wpisy: `<fragment>` albo `<ścieżka strony>|<fragment>`.
Fragment jest dopasowywany jako podciąg trafienia, więc jeden wpis
`/opt/homelab` wycisza wszystkie warianty `/opt/homelab/...`.
`#` zaczyna komentarz w dowolnym miejscu linii — wyjątek bez uzasadnienia
obok siebie szybko staje się wyjątkiem, którego nikt już nie rozumie.
Żaden ze skanowanych wzorców (adresy, porty, ścieżki, tokeny) nie zawiera
`#`, więc obcięcie ogona jest bezpieczne.
"""
if not path.is_file():
return []
entries: list[tuple[str | None, str]] = []
for raw in path.read_text(encoding="utf-8").splitlines():
line = raw.split("#", 1)[0].strip()
if not line:
continue
if "|" in line:
scope, _, fragment = line.partition("|")
entries.append((scope.strip(), fragment.strip()))
else:
entries.append((None, line))
return entries
def whitelisted(entries: list[tuple[str | None, str]], rel: str, value: str) -> bool:
return any(
fragment and fragment in value and (scope is None or scope == rel)
for scope, fragment in entries
)
def check(out_dir: Path, whitelist_path: Path) -> int:
if not out_dir.is_dir():
print(f"BRAK katalogu {out_dir} — najpierw wygeneruj strony (bez --check).")
return 1
entries = load_whitelist(whitelist_path)
files = sorted(out_dir.rglob("*.html"))
print(f"Skan: {out_dir}")
print(f"Whitelist: {whitelist_path} ({len(entries)} wpis(ów))")
print(f"Plików: {len(files)}")
print()
findings: list[tuple[str, int, str, str]] = []
suppressed = 0
for path in files:
rel = path.relative_to(out_dir).as_posix()
for lineno, line in enumerate(path.read_text(encoding="utf-8").splitlines(), 1):
for pattern, value in scan_line(line):
if whitelisted(entries, rel, value):
suppressed += 1
continue
findings.append((rel, lineno, pattern, value))
if not findings:
print(f"CZYSTO — 0 trafień ({suppressed} wyciszonych whitelistą).")
return 0
by_pattern: dict[str, int] = {}
for rel, lineno, pattern, value in findings:
by_pattern[pattern] = by_pattern.get(pattern, 0) + 1
print(f"{rel}:{lineno} [{pattern}] {value}")
print()
print(f"WYCIEK — {len(findings)} trafień ({suppressed} wyciszonych whitelistą):")
for pattern, count in sorted(by_pattern.items(), key=lambda kv: -kv[1]):
print(f" {pattern}: {count}")
print()
print(
"Napraw źródło w kb/ (usuń adres/ścieżkę/token z dokumentu public) albo "
f"dopisz świadomy wyjątek do {whitelist_path.relative_to(REPO_ROOT)}."
)
return 1
# --- CLI --------------------------------------------------------------
def main() -> int:
parser = argparse.ArgumentParser(
description="Generator publicznej warstwy KB (kb.okit.pl) z dokumentów OKF."
)
parser.add_argument(
"--base-url",
default=DEFAULT_BASE_URL,
help=f"publiczny adres wystawki (domyślnie {DEFAULT_BASE_URL})",
)
parser.add_argument(
"--out",
type=Path,
default=DEFAULT_OUT,
help="katalog wyjściowy (domyślnie build/kb-site)",
)
parser.add_argument(
"--check",
action="store_true",
help="nie generuj — przeskanuj wygenerowany HTML pod kątem wycieków",
)
parser.add_argument(
"--whitelist",
type=Path,
default=DEFAULT_WHITELIST,
help="plik świadomych wyjątków dla --check",
)
args = parser.parse_args()
if args.check:
return check(args.out.resolve(), args.whitelist)
build(args.out.resolve(), args.base_url)
return 0
if __name__ == "__main__":
raise SystemExit(main())

View file

@ -0,0 +1,5 @@
# kb-site
Public slice of the knowledge base (`kb.okit.pl`) — static HTML generated from `kb/**/*.md` by `scripts/kb/gen_pages.py`, served by nginx on PIHA.
Dokumentacja: [kb/services/kb-site.md](../../kb/services/kb-site.md)

View file

@ -0,0 +1,28 @@
services:
kb-site:
image: nginx:alpine
container_name: kb-site
restart: unless-stopped
ports:
# PIHA 82x0 static-HTTP block: 8210 paperless, 8220 nextcloud,
# 8230 kb-query, 8240 narty27 -> 8250 is the next free slot.
# Publicly the site is reached only through the npm@PIHA vhost
# kb.okit.pl; this bind is the proxy's upstream.
- "8250:80"
volumes:
# Generated output of scripts/kb/gen_pages.py — never committed, never
# bind-mounted from the repo. Read-only: nginx only serves it; writes go
# through the helper-container procedure in kb/runbooks/kb-site-deploy.md
# (docker cp cannot write into a :ro mount).
- kb-site_content:/usr/share/nginx/html:ro
# busybox wget — nginx:alpine ships no curl. index.html is generated on
# every run, so it is the one file that must always be there.
healthcheck:
test: ["CMD", "wget", "-q", "-O", "/dev/null", "http://127.0.0.1/index.html"]
interval: 30s
timeout: 10s
retries: 5
start_period: 5s
volumes:
kb-site_content:

View file

@ -0,0 +1,10 @@
# kb-site has NO configuration and NO secrets.
#
# The port bind (8250:80) is static and the content lives in the
# kb-site_kb-site_content Docker volume, generated from kb/ by
# scripts/kb/gen_pages.py. This file exists only to keep the
# services/<service>/ layout from CLAUDE.md complete — there is nothing to
# copy to .env.
#
# The public address is a generator argument, not an env var:
# python3 scripts/kb/gen_pages.py --base-url https://kb.okit.pl

21
services/kb-site/healthcheck.sh Executable file
View file

@ -0,0 +1,21 @@
#!/bin/bash
# Healthcheck for kb-site (nginx:alpine serving the generated public KB)
# Container must be running
if ! docker ps --filter "name=kb-site" --filter "status=running" | grep -qw "kb-site"; then
echo "[FAIL] kb-site container is not running"
exit 1
fi
# The index is generated on every run, so both the bare root and /index.html
# must answer. An empty volume means the content was never loaded — see
# kb/runbooks/kb-site-deploy.md.
for path in "" index.html; do
if ! curl -sf -o /dev/null "http://127.0.0.1:8250/${path}"; then
echo "[FAIL] kb-site is not serving /${path} on 127.0.0.1:8250 (content loaded?)"
exit 1
fi
done
echo "[OK] kb-site is healthy"
exit 0

View file

@ -0,0 +1,33 @@
service:
name: kb-site
owner_node: piha
role: static-html-host # public slice of the KB, rendered by scripts/kb/gen_pages.py
exposure: public # public via npm@PIHA (kb.okit.pl). The container itself binds
# 8250 on the LAN; npm is the sole public entry point.
dependencies: [] # nginx serving a local volume — nothing else required at runtime
ports:
- container: 80
host: 8250
protocol: tcp
healthcheck:
type: http
endpoint: http://127.0.0.1:8250/index.html # content must be loaded first (see runbook)
interval: 30s
timeout: 10s
retries: 5
restart_policy: unless-stopped
persistence:
# Docker named volume kb-site_kb-site_content (compose project prefix), NOT a
# bind under /opt/homelab/data. The content is a pure artifact: regenerate it
# from the repo with scripts/kb/gen_pages.py, no backup job needed.
paths:
- kb-site_kb-site_content
runtime:
config_files: [] # no .env — the port bind is static, no secrets
env_vars: []
content:
# Only kb/ documents with `visibility: public` are published; the generator
# is fail-closed (no frontmatter / no visibility field = private).
generator: scripts/kb/gen_pages.py
leak_check: scripts/kb/gen_pages.py --check # must pass before publishing
source: kb/**/*.md