homelab-codex-ws/kb/services/kb-site.md
oskar bb3792d219 docs(kb-site): przekaz ACCESS_TOKEN generatorowi w procedurze publikacji
Poprzedni commit nauczyl gen_pages.py dopisywac token do linkow, ale nic
go nie podawalo — publikacja poszlaby stara sciezka i dalaby build
z golymi linkami, czyli stan sprzed fiksa.

kb-site nie ma skryptu deployu: generator wolany jest wylacznie recznie
z runbooka (kroki 2 i 7), wiec to tam token musi wejsc.

Zrodlo tokenu: /opt/homelab/config/kb-site/.env na wezle GENERUJACYM
(SATURN/SOLARIA), nie na PIHA. To swiadome odstepstwo od konwencji
config/<serwis>/ z CLAUDE.md — plik trzyma zwykle sekrety wezla, ktory
serwis uruchamia, a ten token jest potrzebny tam, gdzie serwis sie
generuje. Kontener nginx dalej nie ma zadnej konfiguracji ani sekretow;
odnotowane w service.yaml i env.example, zeby nikt nie szukal .env na PIHA.

Token idzie zmienna srodowiskowa (set -a; . plik; set +a), nie flaga
--access-token: argument z linii polecen laduje w historii shella i jest
widoczny w ps dla kazdego uzytkownika wezla.

Lancuch publikacji z kroku 7 dostal dwa nowe ogniwa przed scp: test -n
"$ACCESS_TOKEN" (pusty token = build nieklikalny) oraz grep -q 'key='
w index.html (token byl, ale nie dojechal do generatora). Oba zatrzymuja
publikacje tak samo jak --check.

Krok 6 weryfikuje teraz wlasciwa rzecz: wyciaga href ze spisu i pobiera
GO, zamiast recznie sklejac URL — czyli testuje to, co faktycznie bylo
zepsute. Doszedl tez negatywny test bramki (bez tokenu ma NIE byc 200).

Tabela problemow: "index sie otwiera, ale klikniecie daje 403" (build bez
tokenu) i "403 takze z tokenem" (rotacja tokenu w NPM rozjechana z plikiem
— stary build zostaje z martwym tokenem w kazdym linku).

kb/services/kb-site.md (public) — sekcja Access: token siedzi teraz
w tresci kazdej serwowanej strony, wiec jedna zapisana strona wydaje go
w calosci. Model zagrozen bez zmian (URL wejsciowy zawsze go niosl), ale
warto, zeby dokument mowil to wprost obok zdania "to obscurity, not
access control".

Test: sekwencje z krokow 2 i 7 przepuszczone na symulowanym pliku tokenu
(prod /opt/homelab nietkniety) — token obecny: Token: TAK, 4x key=
w index.html, lancuch dochodzi do tar; token pusty: staje na pierwszym
ogniwie, brak tgz; build bez tokenu przy ustawionej zmiennej: staje na
grep, brak tgz. check_okf.py exit 0, gen_pages --check exit 0,
service.yaml parsuje sie.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 22:26:48 +02:00

4.3 KiB

okf type visibility status updated links
0.1 service public active 2026-08-05
../runbooks/kb-site-deploy.md

kb-site

Public slice of this knowledge base, served as static HTML from a deliberately non-obvious subdomain of okit.pl. Plain nginx:alpine on the PIHA node reading one Docker named volume — no build step at runtime, no database, no dependencies.

The site is a rendering, not a source. Every page is generated from the Markdown documents of the knowledge base by scripts/kb/gen_pages.py and copied into the volume; the HTML is never edited by hand and never committed.

What gets published

Only documents that carry an explicit visibility: public field in their frontmatter. The generator is fail-closed: a document with no frontmatter, with unparseable frontmatter, with no visibility field, or with any other value is treated as private and stays out of the build.

The same rule applies to cross-references. A link pointing at a document that was not published is not rendered as a link — only the label survives, marked [private]. A public page therefore never exposes the location of an internal document and never produces a dead link.

Frontmatter shown on a page is deliberately partial: type, status and the last update date. The links field is omitted, because a path to a private document is already a leak of its name.

Structure of the build

build/kb-site/
  index.html            list of all published documents, grouped by type
  <directory>/<name>.html   one page per document, mirroring the source tree

Each page carries a footer with the generation timestamp and the short commit hash of the repository state it was rendered from, so any page can be traced back to an exact revision.

Leak gate

gen_pages.py --check re-reads the generated HTML — not the sources — and fails on anything that looks like infrastructure detail leaking into a public page: private, carrier-grade and public IP addresses (v4 and v6), high service port numbers, absolute host paths, and long hex or base64 strings that look like credentials. Deliberate exceptions live in a whitelist file that starts out empty, so every exception is a recorded decision.

The check is a release gate: content is copied to the host only after it passes.

Access

Three layers keep the site out of casual sight: a query-parameter token enforced in the nginx/NPM layer (advanced config held in NPM only — the secret is not in this repository), an unguessable subdomain, and noindex, nofollow on every page.

A blanket robots.txt was part of this and has been withdrawn. It turned away every client that honours robots.txt — including the fetch tools of AI assistants — even when the caller already held the access token, so it pushed out precisely the readers the token exists to let in. It also worked against its own goal: a crawler blocked before it can read the page never sees the noindex tag, and a search engine may still list a bare URL it was forbidden to fetch.

Because the gate is enforced on the URL, every internal link the generator emits carries the token too — otherwise the entry page would open and every click from it would return 403. The token is therefore a generation-time input, supplied from outside the repository; a build made without it is a local preview, not a publishable site.

This is obscurity, not access control. A URL token is written to access logs, browser history and outbound Referer headers, and — since the links carry it — into the body of every served page, so a single saved page or shared screenshot of the address bar hands it over in full. Anyone who obtains a link keeps it; nothing here resists a deliberate attacker. The leak gate above, not this, is what keeps private material off the site.

Operations

Deployment, content refresh, reverse-proxy and DNS setup are described in the kb-site deployment runbook, which is internal — on this site the reference above is plain text, exactly as described in the previous section.

Content lifetime

The volume holds an artifact, not data. There is no backup job — recovery is a regeneration from the repository. Because the generator writes a fresh tree on every run and the publish step replaces the volume contents wholesale, a document that flips from public to private disappears from the site on the next publish.