Warstwa "nie daj sie przypadkiem znalezc" dla publicznej wystawki KB: - gen_pages.py: <meta name="robots" content="noindex, nofollow"> w <head> kazdej generowanej strony (page_shell, wiec takze index). - gen_pages.py: DEFAULT_BASE_URL -> https://kb-e2a24af3.okit.pl. Slug musi zgadzac sie z rekordem DNS i vhostem w npm@PIHA (runbook kb-site-deploy). - services/kb-site: static/robots.txt (Disallow: /) montowany ro na /usr/share/nginx/html/robots.txt. Plik nie jest dokumentem KB, wiec jedzie z repo, a nie z wolumenu podmienianego przy kazdej publikacji. - kb/services/kb-site.md: sekcja "Access" — token w query paramie na warstwie nginx/NPM (sekret zyje tylko w NPM, nie w repo) + obscure subdomain + noindex. Explicit: to obscurity, nie kontrola dostepu — token w URL laduje w access logach, historii przegladarki i naglowku Referer. Bramka publikacji bez zmian: gen_pages.py --check exit 0 (22 wyciszone whitelista, jak dotad). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
3.4 KiB
| okf | type | visibility | status | updated | links | |
|---|---|---|---|---|---|---|
| 0.1 | service | public | active | 2026-08-05 |
|
kb-site
Public slice of this knowledge base, served as static HTML from a deliberately
non-obvious subdomain of okit.pl.
Plain nginx:alpine on the PIHA node reading one Docker named volume — no
build step at runtime, no database, no dependencies.
The site is a rendering, not a source. Every page is generated from the
Markdown documents of the knowledge base by scripts/kb/gen_pages.py and
copied into the volume; the HTML is never edited by hand and never committed.
What gets published
Only documents that carry an explicit visibility: public field in their
frontmatter. The generator is fail-closed: a document with no frontmatter,
with unparseable frontmatter, with no visibility field, or with any other
value is treated as private and stays out of the build.
The same rule applies to cross-references. A link pointing at a document that
was not published is not rendered as a link — only the label survives, marked
[private]. A public page therefore never exposes the location of an internal
document and never produces a dead link.
Frontmatter shown on a page is deliberately partial: type, status and the last
update date. The links field is omitted, because a path to a private document
is already a leak of its name.
Structure of the build
build/kb-site/
index.html list of all published documents, grouped by type
<directory>/<name>.html one page per document, mirroring the source tree
Each page carries a footer with the generation timestamp and the short commit hash of the repository state it was rendered from, so any page can be traced back to an exact revision.
Leak gate
gen_pages.py --check re-reads the generated HTML — not the sources — and
fails on anything that looks like infrastructure detail leaking into a public
page: private, carrier-grade and public IP addresses (v4 and v6), high service
port numbers, absolute host paths, and long hex or base64 strings that look
like credentials. Deliberate exceptions live in a whitelist file that starts
out empty, so every exception is a recorded decision.
The check is a release gate: content is copied to the host only after it passes.
Access
Three layers keep the site out of casual sight: a query-parameter token enforced
in the nginx/NPM layer (advanced config held in NPM only — the secret is not in
this repository), an unguessable subdomain, and noindex, nofollow on every page
alongside a blanket robots.txt.
This is obscurity, not access control. A URL token is written to access logs,
browser history and outbound Referer headers, so anyone who obtains a link
keeps it; nothing here resists a deliberate attacker. The leak gate above, not
this, is what keeps private material off the site.
Operations
Deployment, content refresh, reverse-proxy and DNS setup are described in the kb-site deployment runbook, which is internal — on this site the reference above is plain text, exactly as described in the previous section.
Content lifetime
The volume holds an artifact, not data. There is no backup job — recovery is a regeneration from the repository. Because the generator writes a fresh tree on every run and the publish step replaces the volume contents wholesale, a document that flips from public to private disappears from the site on the next publish.