"User-agent: * / Disallow: /" odbijalo nie tylko wyszukiwarki, ale kazdego klienta respektujacego robots.txt (m.in. fetch asystentow AI) — takze takiego, ktory znal token dostepu. Token mial WPUSZCZAC znajacych go, a robots.txt ich WYPYCHAL. Dzialalo tez przeciwko wlasnemu celowi: crawler zablokowany przed pobraniem strony nigdy nie widzi meta noindex, a wyszukiwarka i tak potrafi wylistowac sam URL, ktorego nie wolno jej bylo pobrac. Ochrona przed indeksowaniem zostaje bez zmian: <meta name="robots" content="noindex, nofollow"> w <head> kazdej strony (scripts/kb/gen_pages.py). Plik mial jedna linie polityki, wiec bez niej tracil sens — usuniety razem z bind-mountem w compose (katalog static/ zniknal jako pusty). Test: docker compose config (exit 0, zostaje tylko named volume); gen_pages.py --check — CZYSTO, 0 trafien. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
3.9 KiB
| okf | type | visibility | status | updated | links | |
|---|---|---|---|---|---|---|
| 0.1 | service | public | active | 2026-08-05 |
|
kb-site
Public slice of this knowledge base, served as static HTML from a deliberately
non-obvious subdomain of okit.pl.
Plain nginx:alpine on the PIHA node reading one Docker named volume — no
build step at runtime, no database, no dependencies.
The site is a rendering, not a source. Every page is generated from the
Markdown documents of the knowledge base by scripts/kb/gen_pages.py and
copied into the volume; the HTML is never edited by hand and never committed.
What gets published
Only documents that carry an explicit visibility: public field in their
frontmatter. The generator is fail-closed: a document with no frontmatter,
with unparseable frontmatter, with no visibility field, or with any other
value is treated as private and stays out of the build.
The same rule applies to cross-references. A link pointing at a document that
was not published is not rendered as a link — only the label survives, marked
[private]. A public page therefore never exposes the location of an internal
document and never produces a dead link.
Frontmatter shown on a page is deliberately partial: type, status and the last
update date. The links field is omitted, because a path to a private document
is already a leak of its name.
Structure of the build
build/kb-site/
index.html list of all published documents, grouped by type
<directory>/<name>.html one page per document, mirroring the source tree
Each page carries a footer with the generation timestamp and the short commit hash of the repository state it was rendered from, so any page can be traced back to an exact revision.
Leak gate
gen_pages.py --check re-reads the generated HTML — not the sources — and
fails on anything that looks like infrastructure detail leaking into a public
page: private, carrier-grade and public IP addresses (v4 and v6), high service
port numbers, absolute host paths, and long hex or base64 strings that look
like credentials. Deliberate exceptions live in a whitelist file that starts
out empty, so every exception is a recorded decision.
The check is a release gate: content is copied to the host only after it passes.
Access
Three layers keep the site out of casual sight: a query-parameter token enforced
in the nginx/NPM layer (advanced config held in NPM only — the secret is not in
this repository), an unguessable subdomain, and noindex, nofollow on every page.
A blanket robots.txt was part of this and has been withdrawn. It turned away
every client that honours robots.txt — including the fetch tools of AI
assistants — even when the caller already held the access token, so it pushed
out precisely the readers the token exists to let in. It also worked against its
own goal: a crawler blocked before it can read the page never sees the noindex
tag, and a search engine may still list a bare URL it was forbidden to fetch.
This is obscurity, not access control. A URL token is written to access logs,
browser history and outbound Referer headers, so anyone who obtains a link
keeps it; nothing here resists a deliberate attacker. The leak gate above, not
this, is what keeps private material off the site.
Operations
Deployment, content refresh, reverse-proxy and DNS setup are described in the kb-site deployment runbook, which is internal — on this site the reference above is plain text, exactly as described in the previous section.
Content lifetime
The volume holds an artifact, not data. There is no backup job — recovery is a regeneration from the repository. Because the generator writes a fresh tree on every run and the publish step replaces the volume contents wholesale, a document that flips from public to private disappears from the site on the next publish.