homelab-codex-ws/kb/services/kb-site.md
oskar 19548d888e fix(kb-site): wycofaj robots.txt — blokowal legalny fetch z tokenem
"User-agent: * / Disallow: /" odbijalo nie tylko wyszukiwarki, ale kazdego
klienta respektujacego robots.txt (m.in. fetch asystentow AI) — takze takiego,
ktory znal token dostepu. Token mial WPUSZCZAC znajacych go, a robots.txt ich
WYPYCHAL.

Dzialalo tez przeciwko wlasnemu celowi: crawler zablokowany przed pobraniem
strony nigdy nie widzi meta noindex, a wyszukiwarka i tak potrafi wylistowac
sam URL, ktorego nie wolno jej bylo pobrac.

Ochrona przed indeksowaniem zostaje bez zmian: <meta name="robots"
content="noindex, nofollow"> w <head> kazdej strony (scripts/kb/gen_pages.py).

Plik mial jedna linie polityki, wiec bez niej tracil sens — usuniety razem
z bind-mountem w compose (katalog static/ zniknal jako pusty).

Test: docker compose config (exit 0, zostaje tylko named volume);
gen_pages.py --check — CZYSTO, 0 trafien.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 13:11:43 +02:00

94 lines
3.9 KiB
Markdown

---
okf: "0.1"
type: service
visibility: public
status: active
updated: 2026-08-05
links:
- ../runbooks/kb-site-deploy.md
---
# kb-site
Public slice of this knowledge base, served as static HTML from a deliberately
non-obvious subdomain of `okit.pl`.
Plain `nginx:alpine` on the PIHA node reading one Docker named volume — no
build step at runtime, no database, no dependencies.
The site is a rendering, not a source. Every page is generated from the
Markdown documents of the knowledge base by `scripts/kb/gen_pages.py` and
copied into the volume; the HTML is never edited by hand and never committed.
## What gets published
Only documents that carry an explicit `visibility: public` field in their
frontmatter. The generator is **fail-closed**: a document with no frontmatter,
with unparseable frontmatter, with no `visibility` field, or with any other
value is treated as private and stays out of the build.
The same rule applies to cross-references. A link pointing at a document that
was not published is not rendered as a link — only the label survives, marked
`[private]`. A public page therefore never exposes the location of an internal
document and never produces a dead link.
Frontmatter shown on a page is deliberately partial: type, status and the last
update date. The `links` field is omitted, because a path to a private document
is already a leak of its name.
## Structure of the build
```
build/kb-site/
index.html list of all published documents, grouped by type
<directory>/<name>.html one page per document, mirroring the source tree
```
Each page carries a footer with the generation timestamp and the short commit
hash of the repository state it was rendered from, so any page can be traced
back to an exact revision.
## Leak gate
`gen_pages.py --check` re-reads the generated HTML — not the sources — and
fails on anything that looks like infrastructure detail leaking into a public
page: private, carrier-grade and public IP addresses (v4 and v6), high service
port numbers, absolute host paths, and long hex or base64 strings that look
like credentials. Deliberate exceptions live in a whitelist file that starts
out empty, so every exception is a recorded decision.
The check is a release gate: content is copied to the host only after it
passes.
## Access
Three layers keep the site out of casual sight: a query-parameter token enforced
in the nginx/NPM layer (advanced config held in NPM only — the secret is not in
this repository), an unguessable subdomain, and `noindex, nofollow` on every page.
A blanket `robots.txt` was part of this and has been withdrawn. It turned away
every client that honours robots.txt — including the fetch tools of AI
assistants — even when the caller already held the access token, so it pushed
out precisely the readers the token exists to let in. It also worked against its
own goal: a crawler blocked before it can read the page never sees the `noindex`
tag, and a search engine may still list a bare URL it was forbidden to fetch.
This is obscurity, not access control. A URL token is written to access logs,
browser history and outbound `Referer` headers, so anyone who obtains a link
keeps it; nothing here resists a deliberate attacker. The leak gate above, not
this, is what keeps private material off the site.
## Operations
Deployment, content refresh, reverse-proxy and DNS setup are described in the
[kb-site deployment runbook](../runbooks/kb-site-deploy.md), which is internal —
on this site the reference above is plain text, exactly as described in the
previous section.
## Content lifetime
The volume holds an artifact, not data. There is no backup job — recovery is a
regeneration from the repository. Because the generator writes a fresh tree on
every run and the publish step replaces the volume contents wholesale, a
document that flips from public to private disappears from the site on the next
publish.