"User-agent: * / Disallow: /" odbijalo nie tylko wyszukiwarki, ale kazdego klienta respektujacego robots.txt (m.in. fetch asystentow AI) — takze takiego, ktory znal token dostepu. Token mial WPUSZCZAC znajacych go, a robots.txt ich WYPYCHAL. Dzialalo tez przeciwko wlasnemu celowi: crawler zablokowany przed pobraniem strony nigdy nie widzi meta noindex, a wyszukiwarka i tak potrafi wylistowac sam URL, ktorego nie wolno jej bylo pobrac. Ochrona przed indeksowaniem zostaje bez zmian: <meta name="robots" content="noindex, nofollow"> w <head> kazdej strony (scripts/kb/gen_pages.py). Plik mial jedna linie polityki, wiec bez niej tracil sens — usuniety razem z bind-mountem w compose (katalog static/ zniknal jako pusty). Test: docker compose config (exit 0, zostaje tylko named volume); gen_pages.py --check — CZYSTO, 0 trafien. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
94 lines
3.9 KiB
Markdown
94 lines
3.9 KiB
Markdown
---
|
|
okf: "0.1"
|
|
type: service
|
|
visibility: public
|
|
status: active
|
|
updated: 2026-08-05
|
|
links:
|
|
- ../runbooks/kb-site-deploy.md
|
|
---
|
|
|
|
# kb-site
|
|
|
|
Public slice of this knowledge base, served as static HTML from a deliberately
|
|
non-obvious subdomain of `okit.pl`.
|
|
Plain `nginx:alpine` on the PIHA node reading one Docker named volume — no
|
|
build step at runtime, no database, no dependencies.
|
|
|
|
The site is a rendering, not a source. Every page is generated from the
|
|
Markdown documents of the knowledge base by `scripts/kb/gen_pages.py` and
|
|
copied into the volume; the HTML is never edited by hand and never committed.
|
|
|
|
## What gets published
|
|
|
|
Only documents that carry an explicit `visibility: public` field in their
|
|
frontmatter. The generator is **fail-closed**: a document with no frontmatter,
|
|
with unparseable frontmatter, with no `visibility` field, or with any other
|
|
value is treated as private and stays out of the build.
|
|
|
|
The same rule applies to cross-references. A link pointing at a document that
|
|
was not published is not rendered as a link — only the label survives, marked
|
|
`[private]`. A public page therefore never exposes the location of an internal
|
|
document and never produces a dead link.
|
|
|
|
Frontmatter shown on a page is deliberately partial: type, status and the last
|
|
update date. The `links` field is omitted, because a path to a private document
|
|
is already a leak of its name.
|
|
|
|
## Structure of the build
|
|
|
|
```
|
|
build/kb-site/
|
|
index.html list of all published documents, grouped by type
|
|
<directory>/<name>.html one page per document, mirroring the source tree
|
|
```
|
|
|
|
Each page carries a footer with the generation timestamp and the short commit
|
|
hash of the repository state it was rendered from, so any page can be traced
|
|
back to an exact revision.
|
|
|
|
## Leak gate
|
|
|
|
`gen_pages.py --check` re-reads the generated HTML — not the sources — and
|
|
fails on anything that looks like infrastructure detail leaking into a public
|
|
page: private, carrier-grade and public IP addresses (v4 and v6), high service
|
|
port numbers, absolute host paths, and long hex or base64 strings that look
|
|
like credentials. Deliberate exceptions live in a whitelist file that starts
|
|
out empty, so every exception is a recorded decision.
|
|
|
|
The check is a release gate: content is copied to the host only after it
|
|
passes.
|
|
|
|
## Access
|
|
|
|
Three layers keep the site out of casual sight: a query-parameter token enforced
|
|
in the nginx/NPM layer (advanced config held in NPM only — the secret is not in
|
|
this repository), an unguessable subdomain, and `noindex, nofollow` on every page.
|
|
|
|
A blanket `robots.txt` was part of this and has been withdrawn. It turned away
|
|
every client that honours robots.txt — including the fetch tools of AI
|
|
assistants — even when the caller already held the access token, so it pushed
|
|
out precisely the readers the token exists to let in. It also worked against its
|
|
own goal: a crawler blocked before it can read the page never sees the `noindex`
|
|
tag, and a search engine may still list a bare URL it was forbidden to fetch.
|
|
|
|
This is obscurity, not access control. A URL token is written to access logs,
|
|
browser history and outbound `Referer` headers, so anyone who obtains a link
|
|
keeps it; nothing here resists a deliberate attacker. The leak gate above, not
|
|
this, is what keeps private material off the site.
|
|
|
|
## Operations
|
|
|
|
Deployment, content refresh, reverse-proxy and DNS setup are described in the
|
|
[kb-site deployment runbook](../runbooks/kb-site-deploy.md), which is internal —
|
|
on this site the reference above is plain text, exactly as described in the
|
|
previous section.
|
|
|
|
## Content lifetime
|
|
|
|
The volume holds an artifact, not data. There is no backup job — recovery is a
|
|
regeneration from the repository. Because the generator writes a fresh tree on
|
|
every run and the publish step replaces the volume contents wholesale, a
|
|
document that flips from public to private disappears from the site on the next
|
|
publish.
|