From 0a8a668b05c38f25776d8695f9e11b0e61166158 Mon Sep 17 00:00:00 2001 From: oskar Date: Tue, 4 Aug 2026 15:06:40 +0200 Subject: [PATCH] feat(kb): SPLIT fleet-prometheus -> service + decision + runbook kb/services/fleet-prometheus.md (Stack, Placement & exposure, Data, Next steps) kb/decisions/fleet-prometheus-osobny-od-prom.md ("Why separate from the home prom") kb/runbooks/fleet-prometheus-deploy.md (Configuration, Verify) Wzajemne links. Tresc sekcji nietknieta; kontrola: multizbior niepustych linii czesci == oryginal z HEAD. Co-Authored-By: Claude Opus 5 (1M context) --- .../fleet-prometheus-osobny-od-prom.md | 28 +++++ kb/runbooks/fleet-prometheus-deploy.md | 68 +++++++++++ kb/services/fleet-prometheus.md | 50 ++++++++ services/fleet-prometheus/README.md | 109 ------------------ 4 files changed, 146 insertions(+), 109 deletions(-) create mode 100644 kb/decisions/fleet-prometheus-osobny-od-prom.md create mode 100644 kb/runbooks/fleet-prometheus-deploy.md create mode 100644 kb/services/fleet-prometheus.md delete mode 100644 services/fleet-prometheus/README.md diff --git a/kb/decisions/fleet-prometheus-osobny-od-prom.md b/kb/decisions/fleet-prometheus-osobny-od-prom.md new file mode 100644 index 0000000..9eb8ba9 --- /dev/null +++ b/kb/decisions/fleet-prometheus-osobny-od-prom.md @@ -0,0 +1,28 @@ +--- +okf: "0.1" +type: decision +visibility: private +status: active +updated: 2026-06-24 +links: + - ../services/fleet-prometheus.md + - ../runbooks/fleet-prometheus-deploy.md +--- + +# fleet-prometheus — dlaczego osobny od domowego `prom` + +## Why separate from the home `prom` + +The home `prom` on **PIHA** is LAN-only and monitors the home network. This +instance is intentionally a **distinct** Prometheus dedicated to the distributed +fleet — hence the `fleet-` naming. The split is made explicit in every series: + +- `external_labels: { fleet: "homelab-codex" }` +- job names prefixed/scoped for the fleet (`prometheus`, `fleet-node`) + +It also avoids two anti-patterns the home `prom` has: + +- **No `:latest`** — the image is pinned to `prom/prometheus:v3.5.0` (LTS). +- **No inline secrets** — `prometheus.yml` is secret-free by contract (the home + prom embeds a plaintext HAOS token; we do not). + diff --git a/kb/runbooks/fleet-prometheus-deploy.md b/kb/runbooks/fleet-prometheus-deploy.md new file mode 100644 index 0000000..795c56a --- /dev/null +++ b/kb/runbooks/fleet-prometheus-deploy.md @@ -0,0 +1,68 @@ +--- +okf: "0.1" +type: runbook +visibility: private +status: active +updated: 2026-06-24 +links: + - ../services/fleet-prometheus.md + - ../decisions/fleet-prometheus-osobny-od-prom.md +--- + +# fleet-prometheus — konfiguracja i weryfikacja + +## Configuration + +- `prometheus.yml` — committed, **secret-free**. Global config + two scrape jobs. +- `env.example` → copy to `.env` (gitignored). Holds `TAILSCALE_BIND_IP`, the + VPS Tailscale interface IP the listen socket is bound to. + +### Deploy — `.env` is mandatory + +The port is published as `${TAILSCALE_BIND_IP}:9090:9090` so the socket is bound +**only** to the Tailscale interface, never `0.0.0.0`. Compose auto-loads a +co-located `services/fleet-prometheus/.env`, but if that file is **missing** the +variable resolves to a blank string and Compose **silently binds to 0.0.0.0** +(publicly reachable on the Hetzner IP) — it only warns, it does not fail. + +Therefore, before `up -d`: + +1. Create `services/fleet-prometheus/.env` from `env.example` with the real + `TAILSCALE_BIND_IP` (verify with `tailscale ip -4` on the VPS). +2. Deploy with the env file explicit, to be safe: + + ```bash + docker compose -f services/fleet-prometheus/docker-compose.yml \ + --env-file services/fleet-prometheus/.env up -d + ``` + +3. Confirm the bind is Tailscale-only (must NOT show `0.0.0.0`): + + ```bash + docker compose -f services/fleet-prometheus/docker-compose.yml \ + --env-file services/fleet-prometheus/.env config | grep host_ip + ``` + +### Scrape targets (scaffold) + +| Job | Target | Notes | +|--------------|-----------------------------|-----------------------------------------| +| `prometheus` | `localhost:9090` | self | +| `fleet-node` | `host.docker.internal:9100` | VPS `node_exporter` (runs `network_mode: host`, listens on host `:9100`); reached via `host-gateway` | + +## Verify + +```bash +# Config is syntactically valid (uses promtool inside the image): +docker run --rm -v "$PWD/services/fleet-prometheus/prometheus.yml":/etc/prometheus/prometheus.yml:ro \ + prom/prometheus:v3.5.0 promtool check config /etc/prometheus/prometheus.yml + +# Compose renders: +docker compose -f services/fleet-prometheus/docker-compose.yml config + +# After `up -d`: +curl -sf http://localhost:9090/-/healthy +curl -s http://localhost:9090/api/v1/targets | jq '.data.activeTargets[].health' +./services/fleet-prometheus/healthcheck.sh +``` + diff --git a/kb/services/fleet-prometheus.md b/kb/services/fleet-prometheus.md new file mode 100644 index 0000000..cce2c5e --- /dev/null +++ b/kb/services/fleet-prometheus.md @@ -0,0 +1,50 @@ +--- +okf: "0.1" +type: service +visibility: private +status: active +updated: 2026-06-24 +links: + - ../decisions/fleet-prometheus-osobny-od-prom.md + - ../runbooks/fleet-prometheus-deploy.md +--- + +# fleet-prometheus + +Prometheus instance that is the **source of truth for fleet liveness**. Runs on +the **VPS** (ingress / control-plane host) and is reachable only over Tailscale. + +This is a **clean scaffold**: right now it scrapes only itself and the VPS-local +`node_exporter`. The fleet targets, liveness rules, and alerting integration are +deliberately separate, later steps (see *Next steps* below). + +## Stack + +| Container | Image | Purpose | +|--------------------|-------------------------|--------------------------------| +| `fleet-prometheus` | `prom/prometheus:v3.5.0`| Fleet liveness Prometheus (9090) | + +## Placement & exposure + +- **Node:** VPS (`ubuntu-4gb-hel1-1`, Tailscale `100.95.58.48`). +- **Exposure:** `tailscale-internal` — exposed exactly like `control-plane`: a + plain published `9090:9090` port with **no** public npm reverse-proxy entry. + Reachability is constrained to the Tailscale mesh at the VPS firewall layer. + Do **not** add this to npm or public DNS. +- **Memory:** `mem_limit: 512m`. `oom_score_adj: 200` (positive) — fleet + monitoring is less critical than the control-plane (`-900`), so it is a + sacrificial OOM victim *before* the control-plane, never after. +- **Retention:** `15d` / `2GB` (whichever hits first) — tight on purpose. + +## Data + +TSDB lives in the named volume `fleet-prometheus_tsdb` → `/prometheus`. + +## Next steps (NOT done in this scaffold) + +- **Fleet targets:** add the `100.x` Tailscale `node_exporter`s for SATURN, + SOLARIA, PIHA, CHELSTY to the `fleet-node` job. +- **Liveness rules:** `rule_files` with `up == 0 for: 5m` (placeholder is + commented in `prometheus.yml`). +- **Alerting:** Alertmanager wiring, and/or `brain-watchdog` querying the + firing alerts API → Telegram (placeholder commented in `prometheus.yml`). diff --git a/services/fleet-prometheus/README.md b/services/fleet-prometheus/README.md deleted file mode 100644 index 9dbf920..0000000 --- a/services/fleet-prometheus/README.md +++ /dev/null @@ -1,109 +0,0 @@ -# fleet-prometheus - -Prometheus instance that is the **source of truth for fleet liveness**. Runs on -the **VPS** (ingress / control-plane host) and is reachable only over Tailscale. - -This is a **clean scaffold**: right now it scrapes only itself and the VPS-local -`node_exporter`. The fleet targets, liveness rules, and alerting integration are -deliberately separate, later steps (see *Next steps* below). - -## Why separate from the home `prom` - -The home `prom` on **PIHA** is LAN-only and monitors the home network. This -instance is intentionally a **distinct** Prometheus dedicated to the distributed -fleet — hence the `fleet-` naming. The split is made explicit in every series: - -- `external_labels: { fleet: "homelab-codex" }` -- job names prefixed/scoped for the fleet (`prometheus`, `fleet-node`) - -It also avoids two anti-patterns the home `prom` has: - -- **No `:latest`** — the image is pinned to `prom/prometheus:v3.5.0` (LTS). -- **No inline secrets** — `prometheus.yml` is secret-free by contract (the home - prom embeds a plaintext HAOS token; we do not). - -## Stack - -| Container | Image | Purpose | -|--------------------|-------------------------|--------------------------------| -| `fleet-prometheus` | `prom/prometheus:v3.5.0`| Fleet liveness Prometheus (9090) | - -## Placement & exposure - -- **Node:** VPS (`ubuntu-4gb-hel1-1`, Tailscale `100.95.58.48`). -- **Exposure:** `tailscale-internal` — exposed exactly like `control-plane`: a - plain published `9090:9090` port with **no** public npm reverse-proxy entry. - Reachability is constrained to the Tailscale mesh at the VPS firewall layer. - Do **not** add this to npm or public DNS. -- **Memory:** `mem_limit: 512m`. `oom_score_adj: 200` (positive) — fleet - monitoring is less critical than the control-plane (`-900`), so it is a - sacrificial OOM victim *before* the control-plane, never after. -- **Retention:** `15d` / `2GB` (whichever hits first) — tight on purpose. - -## Data - -TSDB lives in the named volume `fleet-prometheus_tsdb` → `/prometheus`. - -## Configuration - -- `prometheus.yml` — committed, **secret-free**. Global config + two scrape jobs. -- `env.example` → copy to `.env` (gitignored). Holds `TAILSCALE_BIND_IP`, the - VPS Tailscale interface IP the listen socket is bound to. - -### Deploy — `.env` is mandatory - -The port is published as `${TAILSCALE_BIND_IP}:9090:9090` so the socket is bound -**only** to the Tailscale interface, never `0.0.0.0`. Compose auto-loads a -co-located `services/fleet-prometheus/.env`, but if that file is **missing** the -variable resolves to a blank string and Compose **silently binds to 0.0.0.0** -(publicly reachable on the Hetzner IP) — it only warns, it does not fail. - -Therefore, before `up -d`: - -1. Create `services/fleet-prometheus/.env` from `env.example` with the real - `TAILSCALE_BIND_IP` (verify with `tailscale ip -4` on the VPS). -2. Deploy with the env file explicit, to be safe: - - ```bash - docker compose -f services/fleet-prometheus/docker-compose.yml \ - --env-file services/fleet-prometheus/.env up -d - ``` - -3. Confirm the bind is Tailscale-only (must NOT show `0.0.0.0`): - - ```bash - docker compose -f services/fleet-prometheus/docker-compose.yml \ - --env-file services/fleet-prometheus/.env config | grep host_ip - ``` - -### Scrape targets (scaffold) - -| Job | Target | Notes | -|--------------|-----------------------------|-----------------------------------------| -| `prometheus` | `localhost:9090` | self | -| `fleet-node` | `host.docker.internal:9100` | VPS `node_exporter` (runs `network_mode: host`, listens on host `:9100`); reached via `host-gateway` | - -## Verify - -```bash -# Config is syntactically valid (uses promtool inside the image): -docker run --rm -v "$PWD/services/fleet-prometheus/prometheus.yml":/etc/prometheus/prometheus.yml:ro \ - prom/prometheus:v3.5.0 promtool check config /etc/prometheus/prometheus.yml - -# Compose renders: -docker compose -f services/fleet-prometheus/docker-compose.yml config - -# After `up -d`: -curl -sf http://localhost:9090/-/healthy -curl -s http://localhost:9090/api/v1/targets | jq '.data.activeTargets[].health' -./services/fleet-prometheus/healthcheck.sh -``` - -## Next steps (NOT done in this scaffold) - -- **Fleet targets:** add the `100.x` Tailscale `node_exporter`s for SATURN, - SOLARIA, PIHA, CHELSTY to the `fleet-node` job. -- **Liveness rules:** `rule_files` with `up == 0 for: 5m` (placeholder is - commented in `prometheus.yml`). -- **Alerting:** Alertmanager wiring, and/or `brain-watchdog` querying the - firing alerts API → Telegram (placeholder commented in `prometheus.yml`).