homelab-codex-ws/kb/phases/kb-m5-documents-ingest-fazy.md
oskar ccb738cddf feat(kb): SPLIT documents-ingest (30 KB) -> service + phase + runbook
kb/services/job-documents-ingest.md — opis jobu i mechanizmow
  (candidate selection, matching, consume/, idempotency, dry-run)
kb/phases/kb-m5-documents-ingest-fazy.md — faza 2, faza 2 krok 6,
  faza 3 krok 4, faza 3 krok 5 (4 sekcje fazowe wtopione w README)
kb/runbooks/documents-ingest-run.md — Usage, Verifying in Paperless, Tests

Najwiekszy README w repo. Tresc sekcji nietknieta; kontrola multizbioru
linii == oryginal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 16:58:46 +02:00

479 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
okf: "0.1"
type: phase
visibility: private
status: active
updated: 2026-07-30
links:
- ../services/job-documents-ingest.md
- ../runbooks/documents-ingest-run.md
---
# documents-ingest — fazy 2 i 3
## Phase 2 — `documents-ingest-paperless` (Paperless -> envelope adapter)
Module 5, phase 2 (`docs/kb/modules/05-faza2-plan.md`, §4.2-4.3, §6 step 5).
Reads documents from the **Paperless REST API** (read-only — GET only, never
writes to Paperless) and inserts them as `source='paperless'` rows into the
`envelope` table on kb-postgres, reusing `kb_mail.envelope.Envelope` /
`kb_mail.db.insert_envelope` from `packages/kb-mail` (untouched by this
change — see plan §1.6). Existing `source='gmail'` rows and `document_chunk`
are never touched; this job only ever `INSERT`s new `paperless` rows.
### Cross-source link (`source_mail`)
Per plan §1.9/§4.2, the deterministic join uses no heuristics: a document's
`original_file_name` (from the Paperless API) is matched against
`consume_name` in this job's **phase-1 registry**
(`/opt/homelab/data/documents-ingest/registry.json`, produced by
`extractor.py` — see above). A match appends a `source_mail` entity pointing
back at the originating mail envelope; no match means the document was added
outside the faktury-1 pipeline, and the entity is simply omitted — not an
error.
### Install
```bash
pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/
```
### Usage
```bash
# Dry run (default) — fetch from Paperless, map, count; no DB writes:
documents-ingest-paperless --dsn postgresql://kb:<pw>@localhost:5433/kb \
--paperless-token <token>
# Real run — insert new envelope rows:
documents-ingest-paperless --dsn ... --paperless-token ... --apply
# Smoke-test slice:
documents-ingest-paperless --dsn ... --paperless-token ... --limit 5
```
`--dsn` can come from `KB_DSN`, `--paperless-token` from `PAPERLESS_API_TOKEN`,
`--paperless-url` from `PAPERLESS_URL` (defaults to Paperless' fixed LAN
address, `http://192.168.31.5:8210`). No `--offset`: unlike the 225 030-row
header backfill, a full re-scan of Paperless' ~186 documents is cheap and
already idempotent, so there is no need for resumable partitioning — `--limit`
exists only to cap a run for smoke-testing.
### Mapping (plan §4.3)
```
id = f"paperless:{document_id}" -- prefixed: Paperless doc-ids are small
-- sequential ints that would otherwise
-- collide with any future source's ids
ts = documents_document.created -- Paperless-detected date (content/filename),
-- not filesystem mtime
geo = NULL
raw_ref = str(document_id) -- REFERENCE — Paperless is the source of truth,
-- no bytes are copied
entities = content, correspondent, tag(s), filename, content_type,
and source_mail when the registry join hits (plan §4.2)
```
`correspondent`/`tag` are resolved from Paperless' `/api/correspondents/` and
`/api/tags/` (fetched once, cached in memory for the run) and kept purely as
informational metadata — nothing in this pipeline depends on them being
non-null (plan decision 4). A document with empty OCR content (Paperless OCR
sometimes produces none) still gets a normal envelope with `"text": ""` — not
skipped, not an error, just counted (`empty_content`).
### Idempotency
A pre-fetched set of existing `source='paperless'` envelope ids (one query at
the start of each run) skips documents already inserted; `insert_envelope`'s
own `ON CONFLICT (id) DO NOTHING` is the second line of defense. Re-running
`--apply` immediately after a successful run reports `inserted: 0` and
`already_in_db` equal to the previous run's `inserted` count.
### Stats must balance
```
fetched = already_in_db + inserted + errors
```
`source_mail_linked` and `empty_content` are informational subsets of
`fetched`, not separate outcome buckets. A per-document mapping failure
(e.g. an unparseable `created` date) is isolated, logged, and counted as
`errors` — it never aborts the run. `main()` exits 1 on non-zero `errors`
or if the balance invariant above doesn't hold (mirrors
`gmail-bulk-import`'s exit-code convention) — a clean run always exits 0.
### Tests
```bash
pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest
```
Pure unit tests, no DB or real HTTP — `run()` is tested by monkeypatching
`asyncpg.connect` (fake connection) and `aiohttp.ClientSession` (fake session
serving canned JSON pages). Covers: mapping shape (content, correspondent,
tag(s), filename, content_type, source_mail), the registry join (hit and
miss), pagination (both the documents list and the correspondents/tags lookup
tables), `--limit`, idempotency (pre-existing ids skipped, a second `--apply`
run inserts nothing new), isolated per-document mapping errors, and the
stats-balance invariant.
### Definition of Done
Per `CLAUDE.md`: smoke run is `documents-ingest-paperless --dsn ...
--paperless-token ... --limit 5` (dry-run first) against kb-postgres@PIHA and
the live Paperless API, over SSH — **not executed as part of this change**
without operator confirmation (this job reads production Paperless data and
writes production envelope rows on `--apply`). `pytest` passes locally before
this commit.
---
## Phase 2 step 6 — `documents-ingest-embed` (chunk + embed)
Module 5, phase 2, plan step 6 (`docs/kb/modules/05-faza2-plan.md`, §6 step 6,
§2 decision 3). Reads `entities[type=content].text` off every `source='paperless'`
envelope, chunks it, calls Ollama (`POST /api/embeddings`, model `bge-m3`) for
each chunk, and inserts the result into `document_chunk`
(`services/kb-postgres/init/002_chunks.sql`). This job only ever `INSERT`s into
`document_chunk``envelope` is read-only here, and `services/ollama/` is
untouched.
### Where it runs
**On SOLARIA** (that's where Ollama lives), against kb-postgres@PIHA over
Tailscale — the reverse of the other jobs in this package, which run on PIHA.
`--ollama-url` defaults to `http://localhost:11434` (Ollama on the same node);
`--dsn` needs PIHA's Tailscale address, e.g.
`postgresql://kb:<pw>@piha:5433/kb`.
### Install
```bash
pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/
```
### Usage
```bash
# Dry run (default) — chunk and count only, no Ollama calls, no DB writes:
documents-ingest-embed --dsn postgresql://kb:<pw>@piha:5433/kb
# Smoke-test slice:
documents-ingest-embed --dsn ... --apply --limit 10
# Full run:
documents-ingest-embed --dsn ... --apply
```
### Chunking (plan §2 decision 3)
Paragraph-preferring: splits on blank-line boundaries, greedily packs
paragraphs up to `--chunk-size` characters (default 2400, ≈600 tokens at a
~4 chars/token heuristic — no local bge-m3 tokenizer available offline),
`--chunk-overlap` characters of trailing context carried into the next chunk
(default 600, ≈150 tokens). A paragraph that alone exceeds `--chunk-size`
falls back to a hard character-based sliding window — Paperless OCR text has
no page-break markers (plan §1.2), so there's nothing else to split large,
unbroken text on. A document with empty OCR content (the 26 `empty_content`
documents from phase 2 step 5) yields zero chunks and is counted separately,
not as an error.
### Idempotency
A pre-fetched set of `(envelope_id, chunk_index)` pairs already embedded with
`--model` skips re-embedding on rerun — no wasted Ollama calls.
`document_chunk`'s own `UNIQUE (envelope_id, chunk_index)` +
`ON CONFLICT DO NOTHING` is the second line of defense; `insert_chunk`'s
command tag is checked so a silently-skipped row is counted as
`chunks_conflict_skipped`, never miscounted as `chunks_inserted`. Note that
uniqueness is on `(envelope_id, chunk_index)` only, not `model`
re-embedding with a *different* model hits this path and that embedding is
discarded (wasted work, correctly reported via `chunks_conflict_skipped`,
but not persisted). Out of scope for this single-model pilot; the real fix
for whoever indexes a second model later is `UNIQUE (envelope_id,
chunk_index, model)` at the schema layer.
A DB write failure for one chunk (dropped connection, unexpected bytes) is
isolated the same way an embed failure is — counted as `chunks_errors`,
never aborting the rest of the run.
### Dimension guard
Every embedding response's length is checked against
`document_chunk.embedding`'s `VECTOR(1024)` column. A mismatch raises
`EmbeddingDimensionError` and aborts the whole run immediately — never
silently indexes vectors of the wrong dimension.
### Chunk size/overlap validation
`--chunk-overlap` must be smaller than `--chunk-size` — the sliding-window
hard-split fallback advances by `chunk_size - chunk_overlap` per step, so an
overlap `>=` size would never advance and hang. `main()` rejects this
combination before opening a DB connection; `hard_split()` itself also
raises `ValueError` as a second line of defense for direct callers.
### Stats must balance
```
documents_fetched = empty_content + documents_chunked
chunks_total = chunks_already_embedded + chunks_inserted
+ chunks_conflict_skipped + chunks_errors
```
`main()` exits 1 on `chunks_errors > 0`, `chunks_conflict_skipped > 0`, or if
either balance breaks. The summary line also reports
`avg_embed_seconds_per_chunk` — CPU-only Ollama timing, the input for
deciding whether/how to scale this to the mail corpus later (plan §7).
### Tests
```bash
pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest
```
Pure unit tests, no DB or real HTTP — `run()` is tested by monkeypatching
`asyncpg.connect` (fake connection) and `aiohttp.ClientSession` (fake session
serving a canned embedding vector, or a 500 for a chosen prompt to exercise
error isolation). Covers: chunking (paragraph boundaries, overlap, empty
document, document shorter than one chunk, oversized paragraph hard-fallback,
the overlap-must-be-smaller-than-size guard), `extract_content`, idempotency
(pre-existing keys skipped, no Ollama calls made for them, a second `--apply`
run embeds nothing new, existing keys are correctly scoped to `--model`),
dimension-mismatch abort, isolated per-chunk embed and insert errors,
`ON CONFLICT` no-ops counted separately from real inserts, and the
stats-balance invariant.
### Known limitation — Ollama context-length rejections on pathological chunks
Ollama's *runtime* context window for a model can be smaller than the
model's advertised max (bge-m3 supports 8192 tokens, but Ollama's default
`num_ctx` is lower) — and some OCR text tokenizes far more densely than the
~4-chars/token heuristic this job uses to size chunks. Concretely: a table-
of-contents page made almost entirely of dot-leader formatting
(`". . . . . . ."`, repeated hundreds of times) hit this on the pilot run —
Ollama returned `500 {"error":"the input length exceeds the context
length"}` for one 2400-char chunk that should have been well within budget
by character count alone. The job isolates this exactly like any other embed
failure (`chunks_errors`, logged, run continues), so it never crashes a
run — but it also never automatically shrinks and retries the offending
chunk. Given how rare this was (1 chunk out of 2684 in the full pilot, all
from one document's dot-leader ToC), it's left as a known gap rather than
fixed here; a real fix would be either a smaller/adaptive chunk size for
low-character-entropy text, or a shrink-and-retry loop on this specific
Ollama error.
### Definition of Done
Per `CLAUDE.md`: `pytest` passes locally (101 tests). Smoke-tested and then
run to completion live on SOLARIA against the real Ollama instance and
kb-postgres@PIHA:
- Dry-run: 186 fetched, 26 `empty_content`, 2684 chunks planned — matches
the known phase-2-step-5 figures exactly.
- `--apply --limit 10`: 64 chunks embedded, 0 errors, avg ≈0.83s/chunk on CPU.
- Re-run of the same slice: fully idempotent — 0 Ollama calls, 0 inserts.
- Full `--apply` (all 186 documents): **2683/2684 chunks inserted, 1 isolated
error** (see "Known limitation" above) — `chunks_errors=1` correctly
produced a non-zero exit rather than silently reporting success.
`document_chunk` ends at 2683 rows across 160 distinct envelopes, matching
`documents_chunked`. A `document_chunk_envelope_idx`-backed count and an
`ORDER BY embedding <=> ...` nearest-neighbor sanity query both look
correct (top match is the reference chunk itself at distance 0; next
nearest are chunks of the same source document).
- **Timing (CPU-only, no GPU driver on SOLARIA)**: ≈0.79s/chunk average
across 2683 real embeddings (2115.8s total embed time), ≈13.2s/document
average across the 160 chunked documents, ≈35 minutes wall-clock for the
full 186-document pilot. This is the real-world input for scaling this
pipeline to the much larger mail corpus later (plan §7 assumed GPU-based
"minutes for the whole pilot"; SOLARIA's Ollama ran CPU-only for this pilot
per the then-disabled GPU reservation). The 186-document pilot's
≈13.2s/document average is dominated by Paperless' long OCR text (≈22k
chars/doc average, per plan §1.2) — 225 030 mail envelopes will have a
very different, likely much shorter, per-envelope chunk count (email
bodies vs. scanned multi-page PDFs), so this number doesn't extrapolate
directly to a mail-corpus estimate. What it does establish: at
≈0.79s/chunk sequential CPU embedding, any corpus with a non-trivial
average chunk count per item will need either a GPU driver fix,
concurrent/batched Ollama calls, or both, before a full mail-corpus run
is practical — flagged for whoever picks up the mail-indexer phase.
- **GPU (RTX 4070 Ti SUPER, driver 595-open, restored 2026-07-16)**: 207ms/embed
(50 sekwencyjnych wywołań /api/embeddings, ~600-tok prompt) vs 790ms/chunk CPU
baseline — ~3.8× szybciej sekwencyjnie; przy pojedynczych requestach dominuje
overhead HTTP/tokenizacji, realny skok da dopiero batching (backlog).
---
## Phase 3 step 4 — retrieval cascade (`documents_ingest.retrieval`) + quality gate
Module 5, phase 3, plan step 4 (`docs/kb/modules/05-faza3-plan.md`, §6). Two retrieval
paths, both `query_text -> chunk hits (dist, source)` — the intended clean API surface for
phase 4's kb-query, not just this eval:
- `flat_query` — baseline: rank every active `document_chunk` row directly. Formalizes the
phase-2 pilot's ad hoc `/tmp/kbq.sh` query into a tested module.
- `cascade_query` — pre-filter to the top-N `document_summary` envelopes (one `model`,
default `claude-haiku-4-5` — plan §2 decision 3, resolved 2026-07-17) before ranking
`document_chunk` within just those envelopes. Both share **one** query embedding call;
the cascade only adds one extra SQL query (stage 1), never an extra Ollama call.
`envelope`, `document_chunk`, and `document_summary` are read-only — this module only ever
`SELECT`s.
### Quality gate
`eval/queries.yaml` — 7 queries transcribed 1:1 from the phase-2 pilot baseline
(`docs/kb/eval/retrieval-pilot-2026-07-16.md`, left untouched — this is its versioned working
copy) with expected envelope / kind (`hit`, `grey_zone`, `negative_control`,
`negative_control_borderline`) per query.
`eval/retrieval_eval.py` — read-only integration script against the live DB + live Ollama,
**not collected by pytest** (same reasoning as the plan: an eval gate against live data isn't
a mocked unit test). Runs every query through both tracks across an N sweep and checks the
plan's three gate criteria (no flat hit degrades, hit@3 cascade ≥ flat, negative controls
stay > 0.55). Exits 0 on PASS, 1 on FAIL.
```bash
pip install -e packages/kb-mail/ -e jobs/documents-ingest/
python eval/retrieval_eval.py --dsn postgresql://kb:<pw>@piha:5433/kb \
--ollama-url http://solaria:11434 --n-sweep 5,10,20
```
`--transport http --base-url http://<kb-query-host>:8230` (module 5 phase 4 plan §2 decision 6 /
§9) calls a live `kb-query`'s `/search` instead of embedding+querying locally — no `--dsn`
needed, `--n-sweep` is ignored (kb-query serves one server-side default N per request). Gate
criterion: `dist` must be **identical** to the same run with `--transport direct` against the
same live SOLARIA (same DB, same retrieval code — HTTP is only a wrapper).
**Result (2026-07-17, live run)**: PASS at N=10, k=5 — see plan §6.3 for the full table,
the N-sweep calibration (N=5 is the measured safety floor; the plan's N=10 default carries a
2× margin), and the cost/improvement analysis. `cascade_query` (N=10, k=5,
`summary_model='claude-haiku-4-5'`) is now the default retrieval path for phase 4's kb-query;
`flat_query` stays as the baseline/fallback.
### Tests
```bash
pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest
```
`tests/test_retrieval.py` — pure unit tests, no DB or real HTTP. Covers: flat ranking across
all envelopes, cascade stage-1-narrows-stage-2, an envelope whose summary exists but has no
active chunks, N larger than the number of summarized envelopes, the no-summaries
short-circuit (stage 2 never queried), and both query entry points embedding exactly once.
---
## Phase 3 step 5 — cyclic ingest (`documents-ingest-cyclic`) + systemd timer
Module 5, phase 3, plan step 5 (`docs/kb/modules/05-faza3-plan.md`, §7). Orchestrates one
run of the recurring ingest pipeline: `paperless_adapter.run()` (new `source='paperless'`
envelopes) → `chunk_embed.run()` (new `document_chunk` rows) →
`summarize.run_summarize(backend='anthropic')` (new `document_summary` rows,
`model='claude-haiku-4-5'` — plan §2 decision 3) → `summarize.run_embed_summaries()`
(embeds those summaries). All four are the same job functions used elsewhere in this
package, called directly — no changes to `paperless_adapter.py` / `chunk_embed.py` /
`summarize.py`, no new CLI flags on them.
### Ollama-offline tolerance
SOLARIA has `availability_target: medium` (planned power-off, plan §1.3). The wrapper
probes `GET {OLLAMA_URL}/api/tags` before the two embed stages (chunk embedding, summary
embedding); unreachable means **skip, not fail** — both embed passes are idempotent, so
new chunks/summaries left unembedded this tick are picked up whole on the next one. A
growing backlog is what `kb_ingest_embed_backlog` + the `KbEmbedBacklogGrowing` alert are
for, not this wrapper's exit code.
Anything else failing **is** a hard failure: Paperless unreachable, a DB error, a non-zero
job error counter, a broken stats-balance invariant, the Anthropic API failing. Each
stage's pass/fail predicate mirrors that job's own `main()` exit check 1:1 (see
`cyclic_ingest.py`'s module docstring). Stages are isolated, not fail-fast — an earlier
stage failing never skips a later one, mirroring the per-row isolation the underlying jobs
already use.
### Usage
```bash
# Dry run (default) — same idempotent counting as every other job in this family, no writes:
documents-ingest-cyclic --dsn postgresql://kb:<pw>@localhost:5433/kb \
--paperless-token <token> --anthropic-api-key <key>
# Real run (what the timer invokes):
documents-ingest-cyclic --dsn ... --paperless-token ... --anthropic-api-key ... --apply
```
`--dsn`/`--paperless-token`/`--anthropic-api-key` also read from `KB_DSN` /
`PAPERLESS_API_TOKEN` / `ANTHROPIC_API_KEY` env vars — never logged. `--ollama-url`
defaults to `http://solaria:11434` (this wrapper always runs on PIHA, unlike
`chunk_embed`/`summarize`'s own CLI defaults which assume co-location with Ollama).
### Metrics (Prometheus textfile collector)
Every run — success or failure — writes `--prom-path`
(default `/opt/homelab/state/node-exporter/kb-ingest.prom`) atomically (tmp + rename):
| Metric | Meaning |
|---|---|
| `kb_ingest_last_run_timestamp` | Unix ts of the last run, success or failure |
| `kb_ingest_last_success_timestamp` | Unix ts of the last run with no hard failure — carried forward from the previous file on a failing run, never reset to 0/now |
| `kb_ingest_last_exit_code` | 0 or 1 |
| `kb_ingest_documents_inserted` | New envelope rows this run |
| `kb_ingest_chunks_inserted` | New document_chunk rows this run (0 if the embed stage was skipped) |
| `kb_ingest_summaries_inserted` | New document_summary rows this run |
| `kb_ingest_embed_skipped` | 1 if Ollama was unreachable this run (both embed stages skipped), 0 otherwise |
| `kb_ingest_embed_backlog` | Active chunks (`excluded_reason IS NULL`) still missing an embedding |
Scraped by fleet-prometheus via node_exporter's textfile collector on PIHA
(`hosts/piha/runtime/node_exporter/docker-compose.override.yml`); alert rules in
`services/fleet-prometheus/rules/kb-ingest.yml`.
### Install (PIHA)
1. Dedicated venv (per plan §7.1 — not the ad hoc rsync-to-`/tmp` pattern used for the
one-shot jobs elsewhere in this README; this is a permanent, recurring installation):
```bash
python3 -m venv /opt/homelab/kb/venv
/opt/homelab/kb/venv/bin/pip install -e packages/kb-mail -e jobs/documents-ingest
```
(run from a checkout of this repo on PIHA — the checkout is used as an install source
only, per CLAUDE.md's "deploy-only" rule; no development happens there).
2. Secrets in `/opt/homelab/kb/.env` (already holds `PAPERLESS_API_TOKEN`; add
`KB_DSN=postgresql://kb:<pw>@localhost:5433/kb` and `ANTHROPIC_API_KEY=<key>`),
`chmod 600`, never in Git.
3. Copy `jobs/documents-ingest/systemd/kb-ingest-run.sh` to `/opt/homelab/kb/` and
`chmod +x` it.
4. Copy (or symlink) `kb-ingest.service` and `kb-ingest.timer` to `/etc/systemd/system/`,
then:
```bash
systemctl daemon-reload
systemctl enable --now kb-ingest.timer
```
5. Verify: `systemctl list-timers kb-ingest.timer`, `journalctl -u kb-ingest.service`,
`/opt/homelab/logs/kb-ingest/run-YYYYMMDD.log`, and
`/opt/homelab/state/node-exporter/kb-ingest.prom` after the first run (manual
`systemctl start kb-ingest.service` to trigger one immediately without waiting for
03:30).
### Tests
```bash
pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest
```
`tests/test_cyclic_ingest.py` — pure unit tests, no DB or real HTTP/Ollama/Anthropic; every
stage function and the Ollama probe are monkeypatched. Covers: each stage's failure
predicate (pinned against its source job's own exit check), the Ollama-down skip path
(chunk_embed/embed_summaries never even called), stage isolation (an earlier stage failing
never skips a later one, whether via a failed predicate or a raised exception), `.prom`
rendering, atomic write, `last_success_timestamp` carry-forward across a failing run, and
`main()`'s CLI guardrails + exit-code propagation.