From 4d69cf7f8c0109820da26077f36910785638ef38 Mon Sep 17 00:00:00 2001 From: oskar Date: Tue, 4 Aug 2026 15:06:52 +0200 Subject: [PATCH] feat(kb): SPLIT documents-ingest (30 KB) -> service + phase + runbook MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit kb/services/job-documents-ingest.md — opis jobu i mechanizmow (candidate selection, matching, consume/, idempotency, dry-run) kb/phases/kb-m5-documents-ingest-fazy.md — faza 2, faza 2 krok 6, faza 3 krok 4, faza 3 krok 5 (4 sekcje fazowe wtopione w README) kb/runbooks/documents-ingest-run.md — Usage, Verifying in Paperless, Tests Najwiekszy README w repo. Tresc sekcji nietknieta; kontrola multizbioru linii == oryginal. Co-Authored-By: Claude Opus 5 (1M context) --- .../phases/kb-m5-documents-ingest-fazy.md | 195 +----------------- kb/runbooks/documents-ingest-run.md | 64 ++++++ kb/services/job-documents-ingest.md | 146 +++++++++++++ 3 files changed, 221 insertions(+), 184 deletions(-) rename jobs/documents-ingest/README.md => kb/phases/kb-m5-documents-ingest-fazy.md (74%) create mode 100644 kb/runbooks/documents-ingest-run.md create mode 100644 kb/services/job-documents-ingest.md diff --git a/jobs/documents-ingest/README.md b/kb/phases/kb-m5-documents-ingest-fazy.md similarity index 74% rename from jobs/documents-ingest/README.md rename to kb/phases/kb-m5-documents-ingest-fazy.md index 624a05f..b8b0f91 100644 --- a/jobs/documents-ingest/README.md +++ b/kb/phases/kb-m5-documents-ingest-fazy.md @@ -1,188 +1,15 @@ -# documents-ingest - -One-shot job CLI, Phase 1 of module 5 (`docs/kb/modules/05-documents-ingest.md`, -"Domkniecie dlugu z maili"). Extracts a **sample** of PDF attachments from the -Gmail `.eml` archive (already indexed in the `envelope` table of kb-postgres) -and drops them into Paperless' `consume/` directory so Paperless does the OCR -and correspondent-detection. This job does **not** write to the `envelope` -table — the Paperless/Nextcloud envelope adapter is a later phase of module 5. - -## Why a sample, not a bulk import - -The Gmail import left ~70k attachments referenced in `envelope.entities` -manifests (bytes live inside the archived `.eml` files, never extracted). -Dumping all of them into Paperless at once would swamp the OCR worker and the -RAG layer isn't built yet to make use of that volume. This job pulls a small, -recent, size-filtered sample (default: 150 envelopes, PDFs >50KB, from the -last year) as a testbed — mass import is a deliberate later decision. - -## Where it runs - -**Locally on PIHA**, as a plain CLI (not a container). It needs simultaneous -filesystem access to three things that all live on PIHA: - -- the mail archive (`/home/oskar/kb/mail/archive`) -- the Paperless `consume/` directory (`/opt/homelab/data/paperless/consume`) -- kb-postgres (`localhost:5433` from PIHA; reachable from elsewhere over - Tailscale, but the archive and consume dir are not — those are local paths) - -Install (from repo root, on PIHA): - -```bash -pip install -e jobs/documents-ingest/ -``` - -(Or reuse the venv already set up for `gmail-bulk-import`, e.g. -`/home/oskar/kb/venv/` — it already has `asyncpg` + `structlog`.) - -## Usage - -```bash -# Dry run (default) — preview only, no writes: -documents-ingest --dsn postgresql://kb:@localhost:5433/kb - -# Or via env var instead of --dsn: -export KB_DSN=postgresql://kb:@localhost:5433/kb -documents-ingest - -# Real run — write files into consume/ and update the registry: -documents-ingest --apply - -# Smaller/larger sample, different window/threshold: -documents-ingest --limit 50 --since-days 180 --min-size 100000 -``` - -Dry-run is the default and does not require `--consume-dir` to exist yet; -`--apply` does (Paperless must already be deployed with its consume dir in -place). See `documents-ingest --help` for all flags. - -## Candidate selection - -```sql -SELECT id, raw_ref, ts, entities FROM envelope -WHERE source = 'gmail' - AND ts > now() - interval '1 year' - AND EXISTS ( - SELECT 1 FROM jsonb_array_elements(entities) AS att - WHERE att->>'content_type' = 'application/pdf' - AND (att->>'size')::numeric > 50000 - ) -ORDER BY ts DESC -LIMIT 150 -``` - -For each matching envelope, every attachment manifest entry that passes the -filter is a separate candidate (one envelope can yield several PDFs). - -## Matching an attachment inside the .eml - -The manifest (`entities[]`) only has metadata — the attachment bytes live -inside the `.eml` (MIME multipart), so each candidate is resolved against the -freshly parsed message: - -1. Parse the `.eml` with `email.policy.default` and collect every - `application/pdf` MIME part (filename + decoded payload). -2. **sha256 is the proof of identity**, not the filename. The manifest was - built by a different parser at import time (`gmail-bulk-import`, using - `mailbox` + compat32 policy) and can still hold the raw RFC 2047 - encoded-word form of a filename (e.g. `=?UTF-8?b?...?=`, sometimes with - header-folding whitespace baked in), while `email.policy.default` decodes - it to real Unicode today. Comparing those byte-for-byte skipped ~10% of - otherwise-good attachments in testing — see `TestFindPdfParts` / - `TestProcessCandidate` in the test suite for the regression case. So: - match by sha256 across all PDF parts in the message; if none match, use a - filename match only to tell "found the named part but its bytes changed" - (`sha_mismatch`, reported and skipped) apart from "not present at all" - (`parse_error`, skipped). -3. The consume/ filename is built from the *decoded* filename (from the MIME - part), not the possibly-garbled manifest one. - -Mismatches and parse errors are never guessed past — they're logged and -skipped. - -## consume/ filenames - -`_.pdf`, date = envelope `ts`. On collision -(same date + sanitized name already used in this run or already present in -`consume/`), an 8-hex sha256 prefix is appended: -`__.pdf`. - -Files are written with a best-effort `chown` to uid:gid `1000:1000` (the -Paperless container's `USERMAP_UID/GID`, see `services/paperless/README.md`) -so Paperless can read them. If the chown fails (e.g. the job isn't running as -root/uid 1000), a warning is logged but the run continues — the write itself -already succeeded; fix ownership/perms on `consume/` separately if needed. -PIHA's uid/gid convention across the fleet is tracked as its own tech-debt -item (see `docs/backlog/`), not solved here. - -## Idempotency — registry - -A JSON file at `/opt/homelab/data/documents-ingest/registry.json` (default, -override with `--registry`), keyed by attachment sha256: - -```json -{ - "": { - "envelope_id": "...", - "filename": "...", - "consume_name": "2026-06-09_invoice.pdf", - "size": 123456, - "ingested_at": "2026-07-13T19:35:16+00:00" - } -} -``` - -**Why a JSON file and not a kb-postgres table:** this is a one-shot sampling -tool for a bootstrapping phase, not a long-running service — a new table -would formalize infrastructure for something temporary. A flat file needs no -migration, is trivial to inspect (`jq`) or reset, and sits under -`/opt/homelab/data/` alongside other node-local state per the repo's runtime -path convention. If/when module 5's real Paperless/Nextcloud adapter phase -starts writing `envelope` rows for `source=paperless`, that's the natural -point to fold this into a proper DB-backed ingest log — re-litigate then, not -now. - -Re-running the job only ever *adds* to the registry (on `--apply`); it's -never consulted or mutated in dry-run mode beyond being read for the preview. - -## Dry-run output - -Logs one line per skip (`skip.duplicate` / `skip.sha_mismatch` / -`skip.parse_error`, with reason), a `summary` line with full counts -(`envelopes_scanned`, `pdf_candidates`, `extracted`, `skipped_duplicate`, -`skipped_sha_mismatch`, `skipped_parse_error`, `errors`), and up to 20 example -`(target_name, size, envelope_id)` rows so you can sanity-check filenames -before running `--apply`. - -## Verifying the result in Paperless - -After `--apply`: - -1. Paperless' consumer picks files up from `consume/` automatically (polling - or inotify, per its own config) — no action needed on this job's side. -2. Watch progress: Paperless UI → Documents (new items appear as OCR - finishes), or `docker logs -f paperless` on PIHA for consumer/OCR activity. -3. Cross-check count: number of new documents in Paperless should equal - `stats["extracted"]` from the `--apply` run's summary line. -4. Confirm idempotency: re-running `--apply` immediately after should report - `extracted: 0` and `skipped_duplicate` equal to the previous run's - `extracted` count — nothing new lands in `consume/`. - -## Tests - -```bash -pip install -e jobs/documents-ingest/ -cd jobs/documents-ingest && pytest -``` - -Pure unit tests, no DB or filesystem outside `tmp_path` required — `run()` is -tested by monkeypatching `asyncpg.connect` with an in-memory fake connection. -Covers: filename sanitization, consume-name collision handling, manifest -filtering, MIME PDF-part extraction (including the RFC 2047 decoding -mismatch), sha256 match/mismatch, duplicate detection, dry-run vs `--apply` -behavior, and multi-attachment envelopes. - --- +okf: "0.1" +type: phase +visibility: private +status: active +updated: 2026-07-30 +links: + - ../services/job-documents-ingest.md + - ../runbooks/documents-ingest-run.md +--- + +# documents-ingest — fazy 2 i 3 ## Phase 2 — `documents-ingest-paperless` (Paperless -> envelope adapter) diff --git a/kb/runbooks/documents-ingest-run.md b/kb/runbooks/documents-ingest-run.md new file mode 100644 index 0000000..755f9a3 --- /dev/null +++ b/kb/runbooks/documents-ingest-run.md @@ -0,0 +1,64 @@ +--- +okf: "0.1" +type: runbook +visibility: private +status: active +updated: 2026-07-30 +links: + - ../services/job-documents-ingest.md + - ../phases/kb-m5-documents-ingest-fazy.md +--- + +# documents-ingest — uruchomienie i weryfikacja + +## Usage + +```bash +# Dry run (default) — preview only, no writes: +documents-ingest --dsn postgresql://kb:@localhost:5433/kb + +# Or via env var instead of --dsn: +export KB_DSN=postgresql://kb:@localhost:5433/kb +documents-ingest + +# Real run — write files into consume/ and update the registry: +documents-ingest --apply + +# Smaller/larger sample, different window/threshold: +documents-ingest --limit 50 --since-days 180 --min-size 100000 +``` + +Dry-run is the default and does not require `--consume-dir` to exist yet; +`--apply` does (Paperless must already be deployed with its consume dir in +place). See `documents-ingest --help` for all flags. + +## Verifying the result in Paperless + +After `--apply`: + +1. Paperless' consumer picks files up from `consume/` automatically (polling + or inotify, per its own config) — no action needed on this job's side. +2. Watch progress: Paperless UI → Documents (new items appear as OCR + finishes), or `docker logs -f paperless` on PIHA for consumer/OCR activity. +3. Cross-check count: number of new documents in Paperless should equal + `stats["extracted"]` from the `--apply` run's summary line. +4. Confirm idempotency: re-running `--apply` immediately after should report + `extracted: 0` and `skipped_duplicate` equal to the previous run's + `extracted` count — nothing new lands in `consume/`. + +## Tests + +```bash +pip install -e jobs/documents-ingest/ +cd jobs/documents-ingest && pytest +``` + +Pure unit tests, no DB or filesystem outside `tmp_path` required — `run()` is +tested by monkeypatching `asyncpg.connect` with an in-memory fake connection. +Covers: filename sanitization, consume-name collision handling, manifest +filtering, MIME PDF-part extraction (including the RFC 2047 decoding +mismatch), sha256 match/mismatch, duplicate detection, dry-run vs `--apply` +behavior, and multi-attachment envelopes. + +--- + diff --git a/kb/services/job-documents-ingest.md b/kb/services/job-documents-ingest.md new file mode 100644 index 0000000..f2f85b3 --- /dev/null +++ b/kb/services/job-documents-ingest.md @@ -0,0 +1,146 @@ +--- +okf: "0.1" +type: service +visibility: private +status: active +updated: 2026-07-30 +links: + - ../phases/kb-m5-documents-ingest-fazy.md + - ../runbooks/documents-ingest-run.md +--- + +# documents-ingest + +One-shot job CLI, Phase 1 of module 5 (`docs/kb/modules/05-documents-ingest.md`, +"Domkniecie dlugu z maili"). Extracts a **sample** of PDF attachments from the +Gmail `.eml` archive (already indexed in the `envelope` table of kb-postgres) +and drops them into Paperless' `consume/` directory so Paperless does the OCR +and correspondent-detection. This job does **not** write to the `envelope` +table — the Paperless/Nextcloud envelope adapter is a later phase of module 5. + +## Why a sample, not a bulk import + +The Gmail import left ~70k attachments referenced in `envelope.entities` +manifests (bytes live inside the archived `.eml` files, never extracted). +Dumping all of them into Paperless at once would swamp the OCR worker and the +RAG layer isn't built yet to make use of that volume. This job pulls a small, +recent, size-filtered sample (default: 150 envelopes, PDFs >50KB, from the +last year) as a testbed — mass import is a deliberate later decision. + +## Where it runs + +**Locally on PIHA**, as a plain CLI (not a container). It needs simultaneous +filesystem access to three things that all live on PIHA: + +- the mail archive (`/home/oskar/kb/mail/archive`) +- the Paperless `consume/` directory (`/opt/homelab/data/paperless/consume`) +- kb-postgres (`localhost:5433` from PIHA; reachable from elsewhere over + Tailscale, but the archive and consume dir are not — those are local paths) + +Install (from repo root, on PIHA): + +```bash +pip install -e jobs/documents-ingest/ +``` + +(Or reuse the venv already set up for `gmail-bulk-import`, e.g. +`/home/oskar/kb/venv/` — it already has `asyncpg` + `structlog`.) + +## Candidate selection + +```sql +SELECT id, raw_ref, ts, entities FROM envelope +WHERE source = 'gmail' + AND ts > now() - interval '1 year' + AND EXISTS ( + SELECT 1 FROM jsonb_array_elements(entities) AS att + WHERE att->>'content_type' = 'application/pdf' + AND (att->>'size')::numeric > 50000 + ) +ORDER BY ts DESC +LIMIT 150 +``` + +For each matching envelope, every attachment manifest entry that passes the +filter is a separate candidate (one envelope can yield several PDFs). + +## Matching an attachment inside the .eml + +The manifest (`entities[]`) only has metadata — the attachment bytes live +inside the `.eml` (MIME multipart), so each candidate is resolved against the +freshly parsed message: + +1. Parse the `.eml` with `email.policy.default` and collect every + `application/pdf` MIME part (filename + decoded payload). +2. **sha256 is the proof of identity**, not the filename. The manifest was + built by a different parser at import time (`gmail-bulk-import`, using + `mailbox` + compat32 policy) and can still hold the raw RFC 2047 + encoded-word form of a filename (e.g. `=?UTF-8?b?...?=`, sometimes with + header-folding whitespace baked in), while `email.policy.default` decodes + it to real Unicode today. Comparing those byte-for-byte skipped ~10% of + otherwise-good attachments in testing — see `TestFindPdfParts` / + `TestProcessCandidate` in the test suite for the regression case. So: + match by sha256 across all PDF parts in the message; if none match, use a + filename match only to tell "found the named part but its bytes changed" + (`sha_mismatch`, reported and skipped) apart from "not present at all" + (`parse_error`, skipped). +3. The consume/ filename is built from the *decoded* filename (from the MIME + part), not the possibly-garbled manifest one. + +Mismatches and parse errors are never guessed past — they're logged and +skipped. + +## consume/ filenames + +`_.pdf`, date = envelope `ts`. On collision +(same date + sanitized name already used in this run or already present in +`consume/`), an 8-hex sha256 prefix is appended: +`__.pdf`. + +Files are written with a best-effort `chown` to uid:gid `1000:1000` (the +Paperless container's `USERMAP_UID/GID`, see `services/paperless/README.md`) +so Paperless can read them. If the chown fails (e.g. the job isn't running as +root/uid 1000), a warning is logged but the run continues — the write itself +already succeeded; fix ownership/perms on `consume/` separately if needed. +PIHA's uid/gid convention across the fleet is tracked as its own tech-debt +item (see `docs/backlog/`), not solved here. + +## Idempotency — registry + +A JSON file at `/opt/homelab/data/documents-ingest/registry.json` (default, +override with `--registry`), keyed by attachment sha256: + +```json +{ + "": { + "envelope_id": "...", + "filename": "...", + "consume_name": "2026-06-09_invoice.pdf", + "size": 123456, + "ingested_at": "2026-07-13T19:35:16+00:00" + } +} +``` + +**Why a JSON file and not a kb-postgres table:** this is a one-shot sampling +tool for a bootstrapping phase, not a long-running service — a new table +would formalize infrastructure for something temporary. A flat file needs no +migration, is trivial to inspect (`jq`) or reset, and sits under +`/opt/homelab/data/` alongside other node-local state per the repo's runtime +path convention. If/when module 5's real Paperless/Nextcloud adapter phase +starts writing `envelope` rows for `source=paperless`, that's the natural +point to fold this into a proper DB-backed ingest log — re-litigate then, not +now. + +Re-running the job only ever *adds* to the registry (on `--apply`); it's +never consulted or mutated in dry-run mode beyond being read for the preview. + +## Dry-run output + +Logs one line per skip (`skip.duplicate` / `skip.sha_mismatch` / +`skip.parse_error`, with reason), a `summary` line with full counts +(`envelopes_scanned`, `pdf_candidates`, `extracted`, `skipped_duplicate`, +`skipped_sha_mismatch`, `skipped_parse_error`, `errors`), and up to 20 example +`(target_name, size, envelope_id)` rows so you can sanity-check filenames +before running `--apply`. +