feat(kb): SPLIT documents-ingest (30 KB) -> service + phase + runbook

kb/services/job-documents-ingest.md — opis jobu i mechanizmow
  (candidate selection, matching, consume/, idempotency, dry-run)
kb/phases/kb-m5-documents-ingest-fazy.md — faza 2, faza 2 krok 6,
  faza 3 krok 4, faza 3 krok 5 (4 sekcje fazowe wtopione w README)
kb/runbooks/documents-ingest-run.md — Usage, Verifying in Paperless, Tests

Najwiekszy README w repo. Tresc sekcji nietknieta; kontrola multizbioru
linii == oryginal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
oskar 2026-08-04 15:06:52 +02:00
parent 5439c00def
commit ccb738cddf
3 changed files with 221 additions and 184 deletions

View file

@ -1,188 +1,15 @@
# documents-ingest
One-shot job CLI, Phase 1 of module 5 (`docs/kb/modules/05-documents-ingest.md`,
"Domkniecie dlugu z maili"). Extracts a **sample** of PDF attachments from the
Gmail `.eml` archive (already indexed in the `envelope` table of kb-postgres)
and drops them into Paperless' `consume/` directory so Paperless does the OCR
and correspondent-detection. This job does **not** write to the `envelope`
table — the Paperless/Nextcloud envelope adapter is a later phase of module 5.
## Why a sample, not a bulk import
The Gmail import left ~70k attachments referenced in `envelope.entities`
manifests (bytes live inside the archived `.eml` files, never extracted).
Dumping all of them into Paperless at once would swamp the OCR worker and the
RAG layer isn't built yet to make use of that volume. This job pulls a small,
recent, size-filtered sample (default: 150 envelopes, PDFs >50KB, from the
last year) as a testbed — mass import is a deliberate later decision.
## Where it runs
**Locally on PIHA**, as a plain CLI (not a container). It needs simultaneous
filesystem access to three things that all live on PIHA:
- the mail archive (`/home/oskar/kb/mail/archive`)
- the Paperless `consume/` directory (`/opt/homelab/data/paperless/consume`)
- kb-postgres (`localhost:5433` from PIHA; reachable from elsewhere over
Tailscale, but the archive and consume dir are not — those are local paths)
Install (from repo root, on PIHA):
```bash
pip install -e jobs/documents-ingest/
```
(Or reuse the venv already set up for `gmail-bulk-import`, e.g.
`/home/oskar/kb/venv/` — it already has `asyncpg` + `structlog`.)
## Usage
```bash
# Dry run (default) — preview only, no writes:
documents-ingest --dsn postgresql://kb:<pw>@localhost:5433/kb
# Or via env var instead of --dsn:
export KB_DSN=postgresql://kb:<pw>@localhost:5433/kb
documents-ingest
# Real run — write files into consume/ and update the registry:
documents-ingest --apply
# Smaller/larger sample, different window/threshold:
documents-ingest --limit 50 --since-days 180 --min-size 100000
```
Dry-run is the default and does not require `--consume-dir` to exist yet;
`--apply` does (Paperless must already be deployed with its consume dir in
place). See `documents-ingest --help` for all flags.
## Candidate selection
```sql
SELECT id, raw_ref, ts, entities FROM envelope
WHERE source = 'gmail'
AND ts > now() - interval '1 year'
AND EXISTS (
SELECT 1 FROM jsonb_array_elements(entities) AS att
WHERE att->>'content_type' = 'application/pdf'
AND (att->>'size')::numeric > 50000
)
ORDER BY ts DESC
LIMIT 150
```
For each matching envelope, every attachment manifest entry that passes the
filter is a separate candidate (one envelope can yield several PDFs).
## Matching an attachment inside the .eml
The manifest (`entities[]`) only has metadata — the attachment bytes live
inside the `.eml` (MIME multipart), so each candidate is resolved against the
freshly parsed message:
1. Parse the `.eml` with `email.policy.default` and collect every
`application/pdf` MIME part (filename + decoded payload).
2. **sha256 is the proof of identity**, not the filename. The manifest was
built by a different parser at import time (`gmail-bulk-import`, using
`mailbox` + compat32 policy) and can still hold the raw RFC 2047
encoded-word form of a filename (e.g. `=?UTF-8?b?...?=`, sometimes with
header-folding whitespace baked in), while `email.policy.default` decodes
it to real Unicode today. Comparing those byte-for-byte skipped ~10% of
otherwise-good attachments in testing — see `TestFindPdfParts` /
`TestProcessCandidate` in the test suite for the regression case. So:
match by sha256 across all PDF parts in the message; if none match, use a
filename match only to tell "found the named part but its bytes changed"
(`sha_mismatch`, reported and skipped) apart from "not present at all"
(`parse_error`, skipped).
3. The consume/ filename is built from the *decoded* filename (from the MIME
part), not the possibly-garbled manifest one.
Mismatches and parse errors are never guessed past — they're logged and
skipped.
## consume/ filenames
`<YYYY-MM-DD>_<sanitized-filename>.pdf`, date = envelope `ts`. On collision
(same date + sanitized name already used in this run or already present in
`consume/`), an 8-hex sha256 prefix is appended:
`<YYYY-MM-DD>_<sanitized-filename>_<hash8>.pdf`.
Files are written with a best-effort `chown` to uid:gid `1000:1000` (the
Paperless container's `USERMAP_UID/GID`, see `services/paperless/README.md`)
so Paperless can read them. If the chown fails (e.g. the job isn't running as
root/uid 1000), a warning is logged but the run continues — the write itself
already succeeded; fix ownership/perms on `consume/` separately if needed.
PIHA's uid/gid convention across the fleet is tracked as its own tech-debt
item (see `docs/backlog/`), not solved here.
## Idempotency — registry
A JSON file at `/opt/homelab/data/documents-ingest/registry.json` (default,
override with `--registry`), keyed by attachment sha256:
```json
{
"<sha256>": {
"envelope_id": "...",
"filename": "...",
"consume_name": "2026-06-09_invoice.pdf",
"size": 123456,
"ingested_at": "2026-07-13T19:35:16+00:00"
}
}
```
**Why a JSON file and not a kb-postgres table:** this is a one-shot sampling
tool for a bootstrapping phase, not a long-running service — a new table
would formalize infrastructure for something temporary. A flat file needs no
migration, is trivial to inspect (`jq`) or reset, and sits under
`/opt/homelab/data/` alongside other node-local state per the repo's runtime
path convention. If/when module 5's real Paperless/Nextcloud adapter phase
starts writing `envelope` rows for `source=paperless`, that's the natural
point to fold this into a proper DB-backed ingest log — re-litigate then, not
now.
Re-running the job only ever *adds* to the registry (on `--apply`); it's
never consulted or mutated in dry-run mode beyond being read for the preview.
## Dry-run output
Logs one line per skip (`skip.duplicate` / `skip.sha_mismatch` /
`skip.parse_error`, with reason), a `summary` line with full counts
(`envelopes_scanned`, `pdf_candidates`, `extracted`, `skipped_duplicate`,
`skipped_sha_mismatch`, `skipped_parse_error`, `errors`), and up to 20 example
`(target_name, size, envelope_id)` rows so you can sanity-check filenames
before running `--apply`.
## Verifying the result in Paperless
After `--apply`:
1. Paperless' consumer picks files up from `consume/` automatically (polling
or inotify, per its own config) — no action needed on this job's side.
2. Watch progress: Paperless UI → Documents (new items appear as OCR
finishes), or `docker logs -f paperless` on PIHA for consumer/OCR activity.
3. Cross-check count: number of new documents in Paperless should equal
`stats["extracted"]` from the `--apply` run's summary line.
4. Confirm idempotency: re-running `--apply` immediately after should report
`extracted: 0` and `skipped_duplicate` equal to the previous run's
`extracted` count — nothing new lands in `consume/`.
## Tests
```bash
pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest
```
Pure unit tests, no DB or filesystem outside `tmp_path` required — `run()` is
tested by monkeypatching `asyncpg.connect` with an in-memory fake connection.
Covers: filename sanitization, consume-name collision handling, manifest
filtering, MIME PDF-part extraction (including the RFC 2047 decoding
mismatch), sha256 match/mismatch, duplicate detection, dry-run vs `--apply`
behavior, and multi-attachment envelopes.
---
okf: "0.1"
type: phase
visibility: private
status: active
updated: 2026-07-30
links:
- ../services/job-documents-ingest.md
- ../runbooks/documents-ingest-run.md
---
# documents-ingest — fazy 2 i 3
## Phase 2 — `documents-ingest-paperless` (Paperless -> envelope adapter)

View file

@ -0,0 +1,64 @@
---
okf: "0.1"
type: runbook
visibility: private
status: active
updated: 2026-07-30
links:
- ../services/job-documents-ingest.md
- ../phases/kb-m5-documents-ingest-fazy.md
---
# documents-ingest — uruchomienie i weryfikacja
## Usage
```bash
# Dry run (default) — preview only, no writes:
documents-ingest --dsn postgresql://kb:<pw>@localhost:5433/kb
# Or via env var instead of --dsn:
export KB_DSN=postgresql://kb:<pw>@localhost:5433/kb
documents-ingest
# Real run — write files into consume/ and update the registry:
documents-ingest --apply
# Smaller/larger sample, different window/threshold:
documents-ingest --limit 50 --since-days 180 --min-size 100000
```
Dry-run is the default and does not require `--consume-dir` to exist yet;
`--apply` does (Paperless must already be deployed with its consume dir in
place). See `documents-ingest --help` for all flags.
## Verifying the result in Paperless
After `--apply`:
1. Paperless' consumer picks files up from `consume/` automatically (polling
or inotify, per its own config) — no action needed on this job's side.
2. Watch progress: Paperless UI → Documents (new items appear as OCR
finishes), or `docker logs -f paperless` on PIHA for consumer/OCR activity.
3. Cross-check count: number of new documents in Paperless should equal
`stats["extracted"]` from the `--apply` run's summary line.
4. Confirm idempotency: re-running `--apply` immediately after should report
`extracted: 0` and `skipped_duplicate` equal to the previous run's
`extracted` count — nothing new lands in `consume/`.
## Tests
```bash
pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest
```
Pure unit tests, no DB or filesystem outside `tmp_path` required — `run()` is
tested by monkeypatching `asyncpg.connect` with an in-memory fake connection.
Covers: filename sanitization, consume-name collision handling, manifest
filtering, MIME PDF-part extraction (including the RFC 2047 decoding
mismatch), sha256 match/mismatch, duplicate detection, dry-run vs `--apply`
behavior, and multi-attachment envelopes.
---

View file

@ -0,0 +1,146 @@
---
okf: "0.1"
type: service
visibility: private
status: active
updated: 2026-07-30
links:
- ../phases/kb-m5-documents-ingest-fazy.md
- ../runbooks/documents-ingest-run.md
---
# documents-ingest
One-shot job CLI, Phase 1 of module 5 (`docs/kb/modules/05-documents-ingest.md`,
"Domkniecie dlugu z maili"). Extracts a **sample** of PDF attachments from the
Gmail `.eml` archive (already indexed in the `envelope` table of kb-postgres)
and drops them into Paperless' `consume/` directory so Paperless does the OCR
and correspondent-detection. This job does **not** write to the `envelope`
table — the Paperless/Nextcloud envelope adapter is a later phase of module 5.
## Why a sample, not a bulk import
The Gmail import left ~70k attachments referenced in `envelope.entities`
manifests (bytes live inside the archived `.eml` files, never extracted).
Dumping all of them into Paperless at once would swamp the OCR worker and the
RAG layer isn't built yet to make use of that volume. This job pulls a small,
recent, size-filtered sample (default: 150 envelopes, PDFs >50KB, from the
last year) as a testbed — mass import is a deliberate later decision.
## Where it runs
**Locally on PIHA**, as a plain CLI (not a container). It needs simultaneous
filesystem access to three things that all live on PIHA:
- the mail archive (`/home/oskar/kb/mail/archive`)
- the Paperless `consume/` directory (`/opt/homelab/data/paperless/consume`)
- kb-postgres (`localhost:5433` from PIHA; reachable from elsewhere over
Tailscale, but the archive and consume dir are not — those are local paths)
Install (from repo root, on PIHA):
```bash
pip install -e jobs/documents-ingest/
```
(Or reuse the venv already set up for `gmail-bulk-import`, e.g.
`/home/oskar/kb/venv/` — it already has `asyncpg` + `structlog`.)
## Candidate selection
```sql
SELECT id, raw_ref, ts, entities FROM envelope
WHERE source = 'gmail'
AND ts > now() - interval '1 year'
AND EXISTS (
SELECT 1 FROM jsonb_array_elements(entities) AS att
WHERE att->>'content_type' = 'application/pdf'
AND (att->>'size')::numeric > 50000
)
ORDER BY ts DESC
LIMIT 150
```
For each matching envelope, every attachment manifest entry that passes the
filter is a separate candidate (one envelope can yield several PDFs).
## Matching an attachment inside the .eml
The manifest (`entities[]`) only has metadata — the attachment bytes live
inside the `.eml` (MIME multipart), so each candidate is resolved against the
freshly parsed message:
1. Parse the `.eml` with `email.policy.default` and collect every
`application/pdf` MIME part (filename + decoded payload).
2. **sha256 is the proof of identity**, not the filename. The manifest was
built by a different parser at import time (`gmail-bulk-import`, using
`mailbox` + compat32 policy) and can still hold the raw RFC 2047
encoded-word form of a filename (e.g. `=?UTF-8?b?...?=`, sometimes with
header-folding whitespace baked in), while `email.policy.default` decodes
it to real Unicode today. Comparing those byte-for-byte skipped ~10% of
otherwise-good attachments in testing — see `TestFindPdfParts` /
`TestProcessCandidate` in the test suite for the regression case. So:
match by sha256 across all PDF parts in the message; if none match, use a
filename match only to tell "found the named part but its bytes changed"
(`sha_mismatch`, reported and skipped) apart from "not present at all"
(`parse_error`, skipped).
3. The consume/ filename is built from the *decoded* filename (from the MIME
part), not the possibly-garbled manifest one.
Mismatches and parse errors are never guessed past — they're logged and
skipped.
## consume/ filenames
`<YYYY-MM-DD>_<sanitized-filename>.pdf`, date = envelope `ts`. On collision
(same date + sanitized name already used in this run or already present in
`consume/`), an 8-hex sha256 prefix is appended:
`<YYYY-MM-DD>_<sanitized-filename>_<hash8>.pdf`.
Files are written with a best-effort `chown` to uid:gid `1000:1000` (the
Paperless container's `USERMAP_UID/GID`, see `services/paperless/README.md`)
so Paperless can read them. If the chown fails (e.g. the job isn't running as
root/uid 1000), a warning is logged but the run continues — the write itself
already succeeded; fix ownership/perms on `consume/` separately if needed.
PIHA's uid/gid convention across the fleet is tracked as its own tech-debt
item (see `docs/backlog/`), not solved here.
## Idempotency — registry
A JSON file at `/opt/homelab/data/documents-ingest/registry.json` (default,
override with `--registry`), keyed by attachment sha256:
```json
{
"<sha256>": {
"envelope_id": "...",
"filename": "...",
"consume_name": "2026-06-09_invoice.pdf",
"size": 123456,
"ingested_at": "2026-07-13T19:35:16+00:00"
}
}
```
**Why a JSON file and not a kb-postgres table:** this is a one-shot sampling
tool for a bootstrapping phase, not a long-running service — a new table
would formalize infrastructure for something temporary. A flat file needs no
migration, is trivial to inspect (`jq`) or reset, and sits under
`/opt/homelab/data/` alongside other node-local state per the repo's runtime
path convention. If/when module 5's real Paperless/Nextcloud adapter phase
starts writing `envelope` rows for `source=paperless`, that's the natural
point to fold this into a proper DB-backed ingest log — re-litigate then, not
now.
Re-running the job only ever *adds* to the registry (on `--apply`); it's
never consulted or mutated in dry-run mode beyond being read for the preview.
## Dry-run output
Logs one line per skip (`skip.duplicate` / `skip.sha_mismatch` /
`skip.parse_error`, with reason), a `summary` line with full counts
(`envelopes_scanned`, `pdf_candidates`, `extracted`, `skipped_duplicate`,
`skipped_sha_mismatch`, `skipped_parse_error`, `errors`), and up to 20 example
`(target_name, size, envelope_id)` rows so you can sanity-check filenames
before running `--apply`.