homelab-codex-ws/jobs/documents-ingest
oskar f7b61f7da9 feat(documents-ingest): PDF attachment extractor — mail archive -> Paperless consume/
Module 5 phase 1 (docs/kb/modules/05-documents-ingest.md): pulls a sample of
PDF attachments (>50KB, last year, LIMIT 150) out of the Gmail .eml archive
and drops them into Paperless consume/ for OCR, so the RAG layer has real
documents to work with before the full Paperless/Nextcloud envelope adapter
is built. sha256-first attachment matching (manifest filenames can carry raw
RFC 2047 encoded-word artifacts that don't byte-match what email.policy.default
decodes today — confirmed against live data, ~10% of candidates were affected).
Idempotent via a sha256-keyed JSON registry; dry-run by default, --apply to write.

Verified end-to-end on PIHA: dry-run + --apply both run against live
kb-postgres/archive, 185/222 candidate PDFs written to consume/ (37 in-run
duplicates correctly deduped), Paperless picked them up and started OCR
immediately.
2026-07-13 20:03:57 +02:00
..
src/documents_ingest feat(documents-ingest): PDF attachment extractor — mail archive -> Paperless consume/ 2026-07-13 20:03:57 +02:00
tests feat(documents-ingest): PDF attachment extractor — mail archive -> Paperless consume/ 2026-07-13 20:03:57 +02:00
pyproject.toml feat(documents-ingest): PDF attachment extractor — mail archive -> Paperless consume/ 2026-07-13 20:03:57 +02:00
README.md feat(documents-ingest): PDF attachment extractor — mail archive -> Paperless consume/ 2026-07-13 20:03:57 +02:00

documents-ingest

One-shot job CLI, Phase 1 of module 5 (docs/kb/modules/05-documents-ingest.md, "Domkniecie dlugu z maili"). Extracts a sample of PDF attachments from the Gmail .eml archive (already indexed in the envelope table of kb-postgres) and drops them into Paperless' consume/ directory so Paperless does the OCR and correspondent-detection. This job does not write to the envelope table — the Paperless/Nextcloud envelope adapter is a later phase of module 5.

Why a sample, not a bulk import

The Gmail import left ~70k attachments referenced in envelope.entities manifests (bytes live inside the archived .eml files, never extracted). Dumping all of them into Paperless at once would swamp the OCR worker and the RAG layer isn't built yet to make use of that volume. This job pulls a small, recent, size-filtered sample (default: 150 envelopes, PDFs >50KB, from the last year) as a testbed — mass import is a deliberate later decision.

Where it runs

Locally on PIHA, as a plain CLI (not a container). It needs simultaneous filesystem access to three things that all live on PIHA:

  • the mail archive (/home/oskar/kb/mail/archive)
  • the Paperless consume/ directory (/opt/homelab/data/paperless/consume)
  • kb-postgres (localhost:5433 from PIHA; reachable from elsewhere over Tailscale, but the archive and consume dir are not — those are local paths)

Install (from repo root, on PIHA):

pip install -e jobs/documents-ingest/

(Or reuse the venv already set up for gmail-bulk-import, e.g. /home/oskar/kb/venv/ — it already has asyncpg + structlog.)

Usage

# Dry run (default) — preview only, no writes:
documents-ingest --dsn postgresql://kb:<pw>@localhost:5433/kb

# Or via env var instead of --dsn:
export KB_DSN=postgresql://kb:<pw>@localhost:5433/kb
documents-ingest

# Real run — write files into consume/ and update the registry:
documents-ingest --apply

# Smaller/larger sample, different window/threshold:
documents-ingest --limit 50 --since-days 180 --min-size 100000

Dry-run is the default and does not require --consume-dir to exist yet; --apply does (Paperless must already be deployed with its consume dir in place). See documents-ingest --help for all flags.

Candidate selection

SELECT id, raw_ref, ts, entities FROM envelope
WHERE source = 'gmail'
  AND ts > now() - interval '1 year'
  AND EXISTS (
      SELECT 1 FROM jsonb_array_elements(entities) AS att
      WHERE att->>'content_type' = 'application/pdf'
        AND (att->>'size')::numeric > 50000
  )
ORDER BY ts DESC
LIMIT 150

For each matching envelope, every attachment manifest entry that passes the filter is a separate candidate (one envelope can yield several PDFs).

Matching an attachment inside the .eml

The manifest (entities[]) only has metadata — the attachment bytes live inside the .eml (MIME multipart), so each candidate is resolved against the freshly parsed message:

  1. Parse the .eml with email.policy.default and collect every application/pdf MIME part (filename + decoded payload).
  2. sha256 is the proof of identity, not the filename. The manifest was built by a different parser at import time (gmail-bulk-import, using mailbox + compat32 policy) and can still hold the raw RFC 2047 encoded-word form of a filename (e.g. =?UTF-8?b?...?=, sometimes with header-folding whitespace baked in), while email.policy.default decodes it to real Unicode today. Comparing those byte-for-byte skipped ~10% of otherwise-good attachments in testing — see TestFindPdfParts / TestProcessCandidate in the test suite for the regression case. So: match by sha256 across all PDF parts in the message; if none match, use a filename match only to tell "found the named part but its bytes changed" (sha_mismatch, reported and skipped) apart from "not present at all" (parse_error, skipped).
  3. The consume/ filename is built from the decoded filename (from the MIME part), not the possibly-garbled manifest one.

Mismatches and parse errors are never guessed past — they're logged and skipped.

consume/ filenames

<YYYY-MM-DD>_<sanitized-filename>.pdf, date = envelope ts. On collision (same date + sanitized name already used in this run or already present in consume/), an 8-hex sha256 prefix is appended: <YYYY-MM-DD>_<sanitized-filename>_<hash8>.pdf.

Files are written with a best-effort chown to uid:gid 1000:1000 (the Paperless container's USERMAP_UID/GID, see services/paperless/README.md) so Paperless can read them. If the chown fails (e.g. the job isn't running as root/uid 1000), a warning is logged but the run continues — the write itself already succeeded; fix ownership/perms on consume/ separately if needed. PIHA's uid/gid convention across the fleet is tracked as its own tech-debt item (see docs/backlog/), not solved here.

Idempotency — registry

A JSON file at /opt/homelab/data/documents-ingest/registry.json (default, override with --registry), keyed by attachment sha256:

{
  "<sha256>": {
    "envelope_id": "...",
    "filename": "...",
    "consume_name": "2026-06-09_invoice.pdf",
    "size": 123456,
    "ingested_at": "2026-07-13T19:35:16+00:00"
  }
}

Why a JSON file and not a kb-postgres table: this is a one-shot sampling tool for a bootstrapping phase, not a long-running service — a new table would formalize infrastructure for something temporary. A flat file needs no migration, is trivial to inspect (jq) or reset, and sits under /opt/homelab/data/ alongside other node-local state per the repo's runtime path convention. If/when module 5's real Paperless/Nextcloud adapter phase starts writing envelope rows for source=paperless, that's the natural point to fold this into a proper DB-backed ingest log — re-litigate then, not now.

Re-running the job only ever adds to the registry (on --apply); it's never consulted or mutated in dry-run mode beyond being read for the preview.

Dry-run output

Logs one line per skip (skip.duplicate / skip.sha_mismatch / skip.parse_error, with reason), a summary line with full counts (envelopes_scanned, pdf_candidates, extracted, skipped_duplicate, skipped_sha_mismatch, skipped_parse_error, errors), and up to 20 example (target_name, size, envelope_id) rows so you can sanity-check filenames before running --apply.

Verifying the result in Paperless

After --apply:

  1. Paperless' consumer picks files up from consume/ automatically (polling or inotify, per its own config) — no action needed on this job's side.
  2. Watch progress: Paperless UI → Documents (new items appear as OCR finishes), or docker logs -f paperless on PIHA for consumer/OCR activity.
  3. Cross-check count: number of new documents in Paperless should equal stats["extracted"] from the --apply run's summary line.
  4. Confirm idempotency: re-running --apply immediately after should report extracted: 0 and skipped_duplicate equal to the previous run's extracted count — nothing new lands in consume/.

Tests

pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest

Pure unit tests, no DB or filesystem outside tmp_path required — run() is tested by monkeypatching asyncpg.connect with an in-memory fake connection. Covers: filename sanitization, consume-name collision handling, manifest filtering, MIME PDF-part extraction (including the RFC 2047 decoding mismatch), sha256 match/mismatch, duplicate detection, dry-run vs --apply behavior, and multi-attachment envelopes.