# documents-ingest One-shot job CLI, Phase 1 of module 5 (`docs/kb/modules/05-documents-ingest.md`, "Domkniecie dlugu z maili"). Extracts a **sample** of PDF attachments from the Gmail `.eml` archive (already indexed in the `envelope` table of kb-postgres) and drops them into Paperless' `consume/` directory so Paperless does the OCR and correspondent-detection. This job does **not** write to the `envelope` table — the Paperless/Nextcloud envelope adapter is a later phase of module 5. ## Why a sample, not a bulk import The Gmail import left ~70k attachments referenced in `envelope.entities` manifests (bytes live inside the archived `.eml` files, never extracted). Dumping all of them into Paperless at once would swamp the OCR worker and the RAG layer isn't built yet to make use of that volume. This job pulls a small, recent, size-filtered sample (default: 150 envelopes, PDFs >50KB, from the last year) as a testbed — mass import is a deliberate later decision. ## Where it runs **Locally on PIHA**, as a plain CLI (not a container). It needs simultaneous filesystem access to three things that all live on PIHA: - the mail archive (`/home/oskar/kb/mail/archive`) - the Paperless `consume/` directory (`/opt/homelab/data/paperless/consume`) - kb-postgres (`localhost:5433` from PIHA; reachable from elsewhere over Tailscale, but the archive and consume dir are not — those are local paths) Install (from repo root, on PIHA): ```bash pip install -e jobs/documents-ingest/ ``` (Or reuse the venv already set up for `gmail-bulk-import`, e.g. `/home/oskar/kb/venv/` — it already has `asyncpg` + `structlog`.) ## Usage ```bash # Dry run (default) — preview only, no writes: documents-ingest --dsn postgresql://kb:@localhost:5433/kb # Or via env var instead of --dsn: export KB_DSN=postgresql://kb:@localhost:5433/kb documents-ingest # Real run — write files into consume/ and update the registry: documents-ingest --apply # Smaller/larger sample, different window/threshold: documents-ingest --limit 50 --since-days 180 --min-size 100000 ``` Dry-run is the default and does not require `--consume-dir` to exist yet; `--apply` does (Paperless must already be deployed with its consume dir in place). See `documents-ingest --help` for all flags. ## Candidate selection ```sql SELECT id, raw_ref, ts, entities FROM envelope WHERE source = 'gmail' AND ts > now() - interval '1 year' AND EXISTS ( SELECT 1 FROM jsonb_array_elements(entities) AS att WHERE att->>'content_type' = 'application/pdf' AND (att->>'size')::numeric > 50000 ) ORDER BY ts DESC LIMIT 150 ``` For each matching envelope, every attachment manifest entry that passes the filter is a separate candidate (one envelope can yield several PDFs). ## Matching an attachment inside the .eml The manifest (`entities[]`) only has metadata — the attachment bytes live inside the `.eml` (MIME multipart), so each candidate is resolved against the freshly parsed message: 1. Parse the `.eml` with `email.policy.default` and collect every `application/pdf` MIME part (filename + decoded payload). 2. **sha256 is the proof of identity**, not the filename. The manifest was built by a different parser at import time (`gmail-bulk-import`, using `mailbox` + compat32 policy) and can still hold the raw RFC 2047 encoded-word form of a filename (e.g. `=?UTF-8?b?...?=`, sometimes with header-folding whitespace baked in), while `email.policy.default` decodes it to real Unicode today. Comparing those byte-for-byte skipped ~10% of otherwise-good attachments in testing — see `TestFindPdfParts` / `TestProcessCandidate` in the test suite for the regression case. So: match by sha256 across all PDF parts in the message; if none match, use a filename match only to tell "found the named part but its bytes changed" (`sha_mismatch`, reported and skipped) apart from "not present at all" (`parse_error`, skipped). 3. The consume/ filename is built from the *decoded* filename (from the MIME part), not the possibly-garbled manifest one. Mismatches and parse errors are never guessed past — they're logged and skipped. ## consume/ filenames `_.pdf`, date = envelope `ts`. On collision (same date + sanitized name already used in this run or already present in `consume/`), an 8-hex sha256 prefix is appended: `__.pdf`. Files are written with a best-effort `chown` to uid:gid `1000:1000` (the Paperless container's `USERMAP_UID/GID`, see `services/paperless/README.md`) so Paperless can read them. If the chown fails (e.g. the job isn't running as root/uid 1000), a warning is logged but the run continues — the write itself already succeeded; fix ownership/perms on `consume/` separately if needed. PIHA's uid/gid convention across the fleet is tracked as its own tech-debt item (see `docs/backlog/`), not solved here. ## Idempotency — registry A JSON file at `/opt/homelab/data/documents-ingest/registry.json` (default, override with `--registry`), keyed by attachment sha256: ```json { "": { "envelope_id": "...", "filename": "...", "consume_name": "2026-06-09_invoice.pdf", "size": 123456, "ingested_at": "2026-07-13T19:35:16+00:00" } } ``` **Why a JSON file and not a kb-postgres table:** this is a one-shot sampling tool for a bootstrapping phase, not a long-running service — a new table would formalize infrastructure for something temporary. A flat file needs no migration, is trivial to inspect (`jq`) or reset, and sits under `/opt/homelab/data/` alongside other node-local state per the repo's runtime path convention. If/when module 5's real Paperless/Nextcloud adapter phase starts writing `envelope` rows for `source=paperless`, that's the natural point to fold this into a proper DB-backed ingest log — re-litigate then, not now. Re-running the job only ever *adds* to the registry (on `--apply`); it's never consulted or mutated in dry-run mode beyond being read for the preview. ## Dry-run output Logs one line per skip (`skip.duplicate` / `skip.sha_mismatch` / `skip.parse_error`, with reason), a `summary` line with full counts (`envelopes_scanned`, `pdf_candidates`, `extracted`, `skipped_duplicate`, `skipped_sha_mismatch`, `skipped_parse_error`, `errors`), and up to 20 example `(target_name, size, envelope_id)` rows so you can sanity-check filenames before running `--apply`. ## Verifying the result in Paperless After `--apply`: 1. Paperless' consumer picks files up from `consume/` automatically (polling or inotify, per its own config) — no action needed on this job's side. 2. Watch progress: Paperless UI → Documents (new items appear as OCR finishes), or `docker logs -f paperless` on PIHA for consumer/OCR activity. 3. Cross-check count: number of new documents in Paperless should equal `stats["extracted"]` from the `--apply` run's summary line. 4. Confirm idempotency: re-running `--apply` immediately after should report `extracted: 0` and `skipped_duplicate` equal to the previous run's `extracted` count — nothing new lands in `consume/`. ## Tests ```bash pip install -e jobs/documents-ingest/ cd jobs/documents-ingest && pytest ``` Pure unit tests, no DB or filesystem outside `tmp_path` required — `run()` is tested by monkeypatching `asyncpg.connect` with an in-memory fake connection. Covers: filename sanitization, consume-name collision handling, manifest filtering, MIME PDF-part extraction (including the RFC 2047 decoding mismatch), sha256 match/mismatch, duplicate detection, dry-run vs `--apply` behavior, and multi-attachment envelopes. --- ## Phase 2 — `documents-ingest-paperless` (Paperless -> envelope adapter) Module 5, phase 2 (`docs/kb/modules/05-faza2-plan.md`, §4.2-4.3, §6 step 5). Reads documents from the **Paperless REST API** (read-only — GET only, never writes to Paperless) and inserts them as `source='paperless'` rows into the `envelope` table on kb-postgres, reusing `kb_mail.envelope.Envelope` / `kb_mail.db.insert_envelope` from `packages/kb-mail` (untouched by this change — see plan §1.6). Existing `source='gmail'` rows and `document_chunk` are never touched; this job only ever `INSERT`s new `paperless` rows. ### Cross-source link (`source_mail`) Per plan §1.9/§4.2, the deterministic join uses no heuristics: a document's `original_file_name` (from the Paperless API) is matched against `consume_name` in this job's **phase-1 registry** (`/opt/homelab/data/documents-ingest/registry.json`, produced by `extractor.py` — see above). A match appends a `source_mail` entity pointing back at the originating mail envelope; no match means the document was added outside the faktury-1 pipeline, and the entity is simply omitted — not an error. ### Install ```bash pip install -e packages/kb-mail/ pip install -e jobs/documents-ingest/ ``` ### Usage ```bash # Dry run (default) — fetch from Paperless, map, count; no DB writes: documents-ingest-paperless --dsn postgresql://kb:@localhost:5433/kb \ --paperless-token # Real run — insert new envelope rows: documents-ingest-paperless --dsn ... --paperless-token ... --apply # Smoke-test slice: documents-ingest-paperless --dsn ... --paperless-token ... --limit 5 ``` `--dsn` can come from `KB_DSN`, `--paperless-token` from `PAPERLESS_API_TOKEN`, `--paperless-url` from `PAPERLESS_URL` (defaults to Paperless' fixed LAN address, `http://192.168.31.5:8210`). No `--offset`: unlike the 225 030-row header backfill, a full re-scan of Paperless' ~186 documents is cheap and already idempotent, so there is no need for resumable partitioning — `--limit` exists only to cap a run for smoke-testing. ### Mapping (plan §4.3) ``` id = f"paperless:{document_id}" -- prefixed: Paperless doc-ids are small -- sequential ints that would otherwise -- collide with any future source's ids ts = documents_document.created -- Paperless-detected date (content/filename), -- not filesystem mtime geo = NULL raw_ref = str(document_id) -- REFERENCE — Paperless is the source of truth, -- no bytes are copied entities = content, correspondent, tag(s), filename, content_type, and source_mail when the registry join hits (plan §4.2) ``` `correspondent`/`tag` are resolved from Paperless' `/api/correspondents/` and `/api/tags/` (fetched once, cached in memory for the run) and kept purely as informational metadata — nothing in this pipeline depends on them being non-null (plan decision 4). A document with empty OCR content (Paperless OCR sometimes produces none) still gets a normal envelope with `"text": ""` — not skipped, not an error, just counted (`empty_content`). ### Idempotency A pre-fetched set of existing `source='paperless'` envelope ids (one query at the start of each run) skips documents already inserted; `insert_envelope`'s own `ON CONFLICT (id) DO NOTHING` is the second line of defense. Re-running `--apply` immediately after a successful run reports `inserted: 0` and `already_in_db` equal to the previous run's `inserted` count. ### Stats must balance ``` fetched = already_in_db + inserted + errors ``` `source_mail_linked` and `empty_content` are informational subsets of `fetched`, not separate outcome buckets. A per-document mapping failure (e.g. an unparseable `created` date) is isolated, logged, and counted as `errors` — it never aborts the run. `main()` exits 1 on non-zero `errors` or if the balance invariant above doesn't hold (mirrors `gmail-bulk-import`'s exit-code convention) — a clean run always exits 0. ### Tests ```bash pip install -e packages/kb-mail/ pip install -e jobs/documents-ingest/ cd jobs/documents-ingest && pytest ``` Pure unit tests, no DB or real HTTP — `run()` is tested by monkeypatching `asyncpg.connect` (fake connection) and `aiohttp.ClientSession` (fake session serving canned JSON pages). Covers: mapping shape (content, correspondent, tag(s), filename, content_type, source_mail), the registry join (hit and miss), pagination (both the documents list and the correspondents/tags lookup tables), `--limit`, idempotency (pre-existing ids skipped, a second `--apply` run inserts nothing new), isolated per-document mapping errors, and the stats-balance invariant. ### Definition of Done Per `CLAUDE.md`: smoke run is `documents-ingest-paperless --dsn ... --paperless-token ... --limit 5` (dry-run first) against kb-postgres@PIHA and the live Paperless API, over SSH — **not executed as part of this change** without operator confirmation (this job reads production Paperless data and writes production envelope rows on `--apply`). `pytest` passes locally before this commit. --- ## Phase 2 step 6 — `documents-ingest-embed` (chunk + embed) Module 5, phase 2, plan step 6 (`docs/kb/modules/05-faza2-plan.md`, §6 step 6, §2 decision 3). Reads `entities[type=content].text` off every `source='paperless'` envelope, chunks it, calls Ollama (`POST /api/embeddings`, model `bge-m3`) for each chunk, and inserts the result into `document_chunk` (`services/kb-postgres/init/002_chunks.sql`). This job only ever `INSERT`s into `document_chunk` — `envelope` is read-only here, and `services/ollama/` is untouched. ### Where it runs **On SOLARIA** (that's where Ollama lives), against kb-postgres@PIHA over Tailscale — the reverse of the other jobs in this package, which run on PIHA. `--ollama-url` defaults to `http://localhost:11434` (Ollama on the same node); `--dsn` needs PIHA's Tailscale address, e.g. `postgresql://kb:@piha:5433/kb`. ### Install ```bash pip install -e packages/kb-mail/ pip install -e jobs/documents-ingest/ ``` ### Usage ```bash # Dry run (default) — chunk and count only, no Ollama calls, no DB writes: documents-ingest-embed --dsn postgresql://kb:@piha:5433/kb # Smoke-test slice: documents-ingest-embed --dsn ... --apply --limit 10 # Full run: documents-ingest-embed --dsn ... --apply ``` ### Chunking (plan §2 decision 3) Paragraph-preferring: splits on blank-line boundaries, greedily packs paragraphs up to `--chunk-size` characters (default 2400, ≈600 tokens at a ~4 chars/token heuristic — no local bge-m3 tokenizer available offline), `--chunk-overlap` characters of trailing context carried into the next chunk (default 600, ≈150 tokens). A paragraph that alone exceeds `--chunk-size` falls back to a hard character-based sliding window — Paperless OCR text has no page-break markers (plan §1.2), so there's nothing else to split large, unbroken text on. A document with empty OCR content (the 26 `empty_content` documents from phase 2 step 5) yields zero chunks and is counted separately, not as an error. ### Idempotency A pre-fetched set of `(envelope_id, chunk_index)` pairs already embedded with `--model` skips re-embedding on rerun — no wasted Ollama calls. `document_chunk`'s own `UNIQUE (envelope_id, chunk_index)` + `ON CONFLICT DO NOTHING` is the second line of defense; `insert_chunk`'s command tag is checked so a silently-skipped row is counted as `chunks_conflict_skipped`, never miscounted as `chunks_inserted`. Note that uniqueness is on `(envelope_id, chunk_index)` only, not `model` — re-embedding with a *different* model hits this path and that embedding is discarded (wasted work, correctly reported via `chunks_conflict_skipped`, but not persisted). Out of scope for this single-model pilot; the real fix for whoever indexes a second model later is `UNIQUE (envelope_id, chunk_index, model)` at the schema layer. A DB write failure for one chunk (dropped connection, unexpected bytes) is isolated the same way an embed failure is — counted as `chunks_errors`, never aborting the rest of the run. ### Dimension guard Every embedding response's length is checked against `document_chunk.embedding`'s `VECTOR(1024)` column. A mismatch raises `EmbeddingDimensionError` and aborts the whole run immediately — never silently indexes vectors of the wrong dimension. ### Chunk size/overlap validation `--chunk-overlap` must be smaller than `--chunk-size` — the sliding-window hard-split fallback advances by `chunk_size - chunk_overlap` per step, so an overlap `>=` size would never advance and hang. `main()` rejects this combination before opening a DB connection; `hard_split()` itself also raises `ValueError` as a second line of defense for direct callers. ### Stats must balance ``` documents_fetched = empty_content + documents_chunked chunks_total = chunks_already_embedded + chunks_inserted + chunks_conflict_skipped + chunks_errors ``` `main()` exits 1 on `chunks_errors > 0`, `chunks_conflict_skipped > 0`, or if either balance breaks. The summary line also reports `avg_embed_seconds_per_chunk` — CPU-only Ollama timing, the input for deciding whether/how to scale this to the mail corpus later (plan §7). ### Tests ```bash pip install -e packages/kb-mail/ pip install -e jobs/documents-ingest/ cd jobs/documents-ingest && pytest ``` Pure unit tests, no DB or real HTTP — `run()` is tested by monkeypatching `asyncpg.connect` (fake connection) and `aiohttp.ClientSession` (fake session serving a canned embedding vector, or a 500 for a chosen prompt to exercise error isolation). Covers: chunking (paragraph boundaries, overlap, empty document, document shorter than one chunk, oversized paragraph hard-fallback, the overlap-must-be-smaller-than-size guard), `extract_content`, idempotency (pre-existing keys skipped, no Ollama calls made for them, a second `--apply` run embeds nothing new, existing keys are correctly scoped to `--model`), dimension-mismatch abort, isolated per-chunk embed and insert errors, `ON CONFLICT` no-ops counted separately from real inserts, and the stats-balance invariant. ### Known limitation — Ollama context-length rejections on pathological chunks Ollama's *runtime* context window for a model can be smaller than the model's advertised max (bge-m3 supports 8192 tokens, but Ollama's default `num_ctx` is lower) — and some OCR text tokenizes far more densely than the ~4-chars/token heuristic this job uses to size chunks. Concretely: a table- of-contents page made almost entirely of dot-leader formatting (`". . . . . . ."`, repeated hundreds of times) hit this on the pilot run — Ollama returned `500 {"error":"the input length exceeds the context length"}` for one 2400-char chunk that should have been well within budget by character count alone. The job isolates this exactly like any other embed failure (`chunks_errors`, logged, run continues), so it never crashes a run — but it also never automatically shrinks and retries the offending chunk. Given how rare this was (1 chunk out of 2684 in the full pilot, all from one document's dot-leader ToC), it's left as a known gap rather than fixed here; a real fix would be either a smaller/adaptive chunk size for low-character-entropy text, or a shrink-and-retry loop on this specific Ollama error. ### Definition of Done Per `CLAUDE.md`: `pytest` passes locally (101 tests). Smoke-tested and then run to completion live on SOLARIA against the real Ollama instance and kb-postgres@PIHA: - Dry-run: 186 fetched, 26 `empty_content`, 2684 chunks planned — matches the known phase-2-step-5 figures exactly. - `--apply --limit 10`: 64 chunks embedded, 0 errors, avg ≈0.83s/chunk on CPU. - Re-run of the same slice: fully idempotent — 0 Ollama calls, 0 inserts. - Full `--apply` (all 186 documents): **2683/2684 chunks inserted, 1 isolated error** (see "Known limitation" above) — `chunks_errors=1` correctly produced a non-zero exit rather than silently reporting success. `document_chunk` ends at 2683 rows across 160 distinct envelopes, matching `documents_chunked`. A `document_chunk_envelope_idx`-backed count and an `ORDER BY embedding <=> ...` nearest-neighbor sanity query both look correct (top match is the reference chunk itself at distance 0; next nearest are chunks of the same source document). - **Timing (CPU-only, no GPU driver on SOLARIA)**: ≈0.79s/chunk average across 2683 real embeddings (2115.8s total embed time), ≈13.2s/document average across the 160 chunked documents, ≈35 minutes wall-clock for the full 186-document pilot. This is the real-world input for scaling this pipeline to the much larger mail corpus later (plan §7 assumed GPU-based "minutes for the whole pilot"; SOLARIA's Ollama ran CPU-only for this pilot per the then-disabled GPU reservation). The 186-document pilot's ≈13.2s/document average is dominated by Paperless' long OCR text (≈22k chars/doc average, per plan §1.2) — 225 030 mail envelopes will have a very different, likely much shorter, per-envelope chunk count (email bodies vs. scanned multi-page PDFs), so this number doesn't extrapolate directly to a mail-corpus estimate. What it does establish: at ≈0.79s/chunk sequential CPU embedding, any corpus with a non-trivial average chunk count per item will need either a GPU driver fix, concurrent/batched Ollama calls, or both, before a full mail-corpus run is practical — flagged for whoever picks up the mail-indexer phase. - **GPU (RTX 4070 Ti SUPER, driver 595-open, restored 2026-07-16)**: TBD po pomiarze 2026-07-16 — sterownik NVIDII naprawiony i GPU reservation przywrócona w `services/ollama/docker-compose.yml` (`docs/backlog.md`, "Ollama SOLARIA: brak sterownika NVIDII"), ale sekwencyjny embedding benchmark vs. ≈0.79s/chunk CPU baseline jeszcze nie zmierzony w tym repo — liczba do uzupełnienia osobnym commitem po pomiarze na żywym Ollamie.