feat(kb): SPLIT documents-ingest (30 KB) -> service + phase + runbook
kb/services/job-documents-ingest.md — opis jobu i mechanizmow (candidate selection, matching, consume/, idempotency, dry-run) kb/phases/kb-m5-documents-ingest-fazy.md — faza 2, faza 2 krok 6, faza 3 krok 4, faza 3 krok 5 (4 sekcje fazowe wtopione w README) kb/runbooks/documents-ingest-run.md — Usage, Verifying in Paperless, Tests Najwiekszy README w repo. Tresc sekcji nietknieta; kontrola multizbioru linii == oryginal. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
5439c00def
commit
ccb738cddf
|
|
@ -1,188 +1,15 @@
|
|||
# documents-ingest
|
||||
|
||||
One-shot job CLI, Phase 1 of module 5 (`docs/kb/modules/05-documents-ingest.md`,
|
||||
"Domkniecie dlugu z maili"). Extracts a **sample** of PDF attachments from the
|
||||
Gmail `.eml` archive (already indexed in the `envelope` table of kb-postgres)
|
||||
and drops them into Paperless' `consume/` directory so Paperless does the OCR
|
||||
and correspondent-detection. This job does **not** write to the `envelope`
|
||||
table — the Paperless/Nextcloud envelope adapter is a later phase of module 5.
|
||||
|
||||
## Why a sample, not a bulk import
|
||||
|
||||
The Gmail import left ~70k attachments referenced in `envelope.entities`
|
||||
manifests (bytes live inside the archived `.eml` files, never extracted).
|
||||
Dumping all of them into Paperless at once would swamp the OCR worker and the
|
||||
RAG layer isn't built yet to make use of that volume. This job pulls a small,
|
||||
recent, size-filtered sample (default: 150 envelopes, PDFs >50KB, from the
|
||||
last year) as a testbed — mass import is a deliberate later decision.
|
||||
|
||||
## Where it runs
|
||||
|
||||
**Locally on PIHA**, as a plain CLI (not a container). It needs simultaneous
|
||||
filesystem access to three things that all live on PIHA:
|
||||
|
||||
- the mail archive (`/home/oskar/kb/mail/archive`)
|
||||
- the Paperless `consume/` directory (`/opt/homelab/data/paperless/consume`)
|
||||
- kb-postgres (`localhost:5433` from PIHA; reachable from elsewhere over
|
||||
Tailscale, but the archive and consume dir are not — those are local paths)
|
||||
|
||||
Install (from repo root, on PIHA):
|
||||
|
||||
```bash
|
||||
pip install -e jobs/documents-ingest/
|
||||
```
|
||||
|
||||
(Or reuse the venv already set up for `gmail-bulk-import`, e.g.
|
||||
`/home/oskar/kb/venv/` — it already has `asyncpg` + `structlog`.)
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
# Dry run (default) — preview only, no writes:
|
||||
documents-ingest --dsn postgresql://kb:<pw>@localhost:5433/kb
|
||||
|
||||
# Or via env var instead of --dsn:
|
||||
export KB_DSN=postgresql://kb:<pw>@localhost:5433/kb
|
||||
documents-ingest
|
||||
|
||||
# Real run — write files into consume/ and update the registry:
|
||||
documents-ingest --apply
|
||||
|
||||
# Smaller/larger sample, different window/threshold:
|
||||
documents-ingest --limit 50 --since-days 180 --min-size 100000
|
||||
```
|
||||
|
||||
Dry-run is the default and does not require `--consume-dir` to exist yet;
|
||||
`--apply` does (Paperless must already be deployed with its consume dir in
|
||||
place). See `documents-ingest --help` for all flags.
|
||||
|
||||
## Candidate selection
|
||||
|
||||
```sql
|
||||
SELECT id, raw_ref, ts, entities FROM envelope
|
||||
WHERE source = 'gmail'
|
||||
AND ts > now() - interval '1 year'
|
||||
AND EXISTS (
|
||||
SELECT 1 FROM jsonb_array_elements(entities) AS att
|
||||
WHERE att->>'content_type' = 'application/pdf'
|
||||
AND (att->>'size')::numeric > 50000
|
||||
)
|
||||
ORDER BY ts DESC
|
||||
LIMIT 150
|
||||
```
|
||||
|
||||
For each matching envelope, every attachment manifest entry that passes the
|
||||
filter is a separate candidate (one envelope can yield several PDFs).
|
||||
|
||||
## Matching an attachment inside the .eml
|
||||
|
||||
The manifest (`entities[]`) only has metadata — the attachment bytes live
|
||||
inside the `.eml` (MIME multipart), so each candidate is resolved against the
|
||||
freshly parsed message:
|
||||
|
||||
1. Parse the `.eml` with `email.policy.default` and collect every
|
||||
`application/pdf` MIME part (filename + decoded payload).
|
||||
2. **sha256 is the proof of identity**, not the filename. The manifest was
|
||||
built by a different parser at import time (`gmail-bulk-import`, using
|
||||
`mailbox` + compat32 policy) and can still hold the raw RFC 2047
|
||||
encoded-word form of a filename (e.g. `=?UTF-8?b?...?=`, sometimes with
|
||||
header-folding whitespace baked in), while `email.policy.default` decodes
|
||||
it to real Unicode today. Comparing those byte-for-byte skipped ~10% of
|
||||
otherwise-good attachments in testing — see `TestFindPdfParts` /
|
||||
`TestProcessCandidate` in the test suite for the regression case. So:
|
||||
match by sha256 across all PDF parts in the message; if none match, use a
|
||||
filename match only to tell "found the named part but its bytes changed"
|
||||
(`sha_mismatch`, reported and skipped) apart from "not present at all"
|
||||
(`parse_error`, skipped).
|
||||
3. The consume/ filename is built from the *decoded* filename (from the MIME
|
||||
part), not the possibly-garbled manifest one.
|
||||
|
||||
Mismatches and parse errors are never guessed past — they're logged and
|
||||
skipped.
|
||||
|
||||
## consume/ filenames
|
||||
|
||||
`<YYYY-MM-DD>_<sanitized-filename>.pdf`, date = envelope `ts`. On collision
|
||||
(same date + sanitized name already used in this run or already present in
|
||||
`consume/`), an 8-hex sha256 prefix is appended:
|
||||
`<YYYY-MM-DD>_<sanitized-filename>_<hash8>.pdf`.
|
||||
|
||||
Files are written with a best-effort `chown` to uid:gid `1000:1000` (the
|
||||
Paperless container's `USERMAP_UID/GID`, see `services/paperless/README.md`)
|
||||
so Paperless can read them. If the chown fails (e.g. the job isn't running as
|
||||
root/uid 1000), a warning is logged but the run continues — the write itself
|
||||
already succeeded; fix ownership/perms on `consume/` separately if needed.
|
||||
PIHA's uid/gid convention across the fleet is tracked as its own tech-debt
|
||||
item (see `docs/backlog/`), not solved here.
|
||||
|
||||
## Idempotency — registry
|
||||
|
||||
A JSON file at `/opt/homelab/data/documents-ingest/registry.json` (default,
|
||||
override with `--registry`), keyed by attachment sha256:
|
||||
|
||||
```json
|
||||
{
|
||||
"<sha256>": {
|
||||
"envelope_id": "...",
|
||||
"filename": "...",
|
||||
"consume_name": "2026-06-09_invoice.pdf",
|
||||
"size": 123456,
|
||||
"ingested_at": "2026-07-13T19:35:16+00:00"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Why a JSON file and not a kb-postgres table:** this is a one-shot sampling
|
||||
tool for a bootstrapping phase, not a long-running service — a new table
|
||||
would formalize infrastructure for something temporary. A flat file needs no
|
||||
migration, is trivial to inspect (`jq`) or reset, and sits under
|
||||
`/opt/homelab/data/` alongside other node-local state per the repo's runtime
|
||||
path convention. If/when module 5's real Paperless/Nextcloud adapter phase
|
||||
starts writing `envelope` rows for `source=paperless`, that's the natural
|
||||
point to fold this into a proper DB-backed ingest log — re-litigate then, not
|
||||
now.
|
||||
|
||||
Re-running the job only ever *adds* to the registry (on `--apply`); it's
|
||||
never consulted or mutated in dry-run mode beyond being read for the preview.
|
||||
|
||||
## Dry-run output
|
||||
|
||||
Logs one line per skip (`skip.duplicate` / `skip.sha_mismatch` /
|
||||
`skip.parse_error`, with reason), a `summary` line with full counts
|
||||
(`envelopes_scanned`, `pdf_candidates`, `extracted`, `skipped_duplicate`,
|
||||
`skipped_sha_mismatch`, `skipped_parse_error`, `errors`), and up to 20 example
|
||||
`(target_name, size, envelope_id)` rows so you can sanity-check filenames
|
||||
before running `--apply`.
|
||||
|
||||
## Verifying the result in Paperless
|
||||
|
||||
After `--apply`:
|
||||
|
||||
1. Paperless' consumer picks files up from `consume/` automatically (polling
|
||||
or inotify, per its own config) — no action needed on this job's side.
|
||||
2. Watch progress: Paperless UI → Documents (new items appear as OCR
|
||||
finishes), or `docker logs -f paperless` on PIHA for consumer/OCR activity.
|
||||
3. Cross-check count: number of new documents in Paperless should equal
|
||||
`stats["extracted"]` from the `--apply` run's summary line.
|
||||
4. Confirm idempotency: re-running `--apply` immediately after should report
|
||||
`extracted: 0` and `skipped_duplicate` equal to the previous run's
|
||||
`extracted` count — nothing new lands in `consume/`.
|
||||
|
||||
## Tests
|
||||
|
||||
```bash
|
||||
pip install -e jobs/documents-ingest/
|
||||
cd jobs/documents-ingest && pytest
|
||||
```
|
||||
|
||||
Pure unit tests, no DB or filesystem outside `tmp_path` required — `run()` is
|
||||
tested by monkeypatching `asyncpg.connect` with an in-memory fake connection.
|
||||
Covers: filename sanitization, consume-name collision handling, manifest
|
||||
filtering, MIME PDF-part extraction (including the RFC 2047 decoding
|
||||
mismatch), sha256 match/mismatch, duplicate detection, dry-run vs `--apply`
|
||||
behavior, and multi-attachment envelopes.
|
||||
|
||||
---
|
||||
okf: "0.1"
|
||||
type: phase
|
||||
visibility: private
|
||||
status: active
|
||||
updated: 2026-07-30
|
||||
links:
|
||||
- ../services/job-documents-ingest.md
|
||||
- ../runbooks/documents-ingest-run.md
|
||||
---
|
||||
|
||||
# documents-ingest — fazy 2 i 3
|
||||
|
||||
## Phase 2 — `documents-ingest-paperless` (Paperless -> envelope adapter)
|
||||
|
||||
64
kb/runbooks/documents-ingest-run.md
Normal file
64
kb/runbooks/documents-ingest-run.md
Normal file
|
|
@ -0,0 +1,64 @@
|
|||
---
|
||||
okf: "0.1"
|
||||
type: runbook
|
||||
visibility: private
|
||||
status: active
|
||||
updated: 2026-07-30
|
||||
links:
|
||||
- ../services/job-documents-ingest.md
|
||||
- ../phases/kb-m5-documents-ingest-fazy.md
|
||||
---
|
||||
|
||||
# documents-ingest — uruchomienie i weryfikacja
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
# Dry run (default) — preview only, no writes:
|
||||
documents-ingest --dsn postgresql://kb:<pw>@localhost:5433/kb
|
||||
|
||||
# Or via env var instead of --dsn:
|
||||
export KB_DSN=postgresql://kb:<pw>@localhost:5433/kb
|
||||
documents-ingest
|
||||
|
||||
# Real run — write files into consume/ and update the registry:
|
||||
documents-ingest --apply
|
||||
|
||||
# Smaller/larger sample, different window/threshold:
|
||||
documents-ingest --limit 50 --since-days 180 --min-size 100000
|
||||
```
|
||||
|
||||
Dry-run is the default and does not require `--consume-dir` to exist yet;
|
||||
`--apply` does (Paperless must already be deployed with its consume dir in
|
||||
place). See `documents-ingest --help` for all flags.
|
||||
|
||||
## Verifying the result in Paperless
|
||||
|
||||
After `--apply`:
|
||||
|
||||
1. Paperless' consumer picks files up from `consume/` automatically (polling
|
||||
or inotify, per its own config) — no action needed on this job's side.
|
||||
2. Watch progress: Paperless UI → Documents (new items appear as OCR
|
||||
finishes), or `docker logs -f paperless` on PIHA for consumer/OCR activity.
|
||||
3. Cross-check count: number of new documents in Paperless should equal
|
||||
`stats["extracted"]` from the `--apply` run's summary line.
|
||||
4. Confirm idempotency: re-running `--apply` immediately after should report
|
||||
`extracted: 0` and `skipped_duplicate` equal to the previous run's
|
||||
`extracted` count — nothing new lands in `consume/`.
|
||||
|
||||
## Tests
|
||||
|
||||
```bash
|
||||
pip install -e jobs/documents-ingest/
|
||||
cd jobs/documents-ingest && pytest
|
||||
```
|
||||
|
||||
Pure unit tests, no DB or filesystem outside `tmp_path` required — `run()` is
|
||||
tested by monkeypatching `asyncpg.connect` with an in-memory fake connection.
|
||||
Covers: filename sanitization, consume-name collision handling, manifest
|
||||
filtering, MIME PDF-part extraction (including the RFC 2047 decoding
|
||||
mismatch), sha256 match/mismatch, duplicate detection, dry-run vs `--apply`
|
||||
behavior, and multi-attachment envelopes.
|
||||
|
||||
---
|
||||
|
||||
146
kb/services/job-documents-ingest.md
Normal file
146
kb/services/job-documents-ingest.md
Normal file
|
|
@ -0,0 +1,146 @@
|
|||
---
|
||||
okf: "0.1"
|
||||
type: service
|
||||
visibility: private
|
||||
status: active
|
||||
updated: 2026-07-30
|
||||
links:
|
||||
- ../phases/kb-m5-documents-ingest-fazy.md
|
||||
- ../runbooks/documents-ingest-run.md
|
||||
---
|
||||
|
||||
# documents-ingest
|
||||
|
||||
One-shot job CLI, Phase 1 of module 5 (`docs/kb/modules/05-documents-ingest.md`,
|
||||
"Domkniecie dlugu z maili"). Extracts a **sample** of PDF attachments from the
|
||||
Gmail `.eml` archive (already indexed in the `envelope` table of kb-postgres)
|
||||
and drops them into Paperless' `consume/` directory so Paperless does the OCR
|
||||
and correspondent-detection. This job does **not** write to the `envelope`
|
||||
table — the Paperless/Nextcloud envelope adapter is a later phase of module 5.
|
||||
|
||||
## Why a sample, not a bulk import
|
||||
|
||||
The Gmail import left ~70k attachments referenced in `envelope.entities`
|
||||
manifests (bytes live inside the archived `.eml` files, never extracted).
|
||||
Dumping all of them into Paperless at once would swamp the OCR worker and the
|
||||
RAG layer isn't built yet to make use of that volume. This job pulls a small,
|
||||
recent, size-filtered sample (default: 150 envelopes, PDFs >50KB, from the
|
||||
last year) as a testbed — mass import is a deliberate later decision.
|
||||
|
||||
## Where it runs
|
||||
|
||||
**Locally on PIHA**, as a plain CLI (not a container). It needs simultaneous
|
||||
filesystem access to three things that all live on PIHA:
|
||||
|
||||
- the mail archive (`/home/oskar/kb/mail/archive`)
|
||||
- the Paperless `consume/` directory (`/opt/homelab/data/paperless/consume`)
|
||||
- kb-postgres (`localhost:5433` from PIHA; reachable from elsewhere over
|
||||
Tailscale, but the archive and consume dir are not — those are local paths)
|
||||
|
||||
Install (from repo root, on PIHA):
|
||||
|
||||
```bash
|
||||
pip install -e jobs/documents-ingest/
|
||||
```
|
||||
|
||||
(Or reuse the venv already set up for `gmail-bulk-import`, e.g.
|
||||
`/home/oskar/kb/venv/` — it already has `asyncpg` + `structlog`.)
|
||||
|
||||
## Candidate selection
|
||||
|
||||
```sql
|
||||
SELECT id, raw_ref, ts, entities FROM envelope
|
||||
WHERE source = 'gmail'
|
||||
AND ts > now() - interval '1 year'
|
||||
AND EXISTS (
|
||||
SELECT 1 FROM jsonb_array_elements(entities) AS att
|
||||
WHERE att->>'content_type' = 'application/pdf'
|
||||
AND (att->>'size')::numeric > 50000
|
||||
)
|
||||
ORDER BY ts DESC
|
||||
LIMIT 150
|
||||
```
|
||||
|
||||
For each matching envelope, every attachment manifest entry that passes the
|
||||
filter is a separate candidate (one envelope can yield several PDFs).
|
||||
|
||||
## Matching an attachment inside the .eml
|
||||
|
||||
The manifest (`entities[]`) only has metadata — the attachment bytes live
|
||||
inside the `.eml` (MIME multipart), so each candidate is resolved against the
|
||||
freshly parsed message:
|
||||
|
||||
1. Parse the `.eml` with `email.policy.default` and collect every
|
||||
`application/pdf` MIME part (filename + decoded payload).
|
||||
2. **sha256 is the proof of identity**, not the filename. The manifest was
|
||||
built by a different parser at import time (`gmail-bulk-import`, using
|
||||
`mailbox` + compat32 policy) and can still hold the raw RFC 2047
|
||||
encoded-word form of a filename (e.g. `=?UTF-8?b?...?=`, sometimes with
|
||||
header-folding whitespace baked in), while `email.policy.default` decodes
|
||||
it to real Unicode today. Comparing those byte-for-byte skipped ~10% of
|
||||
otherwise-good attachments in testing — see `TestFindPdfParts` /
|
||||
`TestProcessCandidate` in the test suite for the regression case. So:
|
||||
match by sha256 across all PDF parts in the message; if none match, use a
|
||||
filename match only to tell "found the named part but its bytes changed"
|
||||
(`sha_mismatch`, reported and skipped) apart from "not present at all"
|
||||
(`parse_error`, skipped).
|
||||
3. The consume/ filename is built from the *decoded* filename (from the MIME
|
||||
part), not the possibly-garbled manifest one.
|
||||
|
||||
Mismatches and parse errors are never guessed past — they're logged and
|
||||
skipped.
|
||||
|
||||
## consume/ filenames
|
||||
|
||||
`<YYYY-MM-DD>_<sanitized-filename>.pdf`, date = envelope `ts`. On collision
|
||||
(same date + sanitized name already used in this run or already present in
|
||||
`consume/`), an 8-hex sha256 prefix is appended:
|
||||
`<YYYY-MM-DD>_<sanitized-filename>_<hash8>.pdf`.
|
||||
|
||||
Files are written with a best-effort `chown` to uid:gid `1000:1000` (the
|
||||
Paperless container's `USERMAP_UID/GID`, see `services/paperless/README.md`)
|
||||
so Paperless can read them. If the chown fails (e.g. the job isn't running as
|
||||
root/uid 1000), a warning is logged but the run continues — the write itself
|
||||
already succeeded; fix ownership/perms on `consume/` separately if needed.
|
||||
PIHA's uid/gid convention across the fleet is tracked as its own tech-debt
|
||||
item (see `docs/backlog/`), not solved here.
|
||||
|
||||
## Idempotency — registry
|
||||
|
||||
A JSON file at `/opt/homelab/data/documents-ingest/registry.json` (default,
|
||||
override with `--registry`), keyed by attachment sha256:
|
||||
|
||||
```json
|
||||
{
|
||||
"<sha256>": {
|
||||
"envelope_id": "...",
|
||||
"filename": "...",
|
||||
"consume_name": "2026-06-09_invoice.pdf",
|
||||
"size": 123456,
|
||||
"ingested_at": "2026-07-13T19:35:16+00:00"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Why a JSON file and not a kb-postgres table:** this is a one-shot sampling
|
||||
tool for a bootstrapping phase, not a long-running service — a new table
|
||||
would formalize infrastructure for something temporary. A flat file needs no
|
||||
migration, is trivial to inspect (`jq`) or reset, and sits under
|
||||
`/opt/homelab/data/` alongside other node-local state per the repo's runtime
|
||||
path convention. If/when module 5's real Paperless/Nextcloud adapter phase
|
||||
starts writing `envelope` rows for `source=paperless`, that's the natural
|
||||
point to fold this into a proper DB-backed ingest log — re-litigate then, not
|
||||
now.
|
||||
|
||||
Re-running the job only ever *adds* to the registry (on `--apply`); it's
|
||||
never consulted or mutated in dry-run mode beyond being read for the preview.
|
||||
|
||||
## Dry-run output
|
||||
|
||||
Logs one line per skip (`skip.duplicate` / `skip.sha_mismatch` /
|
||||
`skip.parse_error`, with reason), a `summary` line with full counts
|
||||
(`envelopes_scanned`, `pdf_candidates`, `extracted`, `skipped_duplicate`,
|
||||
`skipped_sha_mismatch`, `skipped_parse_error`, `errors`), and up to 20 example
|
||||
`(target_name, size, envelope_id)` rows so you can sanity-check filenames
|
||||
before running `--apply`.
|
||||
|
||||
Loading…
Reference in a new issue