homelab-codex-ws/jobs/gmail-header-backfill/pyproject.toml

24 lines
451 B
TOML
Raw Normal View History

feat(gmail-header-backfill): one-shot job — backfill headers into envelope.entities Module 5 phase 2 (docs/kb/modules/05-faza2-plan.md, §5). 225,030 existing source='gmail' envelope rows carry only an attachment manifest — no from/to/cc/delivered_to/subject anywhere in the DB, which blocks answering "is this on me or my wife" / distinguishing aliases (plan §1.8). Separate job from gmail-bulk-import per plan §2 decision 7: UPDATE semantics against production data is a different risk profile than the historical INSERT job. Parses headers only (no MIME-walk of attachments) from the archived .eml files, appends {"type": "headers", ...} (plan §4.1) via the idempotent UPDATE ... WHERE NOT EXISTS from §5.2, batched via executemany. --limit/ --offset (ORDER BY id) give stable, deterministic partitioning so the full backfill can run and be verified in slices instead of one unattended pass. Plain CLI (pip install -e), no Dockerfile — same convention as gmail-bulk-import/documents-ingest, which run directly on PIHA for local filesystem access to the .eml archive. Verified against kb-postgres@PIHA (100-row dry-run + apply, re-run proved idempotent no-op, 1000-row timed slice): 1096/225030 rows backfilled, 996/1000 succeeded on the timed slice (4 parse_errors — a 2014 spam message with an RFC 2047 encoded-word decoding to an embedded newline in the From display name, correctly caught and skipped rather than crashing the batch). Extrapolated full-run time ~16 minutes. Full 225,030-row run is out of scope for this change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-14 15:34:07 +02:00
[build-system]
requires = ["setuptools>=68"]
build-backend = "setuptools.build_meta"
[project]
name = "gmail-header-backfill"
version = "0.1.0"
requires-python = ">=3.11"
dependencies = [
"asyncpg>=0.29",
"structlog>=24.1",
fix(gmail-bulk-import): harden re-import against 8-bit headers + poison batches Four audit findings (2026-07-14), each reproduced on crafted mboxes; the Takeout corpus has proven 8-bit header bytes, so all are real re-import risks. 1. Whole-run crash on 8-bit Message-ID. compat32 .get() returns an email.header.Header (not str) for raw 8-bit bytes; the old Header.strip() raised AttributeError. _message_id ran BEFORE the per-message try, so one bad header killed the entire import. Fix: str() + sanitize_surrogates() before strip; and move _message_id/_parse_date/_parse_attachments INSIDE the per-message try — a broken message is now errors += 1, never run death. 2. Poison batch. pending.clear() ran only AFTER a successful insert, so a failed flush (DB down / bad row) left pending intact and every later message re-flushed the doomed batch; the final flush sat in try/finally with no except and propagated out, losing all stats. Fix: _flush always clears pending and counts a failed insert as db_insert_failed; the run always reaches import_complete. 3. Stats didn't reconcile with the DB. imported counts archive writes, not DB rows, so a partial-insert drift was invisible. Fix: separate db_inserted/db_insert_failed counters; main() exits non-zero on any error, DB drift, or a processed = imported + skipped + errors imbalance. 4. 8-bit Date → needless epoch_fallback. parsedate_to_datetime(Header) raised even when str(header) parses fine. Fix: str() before the epoch fallback. Shared helper: _sanitize moved from gmail-header-backfill into packages/kb-mail (kb_mail.text.sanitize_surrogates) and used by both jobs; gmail-header-backfill now depends on kb-mail. Tests: regression coverage for all four findings in gmail-bulk-import (8-bit id/date, per-message guard, failed-insert non-poisoning, stats balance) plus kb_mail.text unit tests. Full suites green: kb-mail 27, gmail-bulk-import 33, gmail-header-backfill 43. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 19:51:35 +02:00
"kb-mail",
feat(gmail-header-backfill): one-shot job — backfill headers into envelope.entities Module 5 phase 2 (docs/kb/modules/05-faza2-plan.md, §5). 225,030 existing source='gmail' envelope rows carry only an attachment manifest — no from/to/cc/delivered_to/subject anywhere in the DB, which blocks answering "is this on me or my wife" / distinguishing aliases (plan §1.8). Separate job from gmail-bulk-import per plan §2 decision 7: UPDATE semantics against production data is a different risk profile than the historical INSERT job. Parses headers only (no MIME-walk of attachments) from the archived .eml files, appends {"type": "headers", ...} (plan §4.1) via the idempotent UPDATE ... WHERE NOT EXISTS from §5.2, batched via executemany. --limit/ --offset (ORDER BY id) give stable, deterministic partitioning so the full backfill can run and be verified in slices instead of one unattended pass. Plain CLI (pip install -e), no Dockerfile — same convention as gmail-bulk-import/documents-ingest, which run directly on PIHA for local filesystem access to the .eml archive. Verified against kb-postgres@PIHA (100-row dry-run + apply, re-run proved idempotent no-op, 1000-row timed slice): 1096/225030 rows backfilled, 996/1000 succeeded on the timed slice (4 parse_errors — a 2014 spam message with an RFC 2047 encoded-word decoding to an embedded newline in the From display name, correctly caught and skipped rather than crashing the batch). Extrapolated full-run time ~16 minutes. Full 225,030-row run is out of scope for this change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-14 15:34:07 +02:00
]
[project.scripts]
gmail-header-backfill = "gmail_header_backfill.backfill:main"
[tool.setuptools.packages.find]
where = ["src"]
[tool.pytest.ini_options]
asyncio_mode = "auto"
testpaths = ["tests"]