From a95524c363f72c09cc751d21b4cc48a9d45dd68d Mon Sep 17 00:00:00 2001 From: oskar Date: Wed, 22 Jul 2026 19:01:59 +0200 Subject: [PATCH] docs(mail-body-ingest): add job README Usage, pipeline walkthrough, stats bilans, Ollama-offline tolerance, idempotency, and DoD -- mirrors gmail-header-backfill's README structure. --- jobs/mail-body-ingest/README.md | 131 ++++++++++++++++++++++++++++++++ 1 file changed, 131 insertions(+) create mode 100644 jobs/mail-body-ingest/README.md diff --git a/jobs/mail-body-ingest/README.md b/jobs/mail-body-ingest/README.md new file mode 100644 index 0000000..99fc1ce --- /dev/null +++ b/jobs/mail-body-ingest/README.md @@ -0,0 +1,131 @@ +# mail-body-ingest + +Module 5, faza mailowa (`docs/kb/modules/05-faza-mailowa-plan.md`, §5, Krok 2). Second full +pass over the gmail `.eml` archive — `gmail-bulk-import` deliberately skipped inline +`text/plain`/`text/html` parts (`_parse_attachments` does `continue` on them); this job reads +exactly the content that gap left out, chunks it, embeds it, and inserts it into +`document_chunk` alongside the existing paperless chunks. + +## Why a separate job, not an extension of documents-ingest's chunk_embed + +`chunk_embed.py` is wired to `source='paperless'` + `entities[type=content]` (pre-extracted +text already in the DB). Mail content isn't in the DB yet — it has to be read from `.eml` +files, MIME-walked, quote-stripped, and classified, none of which paperless chunks need. The +only thing genuinely shared is the chunker itself, which is why it was extracted to +`kb_mail.chunking` first (Krok 0) instead of being copy-pasted here. + +## Where it runs + +**On SOLARIA** (needs Ollama on `localhost` for `/api/embed`), against `kb-postgres@PIHA` over +Tailscale. The `.eml` archive is rsync'd PIHA -> SOLARIA once (plan §7, Krok 4) rather than +read live over the network — 225k small files over Tailscale would be slow and fragile. + +Install (from repo root, on SOLARIA): + +```bash +pip install -e packages/kb-mail/ +pip install -e packages/kb-retrieval/ +pip install -e jobs/mail-body-ingest/ +``` + +## Usage + +```bash +# Dry run (default) — parse, quote-strip, classify, chunk, count. Zero Ollama calls, zero +# DB writes (including the threading UPDATE): +mail-body-ingest --dsn postgresql://kb:@piha:5433/kb --archive-root /home/oskar/kb/mail/archive + +# Etap A pilot — last 12 months only (plan Decyzja 9): +mail-body-ingest --dsn ... --since 2025-07-01 --apply > mail-ingest-etapA.log 2>&1 + +# Smoke-test slice: +mail-body-ingest --dsn ... --apply --limit 10 +``` + +DSN can also come from `KB_DSN`, Ollama URL from `OLLAMA_URL` (default +`http://localhost:11434` — this job is meant to run where Ollama lives). + +## Pipeline (per envelope) + +1. **Read** `archive_root / raw_ref` — `missing_file`/`read_error` counted like + `gmail-header-backfill`. +2. **Parse**: typed (`policy.default`) with a `compat32` fallback (same ~9/225030 failure mode + `gmail-header-backfill` documents). +3. **Body extraction**: inline `text/plain` preferred; HTML->text via a small stdlib + `HTMLParser` when the mail is HTML-only (plan §1.3: 15% of the corpus) — zero new + dependencies, skips `style`/`script`/`head` content. +4. **Quote-strip** (Decyzja 2): truncate at the earliest reply marker (`On ... wrote:`, + `Dnia ... napisał(a):`, `W dniu ... pisze:`, `-----Original Message-----`, Outlook's + underscore separator), then drop remaining `>`-quoted lines. In HTML, `blockquote` and + `div.gmail_quote` subtrees are skipped before conversion to text. `quoted_chars_stripped` + is tallied for calibration review (plan §7). +5. **Classification**: `newsletter` (`List-Unsubscribe`/`List-Id`/`Precedence: bulk|list`, + read from the same parsed message); `body_empty` (after quote-strip — an empty mail is + still counted, just produces zero chunks). +6. **Prefix** (Decyzya 3): `Temat: ... | Od: ... | Data: YYYY-MM-DD` built from the + already-backfilled `entities[type=headers]` + `envelope.ts` — zero header re-parse. +7. **Chunk**: `kb_mail.chunking.chunk_text` (2400/600 chars, same as paperless). +8. **Embed + insert**: newsletter chunks are inserted immediately with + `excluded_reason='newsletter'`, `embedding=NULL` (no Ollama call, reversible later); + everything else is buffered up to `--batch-size` (default 64) and sent through + `kb_retrieval.embed.embed_batch` (`/api/embed`) before inserting. `ON CONFLICT + (envelope_id, chunk_index, model) DO NOTHING` is checked via the command tag, so a + silent no-op counts as `chunks_conflict_skipped`, never `chunks_inserted`. +9. **Threading append** (Decyzja 10): `In-Reply-To`/`References` (angle brackets stripped, + matching `envelope.id`'s bare-Message-ID convention) appended as + `entities[type=threading]` via the same idempotent `WHERE NOT EXISTS` UPDATE pattern as + `gmail-header-backfill` — done for every parsed envelope regardless of newsletter/ + body_empty status, since it's the same read either way. + +## Stats must balance + +``` +mails_scanned = missing_file + read_errors + parse_errors + body_empty + mails_chunked +chunks_total = chunks_inserted + chunks_newsletter_flagged + chunks_already_embedded + + chunks_conflict_skipped + chunks_errors +``` + +Any non-zero `read_errors`/`parse_errors`/`missing_file`/`chunks_errors`/ +`chunks_conflict_skipped`, or an unbalanced sum, makes the CLI exit 1 — same convention as +`gmail-header-backfill`/`documents-ingest`'s `chunk_embed`. + +## Ollama-offline tolerance + +A failed `embed_batch()` call is caught per-batch (`aiohttp.ClientError` -> the whole batch, +up to `--batch-size` chunks, counts as `chunks_errors`; the run logs a warning and continues). +Those chunks never enter the idempotency set, so a later re-run retries them automatically — +no separate checkpointing needed. Only a wrong embedding dimension +(`EmbeddingDimensionError`) aborts the entire run, since that would otherwise silently index +a vector that doesn't match `document_chunk.embedding VECTOR(1024)`. + +## Idempotency + +Pre-fetched `(envelope_id, chunk_index)` pairs, scoped server-side to `--model` (`WHERE model += $1`), skip chunks already inserted — the pair doesn't need `model` redundantly since the +fetch is already scoped to it. Built this way from the start per the plan's note that +`chunk_embed.py`'s otherwise-equivalent pre-fetch is easy to mis-key across runs that mix +models (not an active bug there today, since one run always uses one model, but worth not +repeating the ambiguity here). + +## Tests + +```bash +pip install -e "jobs/mail-body-ingest[dev]" +cd jobs/mail-body-ingest && pytest +``` + +Pure unit tests (48), no DB/Ollama — `run()` is tested by monkeypatching `asyncpg.connect` +and `aiohttp.ClientSession` with in-memory fakes, `.eml` bytes written to `tmp_path`. Covers: +quote-strip (EN/PL/Outlook markers, bare `>` lines), HTML->text (style/script/blockquote/ +gmail_quote skipping), newsletter classification, threading extraction, prefix building, +body extraction (plain-preferred, HTML fallback, attachment-only), the typed/compat32 parse +fallback, stats balance, idempotency (second run inserts nothing new), newsletter chunks +never reaching Ollama, Ollama-offline batch isolation, and dimension-mismatch abort. + +## Definition of Done + +Per `CLAUDE.md`: `pytest` passes (48/48) + a `--limit 5` dry-run smoke against live +`kb-postgres@PIHA` before committing (confirms DSN/query wiring; a missing local archive +mirror correctly reports `missing_file` rather than crashing). The Etap A pilot (`--since +2025-07-01 --apply`, plan §7) is a separate, explicitly-confirmed run — not part of this +job's DoD, since it's the first real write against production data.