homelab-codex-ws/jobs/mail-body-ingest/README.md
oskar a0c9dbfa52 docs(mail-body-ingest): add job README
Usage, pipeline walkthrough, stats bilans, Ollama-offline tolerance,
idempotency, and DoD -- mirrors gmail-header-backfill's README structure.
2026-07-22 19:01:59 +02:00

6.8 KiB

mail-body-ingest

Module 5, faza mailowa (docs/kb/modules/05-faza-mailowa-plan.md, §5, Krok 2). Second full pass over the gmail .eml archive — gmail-bulk-import deliberately skipped inline text/plain/text/html parts (_parse_attachments does continue on them); this job reads exactly the content that gap left out, chunks it, embeds it, and inserts it into document_chunk alongside the existing paperless chunks.

Why a separate job, not an extension of documents-ingest's chunk_embed

chunk_embed.py is wired to source='paperless' + entities[type=content] (pre-extracted text already in the DB). Mail content isn't in the DB yet — it has to be read from .eml files, MIME-walked, quote-stripped, and classified, none of which paperless chunks need. The only thing genuinely shared is the chunker itself, which is why it was extracted to kb_mail.chunking first (Krok 0) instead of being copy-pasted here.

Where it runs

On SOLARIA (needs Ollama on localhost for /api/embed), against kb-postgres@PIHA over Tailscale. The .eml archive is rsync'd PIHA -> SOLARIA once (plan §7, Krok 4) rather than read live over the network — 225k small files over Tailscale would be slow and fragile.

Install (from repo root, on SOLARIA):

pip install -e packages/kb-mail/
pip install -e packages/kb-retrieval/
pip install -e jobs/mail-body-ingest/

Usage

# Dry run (default) — parse, quote-strip, classify, chunk, count. Zero Ollama calls, zero
# DB writes (including the threading UPDATE):
mail-body-ingest --dsn postgresql://kb:<pw>@piha:5433/kb --archive-root /home/oskar/kb/mail/archive

# Etap A pilot — last 12 months only (plan Decyzja 9):
mail-body-ingest --dsn ... --since 2025-07-01 --apply > mail-ingest-etapA.log 2>&1

# Smoke-test slice:
mail-body-ingest --dsn ... --apply --limit 10

DSN can also come from KB_DSN, Ollama URL from OLLAMA_URL (default http://localhost:11434 — this job is meant to run where Ollama lives).

Pipeline (per envelope)

  1. Read archive_root / raw_refmissing_file/read_error counted like gmail-header-backfill.
  2. Parse: typed (policy.default) with a compat32 fallback (same ~9/225030 failure mode gmail-header-backfill documents).
  3. Body extraction: inline text/plain preferred; HTML->text via a small stdlib HTMLParser when the mail is HTML-only (plan §1.3: 15% of the corpus) — zero new dependencies, skips style/script/head content.
  4. Quote-strip (Decyzja 2): truncate at the earliest reply marker (On ... wrote:, Dnia ... napisał(a):, W dniu ... pisze:, -----Original Message-----, Outlook's underscore separator), then drop remaining >-quoted lines. In HTML, blockquote and div.gmail_quote subtrees are skipped before conversion to text. quoted_chars_stripped is tallied for calibration review (plan §7).
  5. Classification: newsletter (List-Unsubscribe/List-Id/Precedence: bulk|list, read from the same parsed message); body_empty (after quote-strip — an empty mail is still counted, just produces zero chunks).
  6. Prefix (Decyzya 3): Temat: ... | Od: ... | Data: YYYY-MM-DD built from the already-backfilled entities[type=headers] + envelope.ts — zero header re-parse.
  7. Chunk: kb_mail.chunking.chunk_text (2400/600 chars, same as paperless).
  8. Embed + insert: newsletter chunks are inserted immediately with excluded_reason='newsletter', embedding=NULL (no Ollama call, reversible later); everything else is buffered up to --batch-size (default 64) and sent through kb_retrieval.embed.embed_batch (/api/embed) before inserting. ON CONFLICT (envelope_id, chunk_index, model) DO NOTHING is checked via the command tag, so a silent no-op counts as chunks_conflict_skipped, never chunks_inserted.
  9. Threading append (Decyzja 10): In-Reply-To/References (angle brackets stripped, matching envelope.id's bare-Message-ID convention) appended as entities[type=threading] via the same idempotent WHERE NOT EXISTS UPDATE pattern as gmail-header-backfill — done for every parsed envelope regardless of newsletter/ body_empty status, since it's the same read either way.

Stats must balance

mails_scanned = missing_file + read_errors + parse_errors + body_empty + mails_chunked
chunks_total  = chunks_inserted + chunks_newsletter_flagged + chunks_already_embedded
                + chunks_conflict_skipped + chunks_errors

Any non-zero read_errors/parse_errors/missing_file/chunks_errors/ chunks_conflict_skipped, or an unbalanced sum, makes the CLI exit 1 — same convention as gmail-header-backfill/documents-ingest's chunk_embed.

Ollama-offline tolerance

A failed embed_batch() call is caught per-batch (aiohttp.ClientError -> the whole batch, up to --batch-size chunks, counts as chunks_errors; the run logs a warning and continues). Those chunks never enter the idempotency set, so a later re-run retries them automatically — no separate checkpointing needed. Only a wrong embedding dimension (EmbeddingDimensionError) aborts the entire run, since that would otherwise silently index a vector that doesn't match document_chunk.embedding VECTOR(1024).

Idempotency

Pre-fetched (envelope_id, chunk_index) pairs, scoped server-side to --model (WHERE model = $1), skip chunks already inserted — the pair doesn't need model redundantly since the fetch is already scoped to it. Built this way from the start per the plan's note that chunk_embed.py's otherwise-equivalent pre-fetch is easy to mis-key across runs that mix models (not an active bug there today, since one run always uses one model, but worth not repeating the ambiguity here).

Tests

pip install -e "jobs/mail-body-ingest[dev]"
cd jobs/mail-body-ingest && pytest

Pure unit tests (48), no DB/Ollama — run() is tested by monkeypatching asyncpg.connect and aiohttp.ClientSession with in-memory fakes, .eml bytes written to tmp_path. Covers: quote-strip (EN/PL/Outlook markers, bare > lines), HTML->text (style/script/blockquote/ gmail_quote skipping), newsletter classification, threading extraction, prefix building, body extraction (plain-preferred, HTML fallback, attachment-only), the typed/compat32 parse fallback, stats balance, idempotency (second run inserts nothing new), newsletter chunks never reaching Ollama, Ollama-offline batch isolation, and dimension-mismatch abort.

Definition of Done

Per CLAUDE.md: pytest passes (48/48) + a --limit 5 dry-run smoke against live kb-postgres@PIHA before committing (confirms DSN/query wiring; a missing local archive mirror correctly reports missing_file rather than crashing). The Etap A pilot (--since 2025-07-01 --apply, plan §7) is a separate, explicitly-confirmed run — not part of this job's DoD, since it's the first real write against production data.