homelab-codex-ws/kb/services/job-mail-body-ingest.md
oskar 01db57ab82 fix(kb): przepiecie wszystkich odwolan wewnetrznych po migracji
126 plikow (md, yaml, sh, py) odwolywalo sie do sciezek sprzed migracji.

  15  markdown-linkow [..](..) -> policzona sciezka WZGLEDNA wobec pliku
      odsylajacego (wczesniej czesc z nich byla repo-root-relative i nie
      rozwiazywala sie z katalogu, w ktorym lezala)
 200  odwolan tekstowych (backticki, proza, yaml, importy w kodzie)
      -> nowa sciezka repo-root-relative, zgodnie z konwencja repo
   5  linkow rodzenstwa (gole nazwy plikow, np. "](DEPLOY.md)") — dzialaly
      tylko w starym katalogu; przeliczone recznie

Objete m.in.: CLAUDE.md (scripts/onboard/README.md -> kb/runbooks/
node-onboarding-tool.md, docs/backlog.md -> kb/phases/backlog.md),
README.md, .claude/skills/, 20 session logow, kod jobow.

Ostatnie 5 odwolan pochodzi z tresci wciagnietej rebasem z origin/master
(session log 2026-07-31, override node-agenta na SOLARII, dwie pozycje
backlogu) — wskazywaly na docs/incidents/, docs/kb/modules/ i
services/narty27/README.md sprzed migracji.

Dodany wzajemny link miedzy kb/services/control-plane.md (stub kodu)
a kb/subsystems/control-plane.md (opis, deprecated) — dwa dokumenty o tym
samym systemie, latwe do pomylenia.

Weryfikacja na 790 plikach: 0 odwolan do starych sciezek,
0 martwych linkow markdown. Lint OKF: 190/190 plikow ZGODNE.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 16:53:57 +02:00

5.6 KiB

okf type visibility status updated links
0.1 service private active 2026-07-22
../runbooks/mail-body-ingest-run.md

mail-body-ingest

Module 5, faza mailowa (kb/phases/kb-m5-faza-mailowa.md, §5, Krok 2). Second full pass over the gmail .eml archive — gmail-bulk-import deliberately skipped inline text/plain/text/html parts (_parse_attachments does continue on them); this job reads exactly the content that gap left out, chunks it, embeds it, and inserts it into document_chunk alongside the existing paperless chunks.

Why a separate job, not an extension of documents-ingest's chunk_embed

chunk_embed.py is wired to source='paperless' + entities[type=content] (pre-extracted text already in the DB). Mail content isn't in the DB yet — it has to be read from .eml files, MIME-walked, quote-stripped, and classified, none of which paperless chunks need. The only thing genuinely shared is the chunker itself, which is why it was extracted to kb_mail.chunking first (Krok 0) instead of being copy-pasted here.

Where it runs

On SOLARIA (needs Ollama on localhost for /api/embed), against kb-postgres@PIHA over Tailscale. The .eml archive is rsync'd PIHA -> SOLARIA once (plan §7, Krok 4) rather than read live over the network — 225k small files over Tailscale would be slow and fragile.

Install (from repo root, on SOLARIA):

pip install -e packages/kb-mail/
pip install -e packages/kb-retrieval/
pip install -e jobs/mail-body-ingest/

Pipeline (per envelope)

  1. Read archive_root / raw_refmissing_file/read_error counted like gmail-header-backfill.
  2. Parse: typed (policy.default) with a compat32 fallback (same ~9/225030 failure mode gmail-header-backfill documents).
  3. Body extraction: inline text/plain preferred; HTML->text via a small stdlib HTMLParser when the mail is HTML-only (plan §1.3: 15% of the corpus) — zero new dependencies, skips style/script/head content.
  4. Quote-strip (Decyzja 2): truncate at the earliest reply marker (On ... wrote:, Dnia ... napisał(a):, W dniu ... pisze:, -----Original Message-----, Outlook's underscore separator), then drop remaining >-quoted lines. In HTML, blockquote and div.gmail_quote subtrees are skipped before conversion to text. quoted_chars_stripped is tallied for calibration review (plan §7).
  5. Classification: newsletter (List-Unsubscribe/List-Id/Precedence: bulk|list, read from the same parsed message); body_empty (after quote-strip — an empty mail is still counted, just produces zero chunks).
  6. Prefix (Decyzya 3): Temat: ... | Od: ... | Data: YYYY-MM-DD built from the already-backfilled entities[type=headers] + envelope.ts — zero header re-parse.
  7. Chunk: kb_mail.chunking.chunk_text (2400/600 chars, same as paperless).
  8. Embed + insert: newsletter chunks are inserted immediately with excluded_reason='newsletter', embedding=NULL (no Ollama call, reversible later); everything else is buffered up to --batch-size (default 64) and sent through kb_retrieval.embed.embed_batch (/api/embed) before inserting. ON CONFLICT (envelope_id, chunk_index, model) DO NOTHING is checked via the command tag, so a silent no-op counts as chunks_conflict_skipped, never chunks_inserted.
  9. Threading append (Decyzja 10): In-Reply-To/References (angle brackets stripped, matching envelope.id's bare-Message-ID convention) appended as entities[type=threading] via the same idempotent WHERE NOT EXISTS UPDATE pattern as gmail-header-backfill — done for every parsed envelope regardless of newsletter/ body_empty status, since it's the same read either way.

Stats must balance

mails_scanned = missing_file + read_errors + parse_errors + body_empty + mails_chunked
chunks_total  = chunks_inserted + chunks_newsletter_flagged + chunks_already_embedded
                + chunks_conflict_skipped + chunks_errors

Any non-zero read_errors/parse_errors/missing_file/chunks_errors/ chunks_conflict_skipped, or an unbalanced sum, makes the CLI exit 1 — same convention as gmail-header-backfill/documents-ingest's chunk_embed.

Ollama-offline tolerance

A failed embed_batch() call is caught per-batch (aiohttp.ClientError -> the whole batch, up to --batch-size chunks, counts as chunks_errors; the run logs a warning and continues). Those chunks never enter the idempotency set, so a later re-run retries them automatically — no separate checkpointing needed. Only a wrong embedding dimension (EmbeddingDimensionError) aborts the entire run, since that would otherwise silently index a vector that doesn't match document_chunk.embedding VECTOR(1024).

Idempotency

Pre-fetched (envelope_id, chunk_index) pairs, scoped server-side to --model (WHERE model = $1), skip chunks already inserted — the pair doesn't need model redundantly since the fetch is already scoped to it. Built this way from the start per the plan's note that chunk_embed.py's otherwise-equivalent pre-fetch is easy to mis-key across runs that mix models (not an active bug there today, since one run always uses one model, but worth not repeating the ambiguity here).

Definition of Done

Per CLAUDE.md: pytest passes (48/48) + a --limit 5 dry-run smoke against live kb-postgres@PIHA before committing (confirms DSN/query wiring; a missing local archive mirror correctly reports missing_file rather than crashing). The Etap A pilot (--since 2025-07-01 --apply, plan §7) is a separate, explicitly-confirmed run — not part of this job's DoD, since it's the first real write against production data.