--- okf: "0.1" type: phase visibility: private status: active updated: 2026-07-30 links: - ../services/job-documents-ingest.md - ../runbooks/documents-ingest-run.md --- # documents-ingest — fazy 2 i 3 ## Phase 2 — `documents-ingest-paperless` (Paperless -> envelope adapter) Module 5, phase 2 (`kb/phases/kb-m5-faza2.md`, §4.2-4.3, §6 step 5). Reads documents from the **Paperless REST API** (read-only — GET only, never writes to Paperless) and inserts them as `source='paperless'` rows into the `envelope` table on kb-postgres, reusing `kb_mail.envelope.Envelope` / `kb_mail.db.insert_envelope` from `packages/kb-mail` (untouched by this change — see plan §1.6). Existing `source='gmail'` rows and `document_chunk` are never touched; this job only ever `INSERT`s new `paperless` rows. ### Cross-source link (`source_mail`) Per plan §1.9/§4.2, the deterministic join uses no heuristics: a document's `original_file_name` (from the Paperless API) is matched against `consume_name` in this job's **phase-1 registry** (`/opt/homelab/data/documents-ingest/registry.json`, produced by `extractor.py` — see above). A match appends a `source_mail` entity pointing back at the originating mail envelope; no match means the document was added outside the faktury-1 pipeline, and the entity is simply omitted — not an error. ### Install ```bash pip install -e packages/kb-mail/ pip install -e jobs/documents-ingest/ ``` ### Usage ```bash # Dry run (default) — fetch from Paperless, map, count; no DB writes: documents-ingest-paperless --dsn postgresql://kb:@localhost:5433/kb \ --paperless-token # Real run — insert new envelope rows: documents-ingest-paperless --dsn ... --paperless-token ... --apply # Smoke-test slice: documents-ingest-paperless --dsn ... --paperless-token ... --limit 5 ``` `--dsn` can come from `KB_DSN`, `--paperless-token` from `PAPERLESS_API_TOKEN`, `--paperless-url` from `PAPERLESS_URL` (defaults to Paperless' fixed LAN address, `http://192.168.31.5:8210`). No `--offset`: unlike the 225 030-row header backfill, a full re-scan of Paperless' ~186 documents is cheap and already idempotent, so there is no need for resumable partitioning — `--limit` exists only to cap a run for smoke-testing. ### Mapping (plan §4.3) ``` id = f"paperless:{document_id}" -- prefixed: Paperless doc-ids are small -- sequential ints that would otherwise -- collide with any future source's ids ts = documents_document.created -- Paperless-detected date (content/filename), -- not filesystem mtime geo = NULL raw_ref = str(document_id) -- REFERENCE — Paperless is the source of truth, -- no bytes are copied entities = content, correspondent, tag(s), filename, content_type, and source_mail when the registry join hits (plan §4.2) ``` `correspondent`/`tag` are resolved from Paperless' `/api/correspondents/` and `/api/tags/` (fetched once, cached in memory for the run) and kept purely as informational metadata — nothing in this pipeline depends on them being non-null (plan decision 4). A document with empty OCR content (Paperless OCR sometimes produces none) still gets a normal envelope with `"text": ""` — not skipped, not an error, just counted (`empty_content`). ### Idempotency A pre-fetched set of existing `source='paperless'` envelope ids (one query at the start of each run) skips documents already inserted; `insert_envelope`'s own `ON CONFLICT (id) DO NOTHING` is the second line of defense. Re-running `--apply` immediately after a successful run reports `inserted: 0` and `already_in_db` equal to the previous run's `inserted` count. ### Stats must balance ``` fetched = already_in_db + inserted + errors ``` `source_mail_linked` and `empty_content` are informational subsets of `fetched`, not separate outcome buckets. A per-document mapping failure (e.g. an unparseable `created` date) is isolated, logged, and counted as `errors` — it never aborts the run. `main()` exits 1 on non-zero `errors` or if the balance invariant above doesn't hold (mirrors `gmail-bulk-import`'s exit-code convention) — a clean run always exits 0. ### Tests ```bash pip install -e packages/kb-mail/ pip install -e jobs/documents-ingest/ cd jobs/documents-ingest && pytest ``` Pure unit tests, no DB or real HTTP — `run()` is tested by monkeypatching `asyncpg.connect` (fake connection) and `aiohttp.ClientSession` (fake session serving canned JSON pages). Covers: mapping shape (content, correspondent, tag(s), filename, content_type, source_mail), the registry join (hit and miss), pagination (both the documents list and the correspondents/tags lookup tables), `--limit`, idempotency (pre-existing ids skipped, a second `--apply` run inserts nothing new), isolated per-document mapping errors, and the stats-balance invariant. ### Definition of Done Per `CLAUDE.md`: smoke run is `documents-ingest-paperless --dsn ... --paperless-token ... --limit 5` (dry-run first) against kb-postgres@PIHA and the live Paperless API, over SSH — **not executed as part of this change** without operator confirmation (this job reads production Paperless data and writes production envelope rows on `--apply`). `pytest` passes locally before this commit. --- ## Phase 2 step 6 — `documents-ingest-embed` (chunk + embed) Module 5, phase 2, plan step 6 (`kb/phases/kb-m5-faza2.md`, §6 step 6, §2 decision 3). Reads `entities[type=content].text` off every `source='paperless'` envelope, chunks it, calls Ollama (`POST /api/embeddings`, model `bge-m3`) for each chunk, and inserts the result into `document_chunk` (`services/kb-postgres/init/002_chunks.sql`). This job only ever `INSERT`s into `document_chunk` — `envelope` is read-only here, and `services/ollama/` is untouched. ### Where it runs **On SOLARIA** (that's where Ollama lives), against kb-postgres@PIHA over Tailscale — the reverse of the other jobs in this package, which run on PIHA. `--ollama-url` defaults to `http://localhost:11434` (Ollama on the same node); `--dsn` needs PIHA's Tailscale address, e.g. `postgresql://kb:@piha:5433/kb`. ### Install ```bash pip install -e packages/kb-mail/ pip install -e jobs/documents-ingest/ ``` ### Usage ```bash # Dry run (default) — chunk and count only, no Ollama calls, no DB writes: documents-ingest-embed --dsn postgresql://kb:@piha:5433/kb # Smoke-test slice: documents-ingest-embed --dsn ... --apply --limit 10 # Full run: documents-ingest-embed --dsn ... --apply ``` ### Chunking (plan §2 decision 3) Paragraph-preferring: splits on blank-line boundaries, greedily packs paragraphs up to `--chunk-size` characters (default 2400, ≈600 tokens at a ~4 chars/token heuristic — no local bge-m3 tokenizer available offline), `--chunk-overlap` characters of trailing context carried into the next chunk (default 600, ≈150 tokens). A paragraph that alone exceeds `--chunk-size` falls back to a hard character-based sliding window — Paperless OCR text has no page-break markers (plan §1.2), so there's nothing else to split large, unbroken text on. A document with empty OCR content (the 26 `empty_content` documents from phase 2 step 5) yields zero chunks and is counted separately, not as an error. ### Idempotency A pre-fetched set of `(envelope_id, chunk_index)` pairs already embedded with `--model` skips re-embedding on rerun — no wasted Ollama calls. `document_chunk`'s own `UNIQUE (envelope_id, chunk_index)` + `ON CONFLICT DO NOTHING` is the second line of defense; `insert_chunk`'s command tag is checked so a silently-skipped row is counted as `chunks_conflict_skipped`, never miscounted as `chunks_inserted`. Note that uniqueness is on `(envelope_id, chunk_index)` only, not `model` — re-embedding with a *different* model hits this path and that embedding is discarded (wasted work, correctly reported via `chunks_conflict_skipped`, but not persisted). Out of scope for this single-model pilot; the real fix for whoever indexes a second model later is `UNIQUE (envelope_id, chunk_index, model)` at the schema layer. A DB write failure for one chunk (dropped connection, unexpected bytes) is isolated the same way an embed failure is — counted as `chunks_errors`, never aborting the rest of the run. ### Dimension guard Every embedding response's length is checked against `document_chunk.embedding`'s `VECTOR(1024)` column. A mismatch raises `EmbeddingDimensionError` and aborts the whole run immediately — never silently indexes vectors of the wrong dimension. ### Chunk size/overlap validation `--chunk-overlap` must be smaller than `--chunk-size` — the sliding-window hard-split fallback advances by `chunk_size - chunk_overlap` per step, so an overlap `>=` size would never advance and hang. `main()` rejects this combination before opening a DB connection; `hard_split()` itself also raises `ValueError` as a second line of defense for direct callers. ### Stats must balance ``` documents_fetched = empty_content + documents_chunked chunks_total = chunks_already_embedded + chunks_inserted + chunks_conflict_skipped + chunks_errors ``` `main()` exits 1 on `chunks_errors > 0`, `chunks_conflict_skipped > 0`, or if either balance breaks. The summary line also reports `avg_embed_seconds_per_chunk` — CPU-only Ollama timing, the input for deciding whether/how to scale this to the mail corpus later (plan §7). ### Tests ```bash pip install -e packages/kb-mail/ pip install -e jobs/documents-ingest/ cd jobs/documents-ingest && pytest ``` Pure unit tests, no DB or real HTTP — `run()` is tested by monkeypatching `asyncpg.connect` (fake connection) and `aiohttp.ClientSession` (fake session serving a canned embedding vector, or a 500 for a chosen prompt to exercise error isolation). Covers: chunking (paragraph boundaries, overlap, empty document, document shorter than one chunk, oversized paragraph hard-fallback, the overlap-must-be-smaller-than-size guard), `extract_content`, idempotency (pre-existing keys skipped, no Ollama calls made for them, a second `--apply` run embeds nothing new, existing keys are correctly scoped to `--model`), dimension-mismatch abort, isolated per-chunk embed and insert errors, `ON CONFLICT` no-ops counted separately from real inserts, and the stats-balance invariant. ### Known limitation — Ollama context-length rejections on pathological chunks Ollama's *runtime* context window for a model can be smaller than the model's advertised max (bge-m3 supports 8192 tokens, but Ollama's default `num_ctx` is lower) — and some OCR text tokenizes far more densely than the ~4-chars/token heuristic this job uses to size chunks. Concretely: a table- of-contents page made almost entirely of dot-leader formatting (`". . . . . . ."`, repeated hundreds of times) hit this on the pilot run — Ollama returned `500 {"error":"the input length exceeds the context length"}` for one 2400-char chunk that should have been well within budget by character count alone. The job isolates this exactly like any other embed failure (`chunks_errors`, logged, run continues), so it never crashes a run — but it also never automatically shrinks and retries the offending chunk. Given how rare this was (1 chunk out of 2684 in the full pilot, all from one document's dot-leader ToC), it's left as a known gap rather than fixed here; a real fix would be either a smaller/adaptive chunk size for low-character-entropy text, or a shrink-and-retry loop on this specific Ollama error. ### Definition of Done Per `CLAUDE.md`: `pytest` passes locally (101 tests). Smoke-tested and then run to completion live on SOLARIA against the real Ollama instance and kb-postgres@PIHA: - Dry-run: 186 fetched, 26 `empty_content`, 2684 chunks planned — matches the known phase-2-step-5 figures exactly. - `--apply --limit 10`: 64 chunks embedded, 0 errors, avg ≈0.83s/chunk on CPU. - Re-run of the same slice: fully idempotent — 0 Ollama calls, 0 inserts. - Full `--apply` (all 186 documents): **2683/2684 chunks inserted, 1 isolated error** (see "Known limitation" above) — `chunks_errors=1` correctly produced a non-zero exit rather than silently reporting success. `document_chunk` ends at 2683 rows across 160 distinct envelopes, matching `documents_chunked`. A `document_chunk_envelope_idx`-backed count and an `ORDER BY embedding <=> ...` nearest-neighbor sanity query both look correct (top match is the reference chunk itself at distance 0; next nearest are chunks of the same source document). - **Timing (CPU-only, no GPU driver on SOLARIA)**: ≈0.79s/chunk average across 2683 real embeddings (2115.8s total embed time), ≈13.2s/document average across the 160 chunked documents, ≈35 minutes wall-clock for the full 186-document pilot. This is the real-world input for scaling this pipeline to the much larger mail corpus later (plan §7 assumed GPU-based "minutes for the whole pilot"; SOLARIA's Ollama ran CPU-only for this pilot per the then-disabled GPU reservation). The 186-document pilot's ≈13.2s/document average is dominated by Paperless' long OCR text (≈22k chars/doc average, per plan §1.2) — 225 030 mail envelopes will have a very different, likely much shorter, per-envelope chunk count (email bodies vs. scanned multi-page PDFs), so this number doesn't extrapolate directly to a mail-corpus estimate. What it does establish: at ≈0.79s/chunk sequential CPU embedding, any corpus with a non-trivial average chunk count per item will need either a GPU driver fix, concurrent/batched Ollama calls, or both, before a full mail-corpus run is practical — flagged for whoever picks up the mail-indexer phase. - **GPU (RTX 4070 Ti SUPER, driver 595-open, restored 2026-07-16)**: 207ms/embed (50 sekwencyjnych wywołań /api/embeddings, ~600-tok prompt) vs 790ms/chunk CPU baseline — ~3.8× szybciej sekwencyjnie; przy pojedynczych requestach dominuje overhead HTTP/tokenizacji, realny skok da dopiero batching (backlog). --- ## Phase 3 step 4 — retrieval cascade (`documents_ingest.retrieval`) + quality gate Module 5, phase 3, plan step 4 (`kb/phases/kb-m5-faza3.md`, §6). Two retrieval paths, both `query_text -> chunk hits (dist, source)` — the intended clean API surface for phase 4's kb-query, not just this eval: - `flat_query` — baseline: rank every active `document_chunk` row directly. Formalizes the phase-2 pilot's ad hoc `/tmp/kbq.sh` query into a tested module. - `cascade_query` — pre-filter to the top-N `document_summary` envelopes (one `model`, default `claude-haiku-4-5` — plan §2 decision 3, resolved 2026-07-17) before ranking `document_chunk` within just those envelopes. Both share **one** query embedding call; the cascade only adds one extra SQL query (stage 1), never an extra Ollama call. `envelope`, `document_chunk`, and `document_summary` are read-only — this module only ever `SELECT`s. ### Quality gate `eval/queries.yaml` — 7 queries transcribed 1:1 from the phase-2 pilot baseline (`kb/phases/kb-m5-eval-retrieval-pilot.md`, left untouched — this is its versioned working copy) with expected envelope / kind (`hit`, `grey_zone`, `negative_control`, `negative_control_borderline`) per query. `eval/retrieval_eval.py` — read-only integration script against the live DB + live Ollama, **not collected by pytest** (same reasoning as the plan: an eval gate against live data isn't a mocked unit test). Runs every query through both tracks across an N sweep and checks the plan's three gate criteria (no flat hit degrades, hit@3 cascade ≥ flat, negative controls stay > 0.55). Exits 0 on PASS, 1 on FAIL. ```bash pip install -e packages/kb-mail/ -e jobs/documents-ingest/ python eval/retrieval_eval.py --dsn postgresql://kb:@piha:5433/kb \ --ollama-url http://solaria:11434 --n-sweep 5,10,20 ``` `--transport http --base-url http://:8230` (module 5 phase 4 plan §2 decision 6 / §9) calls a live `kb-query`'s `/search` instead of embedding+querying locally — no `--dsn` needed, `--n-sweep` is ignored (kb-query serves one server-side default N per request). Gate criterion: `dist` must be **identical** to the same run with `--transport direct` against the same live SOLARIA (same DB, same retrieval code — HTTP is only a wrapper). **Result (2026-07-17, live run)**: PASS at N=10, k=5 — see plan §6.3 for the full table, the N-sweep calibration (N=5 is the measured safety floor; the plan's N=10 default carries a 2× margin), and the cost/improvement analysis. `cascade_query` (N=10, k=5, `summary_model='claude-haiku-4-5'`) is now the default retrieval path for phase 4's kb-query; `flat_query` stays as the baseline/fallback. ### Tests ```bash pip install -e packages/kb-mail/ pip install -e jobs/documents-ingest/ cd jobs/documents-ingest && pytest ``` `tests/test_retrieval.py` — pure unit tests, no DB or real HTTP. Covers: flat ranking across all envelopes, cascade stage-1-narrows-stage-2, an envelope whose summary exists but has no active chunks, N larger than the number of summarized envelopes, the no-summaries short-circuit (stage 2 never queried), and both query entry points embedding exactly once. --- ## Phase 3 step 5 — cyclic ingest (`documents-ingest-cyclic`) + systemd timer Module 5, phase 3, plan step 5 (`kb/phases/kb-m5-faza3.md`, §7). Orchestrates one run of the recurring ingest pipeline: `paperless_adapter.run()` (new `source='paperless'` envelopes) → `chunk_embed.run()` (new `document_chunk` rows) → `summarize.run_summarize(backend='anthropic')` (new `document_summary` rows, `model='claude-haiku-4-5'` — plan §2 decision 3) → `summarize.run_embed_summaries()` (embeds those summaries). All four are the same job functions used elsewhere in this package, called directly — no changes to `paperless_adapter.py` / `chunk_embed.py` / `summarize.py`, no new CLI flags on them. ### Ollama-offline tolerance SOLARIA has `availability_target: medium` (planned power-off, plan §1.3). The wrapper probes `GET {OLLAMA_URL}/api/tags` before the two embed stages (chunk embedding, summary embedding); unreachable means **skip, not fail** — both embed passes are idempotent, so new chunks/summaries left unembedded this tick are picked up whole on the next one. A growing backlog is what `kb_ingest_embed_backlog` + the `KbEmbedBacklogGrowing` alert are for, not this wrapper's exit code. Anything else failing **is** a hard failure: Paperless unreachable, a DB error, a non-zero job error counter, a broken stats-balance invariant, the Anthropic API failing. Each stage's pass/fail predicate mirrors that job's own `main()` exit check 1:1 (see `cyclic_ingest.py`'s module docstring). Stages are isolated, not fail-fast — an earlier stage failing never skips a later one, mirroring the per-row isolation the underlying jobs already use. ### Usage ```bash # Dry run (default) — same idempotent counting as every other job in this family, no writes: documents-ingest-cyclic --dsn postgresql://kb:@localhost:5433/kb \ --paperless-token --anthropic-api-key # Real run (what the timer invokes): documents-ingest-cyclic --dsn ... --paperless-token ... --anthropic-api-key ... --apply ``` `--dsn`/`--paperless-token`/`--anthropic-api-key` also read from `KB_DSN` / `PAPERLESS_API_TOKEN` / `ANTHROPIC_API_KEY` env vars — never logged. `--ollama-url` defaults to `http://solaria:11434` (this wrapper always runs on PIHA, unlike `chunk_embed`/`summarize`'s own CLI defaults which assume co-location with Ollama). ### Metrics (Prometheus textfile collector) Every run — success or failure — writes `--prom-path` (default `/opt/homelab/state/node-exporter/kb-ingest.prom`) atomically (tmp + rename): | Metric | Meaning | |---|---| | `kb_ingest_last_run_timestamp` | Unix ts of the last run, success or failure | | `kb_ingest_last_success_timestamp` | Unix ts of the last run with no hard failure — carried forward from the previous file on a failing run, never reset to 0/now | | `kb_ingest_last_exit_code` | 0 or 1 | | `kb_ingest_documents_inserted` | New envelope rows this run | | `kb_ingest_chunks_inserted` | New document_chunk rows this run (0 if the embed stage was skipped) | | `kb_ingest_summaries_inserted` | New document_summary rows this run | | `kb_ingest_embed_skipped` | 1 if Ollama was unreachable this run (both embed stages skipped), 0 otherwise | | `kb_ingest_embed_backlog` | Active chunks (`excluded_reason IS NULL`) still missing an embedding | Scraped by fleet-prometheus via node_exporter's textfile collector on PIHA (`hosts/piha/runtime/node_exporter/docker-compose.override.yml`); alert rules in `services/fleet-prometheus/rules/kb-ingest.yml`. ### Install (PIHA) 1. Dedicated venv (per plan §7.1 — not the ad hoc rsync-to-`/tmp` pattern used for the one-shot jobs elsewhere in this README; this is a permanent, recurring installation): ```bash python3 -m venv /opt/homelab/kb/venv /opt/homelab/kb/venv/bin/pip install -e packages/kb-mail -e jobs/documents-ingest ``` (run from a checkout of this repo on PIHA — the checkout is used as an install source only, per CLAUDE.md's "deploy-only" rule; no development happens there). 2. Secrets in `/opt/homelab/kb/.env` (already holds `PAPERLESS_API_TOKEN`; add `KB_DSN=postgresql://kb:@localhost:5433/kb` and `ANTHROPIC_API_KEY=`), `chmod 600`, never in Git. 3. Copy `jobs/documents-ingest/systemd/kb-ingest-run.sh` to `/opt/homelab/kb/` and `chmod +x` it. 4. Copy (or symlink) `kb-ingest.service` and `kb-ingest.timer` to `/etc/systemd/system/`, then: ```bash systemctl daemon-reload systemctl enable --now kb-ingest.timer ``` 5. Verify: `systemctl list-timers kb-ingest.timer`, `journalctl -u kb-ingest.service`, `/opt/homelab/logs/kb-ingest/run-YYYYMMDD.log`, and `/opt/homelab/state/node-exporter/kb-ingest.prom` after the first run (manual `systemctl start kb-ingest.service` to trigger one immediately without waiting for 03:30). ### Tests ```bash pip install -e packages/kb-mail/ pip install -e jobs/documents-ingest/ cd jobs/documents-ingest && pytest ``` `tests/test_cyclic_ingest.py` — pure unit tests, no DB or real HTTP/Ollama/Anthropic; every stage function and the Ollama probe are monkeypatched. Covers: each stage's failure predicate (pinned against its source job's own exit check), the Ollama-down skip path (chunk_embed/embed_summaries never even called), stage isolation (an earlier stage failing never skips a later one, whether via a failed predicate or a raised exception), `.prom` rendering, atomic write, `last_success_timestamp` carry-forward across a failing run, and `main()`'s CLI guardrails + exit-code propagation.