homelab-codex-ws/kb/phases/kb-m5-documents-ingest-fazy.md
oskar 4658089e21 fix(kb): przepiecie wszystkich odwolan wewnetrznych po migracji
126 plikow (md, yaml, sh, py) odwolywalo sie do sciezek sprzed migracji.

  15  markdown-linkow [..](..) -> policzona sciezka WZGLEDNA wobec pliku
      odsylajacego (wczesniej czesc z nich byla repo-root-relative i nie
      rozwiazywala sie z katalogu, w ktorym lezala)
 200  odwolan tekstowych (backticki, proza, yaml, importy w kodzie)
      -> nowa sciezka repo-root-relative, zgodnie z konwencja repo
   5  linkow rodzenstwa (gole nazwy plikow, np. "](DEPLOY.md)") — dzialaly
      tylko w starym katalogu; przeliczone recznie

Objete m.in.: CLAUDE.md (scripts/onboard/README.md -> kb/runbooks/
node-onboarding-tool.md, docs/backlog.md -> kb/phases/backlog.md),
README.md, .claude/skills/, 20 session logow, kod jobow.

Ostatnie 5 odwolan pochodzi z tresci wciagnietej rebasem z origin/master
(session log 2026-07-31, override node-agenta na SOLARII, dwie pozycje
backlogu) — wskazywaly na docs/incidents/, docs/kb/modules/ i
services/narty27/README.md sprzed migracji.

Dodany wzajemny link miedzy kb/services/control-plane.md (stub kodu)
a kb/subsystems/control-plane.md (opis, deprecated) — dwa dokumenty o tym
samym systemie, latwe do pomylenia.

Weryfikacja na 790 plikach: 0 odwolan do starych sciezek,
0 martwych linkow markdown. Lint OKF: 190/190 plikow ZGODNE.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 16:58:46 +02:00

22 KiB
Raw Permalink Blame History

okf type visibility status updated links
0.1 phase private active 2026-07-30
../services/job-documents-ingest.md
../runbooks/documents-ingest-run.md

documents-ingest — fazy 2 i 3

Phase 2 — documents-ingest-paperless (Paperless -> envelope adapter)

Module 5, phase 2 (kb/phases/kb-m5-faza2.md, §4.2-4.3, §6 step 5). Reads documents from the Paperless REST API (read-only — GET only, never writes to Paperless) and inserts them as source='paperless' rows into the envelope table on kb-postgres, reusing kb_mail.envelope.Envelope / kb_mail.db.insert_envelope from packages/kb-mail (untouched by this change — see plan §1.6). Existing source='gmail' rows and document_chunk are never touched; this job only ever INSERTs new paperless rows.

Per plan §1.9/§4.2, the deterministic join uses no heuristics: a document's original_file_name (from the Paperless API) is matched against consume_name in this job's phase-1 registry (/opt/homelab/data/documents-ingest/registry.json, produced by extractor.py — see above). A match appends a source_mail entity pointing back at the originating mail envelope; no match means the document was added outside the faktury-1 pipeline, and the entity is simply omitted — not an error.

Install

pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/

Usage

# Dry run (default) — fetch from Paperless, map, count; no DB writes:
documents-ingest-paperless --dsn postgresql://kb:<pw>@localhost:5433/kb \
    --paperless-token <token>

# Real run — insert new envelope rows:
documents-ingest-paperless --dsn ... --paperless-token ... --apply

# Smoke-test slice:
documents-ingest-paperless --dsn ... --paperless-token ... --limit 5

--dsn can come from KB_DSN, --paperless-token from PAPERLESS_API_TOKEN, --paperless-url from PAPERLESS_URL (defaults to Paperless' fixed LAN address, http://192.168.31.5:8210). No --offset: unlike the 225 030-row header backfill, a full re-scan of Paperless' ~186 documents is cheap and already idempotent, so there is no need for resumable partitioning — --limit exists only to cap a run for smoke-testing.

Mapping (plan §4.3)

id       = f"paperless:{document_id}"   -- prefixed: Paperless doc-ids are small
                                         -- sequential ints that would otherwise
                                         -- collide with any future source's ids
ts       = documents_document.created   -- Paperless-detected date (content/filename),
                                         -- not filesystem mtime
geo      = NULL
raw_ref  = str(document_id)             -- REFERENCE — Paperless is the source of truth,
                                         -- no bytes are copied
entities = content, correspondent, tag(s), filename, content_type,
           and source_mail when the registry join hits (plan §4.2)

correspondent/tag are resolved from Paperless' /api/correspondents/ and /api/tags/ (fetched once, cached in memory for the run) and kept purely as informational metadata — nothing in this pipeline depends on them being non-null (plan decision 4). A document with empty OCR content (Paperless OCR sometimes produces none) still gets a normal envelope with "text": "" — not skipped, not an error, just counted (empty_content).

Idempotency

A pre-fetched set of existing source='paperless' envelope ids (one query at the start of each run) skips documents already inserted; insert_envelope's own ON CONFLICT (id) DO NOTHING is the second line of defense. Re-running --apply immediately after a successful run reports inserted: 0 and already_in_db equal to the previous run's inserted count.

Stats must balance

fetched = already_in_db + inserted + errors

source_mail_linked and empty_content are informational subsets of fetched, not separate outcome buckets. A per-document mapping failure (e.g. an unparseable created date) is isolated, logged, and counted as errors — it never aborts the run. main() exits 1 on non-zero errors or if the balance invariant above doesn't hold (mirrors gmail-bulk-import's exit-code convention) — a clean run always exits 0.

Tests

pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest

Pure unit tests, no DB or real HTTP — run() is tested by monkeypatching asyncpg.connect (fake connection) and aiohttp.ClientSession (fake session serving canned JSON pages). Covers: mapping shape (content, correspondent, tag(s), filename, content_type, source_mail), the registry join (hit and miss), pagination (both the documents list and the correspondents/tags lookup tables), --limit, idempotency (pre-existing ids skipped, a second --apply run inserts nothing new), isolated per-document mapping errors, and the stats-balance invariant.

Definition of Done

Per CLAUDE.md: smoke run is documents-ingest-paperless --dsn ... --paperless-token ... --limit 5 (dry-run first) against kb-postgres@PIHA and the live Paperless API, over SSH — not executed as part of this change without operator confirmation (this job reads production Paperless data and writes production envelope rows on --apply). pytest passes locally before this commit.


Phase 2 step 6 — documents-ingest-embed (chunk + embed)

Module 5, phase 2, plan step 6 (kb/phases/kb-m5-faza2.md, §6 step 6, §2 decision 3). Reads entities[type=content].text off every source='paperless' envelope, chunks it, calls Ollama (POST /api/embeddings, model bge-m3) for each chunk, and inserts the result into document_chunk (services/kb-postgres/init/002_chunks.sql). This job only ever INSERTs into document_chunkenvelope is read-only here, and services/ollama/ is untouched.

Where it runs

On SOLARIA (that's where Ollama lives), against kb-postgres@PIHA over Tailscale — the reverse of the other jobs in this package, which run on PIHA. --ollama-url defaults to http://localhost:11434 (Ollama on the same node); --dsn needs PIHA's Tailscale address, e.g. postgresql://kb:<pw>@piha:5433/kb.

Install

pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/

Usage

# Dry run (default) — chunk and count only, no Ollama calls, no DB writes:
documents-ingest-embed --dsn postgresql://kb:<pw>@piha:5433/kb

# Smoke-test slice:
documents-ingest-embed --dsn ... --apply --limit 10

# Full run:
documents-ingest-embed --dsn ... --apply

Chunking (plan §2 decision 3)

Paragraph-preferring: splits on blank-line boundaries, greedily packs paragraphs up to --chunk-size characters (default 2400, ≈600 tokens at a ~4 chars/token heuristic — no local bge-m3 tokenizer available offline), --chunk-overlap characters of trailing context carried into the next chunk (default 600, ≈150 tokens). A paragraph that alone exceeds --chunk-size falls back to a hard character-based sliding window — Paperless OCR text has no page-break markers (plan §1.2), so there's nothing else to split large, unbroken text on. A document with empty OCR content (the 26 empty_content documents from phase 2 step 5) yields zero chunks and is counted separately, not as an error.

Idempotency

A pre-fetched set of (envelope_id, chunk_index) pairs already embedded with --model skips re-embedding on rerun — no wasted Ollama calls. document_chunk's own UNIQUE (envelope_id, chunk_index) + ON CONFLICT DO NOTHING is the second line of defense; insert_chunk's command tag is checked so a silently-skipped row is counted as chunks_conflict_skipped, never miscounted as chunks_inserted. Note that uniqueness is on (envelope_id, chunk_index) only, not model — re-embedding with a different model hits this path and that embedding is discarded (wasted work, correctly reported via chunks_conflict_skipped, but not persisted). Out of scope for this single-model pilot; the real fix for whoever indexes a second model later is UNIQUE (envelope_id, chunk_index, model) at the schema layer.

A DB write failure for one chunk (dropped connection, unexpected bytes) is isolated the same way an embed failure is — counted as chunks_errors, never aborting the rest of the run.

Dimension guard

Every embedding response's length is checked against document_chunk.embedding's VECTOR(1024) column. A mismatch raises EmbeddingDimensionError and aborts the whole run immediately — never silently indexes vectors of the wrong dimension.

Chunk size/overlap validation

--chunk-overlap must be smaller than --chunk-size — the sliding-window hard-split fallback advances by chunk_size - chunk_overlap per step, so an overlap >= size would never advance and hang. main() rejects this combination before opening a DB connection; hard_split() itself also raises ValueError as a second line of defense for direct callers.

Stats must balance

documents_fetched = empty_content + documents_chunked
chunks_total       = chunks_already_embedded + chunks_inserted
                      + chunks_conflict_skipped + chunks_errors

main() exits 1 on chunks_errors > 0, chunks_conflict_skipped > 0, or if either balance breaks. The summary line also reports avg_embed_seconds_per_chunk — CPU-only Ollama timing, the input for deciding whether/how to scale this to the mail corpus later (plan §7).

Tests

pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest

Pure unit tests, no DB or real HTTP — run() is tested by monkeypatching asyncpg.connect (fake connection) and aiohttp.ClientSession (fake session serving a canned embedding vector, or a 500 for a chosen prompt to exercise error isolation). Covers: chunking (paragraph boundaries, overlap, empty document, document shorter than one chunk, oversized paragraph hard-fallback, the overlap-must-be-smaller-than-size guard), extract_content, idempotency (pre-existing keys skipped, no Ollama calls made for them, a second --apply run embeds nothing new, existing keys are correctly scoped to --model), dimension-mismatch abort, isolated per-chunk embed and insert errors, ON CONFLICT no-ops counted separately from real inserts, and the stats-balance invariant.

Known limitation — Ollama context-length rejections on pathological chunks

Ollama's runtime context window for a model can be smaller than the model's advertised max (bge-m3 supports 8192 tokens, but Ollama's default num_ctx is lower) — and some OCR text tokenizes far more densely than the ~4-chars/token heuristic this job uses to size chunks. Concretely: a table- of-contents page made almost entirely of dot-leader formatting (". . . . . . .", repeated hundreds of times) hit this on the pilot run — Ollama returned 500 {"error":"the input length exceeds the context length"} for one 2400-char chunk that should have been well within budget by character count alone. The job isolates this exactly like any other embed failure (chunks_errors, logged, run continues), so it never crashes a run — but it also never automatically shrinks and retries the offending chunk. Given how rare this was (1 chunk out of 2684 in the full pilot, all from one document's dot-leader ToC), it's left as a known gap rather than fixed here; a real fix would be either a smaller/adaptive chunk size for low-character-entropy text, or a shrink-and-retry loop on this specific Ollama error.

Definition of Done

Per CLAUDE.md: pytest passes locally (101 tests). Smoke-tested and then run to completion live on SOLARIA against the real Ollama instance and kb-postgres@PIHA:

  • Dry-run: 186 fetched, 26 empty_content, 2684 chunks planned — matches the known phase-2-step-5 figures exactly.
  • --apply --limit 10: 64 chunks embedded, 0 errors, avg ≈0.83s/chunk on CPU.
  • Re-run of the same slice: fully idempotent — 0 Ollama calls, 0 inserts.
  • Full --apply (all 186 documents): 2683/2684 chunks inserted, 1 isolated error (see "Known limitation" above) — chunks_errors=1 correctly produced a non-zero exit rather than silently reporting success. document_chunk ends at 2683 rows across 160 distinct envelopes, matching documents_chunked. A document_chunk_envelope_idx-backed count and an ORDER BY embedding <=> ... nearest-neighbor sanity query both look correct (top match is the reference chunk itself at distance 0; next nearest are chunks of the same source document).
  • Timing (CPU-only, no GPU driver on SOLARIA): ≈0.79s/chunk average across 2683 real embeddings (2115.8s total embed time), ≈13.2s/document average across the 160 chunked documents, ≈35 minutes wall-clock for the full 186-document pilot. This is the real-world input for scaling this pipeline to the much larger mail corpus later (plan §7 assumed GPU-based "minutes for the whole pilot"; SOLARIA's Ollama ran CPU-only for this pilot per the then-disabled GPU reservation). The 186-document pilot's ≈13.2s/document average is dominated by Paperless' long OCR text (≈22k chars/doc average, per plan §1.2) — 225 030 mail envelopes will have a very different, likely much shorter, per-envelope chunk count (email bodies vs. scanned multi-page PDFs), so this number doesn't extrapolate directly to a mail-corpus estimate. What it does establish: at ≈0.79s/chunk sequential CPU embedding, any corpus with a non-trivial average chunk count per item will need either a GPU driver fix, concurrent/batched Ollama calls, or both, before a full mail-corpus run is practical — flagged for whoever picks up the mail-indexer phase.
  • GPU (RTX 4070 Ti SUPER, driver 595-open, restored 2026-07-16): 207ms/embed (50 sekwencyjnych wywołań /api/embeddings, ~600-tok prompt) vs 790ms/chunk CPU baseline — ~3.8× szybciej sekwencyjnie; przy pojedynczych requestach dominuje overhead HTTP/tokenizacji, realny skok da dopiero batching (backlog).

Phase 3 step 4 — retrieval cascade (documents_ingest.retrieval) + quality gate

Module 5, phase 3, plan step 4 (kb/phases/kb-m5-faza3.md, §6). Two retrieval paths, both query_text -> chunk hits (dist, source) — the intended clean API surface for phase 4's kb-query, not just this eval:

  • flat_query — baseline: rank every active document_chunk row directly. Formalizes the phase-2 pilot's ad hoc /tmp/kbq.sh query into a tested module.
  • cascade_query — pre-filter to the top-N document_summary envelopes (one model, default claude-haiku-4-5 — plan §2 decision 3, resolved 2026-07-17) before ranking document_chunk within just those envelopes. Both share one query embedding call; the cascade only adds one extra SQL query (stage 1), never an extra Ollama call.

envelope, document_chunk, and document_summary are read-only — this module only ever SELECTs.

Quality gate

eval/queries.yaml — 7 queries transcribed 1:1 from the phase-2 pilot baseline (kb/phases/kb-m5-eval-retrieval-pilot.md, left untouched — this is its versioned working copy) with expected envelope / kind (hit, grey_zone, negative_control, negative_control_borderline) per query.

eval/retrieval_eval.py — read-only integration script against the live DB + live Ollama, not collected by pytest (same reasoning as the plan: an eval gate against live data isn't a mocked unit test). Runs every query through both tracks across an N sweep and checks the plan's three gate criteria (no flat hit degrades, hit@3 cascade ≥ flat, negative controls stay > 0.55). Exits 0 on PASS, 1 on FAIL.

pip install -e packages/kb-mail/ -e jobs/documents-ingest/
python eval/retrieval_eval.py --dsn postgresql://kb:<pw>@piha:5433/kb \
    --ollama-url http://solaria:11434 --n-sweep 5,10,20

--transport http --base-url http://<kb-query-host>:8230 (module 5 phase 4 plan §2 decision 6 / §9) calls a live kb-query's /search instead of embedding+querying locally — no --dsn needed, --n-sweep is ignored (kb-query serves one server-side default N per request). Gate criterion: dist must be identical to the same run with --transport direct against the same live SOLARIA (same DB, same retrieval code — HTTP is only a wrapper).

Result (2026-07-17, live run): PASS at N=10, k=5 — see plan §6.3 for the full table, the N-sweep calibration (N=5 is the measured safety floor; the plan's N=10 default carries a 2× margin), and the cost/improvement analysis. cascade_query (N=10, k=5, summary_model='claude-haiku-4-5') is now the default retrieval path for phase 4's kb-query; flat_query stays as the baseline/fallback.

Tests

pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest

tests/test_retrieval.py — pure unit tests, no DB or real HTTP. Covers: flat ranking across all envelopes, cascade stage-1-narrows-stage-2, an envelope whose summary exists but has no active chunks, N larger than the number of summarized envelopes, the no-summaries short-circuit (stage 2 never queried), and both query entry points embedding exactly once.


Phase 3 step 5 — cyclic ingest (documents-ingest-cyclic) + systemd timer

Module 5, phase 3, plan step 5 (kb/phases/kb-m5-faza3.md, §7). Orchestrates one run of the recurring ingest pipeline: paperless_adapter.run() (new source='paperless' envelopes) → chunk_embed.run() (new document_chunk rows) → summarize.run_summarize(backend='anthropic') (new document_summary rows, model='claude-haiku-4-5' — plan §2 decision 3) → summarize.run_embed_summaries() (embeds those summaries). All four are the same job functions used elsewhere in this package, called directly — no changes to paperless_adapter.py / chunk_embed.py / summarize.py, no new CLI flags on them.

Ollama-offline tolerance

SOLARIA has availability_target: medium (planned power-off, plan §1.3). The wrapper probes GET {OLLAMA_URL}/api/tags before the two embed stages (chunk embedding, summary embedding); unreachable means skip, not fail — both embed passes are idempotent, so new chunks/summaries left unembedded this tick are picked up whole on the next one. A growing backlog is what kb_ingest_embed_backlog + the KbEmbedBacklogGrowing alert are for, not this wrapper's exit code.

Anything else failing is a hard failure: Paperless unreachable, a DB error, a non-zero job error counter, a broken stats-balance invariant, the Anthropic API failing. Each stage's pass/fail predicate mirrors that job's own main() exit check 1:1 (see cyclic_ingest.py's module docstring). Stages are isolated, not fail-fast — an earlier stage failing never skips a later one, mirroring the per-row isolation the underlying jobs already use.

Usage

# Dry run (default) — same idempotent counting as every other job in this family, no writes:
documents-ingest-cyclic --dsn postgresql://kb:<pw>@localhost:5433/kb \
    --paperless-token <token> --anthropic-api-key <key>

# Real run (what the timer invokes):
documents-ingest-cyclic --dsn ... --paperless-token ... --anthropic-api-key ... --apply

--dsn/--paperless-token/--anthropic-api-key also read from KB_DSN / PAPERLESS_API_TOKEN / ANTHROPIC_API_KEY env vars — never logged. --ollama-url defaults to http://solaria:11434 (this wrapper always runs on PIHA, unlike chunk_embed/summarize's own CLI defaults which assume co-location with Ollama).

Metrics (Prometheus textfile collector)

Every run — success or failure — writes --prom-path (default /opt/homelab/state/node-exporter/kb-ingest.prom) atomically (tmp + rename):

Metric Meaning
kb_ingest_last_run_timestamp Unix ts of the last run, success or failure
kb_ingest_last_success_timestamp Unix ts of the last run with no hard failure — carried forward from the previous file on a failing run, never reset to 0/now
kb_ingest_last_exit_code 0 or 1
kb_ingest_documents_inserted New envelope rows this run
kb_ingest_chunks_inserted New document_chunk rows this run (0 if the embed stage was skipped)
kb_ingest_summaries_inserted New document_summary rows this run
kb_ingest_embed_skipped 1 if Ollama was unreachable this run (both embed stages skipped), 0 otherwise
kb_ingest_embed_backlog Active chunks (excluded_reason IS NULL) still missing an embedding

Scraped by fleet-prometheus via node_exporter's textfile collector on PIHA (hosts/piha/runtime/node_exporter/docker-compose.override.yml); alert rules in services/fleet-prometheus/rules/kb-ingest.yml.

Install (PIHA)

  1. Dedicated venv (per plan §7.1 — not the ad hoc rsync-to-/tmp pattern used for the one-shot jobs elsewhere in this README; this is a permanent, recurring installation):
    python3 -m venv /opt/homelab/kb/venv
    /opt/homelab/kb/venv/bin/pip install -e packages/kb-mail -e jobs/documents-ingest
    
    (run from a checkout of this repo on PIHA — the checkout is used as an install source only, per CLAUDE.md's "deploy-only" rule; no development happens there).
  2. Secrets in /opt/homelab/kb/.env (already holds PAPERLESS_API_TOKEN; add KB_DSN=postgresql://kb:<pw>@localhost:5433/kb and ANTHROPIC_API_KEY=<key>), chmod 600, never in Git.
  3. Copy jobs/documents-ingest/systemd/kb-ingest-run.sh to /opt/homelab/kb/ and chmod +x it.
  4. Copy (or symlink) kb-ingest.service and kb-ingest.timer to /etc/systemd/system/, then:
    systemctl daemon-reload
    systemctl enable --now kb-ingest.timer
    
  5. Verify: systemctl list-timers kb-ingest.timer, journalctl -u kb-ingest.service, /opt/homelab/logs/kb-ingest/run-YYYYMMDD.log, and /opt/homelab/state/node-exporter/kb-ingest.prom after the first run (manual systemctl start kb-ingest.service to trigger one immediately without waiting for 03:30).

Tests

pip install -e packages/kb-mail/
pip install -e jobs/documents-ingest/
cd jobs/documents-ingest && pytest

tests/test_cyclic_ingest.py — pure unit tests, no DB or real HTTP/Ollama/Anthropic; every stage function and the Ollama probe are monkeypatched. Covers: each stage's failure predicate (pinned against its source job's own exit check), the Ollama-down skip path (chunk_embed/embed_summaries never even called), stage isolation (an earlier stage failing never skips a later one, whether via a failed predicate or a raised exception), .prom rendering, atomic write, last_success_timestamp carry-forward across a failing run, and main()'s CLI guardrails + exit-code propagation.