Every RAG demo starts with clean text. Every real client starts with 200 PDFs, half of them scanned sideways. The copilot you're about to build is only as good as the pipeline that feeds it.

You'll need: Python 3.10+, pip install pypdf reportlab beautifulsoup4. No API keys — everything in this lesson runs locally, for real.

The customer problem

Dev, sliding a USB stick across the table: "Here are 200 agency procedure PDFs. Sanitation, permits, noise complaints — everything the 'Ask CityOps' copilot needs to know."

You, opening the first file: a scan of a photocopy, slightly rotated. The second: headers and footers on every page, a table split across a page break. The third: won't open at all.

Dev: "So… can we just, like, upload them to the AI?"

No. You cannot "just upload them." Between a folder of PDFs and a copilot that answers questions stands an entire system most tutorials skip: the ingestion pipeline. Extract the text. Clean it. Chop it into chunks. Tag every chunk with metadata so the answer can cite its source. Get any of these wrong and your copilot doesn't just fail — it fails confidently, quoting half a table or a footer as if it were procedure.

This lesson builds that pipeline against a realistic sample document, with every output below produced by a real run.

Clarify the ask

Before writing code, pin down what "searchable" has to mean for the copilot:

  • Chunk boundaries should preserve enough local context that retrieved chunks remain interpretable — while allowing retrieval to combine multiple chunks when an answer genuinely needs them. A chunk that says "escalate to them within 48 hours" without saying who is a wrong answer waiting to happen; but don't bloat chunks trying to make each one independently answerable.
  • Every chunk must carry provenance. Source file, page number, section — because the copilot will cite sources, and "trust me" is not a citation.
  • Unreadable documents must be reported, not silently dropped. If 12 of 200 PDFs are scans, Dev needs a list, not a copilot that quietly knows nothing about weekend coverage.
  • Re-running ingestion must be safe. When procedure #47 gets revised, the new version must replace the old chunks — re-ingesting without duplicating the other 199, and without leaving stale policy searchable.

That's the contract. Now the minimum concept.

The minimum concept

The ingestion pipeline

  • Extract — pull raw text out of each format (PDF, HTML, DOCX). Different formats, different tools, different failure modes.
  • Clean — remove what isn't content: headers, footers, page numbers, navigation chrome, hyphenation artifacts, mojibake.
  • Chunk — split the cleaned text into retrievable units. The chunk is the atomic unit of your search index; its boundaries decide what the model ever gets to see.
  • Tag — attach metadata to every chunk: source, page, section, document hash. Metadata is what turns "a similar paragraph" into "Section 3 of the sanitation procedure, page 1."
  • Validate — check the output before it reaches the index. Did we extract anything? Did expected pages disappear? Are chunks empty? Does every chunk carry provenance? Did every source land in exactly one outcome — indexed, OCR candidate, quarantined, or skipped? Validation is what makes "nothing silently dropped" a checkable claim instead of a hope.
flowchart TD A[200 messy PDFs] --> B[Extract: pypdf / OCR gate] B --> C{Clean: headers, footers, hyphenation} C --> D[Chunk: fixed / overlap / section-aware] D --> E[Tag: source, page, section, hash] E --> V[Validate: provenance, page counts, reconciliation] V --> F[(Search index)] B -->|no text layer| G[OCR queue] B -->|unreadable| H[Quarantine + report]

The uncomfortable truth this diagram hides: retrieval quality is bounded above by ingestion quality. Information destroyed during ingestion can't reliably be recovered later — no embedding model, no reranker, no clever prompt can be counted on to reconstruct a table of escalation contacts that your chunker split in half. The unglamorous work is the work.

Build: from PDF to clean text

Step 1 — A realistic sample document

First, a sample that behaves like Dev's real files: headers and footers on every page, numbered sections, and a table. We generate it with reportlab so you can reproduce this exactly:

from reportlab.lib.pagesizes import letter
from reportlab.platypus import (SimpleDocTemplate, Paragraph, Spacer,
                                Table, TableStyle, PageBreak)
from reportlab.lib.styles import getSampleStyleSheet
from reportlab.lib import colors

HDR = "CITY OF NEW YORK - DEPARTMENT OF SANITATION - INTERNAL USE ONLY"

def header_footer(canvas, doc):
    canvas.saveState()
    canvas.setFont("Helvetica", 8)
    canvas.drawString(72, 750, HDR)
    canvas.drawString(72, 40, f"Page {doc.page} - Confidential - Do Not Distribute")
    canvas.restoreState()

styles = getSampleStyleSheet()
story = [Paragraph("Sanitation Complaint Handling Procedure", styles["Title"]),
         Spacer(1, 12),
         Paragraph("1. INTAKE", styles["Heading2"]),
         Paragraph("All 311 sanitation complaints enter through the central intake "
                   "queue. The intake operator verifies the borough and the complaint "
                   "category before the request is routed.", styles["Normal"]),
         Paragraph("2. CATEGORIZATION", styles["Heading2"]),
         Paragraph("Each complaint is assigned one category: missed pickup, "
                   "overflowing basket, illegal dumping, or street cleaning.",
                   styles["Normal"]),
         Paragraph("3. ESCALATION CONTACTS", styles["Heading2"]),
         Paragraph("If a complaint is unresolved after 48 hours, escalate:", styles["Normal"])]
data = [["Borough", "Escalation contact", "Phone"],
        ["Manhattan", "M. Okafor", "555-0101"],
        ["Brooklyn", "J. Rivera", "555-0102"],
        ["Queens", "S. Patel", "555-0103"]]
t = Table(data)
t.setStyle(TableStyle([("GRID", (0, 0), (-1, -1), 0.5, colors.grey),
                       ("BACKGROUND", (0, 0), (-1, 0), colors.lightgrey)]))
story += [t, PageBreak(),
          Paragraph("4. CLOSURE", styles["Heading2"]),
          Paragraph("A request is closed only after the crew confirms the work "
                    "is complete and the resident is notified.", styles["Normal"])]

SimpleDocTemplate("sanitation_procedure.pdf", pagesize=letter).build(
    story, onFirstPage=header_footer, onLaterPages=header_footer)

Step 2 — Extract, and look at what you actually got

from pypdf import PdfReader

reader = PdfReader("sanitation_procedure.pdf")
pages = [p.extract_text() or "" for p in reader.pages]
print("pages:", len(pages))
print(repr(pages[0][:300]))
pages: 2
'CITY OF NEW YORK - DEPARTMENT OF SANITATION - INTERNAL USE ONLY\nPage 1 - Confidential - Do Not Distribute\n Sanitation Complaint Handling Procedure\n1. INTAKE\nAll 311 sanitation complaints enter through the central intake queue. The intake operator verifies the\nborough and the complaint category before the request is routed...'

There's your first lesson, free of charge: the header and footer are part of the "text" now. Index this raw and every chunk in your system will contain "CITY OF NEW YORK - DEPARTMENT OF SANITATION - INTERNAL USE ONLY" — pure noise that dilutes every search. (This is the Files lesson's messy-real-world discipline, applied to AI.) One provenance nuance for later: the page numbers here are the PDF's own page index, and the printed page label inside the document isn't always the same — cover pages and appendices shift them. If citation accuracy matters, preserve whichever page identifiers your source exposes.

Step 3 — Clean: evict the boilerplate

Headers repeat on every page — so detect repetition instead of hardcoding strings. But watch what happens:

import re
from collections import Counter

def strip_boilerplate(pages):
    # v1: drop lines that appear on EVERY page
    line_pages = Counter()
    for p in pages:
        for line in set(p.split("\n")):
            if line.strip():
                line_pages[line.strip()] += 1
    n = len(pages)
    boilerplate = {l for l, c in line_pages.items() if c >= n and n > 1}
    print("detected boilerplate:", boilerplate)
    return ["\n".join(l for l in p.split("\n")
                       if l.strip() not in boilerplate) for p in pages]

cleaned = strip_boilerplate(pages)
detected boilerplate: {'CITY OF NEW YORK - DEPARTMENT OF SANITATION - INTERNAL USE ONLY'}

It caught the header — but missed the footer. Why? The footer contains the page number, so "Page 1 - Confidential…" and "Page 2 - Confidential…" are different strings. Exact-repeat detection can't see the pattern. This is the kind of thing you only learn by looking at real output:

def strip_boilerplate(pages):
    # v2: ...plus a pattern for numbered footers
    line_pages = Counter()
    for p in pages:
        for line in set(x.strip() for x in p.split("\n") if x.strip()):
            line_pages[line] += 1
    n = len(pages)
    boilerplate = {l for l, c in line_pages.items() if c >= n and n > 1}
    page_num = re.compile(r"^Page \d+\b")
    return ["\n".join(l for l in p.split("\n")
                      if l.strip() and l.strip() not in boilerplate
                      and not page_num.match(l.strip()))
            for p in pages]

def fix_hyphenation(text):
    # "com-\nplaint" -> "complaint" (line-break hyphens from PDF layout)
    return re.sub(r"(\w)-\n(\w)", r"\1\2", text)

def normalize(text):
    return re.sub(r"\s+", " ", text.replace("\n", " ")).strip()

doc = "\n\n".join(normalize(fix_hyphenation(p)) for p in strip_boilerplate(pages))
print("chars:", len(doc))
print("contains 'Confidential'?", "Confidential" in doc)
print(doc[:280])
chars: 653
contains 'Confidential'? False
Sanitation Complaint Handling Procedure 1. INTAKE All 311 sanitation complaints enter through the central intake queue. The intake operator verifies the borough and the complaint category before the request is routed. 2. CATEGORIZATION Each complaint is assigned one category: mis

Don't memorize the regex — memorize the move

The specific patterns (repeat-lines, Page \d+) will differ for every client. The move is: extract, print the raw text, read it like a detective, then write the cleaning rule the evidence demands. Cleaning is never "done" in one pass — you iterate against samples of the real corpus.

Two warnings that come with the move. First, v2 above still demands a line appear on every page — a cover page without the header defeats it. Production detectors look for frequent positional repetition, not necessarily 100%. Second, repeated does not automatically mean boilerplate: a legal instruction repeated on every page can be real content. Frequency is one signal; position and document structure are the others. And remember that cleaning rules are transformations that can destroy information too (after-\nhours → afterhours) — test them against samples, don't bless them.

Step 4 — HTML gets the same treatment

Agency procedures also live on intranet pages, which bring their own boilerplate — nav bars, scripts, footers. For simple HTML, DOM-aware extraction plus removal of known non-content elements is a good starting point (real article pages may need main-content selection or readability heuristics — that's a later lesson):

from bs4 import BeautifulSoup

html = """<html><head><title>DSNY Procedure Update</title>
<style>.x{color:red}</style><script>track();</script></head><body>
<nav>Home | About | Contact</nav>
<h1>Holiday Pickup Schedule</h1>
<p>There is <b>no pickup</b> on Thanksgiving &amp; Christmas.</p>
<footer>Copyright 2026 - Page generated 10:42:11</footer>
</body></html>"""

soup = BeautifulSoup(html, "html.parser")
for tag in soup(["script", "style", "nav", "footer"]):
    tag.decompose()
text = re.sub(r"\s+", " ", soup.get_text(separator=" ")).strip()
print(repr(text))
'DSNY Procedure Update Holiday Pickup Schedule There is no pickup on Thanksgiving & Christmas.'

Scripts, styles, nav, and the timestamped footer are gone; the &amp; entity decoded properly. What remains is content. That's the whole job.

Build continued: chunking — where retrieval is won or lost

The same cleaned document, chunked three ways. Read the outputs, not just the code. One labeling note first: the size=400 below is characters, chosen to keep the example inspectable — production chunk budgets are commonly based on tokens and document structure, and 400 characters is not a recommended RAG chunk size. Chunk size and overlap are retrieval parameters to evaluate on your corpus, not constants to copy (your vector-search lesson already taught you why: measure, don't guess).

def chunk_fixed(text, size=400):
    return [text[i:i+size] for i in range(0, len(text), size)]

def chunk_overlap(text, size=400, overlap=100):
    chunks, i = [], 0
    while i < len(text):
        chunks.append(text[i:i+size])
        i += size - overlap
    return chunks

def chunk_sections(text):
    # split on numbered section headings: "1. INTAKE", "2. CATEGORIZATION", ...
    parts = re.split(r"(?=\d+\.\s+[A-Z][A-Z ]+$|\d+\.\s+[A-Z][A-Z ]+ )", text)
    return [p.strip() for p in parts if p.strip()]

for name, fn in [("fixed-400", chunk_fixed),
                 ("overlap-400/100", chunk_overlap),
                 ("section-aware", chunk_sections)]:
    ch = fn(doc)
    print(f"--- {name}: {len(ch)} chunks ---")
    for c in ch[:2]:
        print(" ", c[:100], "...")
--- fixed-400: 2 chunks ---
  Sanitation Complaint Handling Procedure 1. INTAKE All 311 sanitation complaints enter through the ce ...
   after 48 hours, escalate: Borough Escalation contact Phone Manhattan M. Okafor 555-0101 Brooklyn J. ...
--- overlap-400/100: 3 chunks ---
  Sanitation Complaint Handling Procedure 1. INTAKE All 311 sanitation complaints enter through the ce ...
  ing basket, illegal dumping, or street cleaning. 3. ESCALATION CONTACTS If a complaint is unresolved ...
--- section-aware: 5 chunks ---
  Sanitation Complaint Handling Procedure ...
  1. INTAKE All 311 sanitation complaints enter through the central intake queue. The intake operator  ...
StrategyWhat it buys youWhat it costs you
Fixed-sizeDead simple; uniform chunk sizes keep the index predictableCuts mid-sentence and mid-section — chunk 2 above starts mid-thought
OverlappingBoundary context survives; fewer "the answer was split across two chunks" misses~25% more storage and indexing; overlap is still arbitrary
Section-awareChunks align with the document's own structure — each one carries its section's local contextNeeds detectable structure; a scanned memo with no headings defeats it

The rule of thumb

Prefer structure-aware chunking whenever the documents have structure; fall back to overlapping fixed-size when they don't. And whatever you choose, read a sample of the chunks. The chunk list is the actual product — the model will never see anything else.

One more principle, and it's the load-bearing one for the pipeline below: cleaning should remove noise without erasing structure. Line breaks encode headings, paragraphs, list items, table rows — structure the chunker needs. So detect sections on the structured text, then chunk; don't flatten everything to one line during cleaning and try to rediscover the structure during chunking. (Notice the tension above: normalize() already collapsed the line breaks that chunk_sections' regex is sniffing for. Fine for a demo — fragile for production.)

Related: page boundaries are provenance, not necessarily semantic boundaries. A sentence split across a page break shouldn't be cut in half just because the PDF paginated there — if a chunk spans pages, store a page range, not a single page.

Break: three ways ingestion silently ruins retrieval

1. The chunker that split the table

Remember the escalation-contacts table? Watch a 120-character fixed chunker do surgery on it:

table = "Borough | Contact | Phone\n" + "\n".join(
    f"{b} | {n} | {p}" for b, n, p in
    [("Manhattan", "M. Okafor", "555-0101"), ("Brooklyn", "J. Rivera", "555-0102"),
     ("Queens", "S. Patel", "555-0103"), ("Bronx", "T. Nguyen", "555-0104"),
     ("Staten Island", "A. Kowalski", "555-0105")])

for c in chunk_fixed(table, 120):
    print("--- chunk ---")
    print(c)
--- chunk ---
Borough | Contact | Phone
Manhattan | M. Okafor | 555-0101
Brooklyn | J. Rivera | 555-0102
Queens | S. Patel | 555-0103

--- chunk ---
Bronx | T. Nguyen | 555-0104
Staten Island | A. Kowalski | 555-0105

A resident asks "who do I escalate a Bronx complaint to?" Retrieval returns chunk 2 — which has no header row. The model receives ambiguous context and may misinterpret the fields. Preserve table structure: small tables may stay whole; large tables may need row-group chunking with the header repeated on each chunk (header + Manhattan/Brooklyn/Queens, then header + Bronx/Staten Island) — or structured storage. Never split blindly.

2. The headers that became "knowledge"

Skip cleaning and your first chunk looks like this — real output from the raw extraction:

'CITY OF NEW YORK - DEPARTMENT OF SANITATION - INTERNAL USE ONLY\nPage 1 - Confidential - Do Not Distribute\n Sanitation Complaint Handling Procedure\n1. INTAKE\n...'

Now imagine the copilot citing "Page 1 - Confidential - Do Not Distribute" as part of its answer about intake procedure. It happens. Boilerplate in the index is boilerplate in the answers.

3. The scanned PDF that yielded nothing

Some of Dev's PDFs are scans — photographs of pages, no text layer at all. extract_text() on those returns an honest empty string:

scanned = PdfReader("scanned_form.pdf").pages[0].extract_text()
print(repr(scanned))
''

Empty string is not an error — it's a classification: no usable text extracted — inspect / OCR candidate. Don't conclusively call it "a scan": empty extraction can also mean a blank page, an extraction limitation, or an unusual encoding. This document doesn't need a better parser; it likely needs OCR (a separate tool, a separate queue, a separate conversation with Dev about cost). And the corrupt file?

PdfReader("corrupt.pdf")
# PdfStreamError: Stream has ended unexpectedly

That one is an error — quarantine it, log it, report it. Two different failure modes, two different destinations. Never mix them. And one more gate the pipeline below implements: this check has to run per page. A 12-page PDF with 11 text pages and one scanned form must not silently lose page 9 — the pipeline reports "document has 12 pages, 11 extracted, 1 requires OCR" instead of dropping the page or the document.

✗ Myth: "Ingestion is a solved problem — just call a parser."

✓ Reality: Parsing is often the easy part. Deciding what counts as content, keeping tables whole, detecting scans, quarantining corruption, and tagging provenance — that's the system. The parser is a function call; the pipeline is engineering.

Productionize: the IngestionPipeline

Now assemble it: per-document error handling, the per-page OCR gate, section-aware metadata on every chunk, validation before anything hits the index — and versioning done right. Two concepts carry the design: document identity (which document is this, stable across renames and edits) is separate from content hash (have this document's bytes changed). The principle: a hash tells you that content changed; a stable document identity tells you what that new content replaces. Same document + same hash → skip. Same document + new hash → ingest the new version and retire the old chunks, or retrieval returns both the old and the new policy. (Production systems take identity from a source-system ID, canonical URL, or document-registry ID — filenames get renamed, so the toy below uses the filename stem and says so.)

import hashlib
import re
from collections import Counter
from pypdf import PdfReader

class IngestionPipeline:
    """Toy ingestion pipeline. Production swaps the in-memory dicts for a
    real store, but the stage boundaries are the same:
    extract -> clean -> chunk -> tag -> validate -> index."""

    def __init__(self, chunk_size=400):
        self.chunk_size = chunk_size
        # ---- persistent index state (survives across runs) ----
        self.index = {}     # document_id -> {"content_hash", "source"}
        self.records = []   # every live chunk: what gets embedded + indexed
        self._live = {}     # chunk_key -> record (for version replacement)

    # ---------- identity ----------
    def _document_id(self, path):
        # Toy identity: filename stem. Production needs a STABLE identifier
        # separate from both filename and content hash: a source-system ID,
        # canonical URL, or document-registry ID. Filenames get renamed;
        # hashes change on every edit — neither one is an identity.
        return path.split("/")[-1].rsplit(".", 1)[0]

    def _hash(self, path):
        h = hashlib.sha256()
        with open(path, "rb") as f:
            for block in iter(lambda: f.read(65536), b""):  # stream; don't
                h.update(block)                             # load whole file
        return h.hexdigest()  # full hash stored; truncate only for display

    # ---------- extract ----------
    def _extract(self, path):
        # Per-page extraction. The page number recorded is the PDF's own
        # 1-based index — never renumber after filtering (an empty page 2
        # must not make page 3 become "page 2"). Note: the PDF page index
        # and the printed page label aren't always the same (cover pages,
        # appendices); preserve whichever identifiers your source exposes.
        reader = PdfReader(path)
        return [(i + 1, p.extract_text() or "") for i, p in enumerate(reader.pages)]

    # ---------- clean (structure-preserving) ----------
    def _clean(self, pages):
        """pages: [(page_no, raw_text)] -> [(page_no, cleaned_text)] for pages
        with usable text. Cleaning removes noise (boilerplate lines,
        hyphenation artifacts) but preserves line breaks — headings,
        paragraphs and table rows are structure the chunker needs."""
        line_pages = Counter()
        for _, t in pages:
            for line in set(x.strip() for x in t.split("\n") if x.strip()):
                line_pages[line] += 1
        n = len(pages)
        # frequent positional repetition, not necessarily 100%: a cover page
        # without the header shouldn't save the header. Tune per corpus.
        boilerplate = {l for l, c in line_pages.items()
                       if c >= max(2, int(n * 0.8))}
        pgnum = re.compile(r"^Page \d+\b")
        out = []
        for page_no, t in pages:
            kept = [l for l in (x.strip() for x in t.split("\n"))
                    if l and l not in boilerplate and not pgnum.match(l)]
            # Repeated does not automatically mean boilerplate — a repeated
            # legal instruction can be real content. Frequency is one signal;
            # position and document structure are the others. Inspect samples.
            text = "\n".join(kept)
            text = re.sub(r"(\w)-\n(\w)", r"\1\2", text)  # hyphenation repair
            # ^ cleaning rules are transformations that can destroy information
            # too (after-\nhours -> afterhours). Test them against samples.
            if text.strip():
                out.append((page_no, text))
        return out

    # ---------- chunk + tag ----------
    SECTION_RE = re.compile(r"^(\d+)\.\s+([A-Z][A-Z \-]+)$", re.M)

    def _chunk_page(self, doc_id, digest, source, page_no, text):
        # Detect sections on the STRUCTURED text, before chunking — don't
        # destroy structure during cleaning and rediscover it later.
        spans = [(m.start(), m.group(0))
                 for m in self.SECTION_RE.finditer(text)]
        chunks, i, n = [], 0, 0
        while i < len(text):
            c = text[i:i + self.chunk_size]
            section = None
            for s_off, s_title in spans:
                if s_off <= i:
                    section = s_title
            flat = re.sub(r"\s+", " ", c).strip()
            if flat:
                chunks.append({
                    "chunk_key": f"{doc_id}:p{page_no}:{n}",
                    "document_id": doc_id,
                    "source": source,
                    "content_hash": digest,
                    "page": page_no,
                    "section": section,   # None when no detectable structure
                    "text": flat,
                    
                })
                n += 1
            i += self.chunk_size
        return chunks

    REQUIRED = ("chunk_key", "document_id", "source",
                "content_hash", "page", "text")

    def _validate(self, records):
        # metadata invariant, checked BEFORE anything hits the index.
        # Section is optional when the document has no detectable structure —
        # validate what's required, don't invent what's absent.
        for r in records:
            missing = [k for k in self.REQUIRED if not r.get(k)]
            if missing or not r["text"].strip():
                raise ValueError(f"chunk failed validation: {missing or ['empty text']}")

    def _retire(self, doc_id):
        # version replacement: the old chunks must go, or retrieval returns
        # both the old and the new policy. Atomic here; transactional in prod.
        self._live = {k: v for k, v in self._live.items()
                      if not k.startswith(doc_id + ":")}
        self.records = [r for r in self.records
                        if r["document_id"] != doc_id]

    def _quarantine(self, run, doc_id, path, stage, exc):
        run["quarantined"].append({
            "document_id": doc_id, "source": path.split("/")[-1],
            "stage": stage, "error": type(exc).__name__,
            "detail": str(exc)[:160]})

    # ---------- run ----------
    def ingest(self, paths):
        # run-scoped metrics: fresh every run. self.index / self.records are
        # the persistent store — mixing the two makes the report lie.
        run = {"indexed": [], "replaced": [], "skipped": [], "ocr": [],
               "quarantined": [], "chunks": 0, "pages": []}
        for path in paths:
            doc_id = self._document_id(path)
            source = path.split("/")[-1]
            digest = self._hash(path)
            prev = self.index.get(doc_id)
            if prev and prev["content_hash"] == digest:
                run["skipped"].append(doc_id)   # same document, same content
                continue
            try:
                pages = self._extract(path)
            except Exception as e:              # per-document isolation:
                self._quarantine(run, doc_id, path, "extract", e)
                continue                        # one bad file never kills the run
            text_pages = [(pn, t) for pn, t in pages if t.strip()]
            no_text = [pn for pn, t in pages if not t.strip()]
            # per-page scan gate: a mixed PDF must not silently lose pages.
            # "No usable text" is a classification, not a diagnosis — it can
            # mean a scan, a blank page, or an encoding the extractor can't
            # read. Evidence before diagnosis.
            run["pages"].append((doc_id, len(pages), len(text_pages), no_text))
            if not text_pages:
                run["ocr"].append((doc_id, no_text))
                continue
            new_records = []
            for page_no, ctext in self._clean(text_pages):
                new_records += self._chunk_page(doc_id, digest, source,
                                                page_no, ctext)
            self._validate(new_records)
            if prev:                            # same document, NEW content:
                self._retire(doc_id)            # replace, don't duplicate
                run["replaced"].append(doc_id)
            for r in new_records:
                self._live[r["chunk_key"]] = r
                self.records.append(r)
            self.index[doc_id] = {"content_hash": digest, "source": source}
            run["indexed"].append(doc_id)
            run["chunks"] += len(new_records)
        return self._report(run, len(paths))

    def _report(self, run, n_sources):
        # reconciliation: every source lands in exactly one outcome.
        accounted = (len(run["indexed"]) + len(run["ocr"])
                     + len(run["quarantined"]) + len(run["skipped"]))
        L = ["INGESTION REPORT",
             f"  sources reconciled : {accounted}/{n_sources} "
             f"(indexed {len(run['indexed'])}, ocr-candidate {len(run['ocr'])}, "
             f"quarantined {len(run['quarantined'])}, "
             f"skipped-unchanged {len(run['skipped'])})",
             f"  chunks indexed     : {run['chunks']} (this run; "
             f"{len(self.records)} live in index)"]
        for doc_id, total, got, missing in run["pages"]:
            if missing and got:
                L.append(f"    - {doc_id}: {got}/{total} pages with text; "
                         f"pages {missing} no usable text — OCR candidate")
        for doc_id, pages_ in run["ocr"]:
            L.append(f"    - {doc_id}: no usable text on pages {pages_} — "
                     f"inspect / OCR candidate")
        for q in run["quarantined"]:
            L.append(f"    - {q['source']}: quarantined at {q['stage']} "
                     f"({q['error']}: {q['detail']})")
        if run["replaced"]:
            L.append(f"  replaced versions  : {', '.join(run['replaced'])} "
                     f"(old chunks retired)")
        return "\n".join(L)

Run it against a three-file corpus — one clean PDF, one scan, one corrupt file:

pipe = IngestionPipeline()
print(pipe.ingest(["sanitation_procedure.pdf", "scanned_form.pdf", "corrupt.pdf"]))
print()
r0 = pipe.records[0]
print("sample:", r0["chunk_key"], "| page", r0["page"],
      "| section", repr(r0["section"]), "| sha256 stored (full, 64 hex)")
print()
print("=== second run: nothing changed, nothing duplicated ===")
print(pipe.ingest(["sanitation_procedure.pdf"]))
INGESTION REPORT
  sources reconciled : 3/3 (indexed 1, ocr-candidate 1, quarantined 1, skipped-unchanged 0)
  chunks indexed     : 3 (this run; 3 live in index)
    - scanned_form: no usable text on pages [1] — inspect / OCR candidate
    - corrupt.pdf: quarantined at extract (PdfStreamError: Stream has ended unexpectedly)

sample: sanitation_procedure:p1:0 | page 1 | section None | sha256 stored (full, 64 hex)

=== second run: nothing changed, nothing duplicated ===
INGESTION REPORT
  sources reconciled : 1/1 (indexed 0, ocr-candidate 0, quarantined 0, skipped-unchanged 1)
  chunks indexed     : 0 (this run; 3 live in index)

Note what the report separates: this run's metrics (chunks indexed: 0) from persistent state (3 live in index). Mixing the two is how re-runs start lying. And the reconciliation line — indexed + OCR + quarantined + skipped = sources — is the check that nothing silently disappeared.

Now the critical test — Dev revises procedure #47. Same filename, new content. (Reproduce it: re-run the Step-1 generator with a 5. WEEKEND COVERAGE section added before the PageBreak.) Watch what the pipeline does with the old chunks:

print("=== third run: procedure revised — new version replaces old ===")
print(pipe.ingest(["sanitation_procedure.pdf"]))
print("live chunks:", len(pipe.records))
=== third run: procedure revised — new version replaces old ===
INGESTION REPORT
  sources reconciled : 1/1 (indexed 1, ocr-candidate 0, quarantined 0, skipped-unchanged 0)
  chunks indexed     : 3 (this run; 3 live in index)
  replaced versions  : sanitation_procedure (old chunks retired)
live chunks: 3

Three live chunks, not six. The new version's chunks replaced the old ones — hash-only "idempotency" would have indexed both and let retrieval return superseded policy next to current policy. That's the difference between skipping the seen and versioning the known.

Three design decisions worth naming, because you'll defend each one to a client:

  • Quarantine, don't crash. One corrupt file must never kill a 200-document run — and the quarantine entry carries the document ID, the stage, and the error, so Dev can actually investigate. (The Debugging lesson's quarantine philosophy, applied to documents — broad except is acceptable at this isolation boundary because the diagnostics are preserved.)
  • Separate document identity from content hash. Same document + same hash → skip. Same document + new hash → replace the old chunks, don't just add new ones. A hash tells you content changed; identity tells you what it replaces.
  • Store the chunk text alongside the vector. Embeddings get you to the chunk; the text is what the model actually reads. A vector without its text and metadata is an address with no house.

Useful later

Real OCR (Tesseract, cloud vision APIs), DOCX/Excel extraction, near-duplicate detection across the corpus, and page-range chunks for content that spans page boundaries. You don't need them for the milestone — but the pipeline above has a labeled place for each.

Communicate: the report Dev actually needs

Dev doesn't want to hear about your regexes. Dev wants to know what the copilot knows and what it doesn't. The shape of the report comes straight from the real run above — scaled here to 200 documents to show what it looks like at full size (the 200-document numbers below are illustrative, not from the 3-file run):

Subject: Ingestion report — agency procedure corpus (run 2026-10-03)

200 documents processed: 184 fully indexed (6,412 chunks), 12 need OCR (scanned forms, listed below — I need a call on whether we OCR these or you re-supply them), 4 quarantined (2 corrupt files, 2 password-protected — I don't have the passwords).

Every chunk carries source file and page (plus section where the document has detectable structure), so the copilot can cite where each answer comes from. Re-running ingestion is safe — unchanged files are skipped by content hash, and revised files replace their old chunks instead of duplicating them.

The 12 OCR-needed files are attached as a list. Reconciliation check: 184 + 12 + 4 + 0 skipped = 200 — every source landed in exactly one outcome. Nothing was silently dropped: if a document isn't in the index, it's in this report with a reason.

That last sentence is the whole ethic. "Validate at the boundary; trust inside" — the Pydantic lesson's principle — applies to documents too. Ingestion is the boundary between the client's messy world and your clean index.

CityOps: the corpus becomes the knowledge base

Those 184 indexed procedures — intake rules, escalation contacts, weekend coverage — are the approved knowledge corpus the "Ask CityOps" copilot will reason over. When it answers "who handles a Bronx escalation after 48 hours," the system can retrieve and cite the approved source chunk (source: sanitation_procedure.pdf, page: 1, section: 3. ESCALATION CONTACTS) instead of relying solely on the model's prior knowledge.

Next, that corpus gets embedded and searched — but the search will only ever be as honest as this pipeline. You just built the foundation everything else stands on.

Field check

  1. Why is boilerplate (headers, footers, nav bars) dangerous in a search index, beyond just wasting space?
  2. Your repeat-line detector misses footers containing page numbers. What's the general fix, and why can't exact matching ever be enough?
  3. A 120-character fixed chunker splits the escalation table between Queens and Bronx. Name two concrete harms this causes at answer time.
  4. extract_text() returns "" for a scanned PDF and raises PdfStreamError for a corrupt one. Why do these go to different queues?
  5. What does hashing the source file buy you on re-ingestion day, when procedure #47 is revised and the other 199 aren't?
Answer 1

Boilerplate repeats in every chunk, so it dilutes similarity scores — every chunk looks a little bit like every other chunk — and worse, the model can quote it as if it were content, e.g. citing "Page 1 - Confidential - Do Not Distribute" inside an answer about intake procedure.

Answer 2

Match the pattern, not the string — e.g. a regex like ^Page \d+. Exact matching fails whenever the repeated element contains a variable part (page numbers, timestamps, dates). The general rule: repeated structure needs pattern detection; repeated literals need set detection. Real corpora need both.

Answer 3

First, the second chunk loses the header row, so the model receives unlabeled name/phone pairs and may misinterpret the fields. Second, a query about Bronx escalation retrieves a chunk with no borough labels at all, so even a correct model can't ground its answer in the retrieved text — retrieval returned evidence that can't support the answer.

Answer 4

They're different failure modes needing different actions. Empty text means "no usable text extracted" — likely a scan, but possibly a blank page or an encoding the extractor can't read — so it's an inspect/OCR candidate, which is a cost/scheduling conversation with the client. A corrupt file means "this is broken" — it needs quarantine, logging, and a re-supply request. Mixing them would either waste OCR money on garbage or silently drop readable-after-OCR documents.

Answer 5

The hash identifies unchanged files, so re-ingestion skips them — no duplicate chunks, no wasted embedding cost, no risk of double-counting. But the hash alone isn't enough: the revised file gets a new hash, and the pipeline must also know it's the same document (document identity, separate from the hash) so it can retire the old version's chunks when it ingests the new ones. Otherwise both the old and the new policy remain searchable — and the copilot can cite superseded procedure with full confidence. It turns a scary "re-run everything" operation into a boring, safe one.