A chatbot that answers from your documents sounds like magic. It isn't — it's plumbing. And the plumbing matters more than the model.

You'll need: Python + numpy + pydantic, the vector-search intuition from the Embeddings lesson (embeddings turn "find similar" into arithmetic), and the contract discipline from Trust No Input: Data Validation with Pydantic. No API keys needed — everything below runs locally.

The customer problem

Lisa drops into your standup with a printout. It's the 311 call log from last week — forty pages of residents asking the same twenty questions.

"When is my recycling collected?" "There's a pothole on 4th — who fixes it?" "My garbage didn't get picked up, what now?"

"Can we put something on the website," she asks, "that answers these from our actual procedures — not from whatever the AI happens to remember?"

That last clause is doing all the work. Lisa has seen a demo chatbot confidently invent a recycling schedule. She doesn't want clever. She wants traceable.

Clarify the ask

Before writing code, pin down what "answer from our actual procedures" means:

QuestionLisa's answer
Where do answers come from?The six procedure documents the agency maintains. Nothing else.
What if the docs don't cover it?Say "I don't know" — never invent.
How do we trust an answer?Every answer shows which document it came from.
Who updates it?Lisa's team edits the docs; the system picks up changes without retraining anything.

That is a spec for RAG — retrieval-augmented generation. Not "an AI that knows things." A retrieval system with a language model attached to it. The order of those words is the whole lesson.

The principle

"Retrieval quality comes before prompt cleverness." A brilliant prompt over the wrong documents produces a confident wrong answer. A plain prompt over the right documents produces a useful one. When a RAG answer is wrong, inspect retrieval before changing the prompt — if the right evidence never reached the model, generation cannot recover it reliably.

The minimum concept

RAG is four steps, and you already understand three of them:

flowchart LR A[Agency documents] --> B[INGEST: extract, clean, chunk, tag, embed, index] B --> C[(Vector index)] D[Resident question] --> E[RETRIEVE: embed question, find nearest chunks] C --> E E --> F[AUGMENT: stuff chunks into the prompt] F --> G[GENERATE: model answers from the chunks] G --> H[Answer + citations]

Ingest happens once — and again whenever docs change. It extracts, cleans, chunks, tags, embeds and indexes the approved documents: the same pipeline from the ingestion lesson, with an embedding step added per chunk. Retrieve happens per question: embed the question, find the nearest chunks by cosine similarity. Augment means the prompt you send the model contains the retrieved chunks — "here is what we know, answer from this." Generate is the model call, which you already know how to make reliable.

Notice what RAG does not do: it never teaches the model anything. The model's weights don't change. That is the point — when Lisa's team edits a procedure, you re-index the document and the answers update. No retraining, no fine-tuning, no redeploy. For applications that need answers grounded in frequently changing documents, RAG is often a better fit than fine-tuning, because the knowledge can be updated without changing model weights. Note the two solve different jobs: RAG supplies current, external knowledge; fine-tuning adapts the model's behavior, style, or task performance. They aren't competitors — a system can use both.

Build: the loop, running for real

We'll build the whole thing locally over six real CityOps procedure documents — missed garbage collection, noise complaints, pothole repair, water main breaks, recycling contamination, graffiti removal. One honesty note before we start:

What runs and what's a stand-in

The retrieval plumbing — chunking, embedding, cosine ranking, top-k selection, the refusal gate, citation assembly — all runs for real below, with true outputs. Two pieces are stand-ins, clearly marked: the embedding function (this environment has no embedding API, so a small deterministic TF-IDF vectorizer stands in for a real embedding model — the cosine math is identical, but it scores word overlap, not meaning) and the generator (a stubbed function stands in for the LLM call, so what you're actually exercising is the retrieval and answer-assembly logic). In production you swap both for real API calls. The architecture stays the same, although thresholds, evaluation results and operational handling must be recalibrated for the real models.

Everything below runs against six real procedure documents and a small honest stand-in for the embedding model. Here is the full setup — corpus, chunk type, vectorizer, and similarity — so every output that follows is reproducible:

import re
import numpy as np
from dataclasses import dataclass

@dataclass
class Chunk:
    doc_id: str
    title: str
    idx: int
    text: str
    page: int = 1                 # provenance, per the ingestion lesson
    section: str | None = None

DOCS = [
    {"id": "PROC-SAN-001", "title": "Missed Garbage Collection", "text":
     "If a resident reports a missed garbage collection, first verify the route and the scheduled "
     "pickup day in the dispatch system. Remind residents to place bins at the curb by 6:00 AM on "
     "their scheduled day, with lids fully closed. If the crew missed a properly set out bin, dispatch "
     "a contractor crew within the 48-hour contractor window and promptly notify the resident of the new pickup "
     "time by text message. Escalate repeat misses to the borough supervisor for review. "
     "Log the incident in the 311 system with priority flag."},
    {"id": "PROC-NOI-002", "title": "Noise Complaints", "text":
     "Construction noise is permitted on weekdays between 7:00 AM and 6:00 PM only. Work outside "
     "these hours, including Sunday morning construction, requires a special after-hours permit from "
     "the borough office. When a resident reports a noise violation, log the address, the time, and "
     "the type of equipment involved for enforcement follow-up."},
    {"id": "PROC-POT-003", "title": "Pothole Repair", "text":
     "Pothole reports require the nearest cross streets and an estimate of the pothole size before a "
     "crew is dispatched to the location. Potholes on arterial roads are repaired within 72 hours of the "
     "report; potholes on residential streets are repaired within 5 business days. If the pothole caused "
     "vehicle damage, the resident must file a damage claim with the comptroller office within 30 days "
     "for review and reimbursement processing."},
    {"id": "PROC-WAT-004", "title": "Water Main Breaks", "text":
     "When a water main break floods the street, first shut the nearest upstream valve and call for a "
     "supervisor. Evacuate nearby basements if flooding is severe and notify the utility crew. Crews "
     "must restore service within 12 hours and document the repair for the weekly report."},
    {"id": "PROC-REC-005", "title": "Recycling Contamination", "text":
     "Recycling bins contaminated with food waste or plastic bags are rejected at pickup and tagged "
     "with a contamination notice explaining what was wrong. Residents may put the bin out again on the "
     "next scheduled day if corrected, with the contaminating items removed. Three contaminated pickups "
     "in a quarter trigger a mandatory recycling education visit from the outreach team for the household, "
     "scheduled within two weeks of the third violation."},
    {"id": "PROC-GRA-006", "title": "Graffiti Removal", "text":
     "Graffiti on public property is removed by the city cleanup crew within 10 business days of a "
     "report. Graffiti on private property is the owner's responsibility; the city provides a list of "
     "approved contractors but does not perform the removal for private buildings."},
]

def _tokens(s):
    return re.findall(r"[a-z0-9]+", s.lower())

_VOCAB = sorted({t for d in DOCS for t in _tokens(d["text"])})
_VIDX = {t: i for i, t in enumerate(_VOCAB)}
# IDF weights: common words ("the") count for almost nothing,
# rare content words count more. Still word overlap, not meaning.
_DF = np.zeros(len(_VOCAB))
for _d in DOCS:
    for _t in set(_tokens(_d["text"])):
        _DF[_VIDX[_t]] += 1.0
_IDF = np.log(len(DOCS) / np.maximum(_DF, 1.0))

def embed(s):
    """STAND-IN for a real embedding model: deterministic TF-IDF word-overlap
    vector. The cosine math is identical to production; but this scores shared
    words, not meaning. A real model is swapped in later."""
    v = np.zeros(len(_VOCAB))
    for t in _tokens(s):
        j = _VIDX.get(t)
        if j is not None:
            v[j] += 1.0
    return v * _IDF

def cosine(a, b):
    na, nb = np.linalg.norm(a), np.linalg.norm(b)
    return float(a @ b / (na * nb)) if na and nb else 0.0

def generate_stub(question, hits):
    """STAND-IN for the LLM call: drafts an answer from the top hit's text."""
    top = hits[0][1]
    return f"Based on {top.doc_id}: {top.text[:160]}..."

Step 1: Chunk the documents

Chunking is the first retrieval decision you'll ever make, and it has consequences. Too big, and each chunk's embedding is a blurry average — the specific fact drowns. Too small, and a fact gets split across chunks — retrieval finds half an answer. Watch:

def chunk_text(text, size, overlap=0):
    """Naive fixed-size word chunker with optional overlap."""
    words = text.split()
    chunks, i = [], 0
    while i < len(words):
        chunks.append(" ".join(words[i:i + size]))
        i += size - overlap if overlap else size
    return chunks

doc = DOCS[0]["text"]  # missed garbage collection procedure, 90 words
for size, ov in [(200, 0), (40, 0), (40, 10)]:
    cs = chunk_text(doc, size, ov)
    print(f"size={size} overlap={ov}: {len(cs)} chunk(s), "
          f"first={len(cs[0].split())} words, last={len(cs[-1].split())} words")
size=200 overlap=0: 1 chunk(s), first=90 words, last=90 words
size=40 overlap=0: 3 chunk(s), first=40 words, last=10 words
size=40 overlap=10: 3 chunk(s), first=40 words, last=30 words

Size 200 swallows the whole procedure into one vector — a single embedding trying to mean "verify the route and the 6 AM rule and the 48-hour contractor window and escalation." Size 40 with no overlap leaves a 10-word orphan ("Log the incident in the 311 system with priority flag") with no context about what incident. Size 40 with 10 words of overlap keeps every chunk self-contained enough to stand alone. Overlap is cheap insurance against facts that straddle a boundary — but it isn't free: overlap duplicates text, grows the index, and can surface near-duplicate chunks in retrieval. More overlap isn't automatically better. We'll index at 40/10 — 15 chunks from 6 documents.

Chunking strategyRetrieval consequence
Whole document, one chunkDiluted embedding — the specific fact drowns in the average.
Tiny chunks, no overlapFragmented facts — retrieval finds "the 48-hour rule" with no idea what it's for.
~40 words + overlapEach chunk self-contained. 40/10 is convenient for this tiny teaching corpus; production chunk size should be selected through retrieval evaluation.
Section-aware (headings, paragraphs)Better when docs have real structure — the upgrade path.

Step 2: Embed and index

Each chunk becomes a vector; the vectors go into an index. In production this is a vector database. For 15 chunks, storing vectors in a NumPy array and brute-force scoring them is enough — no vector database needed, and no infrastructure before scale demands it:

def build_index(chunk_size=40, overlap=10):
    chunks, vecs = [], []
    for d in DOCS:
        for i, c in enumerate(chunk_text(d["text"], chunk_size, overlap)):
            chunks.append(Chunk(doc_id=d["id"], title=d["title"], idx=i, text=c))
            vecs.append(embed(c))          # embed(): the embedding model
    return chunks, np.array(vecs)

chunks, vecs = build_index()
print(f"index: {len(chunks)} chunks from {len(DOCS)} docs")
index: 15 chunks from 6 docs

Step 3: Retrieve — the moment of truth

Embed the question, rank every chunk by cosine similarity, take the top-k. Everything downstream depends on this step being right:

def retrieve(query, chunks, vecs, k=2, min_score=0.0):
    qv = embed(query)
    ranked = sorted(((cosine(qv, v), c) for v, c in zip(vecs, chunks)),
                    key=lambda s: -s[0])
    return [(s, c) for s, c in ranked[:k] if s >= min_score]

for q in ["huge pothole damaged my car",
          "is my recycling bin contaminated",
          "missed garbage collection what do i do"]:
    print(f"Q: {q}")
    for s, c in retrieve(q, chunks, vecs, k=2):
        print(f"  {s:.3f} [{c.doc_id} c{c.idx}] {c.text[:60]}...")
    print()
Q: huge pothole damaged my car
  0.360 [PROC-POT-003 c0] Pothole reports require the nearest cross streets and an est...
  0.213 [PROC-POT-003 c1] hours of the report; potholes on residential streets are rep...

Q: is my recycling bin contaminated
  0.300 [PROC-REC-005 c0] Recycling bins contaminated with food waste or plastic bags ...
  0.270 [PROC-REC-005 c1] on the next scheduled day if corrected, with the contaminati...

Q: missed garbage collection what do i do
  0.322 [PROC-SAN-001 c0] If a resident reports a missed garbage collection, first ver...
  0.105 [PROC-SAN-001 c1] 6:00 AM on their scheduled day, with lids fully closed. If t...

The top hit is right all three times — but notice the scores are modest (0.30–0.36), and the sanitation query's second hit is a recycling chunk at 0.105: mild noise that shared a word or two. This is normal. Retrieval is a ranking, not a verdict. Only pass sufficiently relevant evidence downstream; rank position and scores are retrieval signals, not proof.

Step 4: Augment — build the prompt

The retrieved chunks become the model's context, with an explicit instruction to stay inside it. Here is the actual prompt our pipeline builds for the pothole question (truncated):

"Answer ONLY from the procedures below. If they don't cover the question, say you don't know.

[1] (PROC-POT-003) Pothole reports require the nearest cross streets and an estimate
of the pothole size before a crew is dispatched to the location. Potholes on
arterial roads are repaired within 72 hours of the report

Q: huge pothole damaged my car"

That first sentence is the grounding instruction — the answering policy for this system: permission to say "I don't know" is a feature, not a failure mode. The instruction communicates the desired behavior; validation and evaluation determine whether the model actually follows it. Combined with retrieval, refusal logic, citation requirements and validation, this creates a more grounded answering system than an unconstrained chatbot.

Step 5: Generate — and cite, or refuse

The generator (stubbed here) produces the answer; then we enforce Lisa's traceability requirement with the Pydantic discipline from the validation lesson. The business rule isn't "every answer has at least one citation" — it's conditional: answered → at least one citation; refused → zero citations. The schema should express that truth, not force fake values:

from pydantic import BaseModel, model_validator

class Answer(BaseModel):
    text: str
    citations: list[int] = []   # 1-based indices into the retrieved hits
    refused: bool = False
    sources: list[str] = []     # resolved provenance, filled by the pipeline

    @model_validator(mode="after")
    def _check_citation_contract(self):
        if self.refused and self.citations:
            raise ValueError("refused answers must carry zero citations")
        if not self.refused and not self.citations:
            raise ValueError("answered responses need at least one citation")
        return self

q = "huge pothole damaged my car"
hits = retrieve(q, chunks, vecs, k=2, min_score=0.28)
text = generate_stub(q, hits)   # STAND-IN for the model call
ans = Answer(text=text, citations=[1],
             sources=[f"{hits[0][1].doc_id} — {hits[0][1].title} (p{hits[0][1].page})"])
print("citations:", ans.citations, "| refused:", ans.refused)
print("sources:", ans.sources)
citations: [1] | refused: False
sources: ['PROC-POT-003 — Pothole Repair (p1)']
# a refusal carries zero citations — no fake [0] to please the validator:
Answer(text="I don't know — no CityOps procedure covers this.", refused=True)

# and an uncited answer never leaves the building:
Answer(text="Potholes are fixed quickly.", citations=[])
ValidationError: Value error, answered responses need at least one citation

Notice what we didn't do: the refusal doesn't smuggle in a fake citations=[0] to satisfy a blanket minimum. That would corrupt the data contract to please the validator. Instead the schema expresses the actual business rule — a direct callback to the Pydantic lesson.

Citation numbers also resolve to real provenance — the document ID, title, and page the ingestion lesson taught us to preserve — so Lisa's "which document did this come from" is answered structurally, not by a bare [1]. But note the limit: the schema guarantees that answered responses include citation references. It does not guarantee the cited chunk supports the answer. A cited answer is not necessarily a grounded answer; the cited evidence must actually support the claim.

The refusal gate: tuning min_score

What happens when nothing relevant exists? The threshold decides. Watch the same pipeline at two settings — first against an unsupported question, then against a legitimate one:

RAGPipeline(chunks, vecs, min_score=0.15).ask("when does the library close")
#   refused=False -> "Based on PROC-GRA-006: list of approved contractors but
#                     does not perform the removal for private buildings...."
#                     <- answers about graffiti contractors. Wrong: no procedure
#                     covers library hours. The query scored 0.269 on a stray
#                     word overlap ("does") and slipped through.

RAGPipeline(chunks, vecs, min_score=0.35).ask("when does the library close")
#   refused=True -> "I don't know — no CityOps procedure covers this."

RAGPipeline(chunks, vecs, min_score=0.35).ask("is my recycling bin contaminated")
#   refused=True -> "I don't know..."  <- a WRONG refusal: the procedure exists,
#                     it just scored 0.300, below the 0.35 bar.

At 0.15 the gate is decorative: a weak 0.269 gets presented as an answer about graffiti contractors. At 0.35 the library nonsense is refused — but so is a legitimate recycling question. The threshold is a tuning decision, not a constant: too low → false answers; too high → unnecessary refusals. Set it from a golden set containing both answerable and unanswerable questions (we'll build exactly that below), and re-evaluate it when the corpus grows or the embedding model changes — the model, the index, and the threshold belong together as retrieval configuration. Never ship the default. And remember: a similarity threshold is one refusal signal, not proof that the retrieved evidence answers the question.

Break: three ways RAG fails

flowchart TD Q[Resident question] --> R[Retrieve top-k] R -->|nothing above threshold| REF[I don't know] R -->|wrong chunks| W[Confident wrong answer] R -->|right chunks| G[Generate from context] G -->|ignores context| M[Hallucination from memory] G -->|uses context| OK[Cited answer]

Failure 1: Retrieval returns the wrong chunks

Remember our stand-in embedder scores word overlap, not meaning. Ask it about trash:

for s, c in retrieve("my trash wasn't collected", chunks, vecs, k=2):
    print(f"  {s:.3f} [{c.doc_id} c{c.idx}] {c.text[:60]}...")
  0.000 [PROC-SAN-001 c0] If a resident reports a missed garbage collection, first ver...
  0.000 [PROC-SAN-001 c1] 6:00 AM on their scheduled day, with lids fully closed. If t...

"My trash wasn't collected" scores 0.000 across the board — the stand-in finds nothing. The word "trash" never appears in the corpus, and the stand-in has no notion that trash ≈ garbage, so the sanitation procedure is invisible to it. A capable embedding model should represent common paraphrases such as trash/garbage more usefully than this word-overlap stand-in — but verify that on your retrieval eval. Vocabulary mismatch is one of the core problems semantic retrieval is meant to solve — paraphrase, domain terminology, abbreviations, and concept relationships all live here. And when retrieval goes wrong, don't jump straight to swapping the embedding model: inspect the retrieved chunks and the query first, then check ingestion and chunking, metadata filters, embedding behavior, and ranking configuration — in that order.

Failure 2: The model ignores the context

Right chunks in, wrong answer out. The model answers from its training memory instead of the provided procedures — the context adherence failure:

def rogue_generate(question, hits):
    """SCRIPTED stand-in for a model answering from memory, not context."""
    return "Potholes are repaired within 24 hours guaranteed."

bad = rogue_generate(q, hits)
# model said: "Potholes are repaired within 24 hours guaranteed."
# check: FAIL — "24 hours guaranteed" appears in 0 retrieved chunks

No procedure promises 24 hours — the real answer is 72 hours / 5 business days. The defense is mechanical, not hopeful: require citations, then verify that the cited evidence semantically supports the claims being made. Support can involve paraphrase or span multiple chunks, so this isn't substring matching — and it's a check you'll eventually automate, not hand-wave. A citation pointing at a chunk that doesn't support the sentence is the tell. Citation presence improves traceability; citation support checking improves grounding; neither replaces end-to-end evaluation. This is also why the refusal instruction matters — a model told "answer only from below" fails closed instead of reaching into memory.

Failure 3: Top-k too large drowns the answer

More context feels safer. It isn't. Watch what k=6 retrieves for the pothole question:

  0.360 [PROC-POT-003 c0]     <- the answer
  0.213 [PROC-POT-003 c1]
  0.000 [PROC-SAN-001 c0]     <- zero similarity. pure filler.
  0.000 [PROC-SAN-001 c1]     <- zero similarity. pure filler.
  0.000 [PROC-SAN-001 c2]     <- zero similarity. pure filler.
  0.000 [PROC-NOI-002 c0]     <- zero similarity. pure filler.

Irrelevant chunks consume context budget without adding useful evidence and may distract generation: models are suggestible, and filler context dilutes the real signal, inviting the model to stitch nonsense from irrelevant chunks. Top-k is a relevance budget. Spend it on the chunks that earned their place, and let the threshold cut the rest. The right k isn't "small" as a universal rule — it's enough evidence to answer, but not indiscriminate context. Some questions need one chunk; some need ten. Treat top-k and the similarity threshold as the two tunable retrieval parameters they are.

Productionize: the pipeline as a system

The demo pieces compose into one class — ingest once, ask many times, refuse cleanly:

class RAGPipeline:
    def __init__(self, chunks, vecs, k=2, min_score=0.28):
        self.chunks, self.vecs = chunks, vecs
        self.k, self.min_score = k, min_score

    def ask(self, question):
        hits = retrieve(question, self.chunks, self.vecs,
                        k=self.k, min_score=self.min_score)
        if not hits:
            return Answer(text="I don't know — no CityOps procedure covers this.",
                          refused=True)
        prompt = self._build_prompt(question, hits)
        text = generate_stub(question, hits)  # STAND-IN for the model call
        cites = list(range(1, len(hits) + 1))
        ans = Answer(text=text, citations=cites)
        ans.sources = self._resolve(hits, cites)  # numbers -> provenance
        return ans

    def _resolve(self, hits, citations):
        out = []
        for n in citations:
            c = hits[n - 1][1]
            sec = f", {c.section}" if c.section else ""
            out.append(f"{c.doc_id} — {c.title} (p{c.page}{sec})")
        return out

    def _build_prompt(self, question, hits):
        ctx = "\n".join(f"[{i+1}] ({c.doc_id}) {c.text[:200]}"
                        for i, (_, c) in enumerate(hits))
        return (f"Answer ONLY from the procedures below. If they don't cover "
                f"the question, say you don't know.\n\n{ctx}\n\nQ: {question}")

Note the refusal path: no fake citations=[0] — a refused answer carries zero citations, and the validator enforces it. And 0.28 isn't a magic number; it's the value that separates our golden positives (0.30–0.54) from our golden negatives (0.00–0.27), chosen after measuring both. Here's the production mental model this pipeline implements — the architecture to carry into every RAG system you build:

flowchart TD Q[Resident question] --> R[Retrieve candidates] R --> F[Apply exact filters / rank] F --> E{Enough evidence?} E -->|No| REF[Refuse: I don't know] E -->|Yes| B[Build grounded context] B --> G[Generate] G --> V[Validate structure] V --> S[Verify citation support] S --> A[Answer]

Retrieve candidates, apply exact filters and rank, check whether the evidence is sufficient — refusing when it isn't — then build grounded context, generate, validate structure, and verify citation support before answering. Every production upgrade below slots into this loop without changing it.

Three production upgrades, in the order you'd actually need them:

1. Hybrid retrieval: vectors for meaning, keywords for exactness

Vector search wins on paraphrase ("my street smells" → garbage complaints). Keyword search wins on exactness: procedure IDs like PROC-SAN-001, dates, dollar amounts, street names. A production RAG system runs both and merges the rankings — reciprocal rank fusion (RRF) is a common way to combine the ranked lists. The rule of thumb:

Use vector search whenUse keyword / filters when
The resident paraphrases ("trash" vs "garbage")They quote an ID, date, or exact term
"Find me things like this""Find me this exact thing"
Meaning matters more than wordingWording matters more than meaning

And structured attributes — borough, status, date ranges — don't belong in vectors at all. Apply exact structured constraints as part of retrieval, with the SQL discipline from SQL for Field Work, then use semantic ranking among the eligible records. Vector search over "all complaints" when you mean "open complaints in Queens" is a bug, not a strategy.

2. Evaluate retrieval separately from generation

When the system gives a bad answer, you need to know which half broke. Keep a golden set of questions with known-correct documents and score retrieval alone — no model involved:

golden = [
    ("huge pothole damaged my car",            "PROC-POT-003"),
    ("is my recycling bin contaminated",       "PROC-REC-005"),
    ("missed garbage collection what do i do", "PROC-SAN-001"),
    ("construction noise on sunday morning",   "PROC-NOI-002"),
    ("water flooding the street",              "PROC-WAT-004"),
    ("graffiti on the bus stop",               "PROC-GRA-006"),
]
negatives = [
    "zebra migration patterns",        # no procedure covers this
    "when does the library close",     # no procedure covers this
    "can I dispute a parking ticket",  # no procedure covers this
]

def doc_hit(q, want):
    # no hits above threshold counts as a missed retrieval — not a crash
    hits = retrieve(q, chunks, vecs, k=1)
    return bool(hits) and hits[0][1].doc_id == want

supported = sum(doc_hit(q, want) for q, want in golden)
print(f"supported hit-rate@1: {supported}/{len(golden)}")

pipe = RAGPipeline(chunks, vecs)
refused = sum(pipe.ask(q).refused for q in negatives)
print(f"unsupported refusal rate: {refused}/{len(negatives)}")
supported hit-rate@1: 6/6
unsupported refusal rate: 3/3

Two honest caveats before you celebrate: this is six questions on six documents, and the queries share vocabulary with the docs — a cooperative test, not an adversarial one. The number isn't the point; the method is. A retrieval eval that runs in seconds, separate from the model, tells you whether a bad answer is a retrieval problem (fix the chunks, the embeddings, the threshold) or a generation problem (fix the prompt, the adherence checks). That separation is what makes RAG debuggable — and it's the seed of the full evals discipline coming later in this stage.

Note what the negatives buy you: refusal is a product requirement, so the eval must measure it. Without unsupported queries in the golden set, you'd tune only for recall — and never notice a threshold so low it answers everything. This same set is what chose our 0.28: positives score 0.30–0.54, negatives 0.00–0.27, and the threshold has to separate them.

One more step for later: document hit-rate is the beginner metric. If a document holds thirty sections, retrieving the right document but the wrong section isn't enough — production evals should judge whether the retrieved chunks actually contain the needed evidence. The ingestion lesson's chunk/page/section provenance exists precisely so you can score at that granularity when you're ready.

3. Re-index on doc changes, or answers go stale

RAG's superpower — answers update without retraining — only works if ingest actually re-runs. The production checklist: version the document set, build a new index version on every edit, validate it against the golden set, then atomically switch traffic to it — replacing or deactivating the old document's chunks so stale text can't stay searchable. (The ingestion lesson showed why: a new content hash alone can leave old chunks live. Index versioning plus old-chunk retirement is the fix.) A RAG system answering from last quarter's procedures is worse than no system — it's authoritative and wrong.

Useful later

Real deployments add: a vector database (pgvector, Qdrant, Pinecone) once the corpus outgrows memory; reranking (a second, slower model re-scores the top-k); query rewriting ("it" → the thing "it" refers to); and per-tenant indexes so one client's documents never leak into another's answers — tenant isolation must be enforced by architecture and access controls, not merely by telling the model which client's documents to use. None of it changes the four-step loop — it just scales it.

Communicate: the doc you'd give Lisa

Lisa doesn't need the architecture diagram. She needs the operating contract:

311 Answers — what it can and can't do (v1)

Coverage: 6 procedures — garbage, noise, potholes, water mains, recycling, graffiti. Ask about anything else and it says "I don't know" instead of guessing.

Trust: every answer names the procedure it came from. If an answer ever lacks a source, that's a bug — report it.

Updates: edit a procedure doc and tell engineering; answers pick it up at the next successful re-index.

Measured quality: on 6 representative resident questions, retrieval finds the right procedure 6/6. We'll grow this test set as residents ask new things.

Known limits: semantic retrieval can match related phrasing even when exact words differ — usually a strength, occasionally surprising. It can't answer questions the procedures don't cover, and it won't try.

Notice what's doing the heavy lifting: coverage stated up front, refusal framed as a feature, a quality number with its limits disclosed. That's how you sell a system whose best behavior is sometimes silence.

CityOps: the retrieval backbone

This pipeline is the foundation the rest of the stage builds on. The procedures index you just built is exactly what the "Ask CityOps" milestone will query — except there, the retrieved context feeds a copilot that can also act: look up complaints, draft escalations, and pause for human approval. Retrieval stays the same four-step loop; what changes is what the system is allowed to do with what it finds. Get the retrieval right here and the milestone inherits it. Get it wrong and no amount of agent cleverness will save it.

Field check

  1. Why does this lesson insist retrieval quality comes before prompt cleverness?
  2. Your RAG system confidently answers from a procedure that changed last week. What broke?
  3. Retrieval returns the right chunk but the answer is still wrong. Name two suspects.
  4. A resident asks for "procedure PROC-SAN-001". Vector search, keyword search, or metadata filter — and why?
  5. A resident asks something no procedure covers. Name the two gates that stop a hallucination.
Reveal answers

1. The model can only answer from what retrieval hands it. A perfect prompt over the wrong chunks produces a confident wrong answer — the failure is upstream of the model, so check retrieval first.

2. The index is stale — ingest didn't re-run after the doc edit. RAG's freshness depends on re-indexing: version the document set, build and validate a new index version against the golden set, then switch traffic to it while retiring the old document's chunks.

3. (a) Generation / context-adherence failure — the model didn't faithfully use the evidence; check whether the cited chunk semantically supports the claim. (b) Context construction / answer assembly failure — the right chunk was truncated, omitted, mislabeled, or the citation mapping broke. (A third possibility: the "right" chunk doesn't actually contain enough evidence to answer — but don't revert to blaming retrieval when the question stipulates it succeeded.)

4. Keyword/metadata — "PROC-SAN-001" is an exact identifier, not a meaning. Vector search is for paraphrase; exact lookups belong to keyword search or a metadata filter.

5. Gate 1 — before generation: the refusal gate. When retrieval finds no adequate evidence (nothing above min_score), the pipeline returns "I don't know" without ever calling the model. Gate 2 — after generation: citation validation and support checking. The schema requires every answered response to carry citations, and support verification checks the cited evidence actually backs the claims — making unsupported output detectable. Keep each piece honest: the threshold is one refusal signal, not proof; the grounding instruction communicates intent but doesn't enforce it; and citation presence is not citation truth.