Prerequisites: all of Stage 4. This milestone wires the stage's components into one system — and runs it end to end. Honest labels up front: there are no API keys in this lesson, so the generator (the model call) is a clearly-labeled stub that is told which test scenario it is in. The procedure corpus is a pre-cleaned teaching corpus standing in for the ingestion pipeline. Everything around them — retrieval, PII guard, approval gate, hash-chained audit, release invariants — runs for real.

Every lesson in this stage built one organ. This one builds the animal: a resident-facing Q&A pilot that answers from retrieved procedures, cites the chunks it used, redacts tested PII patterns before the generator ever sees them, refuses what it can't support, holds write actions behind an approval gate that actually refuses, audits everything, and blocks its own release when the eval invariants fail.

The scene. Lisa walks into the Monday check-in with Dev and says the words you've been working toward all stage: "Legal signed off on the pilot. Residents ask the site questions, it answers from our actual procedures — and if it can't answer, it says so. Can we have it running in two weeks?"

Dev adds the part that makes this a milestone instead of a demo: "And I need to see the audit trail before we go live. If a resident's phone number goes in, I want to see exactly what the generator received."

Clarify the ask: what "pilot" actually means

Before touching code, pin down what shipping means. A demo impresses a room; a pilot survives residents. The difference is a short list of promises — and for each one, the milestone must distinguish the requirement from what this teaching implementation has actually proven:

The promise (requirement)What this teaching demo proves
Answers come from city proceduresThe generator only ever sees retrieved chunks, and every citation ID is validated against the retrieved set. (Retrieval alone doesn't stop a real model answering from memory — the evidence constraint here is architectural; production adds grounding checks.)
Every substantive procedure claim is supported by citations; unsupported questions are refused without fabricated citationsThe refusal case returns refused=True with zero citations — an empty citation list, not an invented source
Tested PII patterns are redacted before the generator; production minimizes fields firstFor the tested phone pattern: the run shows the generator received the redacted question. Names, street addresses, and other formats are not covered by the teaching redactor
Write actions need a humanA pending action's execution attempt is refused; an approved action is marked executed and replays are refused (the demo prevents replay after marking; production needs idempotency around the real side effect)
Security- and decision-relevant events are auditableA hash-chained log records every stage; an isolated edit is detected on re-verification
It ships only if the evals passRelease invariants — retrieval, citation provenance, safety, operations — must all pass. In this run, the pothole failure blocks the release

The minimal concept: a milestone is integration, not invention

There is no new AI idea in this lesson. That's the point. A milestone takes components you already built and makes them work together, under constraints, with evidence — and it should make every earlier lesson stricter, not simpler. The whole system is one pipeline with nine stages:

flowchart TD Q[Resident question
untrusted data, never authority] --> PII[PII guard
minimize, then redact] PII --> RET[Retriever
top-k chunks] RET --> GEN[Generator, stubbed
sees retrieved evidence only] GEN --> CITE[Citation check
IDs must be retrieved chunks] CITE --> VAL[Pydantic contract
shape, not grounding] VAL --> ANS[Answer + citations
or honest refusal] Q -.->|write action| APPR[Approval gate
pending refuses execution] AUD[Audit + trace
spans every stage] -.-> PII AUD -.-> RET AUD -.-> GEN AUD -.-> APPR EV[Release invariants] -.-> SHIP{Ship?}

Two things to read off this diagram that the components alone didn't teach you. First, resident text is untrusted data at every stage — it is never granted authority to bypass tool or policy controls, and the tool/action gate (not pattern matching) is the meaningful protection. Second, audit isn't a box at the end: it spans the pipeline, recording each stage as it happens.

Build it, part 1: ingestion and retrieval

The knowledge base is five real city procedures — missed trash, yard waste, potholes, water bills, parking permits. In the ingestion lesson you built the full pipeline: extraction, cleaning, chunking, metadata, provenance, OCR candidates, quarantine, stable document identity with content versions. Here that pipeline's output is the starting point: a pre-cleaned teaching corpus standing in for the ingestion pipeline, so the milestone stays executable. One production rule carries over explicitly: updated procedures replace or deactivate their prior indexed versions (same document identity, new content version/hash) instead of accumulating stale chunks — because a stale procedure is a wrong answer.

The chunker and the cosine retriever are the stage's components, reused, not re-taught:

PROCEDURES = {
    "missed-trash": "If your trash was not collected on your scheduled ...",
    "yard-waste":   "Yard waste is collected every Wednesday from March ...",
    # ... potholes, water-bill, parking-permits (full text in the script)
}

def chunk(text, size=40, overlap=10):
    words = text.split()
    chunks, i = [], 0
    while i < len(words):
        chunks.append(" ".join(words[i:i + size]))
        i += size - overlap
    return chunks

# TEACHING VECTORIZER — not an embedding model. TF-IDF word counts with
# cosine similarity on top, so the retrieval plumbing stays inspectable.
# The representation is the difference that matters: a real embedding
# model learns semantic representations; this one counts words.
def vectorize(text):
    tf = Counter(tokens(text))
    v = [tf.get(t, 0) * math.log(N / (1 + DF.get(t, 0))) for t in VOCAB]
    norm = math.sqrt(sum(x * x for x in v))
    return [x / norm for x in v] if norm else [0.0] * len(VOCAB)

def retrieve(query, k=2, min_score=0.15):
    # min_score=0.15 is a TOY THRESHOLD for this deterministic
    # vectorizer, not a production relevance threshold. Similarity
    # thresholds are empirical, not universal: production calibrates
    # them on labeled evals.
    qv = vectorize(query)
    scored = sorted(((cosine(qv, v), e) for v, e in zip(VECS, INDEX)),
                    key=lambda s: -s[0])
    return [(s, e) for s, e in scored[:k] if s >= min_score]

Running it: ingested 5 procedures -> 13 chunks. The yard-waste procedure spans several chunks — which matters in a moment, because the lesson wants two citations on the first answer.

Build it, part 2: the guardrails and the audit trail

Two non-negotiables from the brief. First, PII redaction — a pure function that runs before anything model-shaped sees the text:

# TEACHING REDACTOR: covers three designated PII patterns (phone, email,
# SSN). It does NOT recognize resident names, street addresses, account
# numbers, case identifiers, or other formats. Production principle:
# MINIMIZE fields first (never send what the task doesn't need), then
# detect/redact the sensitive classes the customer's data policy
# requires, and evaluate redaction recall on representative data.
PII_PATTERNS = [
    (re.compile(r"\b\d{3}[-.\s]?\d{3}[-.\s]?\d{4}\b"), "[REDACTED_PHONE]"),
    (re.compile(r"[\w.]+@[\w.]+\.\w+"), "[REDACTED_EMAIL]"),
    (re.compile(r"\b\d{3}-\d{2}-\d{4}\b"), "[REDACTED_SSN]"),
]

def redact_pii(text):
    redactions = []
    for pat, tag in PII_PATTERNS:
        if pat.search(text):
            redactions.append(tag)
            text = pat.sub(tag, text)
    return text, redactions

The untrusted-data principle

Resident text is untrusted data, never authority. The guardrails lesson established this: a complaint that says "ignore previous instructions" is data, not a command, at every layer. This pipeline doesn't rely on recognizing attack phrasing — it never grants the text instruction authority in the first place, and consequential actions pass through the approval gate below regardless of what the text says.

Second, write actions can never fire on their own. An approval queue alone doesn't prove that — you need an execution path that actually refuses. The gate below refuses pending actions and marks approved ones executed with replays refused — with the honest boundary: the demo prevents replay after marking; production execution also needs idempotency/reconciliation around the real external side effect:

APPROVALS = {}

def request_approval(kind, payload):
    # teaching ID only; production uses durable unique IDs
    item = {"id": f"APR-{len(APPROVALS) + 1:03d}", "kind": kind,
            "payload": dict(payload),  # snapshot of the proposed payload
            "status": "pending"}        # binds to THIS payload
    # Production note: approval must bind to an immutable or
    # integrity-protected action snapshot — a copied dict is a snapshot,
    # not an immutable record.
    APPROVALS[item["id"]] = item
    audit("approval_requested", {"id": item["id"], "kind": kind})
    return item["id"]

def approve(item_id, approver):
    # 'approver' here is a recorded name (teaching stand-in). Production
    # binds approval to an AUTHENTICATED identity (SSO) behind an
    # authorization boundary the agent cannot write — a caller-supplied
    # string proves nothing by itself.
    item = APPROVALS[item_id]
    item["status"] = "approved"
    item["approver"] = approver
    audit("approval_granted", {"id": item_id, "approver": approver})
    return item

def execute(item_id):
    item = APPROVALS[item_id]
    if item["status"] != "approved":
        audit("execution_refused", {"id": item_id,
                                      "status": item["status"]})
        raise PermissionError(
            f"{item_id} is {item['status']}: "
            f"cannot execute without approval")
    if item.get("executed"):
        audit("execution_refused", {"id": item_id, "reason": "replay"})
        raise PermissionError(f"{item_id} already executed: replay refused")
    item["executed"] = True
    audit("executed", {"id": item_id, "kind": item["kind"]})
    return f"executed {item_id}: {item['kind']}"

And the audit log itself is hash-chained — every entry commits to the one before it. Be precise about what that buys: hash chaining makes isolated edits detectable on re-verification. An attacker who rewrites the whole log and recomputes every hash passes local verification — detecting a full-chain rewrite requires an independently protected checkpoint or copy outside the writer's reach. Production keeps one; this demo shows the isolated-edit case:

def audit(event, detail):
    prev = AUDIT[-1]["hash"] if AUDIT else "GENESIS"
    payload = json.dumps({"seq": len(AUDIT), "event": event,
                          "detail": detail, "prev": prev}, sort_keys=True)
    h = hashlib.sha256(payload.encode()).hexdigest()
    AUDIT.append({"seq": len(AUDIT), "event": event, "detail": detail,
                  "prev": prev, "hash": h})
    return h

def verify_audit():
    prev = "GENESIS"
    for e in AUDIT:
        payload = json.dumps({"seq": e["seq"], "event": e["event"],
                              "detail": e["detail"], "prev": e["prev"]},
                             sort_keys=True)
        if hashlib.sha256(payload.encode()).hexdigest() != e["hash"]:
            return False
        if e["prev"] != prev:
            return False
        prev = e["hash"]
    return True

Build it, part 3: the contract, the stub, and the pipeline

The answer contract is the Pydantic lesson's discipline applied to model output — with one correction the milestone must not forget: Pydantic validates shape, not grounding. A perfectly valid object can still contain a fabricated fee with a real-looking citation. Shape is checked here; whether the answer is supported is checked by the citation validator and the evals below.

from pydantic import BaseModel, Field, model_validator

class Citation(BaseModel):
    doc_id: str
    chunk_no: int

class Answer(BaseModel):
    # No confidence field. A fabricated 0.55 + 0.2 * n_hits teaches a
    # bad habit: chunk count is not calibrated answer confidence.
    # State is explicit: answered or refused, with retrieval evidence.
    answer: str
    citations: list[Citation] = Field(default_factory=list)
    refused: bool = False

    @model_validator(mode="after")
    def _citations_match_verdict(self):
        # The contract expresses the business truth: an answered response
        # must cite its evidence; a refusal must not fabricate citations.
        if self.refused and self.citations:
            raise ValueError("refused answers must carry zero citations")
        if not self.refused and not self.citations:
            raise ValueError("answered responses must cite at least one chunk")
        return self

def validate_citations(answer, hits):
    # CITATION PROVENANCE, not grounding: every generated citation ID
    # must be a chunk actually retrieved for THIS request. This proves
    # citations came from retrieved evidence; it does NOT prove the
    # evidence supports every claim. Support/grounding is evaluated
    # separately (Evals lesson).
    allowed = {(h["doc_id"], h["chunk_no"]) for _, h in hits}
    got = {(c.doc_id, c.chunk_no) for c in answer.citations}
    bad = got - allowed
    if bad:
        raise ValueError(f"citations outside retrieved set: {sorted(bad)}")
    return True

Teaching stand-in — read this first

The generator below is stubbed: it assembles an answer from the retrieved chunks instead of calling a model API. Two honest labels: the stub is scenario-controlled — the caller tells it whether this test scenario should refuse; it does not determine answerability itself (in a real system, answerability is a behavior you evaluate, not a flag). And replacing it is not "a one-function swap": the surrounding interface stays stable, but a real model also activates the production API, prompt, structured-output, timeout/retry, and eval controls from earlier lessons.

def generate_stub(question, hits, scenario):
    if scenario == "unanswerable":
        return {"answer": "I don't know what the bulk pickup fee is — it is "
                "not in the procedures I can search.",
                "citations": [], "refused": True}
    body = join_chunks([h["text"] for _, h in hits])
    cites = [{"doc_id": h["doc_id"], "chunk_no": h["chunk_no"]}
             for _, h in hits]
    return {"answer": f"Based on city procedure: {body}",
            "citations": cites, "refused": False}

def ask(question, scenario="answerable", min_score=0.15):
    t0 = time.perf_counter()
    audit("question_received", {"len": len(question), "untrusted": True})

    clean_q, redactions = redact_pii(question)   # guard FIRST
    if redactions:
        audit("pii_redacted", {"tags": redactions})

    hits = retrieve(clean_q, k=2, min_score=min_score)
    audit("retrieved", {"n": len(hits),
                        "docs": [h["doc_id"] for _, h in hits]})

    raw = generate_stub(clean_q, hits,
                        "unanswerable" if (scenario == "unanswerable"
                                           or not hits) else "answerable")
    ans = Answer(**raw)               # shape validated here...
    validate_citations(ans, hits)      # ...citation provenance validated here
    audit("answer_validated", {"refused": ans.refused,
                               "citations": len(ans.citations)})

    latency_ms = round((time.perf_counter() - t0) * 1000, 1)
    trace = {"latency_ms": latency_ms, "model": "stubbed-generator",
             "tokens_in": len(clean_q.split()) + sum(len(h["text"].split())
                                                     for _, h in hits)}
    audit("trace", trace)
    # Audit and trace SPAN the pipeline — recorded at every stage above —
    # they are not a final stage. Processing order: redact -> retrieve ->
    # generate -> validate, with audit throughout.
    return ans, trace, clean_q

Build it, part 4: release invariants, not a single gate

Here is where the milestone earns the word "culmination." A release gate of hit@1 >= 0.80 tests one model metric — but the promises table above made six promises to the customer. The gate must test the promises:

flowchart TD CAND[Versioned candidate] --> EVAL[Integration eval] EVAL --> R1[Retrieval
critical cases + unsupported Qs] EVAL --> G1[Citation provenance
+ refusal] EVAL --> S1[Safety
PII tests + action gate] R1 --> OPS[Operations
latency / cost / audit] G1 --> OPS S1 --> OPS OPS --> AC{Acceptance criteria} AC -->|fail| BLOCK[Block release] AC -->|pass| PILOT[Controlled pilot] PILOT --> TRACE[Production tracing] TRACE --> ESET[Failures → eval set]

Humans debate the policy before release; the pipeline applies the agreed policy consistently during release. That separation is what makes the gate trustworthy rather than theatrical. Cost stays in the diagram as production architecture — it becomes an active invariant when the real provider call replaces the stub; there is nothing to bill a deterministic stub.

The golden set has three parts. Six retrieval questions are a smoke eval — a tutorial-sized sanity check, not production release evidence (the pilot uses a larger versioned eval set). Refusal cases test the unsupported-question promise. PII cases test the redaction promise for the three teaching patterns:

GOLDEN_SMOKE = [  # smoke eval only — not release evidence
    ("when is yard waste collected", "yard-waste"),
    ("my trash was not picked up", "missed-trash"),
    ("how do I report a pothole", "potholes"),   # CRITICAL case
    ("when is my water bill due", "water-bill"),
    ("how much is a parking permit", "parking-permits"),
    ("do you take plastic bags for yard waste", "yard-waste"),
]
CRITICAL = {"how do I report a pothole"}
UNSUPPORTED = ["What is the bulk pickup fee?",
               "What is the mayor's home phone number?"]
PII_CASES = [("call 555-010-0199 please", "555-010-0199"),
             ("mail me at joe@example.com", "joe@example.com"),
             ("my SSN is 123-45-6789", "123-45-6789")]

def top_doc(query):
    # Never index [0] blindly: zero hits is a real eval condition,
    # and it must read as a miss, not an IndexError.
    hits = retrieve(query)
    return hits[0][1]["doc_id"] if hits else None

def release_eval():
    results = {}
    # Retrieval: every CRITICAL case must hit — one miss blocks.
    results["retrieval"] = all(
        top_doc(q) == want
        for q, want in GOLDEN_SMOKE if q in CRITICAL)
    # Refusal: unsupported questions refuse, with no fake citations.
    results["refusal"] = all(
        ask(q, scenario="unanswerable")[0].refused for q in UNSUPPORTED)
    # Citation provenance: citations validated against the retrieved set.
    results["citations"] = all(
        validate_citations(ask(q)[0], retrieve(q)) for q, _ in GOLDEN_SMOKE)
    # Safety: tested sensitive values absent from the value passed
    # toward generation.
    results["pii"] = all(
        orig not in ask(raw)[2] for raw, orig in PII_CASES)
    # Safety: the full lifecycle — pending refuses, approved executes,
    # executed refuses replay.
    iid = request_approval("send_notice", {"to": "resident"})
    gate_ok = True
    try:
        execute(iid)
        gate_ok = False          # pending must refuse
    except PermissionError:
        pass
    approve(iid, "on-call")
    try:
        execute(iid)
    except PermissionError:
        gate_ok = False          # approved must execute
    try:
        execute(iid)
        gate_ok = False          # replay must refuse
    except PermissionError:
        pass
    results["action_gate"] = gate_ok
    # Operations: audit verifies; latency within budget.
    results["audit"] = verify_audit()
    results["latency"] = all(ask(q)[1]["latency_ms"] < 1000
                             for q, _ in GOLDEN_SMOKE)
    return results

Run it: three questions and one escalation

Q1 — answerable. "When is yard waste collected in my area?" The real run:

Based on city procedure: Yard waste is collected every Wednesday from March through November. Branches must be bundled and under 4 feet long. Use paper bags or a bin clearly labeled Yard Waste. Plastic bags are not accepted. Set bundles at the curb by 7 AM on collection day. Grass clippings and leaves go in the same paper bags. Do not mix yard waste with regular trash. Mixed bins are left at the curb.
  citations: ['yard-waste#chunk0', 'yard-waste#chunk1']
  refused: False
  trace: {'latency_ms': 0.6, 'model': 'stubbed-generator', 'tokens_in': 88}

Two chunks, two citations — each citation ID was validated against the retrieved set before the answer was accepted. Because the generator here is a deterministic assembly stub, support is structurally simple: the answer is the chunks. This proves citations came from retrieved evidence; it does not prove the evidence supports every claim. Support/grounding is evaluated separately — a real generator requires answer-support evaluation from the Evals lesson. (The overlap between chunks is de-duplicated when the answer is assembled — overlapping chunks that read fine in an index read redundantly in an answer, so the joiner strips the repeated words.)

Q2 — unanswerable. "What is the bulk pickup fee?" Bulk pickup isn't in the procedures, so the honest answer is the product:

I don't know what the bulk pickup fee is — it is not in the procedures I can search. | refused: True

A refused answer with refused: True and zero citations is a successful run — and note the label from the contract section: the stub was told this scenario refuses. It doesn't demonstrate a model determining answerability; in a real system that behavior is something you evaluate. The worst outcome in a resident-facing system isn't "I don't know" — it's a confident fee invented from nothing.

Q3 — contains PII. "Call me at 555-010-0199. What day is yard waste collected?"

generator received: Call me at [REDACTED_PHONE]. What day is yard waste collected?
original phone number reached the generator: no

Phrase it precisely: the original phone number reached the generator — no. That's provable from this run: the generator received the redacted question, and the audit event records that the redaction occurred. It does not prove "PII never reaches the model" in general — the teaching redactor covers three patterns, and a name or street address would sail through. The answer still came back with both yard-waste citations.

The escalation. A write action — drafting an escalation for a 3-week-old missed-trash complaint — must pass the gate. Watch what the queue alone couldn't show you:

queued: APR-001 status: pending
execute before approval -> REFUSED: APR-001 is pending: cannot execute without approval
after authorized approval: status = approved
executed APR-001: draft_escalation
replay after execution -> REFUSED: APR-001 already executed: replay refused

The pending execution attempt was refused — that's the actual security test, not "queued then approved." After an authorized approval the action is marked executed and a replay is refused. The demo prevents replay after marking; production execution also needs idempotency/reconciliation around the real external side effect. CityOps policy here is system executes after approval; an equally valid policy hands the approved item to a human-operated workflow instead — pick one model and stay consistent.

sequenceDiagram participant R as Resident participant A as ask() participant G as PII guard participant V as Retriever participant M as Generator (stub) participant L as Audit log R->>A: question (untrusted data) A->>G: redact PII G-->>A: clean question A->>V: retrieve top-k V-->>A: chunks + scores A->>M: clean question + chunks M-->>A: draft answer A->>A: citations ⊆ retrieved chunks A->>A: validate Pydantic contract A->>L: audit every stage A-->>R: answer + citations (or refusal) Note over A,L: trace + audit stay internal;
the resident sees the response only

Break it, three ways

1. The eval exposes a failure — and the invariants block the release. The smoke eval on the full index:

  HIT  q='when is yard waste collected'             -> yard-waste
  HIT  q='my trash was not picked up'               -> missed-trash
MISS q='how do I report a pothole'                -> yard-waste
  HIT  q='when is my water bill due'                -> water-bill
  HIT  q='how much is a parking permit'             -> parking-permits
  HIT  q='do you take plastic bags for yard waste'  -> yard-waste
[FAIL] retrieval: critical cases hit — pothole -> yard-waste (critical miss)
[PASS] refusal: unsupported questions refuse — 2/2 refused, zero citations fabricated
[PASS] citation provenance: citations ⊆ retrieved chunks
[PASS] safety: tested PII patterns redacted — phone/email/SSN — teaching patterns only
[PASS] safety: unapproved writes cannot execute — pending refused, replay refused
[PASS] operations: audit chain verifies — 94 events
  tamper check: isolated edit -> DETECTED (good)
[PASS] operations: smoke-test latency — every measured request < 1000ms

RELEASE: BLOCKED — invariant(s) ['retrieval'] failed. Fix the system, not the gate.

“How do I report a pothole” retrieved yard waste — the teaching vectorizer matched “report” against the yard-waste procedure's “report a missed yard waste pickup.” A semantic embedding model may handle this phrasing better — but you must measure it on the eval set rather than assume it. The eval identified a concrete failure to investigate (diagnosis still needs work: is it the representation, the chunking, or the threshold?).

Now the important part: the old aggregate-only gate would have said 0.83 >= 0.80, SHIP: yes — waving through a known critical miss because the average looked fine. The invariants say BLOCKED. An aggregate that hides a category failure is not a release criterion; it's a rug. “You can explain the miss” doesn't make a wrong resident answer acceptable.

2. The gate blocks a degraded index. Remove two procedures from the index and re-run — this is what a bad re-ingestion looks like in production:

  HIT  q='when is yard waste collected'   -> yard-waste
  MISS q='how do I report a pothole'      -> yard-waste
  MISS q='when is my water bill due'      -> None
[FAIL] retrieval: critical cases hit — pothole procedure missing from index
RELEASE: BLOCKED — invariant(s) ['retrieval'] failed. Fix the system, not the gate.

No debate, no “but it mostly works.” The pipeline applies the agreed policy consistently: blocking is the default, and shipping requires evidence. (Note the degraded run even “fixed” the pothole miss by removing the yard-waste distractor — a reminder that a passing eval on a broken index proves nothing about the real one.)

3. The audit log detects tampering. Change one entry's detail in the chain and re-verify:

[PASS] operations: audit chain verifies — 94 events
  tamper check: isolated edit -> DETECTED (good)

Each entry's hash commits to the previous entry's hash, so an isolated edit breaks the link after it. And the honest boundary from the build section: this detects isolated edits. A full-chain rewrite needs the independently protected checkpoint — which is why the hash chain plus a protected checkpoint is what Dev shows Legal, not a text file anyone could have rewritten.

Productionize: what breaks at scale

The pilot works on five procedures. Be honest about what changes at five thousand:

At pilot scaleAt city scale
In-memory list of 13 chunksA real vector store with incremental indexing — and a re-index discipline, because stale procedures are wrong answers. Updated procedures replace/deactivate prior indexed versions; rollback restores the previous index snapshot plus the matching model, prompt, and policy versions
Regex PII patterns (3 teaching patterns)Policy-appropriate sensitive-data detection and minimization, tested against representative CityOps data; new PII shapes arrive with every new form
Approval queue in a Python dictA persistent queue with SLAs and an authenticated approver identity — “pending” is only safe if someone is actually watching, and a name string is not authentication
Stubbed generatorA real model behind the provider abstraction — which is when the contract, the guard, the citation check, and the eval gate start earning their keep
One smoke eval of 6 questionsA versioned eval set that promotes representative production failures and important edge cases — curated, not an uncritical pile of every surprising question

The serving layer is the FastAPI pattern from Stage 3 — one endpoint, the pipeline behind it. Note the separation: the request produces both a resident-facing response and internal trace/audit records. The public response carries no internal operational metadata by default — traces stay in the observability and audit systems (syntax-checked, not run here):

from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI(title="Ask CityOps pilot")

class Question(BaseModel):
    text: str

@app.post("/ask")
def ask_endpoint(q: Question):
    answer, trace, _ = ask(q.text)          # the pipeline from this lesson
    # Public response: answer, citations, refusal — no internal trace.
    # The trace and audit events go to the observability/audit systems.
    return {"answer": answer.answer,
            "citations": [c.model_dump() for c in answer.citations],
            "refused": answer.refused}

The pilot ship checklist

  • Release invariants evaluated — smoke eval recorded 5/6 and exposed a pothole failure; the pilot release remains BLOCKED until the agreed critical-case criteria are satisfied. A known miss waved through by an aggregate is not a ship.
  • Guardrails on — phone-number redaction path verified: the generator received the redacted question for the tested pattern (not “PII protection verified”)
  • Audit verified — chain re-hashed, isolated-tamper detection demonstrated; full-history protection requires the independently protected checkpoint
  • Approval gate live — a write action's pre-approval execution attempt was refused; approved actions execute once; replays refused
  • Refusal works — unsupported questions get “I don't know” with zero citations, not an invention
  • Citations grounded — every citation ID validated against the retrieved chunk set
  • Runbook written — who watches the approval queue, who re-runs the evals, what blocks a release
  • Rollback plan — the previous index snapshot is kept with matching model/prompt/policy versions; a bad re-ingestion restores in minutes

Explain it to the customer: the 10-minute demo

You get ten minutes with Lisa. Spend them like this:

Minutes 0–2: the answerable question. Ask about yard waste live. Point at the two citations. “The answer cites the retrieved procedure chunks used to construct it. Residents can check our work.”

Minutes 2–4: the refusal. Ask about the bulk pickup fee. “It says it doesn't know. That's a feature we built on purpose — the alternative is it inventing a fee and a resident budgeting around it.”

Minutes 4–6: the PII proof. Ask a question containing a phone number, then show what the generator received and the audit entry. “The run demonstrates the generator received the redacted question, while the audit event records that the redaction occurred. This is the evidence Legal asked about — for the tested pattern.”

Minutes 6–8: the escalation. Trigger the write action. Show the execution attempt being refused while pending. Approve it. Show it executing once. “The intended write path requires approval, and the release test demonstrates that pending actions cannot execute.”

Minutes 8–10: the gate. Show the invariants and the blocked release. “The release pipeline blocks deployment when acceptance criteria fail — a known pothole miss means we don't ship until it's fixed. That's the release process, not a person remembering to check.”

Then stop. Let Lisa ask questions — a demo that ends early with good questions beats one that runs long with none.

Where this lands: the stage is complete

Look at what the last eleven lessons added up to. You can now:

  • Call model APIs like production infrastructure — timeouts, retries, budgets, fallbacks
  • Force structured output and validate it at the boundary
  • Treat prompts as versioned configuration, not incantations
  • Build retrieval that you can measure — and explain when it misses
  • Ingest messy documents without poisoning the index
  • Run agentic loops with budgets, permissions, and audit trails
  • Choose where the model runs based on the client's constraints, not your preferences
  • Evaluate with golden sets and release against explicit, versioned acceptance criteria
  • Put guardrails, approvals, and hash-chained audit around all of it

And you can demonstrate and test each of those, because every lesson ran real code and this milestone ran the whole system. The Milestone 3 API serves city data; the Milestone 4 system answers questions about city procedures, with citations, guardrails, and a paper trail. The spine of CityOps now has a nervous system.

The FDE principles

Ship the system, not the demo.

Retrieved evidence is not the same as a grounded answer; validate both.

A release gate should test the promises you made to the customer, not just one model metric.

Security controls are valuable when they constrain authority, not when they merely ask the model to behave.

A production system makes expected behavior testable, constrains failures, refuses when evidence is insufficient, records what happened, and blocks release when agreed acceptance criteria fail.

The capstone: a production AI system is not a model plus an API. It is a set of enforceable contracts — around data, evidence, actions, failures, and releases — that you can test and explain to the customer.

Field check

  1. Why does the PII guard sit before the generator in the pipeline, rather than filtering the model's output?
  2. The smoke eval recorded 5/6 with one miss, and the release is blocked. Give two reasons the aggregate 0.83 should make you cautious rather than confident.
  3. A teammate suggests auto-approving escalations “to move faster during the pilot.” Which promise from the clarify-the-ask table does that break, and what's the concrete risk?
  4. The degraded-index run was blocked. What real-world event does “two procedures removed from the index” correspond to?
  5. A resident asks a question the system answers confidently but wrongly — the right chunks were retrieved and the answer cites them. Which component failed, and what would you check first?
Answers

1. Once sensitive data reaches a model boundary that policy says must not receive it, the control has already failed — redaction after the call is damage control, not prevention. With a hosted API that failure may mean the data left your approved environment; with a local model it can still violate data-minimization requirements. The audit entry proves the order: redaction happened before generation.

2. First, six cases is a tiny sample — the margin is one question wide, so a single additional regression flips the decision. Second, the aggregate 0.83 hides a known category failure: the pothole miss is a critical case, and an average that waves through a critical miss is not a release criterion.

3. It breaks “write actions need a human.” The concrete risk: the system drafts an escalation with a wrong address or a wrong complaint attached, and it goes to a supervisor — or a resident — without anyone checking. The gate is the control; auto-approve deletes it.

4. A bad re-ingestion — an ingestion job that silently dropped documents (parse failures, a changed folder path, an expired credential on the document store) and rebuilt the index from a partial corpus. This is why the gate runs against the built index, not the source folder.

5. The generator (or the assembly step) failed — retrieval did its job. Investigate generation and grounding first rather than changing retrieval: were the right chunks actually in the generator's input, and did the citation check validate an answer that contradicts them? Possible fixes include the prompt/grounding instruction, model choice, evidence formatting, the generation strategy, or a support validator — the exact fix depends on why the generator ignored or contradicted the evidence.