Prerequisites: Stages 1–5, and the CityOps spine through Milestone 5: CityOps Goes Live — the Compose stack, the intake worker, the CityOps API, the Ask CityOps copilot from Milestone 4, and the launch-day postmortem. This chapter attacks all of it.

CityOps is live — Compose stack, intake pipeline, CityOps API, the Ask CityOps copilot, audit trail, runbook. This is the week it all gets attacked: by its own suppliers, its own growth, and its own users. Six incidents, one trial. Triage, cut scope, redesign, communicate, ship — and at the end, a scorecard that grades judgment, not heroics. The one principle for the whole chapter: production green is not success. Adoption is the next test; customer outcome is the destination.

Dev (Monday, 8:04 AM): "Morning. Six things are going to break this week. Some of them already have. Your job: run the triage protocol on each one, decide what done looks like before you touch anything, and ship the disciplined fix — not the fancy one. Friday, you brief Maria. The scorecard at the end grades all of it."

Maria: "And no, you don't get extra time. The borough meeting is Friday, the stakeholder demo is Thursday, and the numbers are what they are."

How to read this chapter

This chapter is a simulation. Every number in it — 48 hours, 3 days, a 10% rate limit, 18% adoption, 214 quarantined records — is a scenario constant, not a measured fact. The incidents are teaching fiction; the machinery is real. Every fix below runs against the actual CityOps components from the milestones: the intake worker, the FastAPI service, the copilot pipeline, the audit trail. The code is written to be read and adapted, not copy-pasted into production.

The rules of the trial

Before the first inject, the rules — because a trial without rules is just suffering. Each incident below follows the same shape:

  • The inject — the stakeholder message, verbatim. Read it the way you'd read it at 6 AM: what is actually being asked?
  • What "done" looks like — the acceptance line, written before any fix. If you can't state done, you can't stop.
  • The worked example — a strong response: triage commands, the fix, the output. This is what disciplined looks like.
  • The trap — the fancy fix that loses to the disciplined one. Every incident has one.

The week, as scheduled:

DayIncidentThe clock
MondayLisa blocks raw complaint text from reaching the external model48 hours to the stakeholder demo
TuesdayTom needs the borough meeting pack3 days to the borough meeting
Wednesday AMThe weather API collapses to 10% of its rate limitUnannounced; enrichment degrades now
Wednesday PMRAG evals fall below the acceptance line Maria signedDiagnose, fix, or renegotiate — with evidence
ThursdaySchema change #2: the quarantine queue spikes at dawnTriage at dawn, noon note, demo at 2 PM
Week 3"The system works. Nobody uses it." — 18% adoptionInvestigate, then change one thing

Grading follows the curriculum's milestone scorecard — technical correctness, data reconciliation, reliability, security & compliance, time-to-value, customer outcome, scope discipline, communication, operability, adoption. The full rubric is at the end of this chapter. One rule of the trial, stated up front: the technically fanciest solution sometimes scores worse than the boring one. The scorecard grades the customer outcome, not the architecture.

The minimum concept: the triage protocol

Every incident in this chapter — and, Dev would argue, every incident in the job — runs through one protocol, in this order. The order is the concept:

flowchart TD IN[Inject arrives] --> SC[1. Scope
what's broken, blast radius,
what still works] SC --> CT{Containment needed?
data crossing a boundary
it shouldn't?} CT -- Yes --> C[Contain first
stop the calls,
investigate after] CT -- No --> CA C --> CA[2. Cause
evidence first,
theories later] CA --> DS[3. Data safety
lost, late, or wrong?] DS --> FX[4. Fix
smallest safe change,
through the pipeline] FX --> CM[5. Communicate
what happened, what's true now,
what changes permanently] CM --> OUT[Done = the acceptance line
you wrote before fixing]

Three things about this order that the week will keep proving. First, scope before cause: knowing "the API is fine, only new ingestion is delayed" changes what Maria says at the demo more than knowing why it's delayed. Second, data safety before fix: "nothing is lost, it's late" and "records are corrupted" demand completely different responses, and you cannot choose the response until you know which one you're in. Third, containment can precede cause: when data is actively crossing a boundary it shouldn't — PII reaching an external model, corrupt writes landing in a table — you stop the calls first and investigate after. The protocol orders the thinking, not the hands: containment is a reflex, cause is the diagnosis. Don't read the sequence as "never mitigate until the root cause is known." The protocol is a checklist against panic — triage is a protocol, not a personality.

The one question that starts every incident

Before the first command, ask: "What does done look like, and who decides?" For Lisa's block, Maria and Lisa decide. For Tom's export, Tom decides. For the evals, the acceptance criteria decide. Write the done-line down. Everything after it is execution.

Incident 1: Lisa blocks the model boundary (Monday)

Lisa (Monday, 9:12 AM): "Legal reviewed the copilot's data flow for the demo. Complaint text goes to the external model API — and our spot check found a complaint containing a resident's name and phone number in the provider's request logs. The teaching redactor you demoed covers three patterns. Names, parenthesized phone formats, and addresses sail through. Raw complaint text stops crossing the model boundary today. The stakeholder demo is Thursday. Fix it properly — I will spot-check again."

The spot check, reproduced. This is the input Lisa's team sent:

spot = "Call me at (555) 010-0199 — Maria Santos, 44 Elm St, about my trash."
clean, tags = redact_pii(spot)
# clean: "Call me at (555) 010-0199 — Maria Santos, 44 Elm St, about my trash."
# tags:  []   <-- nothing redacted. The parens format, the name, and the
#              address all sailed through the three-pattern teaching redactor.

The honest reading: the Milestone 4 redactor was always labeled a teaching redactor — three patterns, minimize-first as the production principle. The copilot has since moved from a stubbed generator to a real external model API, and the boundary it crosses is now a legal boundary, not just an architectural one. Lisa isn't asking for better regexes. She's asking for a boundary that fails closed.

What "done" looks like (written before any code, agreed with Lisa): only minimized text that passes the approved boundary policy crosses the external model boundary; detected or uncertain sensitive content is refused and queued for human review; every crossing and refusal is audited with a payload hash; the Thursday demo runs on the new boundary. This reduces risk but does not prove an imperfect detector catches every possible identifier. Who decides: Lisa, by spot-check.

Worked example: minimize, then gate, then audit

The redesign has three layers, in priority order — and the order matters more than any one layer:

Layer 1 — minimize. The model needs the question and the retrieved procedures. It does not need the resident's name, address, or the raw complaint thread. The first fix is architectural, not textual: stop sending the complaint. Send only the redacted question plus retrieved chunks:

def prepare_model_input(question, chunks):
    # MINIMIZE first: the external model receives the question and the
    # retrieved procedures — never the raw complaint thread. What never
    # crosses the boundary can't leak across it.
    clean_q, tags = redact_pii(question)
    return clean_q, chunks, tags

Layer 2 — fail closed. Whatever the redactor can't clear must not cross. A broader detector (names, addresses, the phone formats Lisa found) runs after redaction; anything still PII-shaped is refused and queued for human review:

def model_boundary_check(clean_q, chunks, tags):
    leftover = find_pii_shapes(clean_q)  # broader detector: names,
                                         # addresses, phone variants
    if leftover:
        audit("model_boundary_refused",
              {"shapes": sorted(leftover), "redactions": tags})
        queue_for_review(clean_q, leftover)
        raise BoundaryRefused(
            "PII-shaped content blocked from the model boundary; "
            "queued for human review")
    payload = {"question": clean_q, "chunks": chunks}
    audit("model_boundary_crossed",
          {"payload_sha": sha256_of(payload), "redactions": tags})
    return payload

One boundary subtlety worth stating: queue_for_review(clean_q, leftover) deliberately places sensitive text into another system. That's acceptable only because the review queue is an authorized sensitive-data store with access and retention controls — staffed, with an SLA, per Lisa's line at the end of this incident. Refusing model transmission does not mean the data stops being sensitive; it means the data moved to a boundary built to hold it. This is the data-boundary thinking from the security lessons, applied to your own tooling.

The run against Lisa's spot check:

>>> model_boundary_check(*prepare_model_input(spot, chunks))
audit: model_boundary_refused  shapes=['phone_variant', 'person_name', 'address']
BoundaryRefused: PII-shaped content blocked from the model boundary;
                 queued for human review
# And the clean question — "about my trash", no PII — crosses normally:
>>> model_boundary_check(*prepare_model_input("When is yard waste collected?", chunks))
audit: model_boundary_crossed  payload_sha=9f2c…a41d  redactions=[]

Layer 3 — extend the redactor with the formats Lisa's team actually found (parenthesized phones, +1 prefixes), so fewer legitimate questions hit the review queue. This is layer 3, not layer 1: detection improves coverage and throughput; minimization shrinks the problem the detector must solve, and the gate enforces the policy decision on whatever the detector sees.

Be precise about what this proves, because Lisa will ask: the refusal path is proven (the audit event shows the block and the queue entry). The detection is not proven complete — a novel PII shape could still evade the detector and be labeled clean, in which case the gate opens. The gate enforces the decision; it cannot make an imperfect detector omniscient. And the payload hash is evidence of identity, not safety: it proves this exact payload was the one audited — it says nothing about whether the payload contained PII, and a hash is only as durable as the audit store behind it. Immutability comes from the store's protection, not from the hash alone.

The trap: "just add more regexes." The fancy fix is a bigger pattern list — or worse, a second model that "detects PII" (a model guarding a model, evaluated by nobody, 48 hours before a demo). Both keep the same architecture: send everything, hope the filter is perfect. The disciplined fix changes the architecture: minimize first, detect what remains, block what policy says cannot cross, test the boundary continuously. Under a 48-hour clock, you build the gate first — because the gate is the part that enforces a decision even when detection is uncertain.

Thursday's demo runs on the new boundary. Lisa spot-checks with five new complaints — two refuse-and-queue, three cross clean. She signs off, with one line for the permanent record: "The review queue is a control, not a backlog — staffed, or it's theater." The queue gets an owner and an SLA before the demo ends.

Incident 2: Tom needs the borough pack in 3 days (Tuesday)

Tom (Tuesday, 10:40 AM): "Borough meeting Friday. The president's staff wants the numbers — open reports by borough and category, anything aging past 7 and 30 days, top streets. They asked for CSVs. The export endpoint we have gives them raw filtered rows; they need the rollup. Can we get it by Thursday EOD?"

Before designing anything, clarify the ask — because "CSVs for the meeting" hides three different products. Tom is asked one question: who opens this file, and what do they do with it? Answer: the borough president's staff, who paste the numbers into a slide deck. They need meeting-ready rollups, not raw rows and not a new dashboard. That answer kills two scope-creep candidates on the spot: no interactive report builder, no scheduled email delivery. One file, the right columns, generated from a repeatable endpoint.

What "done" looks like: one CSV (plus a small detail file) with exactly the columns the meeting needs, produced by a versioned endpoint that reuses the existing filtered-query layer; a tradeoff memo listing what was explicitly refused. Who decides: Tom, by opening the file Thursday.

Worked example: one query, one endpoint, one memo

The pack is two queries over the reports table — the existing query layer supplies the authorized filters, the new parts are only the grouping and the ranking:

-- borough_pack.sql: the meeting rollup in one query.
-- Runs INSIDE the authorized borough scope: the query layer applies the
-- session's borough filter before this aggregation — the SQL below is the
-- aggregation portion only. Aging buckets are cumulative thresholds:
-- over_30d is a subset of over_7d ("older than 7", "older than 30"),
-- which is what the meeting wants — not 8–30 / 31+ bands.
SELECT b.borough,
       r.service_code,
       COUNT(*) FILTER (WHERE r.status = 'open')                 AS open_reports,
       COUNT(*) FILTER (WHERE r.status = 'open'
                         AND r.age_days > 30)                  AS over_30d,
       COUNT(*) FILTER (WHERE r.status = 'open'
                         AND r.age_days > 7)                   AS over_7d
FROM reports r
JOIN boroughs b ON r.borough_id = b.id
GROUP BY b.borough, r.service_code
ORDER BY b.borough, open_reports DESC;

-- borough_pack_top_streets.sql: top-5 streets per borough, as the small
-- detail file. "Top" means most reports — so rank by count first, then
-- take five. Kept separate from the rollup so the main query stays
-- readable and the ranking stays testable on its own.
SELECT borough, street, reports FROM (
  SELECT b.borough AS borough,
         r.street  AS street,
         COUNT(*)  AS reports,
         ROW_NUMBER() OVER (PARTITION BY b.borough
                            ORDER BY COUNT(*) DESC) AS rn
  FROM reports r
  JOIN boroughs b ON r.borough_id = b.id
  WHERE r.status = 'open'
  GROUP BY b.borough, r.street
) ranked
WHERE rn <= 5
ORDER BY borough, reports DESC;

The endpoints reuse the Milestone 3 export machinery (same authorization, same streaming response) — the new code is the queries and the routes:

@app.get("/reports/borough-pack")
def borough_pack(auth=Depends(require_authorization())):
    # Reuse the existing authorization dependency and authorized query
    # scope: borough_ids come from the session, never from the request.
    # The SQL above is the aggregation portion, applied inside that scope.
    rows = run_query(BOROUGH_PACK_SQL, {"auth": auth})
    audit("report_generated", {"report": "borough-pack",
                               "rows": len(rows), "by": auth.user})
    return stream_csv(rows, filename="borough-pack.csv")

@app.get("/reports/borough-pack/top-streets")
def borough_pack_top_streets(auth=Depends(require_authorization())):
    # Same authorized scope, second file: the detail Tom pastes under
    # each borough's headline numbers.
    rows = run_query(BOROUGH_PACK_TOP_STREETS_SQL, {"auth": auth})
    return stream_csv(rows, filename="borough-pack-top-streets.csv")

The output Tom opens Thursday:

borough,service_code,open_reports,over_30d,over_7d
North,POTHOLE,412,88,201
North,MISSED_TRASH,305,41,133
South,POTHOLE,388,102,176
...
# borough-pack-top-streets.csv (the detail file):
borough,street,reports
North,Elm St,96
North,3rd Ave,71
...

And the tradeoff memo — the artifact that makes "no" a deliverable. It lists what was asked for or imagined and explicitly refused, with the reason:

The borough-pack tradeoff memo (excerpt)

  • Built: two CSV endpoints — borough × category rollup with cumulative aging buckets, top-5 streets as a small detail file. Reuses the existing authorization dependency and export machinery — no new auth code under the 3-day clock.
  • Refused — scheduled weekly email: a new delivery mechanism with its own failure modes, for a meeting that happens Friday. Tom downloads it Thursday.
  • Refused — Excel formatting / charts in the file: staff paste into their own deck; formatting is their tool's job, not ours.
  • Refused — live "refresh" button on the dashboard: a new interactive surface 3 days before a meeting is how demos die.
The trap: the export framework. The fancy fix is a generic "report builder" — configurable columns, saved reports, a UI. It would take three weeks, miss the meeting, and ship untested. The disciplined fix is one query and one route, built in a day, with the refusals written down. Note the asymmetry: the memo's "refused" list is what protects Friday. Scope discipline isn't saying no to the customer — it's saying no to the second, third, and fourth products hiding inside the first request.

Thursday EOD, Tom downloads the file, pastes the North-borough pothole line into his deck, and the meeting runs on CityOps numbers for the first time. Keep that fact — it matters in Incident 6.

Incident 3: the weather API collapses to 10% (Wednesday AM)

Tom (Wednesday, 7:18 AM): "Weather enrichment is throwing 429s. Provider cut our key to 10% of the quota overnight — no notice, no email, just rate limits. The intake worker is still green, but every report since midnight is missing weather context. Do we pause ingestion until they fix the quota?"

Run the protocol. Scope: the API is fine, 311 ingestion is fine — only the enrichment step is failing. Cause: 429s from the weather provider; quota cut to a tenth, unannounced. Data safety: nothing is lost or wrong — reports are complete except for one advisory field. Which answers Tom's question before it's fully asked: no, you don't pause ingestion. Pausing the product because a garnish failed is the incident, not the response.

What "done" looks like: every 311 report still ingests on schedule; reports carry weather with explicit provenance — weather_status ("fresh", "cached", or "unavailable") plus weather_observed_at — so no consumer can mistake degraded context for clean data; no retry storm burns the remaining 10% quota; the remaining quota is reserved for the workloads the customer has explicitly prioritized (Maria chose the Thursday demo queries); an alert fires if degradation lasts past the provider's stated recovery window. Who decides: the data contract — downstream consumers must be able to tell "no weather data" from "clear skies."

Worked example: degrade on purpose

The enrichment worker gets a degradation policy — three branches, no surprises:

def enrich_with_weather(report, weather_client, cache):
    # Enrichment is ADVISORY. The 311 record is the product; weather is
    # context. If the API is sick, the record still ships — with provenance.
    try:
        w, observed_at = weather_client.get(report["date"], report["borough"])
        return {**report, "weather": w, "weather_status": "fresh",
                "weather_observed_at": observed_at}
    except RateLimited:
        # 10% quota: do NOT burn it retrying. Use a cached observation
        # only if it actually corresponds to this report's date/borough —
        # yesterday's weather is not today's weather, and a stale
        # substitution would make the data wrong, not just degraded.
        hit = cache.matching(report["date"], report["borough"])
        if hit is not None:
            audit("weather_degraded",
                  {"report": report["id"], "fallback": "cached"})
            return {**report, "weather": hit.value,
                    "weather_status": "cached",
                    "weather_observed_at": hit.observed_at}
        audit("weather_degraded",
              {"report": report["id"], "fallback": "unavailable"})
        return {**report, "weather": None, "weather_status": "unavailable",
                "weather_observed_at": None}

The worker log at 7:30, after the deploy:

07:31 enrich: 429 from weather API (quota 10%) -> degraded mode
07:31 enrich: report 10412 -> weather_status=cached, observed_at=2026-10-06
07:31 ingest: 47/47 reports stored, 47 with weather_status != fresh
07:31 alert: weather degradation > 60min -> page Tom (sustained, not transient)

The load-bearing detail is the provenance. A boolean flag can't carry what the consumer needs to know: weather_status distinguishes fresh from cached (with when it was observed) from unavailable — four different facts a single bit would collapse into one. A downstream query asking "pothole reports on rainy days" can now decide: fresh rows for the correlation, cached rows with their observation dates for context, unavailable rows excluded. Without that distinction, "no weather data" silently becomes "clear skies" and someone briefs the borough president on a correlation that doesn't exist. Degraded data must be distinguishable from clean data, or the degradation is a corruption. That's a data-contract rule, and it goes into the contract the Pydantic lesson taught: the schema now documents what each status value means.

The trap: retrying harder. The fancy fix is a clever retry policy — exponential backoff with jitter, a queue, a heroic effort to squeeze the full enrichment out of 10% quota. It burns the remaining quota by 9 AM and teaches the provider's rate limiter your IP address. The disciplined fix accepts the degraded state, reserves the remaining quota for the workloads the customer has explicitly prioritized (Maria chose the Thursday demo queries — quota follows business priority, not demo optics), and makes the degradation visible instead of heroic. Graceful degradation is a design decision you make on Wednesday morning, not a retry loop you tune.

Productionize note: the provider restores the quota Thursday. The permanent changes: the weather call gets a quota budget (the limiter that controls the allowed request rate), a cache (fewer calls), and a circuit breaker (stop calling when the dependency is failing — a different job from the budget, not a replacement for it); the weather_status and weather_observed_at columns become permanent, with contract documentation; and the runbook gains a "third-party degradation" page — because the next 10% collapse won't announce itself either.

Incident 4: the evals fall below Maria's line (Wednesday PM)

Dev (Wednesday, 2:05 PM): "Bad news from CI. The RAG evals dropped below the acceptance line — retrieval hit-rate 0.81 against the 0.85 floor Maria signed off on. Nothing in the copilot code changed this week. The demo is tomorrow. Options: I can tune the retriever tonight, or we talk to Maria about the threshold. I don't love either."

This is the incident the whole evals lesson was building toward. The instinct — Dev's stated options — is to change the system or change the bar, tonight, under pressure. The protocol says: scope, cause, data safety first. Scope: which eval dimension failed? Only retrieval hit-rate; refusal, citations, safety, and latency are green. Cause: unknown — nothing in the system changed. Data safety: no production impact; the gate did its job and blocked the release.

What "done" looks like: the cause of the regression is identified with evidence (system, eval set, or measuring stick); either the system is fixed through the normal pipeline, or the threshold is renegotiated with Maria with the per-slice evidence on the table — never silently. Who decides: Maria, as the domain owner who signed the original line.

Worked example: version all three, then talk

The evals lesson's durable rule — version the system, the eval set, and the measuring stick — is the diagnosis procedure. Run it:

$ git log --oneline -- evals/golden.json | head -3
a91f3c2  add 40 production-log questions to golden set   # <-- the change
7d44e10  release v3: acceptance 0.85 (Maria sign-off)
$ eval_report --slice original-120
  retrieval hit-rate: 0.88   (floor 0.85 — PASSES)
$ eval_report --slice new-40
  retrieval hit-rate: 0.62   (never reviewed, never signed off)

No regression is observed on the original signed-off slice. The measuring stick changed: 40 new questions, pulled from production logs, were added to the golden set without domain-owner review — and they're harder in a specific way. Initial inspection suggests they use long, conversational phrasings the TF-IDF-era chunker was never tuned for; root-cause work is scheduled after the domain review, because evals detect failure and diagnosis locates cause. The original 120 questions — the set Maria actually signed off on — still pass at 0.88. And note what the new slice might be telling us: production-derived questions could reveal a real capability gap in current production traffic. That's a separate, unresolved question — not a cleared one.

So Dev doesn't tune the retriever tonight (an untested change before a demo, chasing an unreviewed bar), and nobody edits the threshold in CI to make it green. The worked response is a conversation with evidence. Dev brings Maria the slice breakdown and two honest options:

Dev to Maria: "The 0.81 is two different stories. The 120 questions you signed off on: 0.88, still passing. The 40 new ones from production logs: 0.62 — and nobody reviewed whether those questions are fair or what 'correct' means for them yet. My proposal: tomorrow's demo runs on the signed-off set, the gate stays at 0.85 on that set, and we schedule a review session where you decide which of the 40 belong in the golden set and what the bar should be. If the review says the bar moves, it moves — with your sign-off, not my commit."

Maria: "And if the review says 0.85 was too generous?"

Dev: "Then we change the system until it earns the number, through the pipeline, with the evals watching. The number follows the evidence — never the deadline."

Maria agrees. The demo runs on the signed-off set. The review session goes on the calendar for next week, owned by Maria — because the golden set's definition of correct behavior is owned by the domain expert (she decides which questions and expected outcomes are representative and acceptable), and engineering owns the machinery that maps retrieval metrics onto that contract. That's the evals lesson's split, applied under pressure instead of in a textbook.

And the 40 questions don't disappear: the dashboard gains a second eval line that week — the signed-off set (gating, 0.88) and the production slice (reported, 0.62, unresolved). The 0.62 doesn't gate the release, but it doesn't get to hide either. Running the demo on the signed-off set while reporting the new slice separately is governance; letting the bad number vanish behind a redefined denominator would be the midnight-commit behavior the trap section condemns.

The trap: grade inflation. The fancy fix is a midnight commit — a tweaked threshold, a "temporarily" excluded slice, a retriever tuned against the test set until the number turns green. It ships a demo on a lie and teaches the team that the gate is negotiable by whoever is most tired. The disciplined fix keeps the gate sacred and moves the conversation: renegotiate the bar with evidence, in daylight, with the person who owns it. A threshold changed with evidence and a sign-off is governance. A threshold changed in CI at midnight is fraud with extra steps.

Incident 5: dawn triage — the quarantine spike (Thursday)

Tom (Thursday, 6:40 AM): "Quarantine queue jumped overnight — 214 records, all 'unmappable service_code'. The contract test is green. The demo is at 2 PM. Is this the 2:14 AM thing again?"

It is and it isn't. The Milestone 5 incident was a field rename — category became service_code — and the machinery built after it (tolerant parser, quarantine queue, hourly contract test) is exactly what's working now: nothing crashed, the records were set aside instead of lost, and Tom saw a count instead of a flatline. But the contract test asserts field presence, and this time the field is present. Its values changed.

Run the protocol. Scope: API fine, dashboard fine, 214 reports quarantined — the borough pack Tom built Tuesday used yesterday's data, so the meeting numbers are unaffected. Cause (one sample):

$ psql -c "SELECT raw->>'service_code' AS code, COUNT(*)
           FROM quarantine
           WHERE created_at > now() - interval '12 hours'
           GROUP BY 1 ORDER BY 2 DESC LIMIT 5;"
 code | count
------+-------
 312  |   214
# Yesterday: 'POTHOLE', 'MISSED_TRASH', ...  Today: numeric codes.
# Same field name. Different language. The contract test checks names.

Data safety: late, not lost, not wrong — quarantine did its job. The fix is a mapping, not a rescue.

What "done" looks like: the 214 records mapped and backfilled before the 2 PM demo — 214 quarantined, 214 mapped, 214 accepted, 0 unresolved; the mapping lives in a versioned table with effective dates, not in parser code; the contract test gains a value-domain monitor (a drift signal, not a presence assertion); a noon note (not a war-room postmortem — the machinery worked) on Maria's desk. Who decides: the demo clock.

Worked example: map it, version it, monitor the value domain

The fix is small because the architecture is right. A versioned mapping table — with effective dates, because code 312 means POTHOLE only under codebook v3. If the city ships v4 with 312 reassigned, each record must map under the codebook version in force when it arrived; without effective dates, a versioned table exists but the application can't know which version applies. And the next codebook change shouldn't require a code deploy to understand:

-- service_code_map: the city's vocabulary, versioned with effective dates.
-- v3 added 2026-10-08 after the numeric-code switch; source is the official
-- codebook, confirmed with the named data owner — a phone call alone is not
-- durable mapping authority.
INSERT INTO service_code_map (code, name, version, effective_from, effective_to, source) VALUES
  ('312', 'POTHOLE',      3, '2026-10-08', NULL, 'city codebook 2026-10; confirmed with the data owner'),
  ('318', 'MISSED_TRASH', 3, '2026-10-08', NULL, 'city codebook 2026-10; confirmed with the data owner');
-- Backfill with reconciliation: 214 quarantined → 214 mapped →
-- 214 re-ingested and accepted → 0 unresolved. Queue zero alone proves
-- nothing; the counts must tie.
$ psql -c "SELECT COUNT(*) AS quarantined FROM quarantine
           WHERE reason='unmappable service_code';"
 quarantined
-------------
         214
-- (map + re-ingest)
$ psql -c "SELECT COUNT(*) AS accepted FROM reports
           WHERE service_code_origin='quarantine-backfill-2026-10-08';"
 accepted
----------
      214
$ psql -c "SELECT COUNT(*) AS unresolved FROM quarantine
           WHERE reason='unmappable service_code';"
 unresolved
------------
          0

And the contract test grows a second sense — a value-domain monitor. Presence checks catch renames; the monitor watches the value domain for semantic drift. It raises for investigation; it doesn't assert. A statistical distribution change is a drift signal, not a contract violation:

def value_domain_monitor():
    # A DRIFT SIGNAL, not a contract. Asserting "yesterday's top-10 values
    # must all appear today" would false-positive on normal variation —
    # a quiet graffiti day is not a schema change.
    vals = today_values("service_code")
    unknown_rate = sum(1 for v in vals if v not in KNOWN_DOMAIN) / len(vals)
    coverage = len(set(KNOWN_DOMAIN) & set(vals)) / len(KNOWN_DOMAIN)
    if unknown_rate > UNKNOWN_RATE_ALERT or coverage < COVERAGE_FLOOR:
        raise_for_investigation(
            f"value-domain drift: unknown_rate={unknown_rate:.3f}, "
            f"known_coverage={coverage:.2f}")

By 11:30 the backfill is done and the dashboard is current. The noon note is one page — what changed, what the mapping table is, the new monitor — and Maria forwards it without editing. Note the difference from Milestone 5's postmortem: that one documented a system being brittle to a normal event. This one documents a system absorbing a normal event. The quarantine queue, the versioned mapping table, the drift monitor — that's what "productionize" bought. And one sentence in the noon note is worth its weight: quarantine converted a potentially wrong interpretation into merely late data. That's why the queue matters.

The trap: hardcoding the mapping. The fancy fix is a dict in the parser — {"312": "POTHOLE"} — deployed by 8 AM. It works until the city ships codebook v4 and the mapping needs a deploy, a review, and a developer who remembers where the dict lives. The disciplined fix puts the vocabulary in a versioned table with a source column, because the mapping is data about the world, not logic. Logic changes with releases; world-knowledge changes with the world.

The 2 PM demo runs on current data. Nobody in the room hears about the 214 records — and that's the point worth naming: the best incident response is the one the customer never needs to hear about. Not because it was hidden, but because the system absorbed it before it became their problem.

Incident 6: "The system works. Nobody uses it." (Week 3)

Maria (two weeks after launch): "I pulled the usage numbers for the ops review. Of 128 operations staff, 23 touched CityOps last week — dashboard, copilot, or export. That's 18%. The system is green, the data is current, the demo went well — and four out of five of the people it was built for aren't in it. I need to know why, and I need a plan that isn't 'more features.'"

This is the incident with no pager and no stack trace, which is why it's the hardest one in the chapter — and the most valuable. Everything is working. That's the problem. The triage protocol still applies, but the instruments change: instead of logs, you investigate with people.

What "done" looks like: a written diagnosis naming the specific frictions (ranked by evidence, not opinion), one change shipped to remove the top friction, and a re-measurement date. Explicitly not done: a feature roadmap. Who decides: Maria, against the next usage pull.

Worked example: measure the funnel, then ask humans

Start with what the audit trail already knows — the funnel, not the total:

-- Scenario constants: audit-derived, week of 2026-10-19.
-- logins: 41 | dashboard_view: 31 | ask_or_export: 23  (of 128 staff)
SELECT event, COUNT(DISTINCT user_id) AS users
FROM audit
WHERE event IN ('login','dashboard_view','ask','export')
  AND ts > now() - interval '7 days'
GROUP BY 1;
-- The drop is steepest at login -> dashboard_view. 87 of 128 had no
-- recorded login that week: the metric is absence, not friction.
-- The interviews identify the friction.

Then the human instrument: three 20-minute interviews and one shadowed shift. Initial qualitative findings, prioritized using funnel size plus interview evidence:

  1. The door is heavy. CityOps needs VPN plus a separate login; the old workflow was a spreadsheet on the shared drive. 87 of 128 staff had no recorded login that week — they may have been off shift, on leave, using another authorized workflow, or simply unaware of CityOps. The metric records absence, not friction. The interviews then identified VPN + separate login as a major access friction for those who did try — and "I'll check it later" is how 18% happens. Evidence first.
  2. The meeting runs on slides. The weekly ops meeting — the one place these numbers matter — runs on a slide deck. The borough pack exists (Incident 2 built it), but nobody showed the ops leads the one-click path from dashboard to deck. The export they need is three clicks away and nobody knows it. Workflow mismatch, not missing feature.
  3. The copilot's first impression was the pothole miss. The earliest story anyone heard about Ask CityOps was the eval failure from the pilot — "it couldn't answer a pothole question." Nobody saw the refusal behavior, the citations, or the blocked release that followed. Trust debt, not a trust problem.

Notice what the investigation didn't find: nobody asked for a new feature. Every friction is about the path to the existing value, not the value itself. So the plan is one change, not a roadmap: a 20-minute floor session — Dev walks the ops leads through the one-click borough pack (their meeting, their deck, their numbers) and live-demos the copilot refusing an unanswerable question with citations on the answerable ones. The VPN+login simplification goes on the backlog as the next friction, owned and dated — because the door is the biggest drop, but the session is the fastest win.

The sequencing rule worth naming: prioritize not only by impact, but by evidence, reversibility, effort, and speed to learning. The door is the biggest drop in the funnel; the floor session is the fastest validated learning — it attacks the second and third frictions immediately, and a visible win builds the political capital for the slower, riskier auth work. Biggest drop ≠ first change.

The re-measurement is set for two weeks out, with the honest caveat stated up front: a before/after comparison can't isolate the session from everything else that happens in two weeks. The team records the confounders they know about (the borough meeting cycle, a staffing change in the South office) alongside the number. Measure anyway, and stay honest about what the measurement can't prove — that's the difference between evidence and theater.

One more honesty upgrade: "23 touched CityOps" is reach, not adoption. The ladder is reach → activation → meaningful workflow use → repeat use → customer outcome. A login isn't adoption; neither is one dashboard view. Tom's meeting running on CityOps numbers back in Incident 2 is the first genuine customer outcome in this chapter — that's the rung that counts.

The trap: "they'll come when we add X." The fancy fix is a feature — a mobile view, a Slack bot, a redesigned dashboard. It assumes the problem is the product. The evidence says the problem is the path: the door, the deck, the first impression. Features don't fix friction; they add surface area to it. The disciplined move is subtraction and demonstration — remove one obstacle, show the value that was already there. Production green is not success. Adoption is the next test; customer outcome is the destination.

Communicate: the Friday briefing

Friday, 4 PM. Dev gets five minutes with Maria — the whole week, compressed into the format a customer can actually use. Not a war story: a status. What broke, what's true now, what changed permanently, what's next:

Dev to Maria: "Four things. One — the copilot boundary: Lisa's spot check found PII reaching the external model. We now minimize what crosses the boundary, refuse-and-queue anything detected or uncertain, and audit every crossing and refusal. Lisa re-checked and signed off. Two — the borough pack shipped Thursday; Tom's meeting ran on our numbers. The tradeoff memo lists what we refused. Three — the weather API cut us to 10% quota Wednesday; ingestion never stopped, reports carry explicit provenance, and the provenance columns are now permanent contract fields. Four — the evals: the drop was a measuring-stick change, not a system regression. The gate stays at 0.85 on the set you signed; the 40 new questions are reported as a separate, unresolved risk and go through your review next week before they gate anything."

Maria: "And the 214 records Thursday morning?"

Dev: "Mapped and backfilled before the demo — the quarantine queue worked the way we designed it. The noon note has the details. And the adoption number: 18%. We know the three frictions, we're running the floor session Tuesday, and we re-measure in two weeks. That's the whole week."

Study the shape of this briefing, because it's the skill the chapter is really teaching. Every item follows the same grammar: what happened, what's true now, what changed permanently. No item ends at "we fixed it" — each one ends at the mechanism that makes it cheaper next time. And the adoption number is reported with the same neutrality as the 429s: bad news, delivered early, with a plan attached. The briefing is a deliverable — Maria forwards its shape to her leadership, the way she forwarded the postmortem.

The scorecard: grading the trial

The curriculum's milestone scorecard, applied to the week. Each dimension names what "good" looked like this week — and the trap it punishes:

DimensionWhat "good" looked likeThe trap it punishes
Technical correctnessFixes handle the general case (tolerant parser, versioned mapping table), edge cases quarantined not crashedThe hardcoded dict; the regex-only PII fix
Data reconciliation214 quarantined → 214 mapped → 214 accepted → 0 unresolved; every record accounted for"Mostly backfilled"; uncounted losses
ReliabilityDegradation by design (weather provenance, circuit breaker); the contract test's new value-domain monitorRetry heroics; silent degradation
Security & complianceModel boundary minimizes input, blocks detected/uncertain sensitive content, audits crossings and refusals; Lisa's re-check signed offSend-and-hope; the detector nobody evaluated
Time-to-valueBorough pack in 3 days; PII gate in 48 hours — smallest slice that meets the clockThe report builder; the re-platform
Customer outcomeTom's meeting ran on CityOps numbers; Maria could forward every document uneditedA green dashboard nobody opens
Scope disciplineThe tradeoff memo's refused list; the demo on the signed-off eval setThe second, third, and fourth products inside the request
CommunicationThe Friday briefing grammar: what happened, what's true now, what changed permanentlyThe war story; the softened postmortem
OperabilityRunbook followed at dawn; quarantine review SLA owned; on-call never improvisesHeroics as a process
Adoption18% measured honestly; one friction removed; re-measurement dated"They'll come when we add X"

Read the right column as a whole: every trap is the same mistake wearing different clothes — optimizing for impressiveness over outcomes. The scorecard's central lesson, stated one final time: the best solution isn't the most impressive architecture — it's the smallest reliable system that creates the required customer outcome and survives production. This week, that meant a gate instead of a detector, a query instead of a framework, a conversation instead of a midnight commit, a mapping table instead of a dict, and a floor session instead of a feature.

CityOps: the job, experienced

Step back and name what the trial proved. Not that CityOps works — the milestones proved that. That it survives: its suppliers (the weather quota, the codebook), its own growth (the eval set, the model boundary), and its users (the 18%). Each incident converted a lesson into a system property — the boundary gate, the tradeoff memo, the degradation provenance, the versioned golden set, the value-domain monitor, the adoption funnel. Productionize, across the whole chapter, meant the same thing it meant in every lesson: turn the lesson into a mechanism.

And the arc of the full curriculum lands here. Stage 1 taught you to scope the vague ask. Stage 2 taught you to move data without losing it. Stage 3 taught you to integrate without trusting. Stage 4 taught you to put contracts around intelligence. Stage 5 taught you to operate what you shipped. Stage 6 taught you to present the evidence. This chapter was the exam for all of it — and the exam's real subject was never the incidents. It was the judgment: what to fix, what to refuse, what to measure, and what to tell the customer on Friday at 4 PM.

If the whole curriculum had to fit on one card, it would be this — the synthesis of every incident in the chapter:

flowchart TD REQ[Incident / request] --> DONE[Define done] DONE --> SC[Scope] SC --> CT{Immediate containment needed?} CT -- Yes --> C[Contain] CT -- No --> CA C --> CA[Cause] CA --> DS{Data safety} DS -- Lost --> L[Recover] DS -- Late --> LT[Backfill / wait] DS -- Wrong --> W[Quarantine / correct] L --> FX LT --> FX W --> FX[Smallest safe fix] FX --> VF[Verify / reconcile] VF --> CM[Communicate] CM --> PM[Permanent mechanism] PM --> USE[Customer use] USE --> OUT[Customer outcome]

Five lines to leave with. Triage is a protocol, not a personality. Define "done" before you start fixing. Under pressure, the smallest safe change usually beats the most impressive architecture. Every incident should leave behind a mechanism that makes the next occurrence cheaper. And the last one, upgraded: production green is not success — adoption is the next test, customer outcome is the destination.

The chapter principle

Production green is not success. Adoption is the next test; customer outcome is the destination. The hierarchy: the system works → people use it → the workflow improves → the customer outcome improves. A harmful or inefficient system can have high adoption — green measures the system, adoption measures the usage, and only the outcome measures success. Build for the last one.

Field check

  1. Lisa's spot check found a phone number the redactor missed. Why is "extend the redactor" layer 3 of the fix instead of layer 1 — and what would go wrong if it were layer 1?
  2. Tom's borough pack reused the Milestone 3 export machinery. Name two things the reuse bought the team under the 3-day clock, beyond "less code."
  3. The weather enrichment degraded instead of pausing ingestion. A teammate argues the flagged records are "dirty data" and should have been held back. What's the honest counter-argument — and what makes the flag load-bearing?
  4. The eval regression turned out to be a measuring-stick change. Give two reasons Dev was right not to tune the retriever that night, even though tuning might have pushed the number back over 0.85.
  5. Incident 5's contract test stayed green while values drifted. What does that teach you about what contract tests actually assert — and what's the general fix?
  6. The adoption investigation found three frictions and zero feature requests. Why is "the door is heavy" (VPN + login) the biggest drop in the funnel, yet the floor session the first change shipped?
Answers

1. Because detection can always miss the next format — layer 1 (minimize) shrinks what can leak, and layer 2 (fail closed) makes "uncertain" mean "blocked," regardless of what the next format looks like. If redaction were layer 1, the architecture would still be "send everything, hope the filter is perfect" — and the next spot check would find the next format. Minimize first. Detect what remains. Block what policy says cannot cross. Test the boundary continuously.

2. First, the auth model came free and correct — borough staff see only their borough, with no new authorization code written under pressure. Second, the audit event came free — every generated report is logged with the row count and the requester, so the pack is traceable without extra work. Reuse under a deadline isn't about typing less; it's about inheriting tested properties.

3. Holding the records back would have turned a degraded enrichment into missing 311 data — trading an honest partial record for a gap. The counter-argument: the 311 report is the product and it was complete; only the advisory context was degraded. The flag is load-bearing because it keeps "no weather data" distinguishable from "clear skies" — without it, the degradation becomes silent corruption of every downstream weather correlation.

4. First, tuning against an unreviewed bar optimizes for the wrong target — the 40 new questions had no agreed definition of correct, so "passing" them proves nothing. Second, a night-before change ships without the pipeline's normal scrutiny (no eval-gated release, no review), trading a known-good system for an unknown one hours before a demo. The disciplined move was the conversation, not the commit.

5. Contract tests assert what you wrote down — this one asserted field presence because the last incident was a rename. The general fix: assert the properties the consumer actually depends on. If downstream logic branches on values, the contract must watch the value domain (the drift monitor). A distribution change is a drift signal for investigation, not a contract violation. Every incident teaches the next assertion; contracts grow the way checklists do.

6. The door is the biggest drop in the funnel (87 of 128 with no recorded login that week), but fixing auth is a slower, riskier change involving security review — while the floor session is shippable Tuesday and attacks the second and third frictions (the unknown one-click path, the trust debt) immediately. Sequence by reversibility and speed: demonstrate value first (which also builds the political capital for the auth work), then fix the door. Biggest drop ≠ first change; fastest validated learning does.