Prerequisites: Stage 5 posts 1–7 (Compose and container basics, CI/CD, health checks and monitoring, secrets management, RBAC, the security-review evidence pack) — referenced by title, as they ship in this stage — plus the Milestone 3 CityOps API and the Stage 4 model-operations standards (liveness vs. readiness, deployment identity, timeouts from SLOs).

Tom (ops, 6:12 AM, launch day): "The intake pipeline broke overnight. Zero records since 2:14 AM. The city's 311 feed changed something — I don't know what yet."

Maria (customer): "The mayor's office demo is at 9:00. They open the dashboard, they expect to see this morning's reports. What do I tell them?"

Dev: "Give us until 8. We'll have it fixed, verified, and written up. And Maria — the write-up is part of the deliverable, not an apology note."

Milestone 5 is the production cutover: CityOps goes live on real infrastructure, deployed by pipeline, watched by monitors, guarded by secrets and roles — and then, on launch morning, it breaks. This is the capstone because it integrates everything: the Compose stack, the CI/CD pipeline, the health checks, the audit trail, the evidence pack. And it proves the hardest lesson of the stage: the incident is part of the product.

The customer problem

Two problems arrive at once, which is exactly how production works. The planned problem: cut CityOps over from the pilot environment to production — real Compose stack, real CI/CD, real secrets, real monitoring, real runbook. The unplanned problem: at 2:14 AM on launch day, the city's 311 feed ships a schema change, the intake pipeline crashes on it, and by 6 AM the dashboard is showing a flat line where this morning's reports should be.

Maria's 9 AM demo is the forcing function. She doesn't need a root-cause analysis by 9 — she needs a working dashboard and an honest story. Dev's team has until 8 AM to do three things: triage (what broke, what's the blast radius), fix (restore ingestion without losing data), and communicate (a blameless, client-facing postmortem on Maria's desk by noon).

This is why the milestone is the right end to Stage 5. "Deploy Like a Pro" was never about YAML — it was about what happens when the YAML meets reality. The launch-day break is the exam.

Clarify the ask

In the war-room call at 6:30, the team separates the two asks before touching anything — because conflating "fix the incident" with "finish the cutover" is how you make both worse.

Ask 1 — the cutover (planned): CityOps runs in production behind the Compose stack built across this stage: containers for the API, Postgres, and the intake worker; CI/CD deploying on merge; health checks gating traffic; secrets from the manager, not env files; RBAC on every endpoint; dashboards on metrics; the audit trail from the security-review post recording every sensitive action. Done means Maria can point the mayor's office at a URL and Tom can go back to sleep.

Ask 2 — the incident (unplanned): restore the 311 feed ingestion, backfill the missed records, and deliver the postmortem by noon. Done means the dashboard is current and Maria has a document she can forward to the mayor's office without editing it.

Dev states the principle that governs the whole milestone: a milestone is integration, not invention. Nothing in this post is new technology. The Compose file, the pipeline, the health checks, the audit middleware, the evidence pack — all built in posts 1–7. The milestone's job is to compose them into a system that survives contact with a real Tuesday morning. If you find yourself inventing a new mechanism during a milestone, stop: you're papering over a gap in the earlier posts.

Must know

  • A milestone is integration, not invention. Compose what the stage built; don't invent new machinery under pressure.
  • Separate the planned work (cutover) from the unplanned work (incident) before you touch anything. Different goals, different clocks.
  • The postmortem is a deliverable, not an apology. Maria should be able to forward it unedited.

The minimum concept

The cutover needs one concept, and the incident needs one concept. Both are small.

Cutover concept: the deployment is a closed loop. Code moves through CI (build, test, scan) into images; images move through CD into the Compose stack; the stack reports health; health gates traffic; traffic generates metrics and audit records; metrics feed monitors; monitors page Tom. Every arrow in that loop was built in an earlier post. The milestone just closes it and proves it runs end to end.

Incident concept: the postmortem is the product. The fix restores service; the postmortem restores trust. A blameless postmortem answers five questions — what happened, what was the impact, what was the root cause (a process gap, never a person), what did we learn, what changes now — in language a non-engineer can forward. Maria's demo audience doesn't need to know what a KeyError is. They need to know it won't silently eat their data again.

The integration map: where posts 1–7 live in the final system

"Integration, not invention" is checkable. Every component of the production system below was built in an earlier Stage 5 post — if you can't point at the post, you're inventing:

Stage 5 post (by title)What it builtWhere it runs at launch
Containers & ComposeImages, Compose services, networksThe three-service prod stack
CI/CD pipelinesTest → scan → build → deploy loopDeploys the 7:40 incident fix
Health checks & monitoringLiveness/readiness, metrics, alertsGates traffic; pages Tom at 6:12
Secrets managementManager-backed, file-mounted secretsDB passwords, signing keys
RBAC & access controlRoles, least-privilege service accountsEvery API endpoint; the insert-only worker
Security review evidenceAudit trail, PII inventory, access matrixEvery sensitive action logged; Lisa's pack
This milestoneComposition + the runbook + the postmortem — the only new artifacts

The milestone's genuinely new content is small on purpose: the runbook (an operations document) and the postmortem (a communication document). Everything else is wiring. When a milestone's "new" column is longer than its "integrated" column, the stage plan had a hole.

graph TD A[Git push to main] --> B[CI: test + PII log-scan
+ image build] B --> C[CD: deploy Compose stack] C --> D{Health checks
liveness + readiness} D -->|pass| E[Traffic to new containers] D -->|fail| F[Automatic rollback
previous images stay] E --> G[Metrics → dashboards
Audit → immutable log] G --> H[Monitors → page Tom] H -.->|incident| I[Postmortem
by noon]

Build it: the production cutover

The cutover is assembly, not invention. Each piece below comes from its Stage 5 post (referenced by title); what's new is that they now run together, against production data, behind one URL.

The Compose stack

Three services, one network, no exposed database port. The API serves traffic, the worker ingests the 311 feed, Postgres holds the reports. Secrets arrive as files from the secrets manager (Stage 5's secrets post), never as environment variables in the compose file:

# docker-compose.prod.yml — the production cutover
services:
  api:
    image: registry.internal/cityops/api:${IMAGE_TAG:?IMAGE_TAG required}
    depends_on:
      db:
        condition: service_healthy
    secrets:
      - db_password
      - api_signing_key
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/ready"]
      interval: 15s
      timeout: 5s
      retries: 3
      start_period: 30s
    deploy:
      resources:
        limits: {cpus: "2", memory: 1G}
    networks: [cityops]

  worker:
    image: registry.internal/cityops/worker:${IMAGE_TAG:?IMAGE_TAG required}
    depends_on:
      db:
        condition: service_healthy
    secrets:
      - db_password
    command: ["python", "-m", "ingest.poll_311", "--interval", "300"]
    networks: [cityops]

  db:
    image: postgres:16
    volumes:
      - pgdata:/var/lib/postgresql/data
    secrets:
      - db_password
    environment:
      POSTGRES_PASSWORD_FILE: /run/secrets/db_password
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U cityops"]
      interval: 10s
      retries: 5
    networks: [cityops]
    # no ports: published — the database is not reachable from outside
    # this Compose network. That absence is a security control.

volumes:
  pgdata:

secrets:
  db_password:
    external: true
  api_signing_key:
    external: true

networks:
  cityops:

Read the security properties out of the YAML: the database has no published port (least privilege at the network layer), secrets are external (rotation doesn't require redeploys), every service has a health check, and the image tag is mandatory — no latest, because "pin the exact tag" may not be enough for models (Stage 4), but for containers an immutable digest-pinned tag is exactly the deployment identity you want.

Health checks: liveness vs. readiness, correctly

The Stage 4 model-operations standard applies verbatim here: liveness proves the process responds; readiness proves it can serve. A health GET that returns 200 only proves the server is up. The /ready endpoint the Compose file hits must prove the dependencies are actually usable:

@app.get("/live")
def liveness():
    return {"status": "ok"}  # the process is alive. That's all this claims.

@app.get("/ready")
def readiness(db=Depends(get_db)):
    # Readiness proves the API can actually serve requests:
    try:
        db.execute("SELECT 1")          # database reachable?
        model_registry.ping()           # required model artifact loadable?
    except Exception as e:
        raise HTTPException(503, f"not ready: {type(e).__name__}")
    return {"status": "ready",
            "model": model_registry.identity()}  # immutable artifact id

This distinction is load-bearing in the incident below: during the outage, /live stayed green while /ready would have caught a sick dependency. Monitors watch readiness; pagers fire on readiness. Liveness is for the orchestrator's restart decisions only.

The CI/CD pipeline

Every merge to main runs the same loop — test, scan, build, deploy — so the cutover is a non-event rather than a ceremony:

# .github/workflows/deploy-prod.yml
name: deploy-prod
on:
  push:
    branches: [main]
jobs:
  ci:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: pytest -q                          # unit + contract tests
      - run: python scripts/pii_log_scan.py      # fail on PII in test logs
      - run: pip-audit -r requirements.txt      # dependency vulnerabilities
      - run: docker build -t registry.internal/cityops/api:${{ github.sha }} ./api
      - run: docker push registry.internal/cityops/api:${{ github.sha }}
  cd:
    needs: ci
    runs-on: [self-hosted, prod]
    steps:
      - run: |
          IMAGE_TAG=${{ github.sha }} docker compose \
            -f docker-compose.prod.yml up -d --no-deps api worker
          # Compose healthchecks gate traffic: unhealthy containers
          # never join the network. Previous images stay until the
          # new ones pass /ready — that IS the rollback plan.

Note the rollback story, because Maria will ask: there isn't a separate rollback procedure. The previous containers keep running until the new ones pass their health checks; if the new ones never pass, traffic never moves. Rollback is the absence of rollforward — the safest kind, because it has no new code paths to break at 2 AM.

Deployment identity: knowing exactly what is running

The Stage 4 standard — record an immutable deployment identity for every deployment — applies to containers too. "We deployed the fix" is not an identity. The team writes a deploy record on every CD run, and Tom checks it during triage instead of asking Dev "did the fix go out?":

# /var/lib/cityops/deploys/2026-10-04-0740.json — written by the CD job
{
  "deployed_at": "2026-10-04T07:40:12-07:00",
  "api_image": "registry.internal/cityops/api@sha256:9f2c…a41d",
  "worker_image": "registry.internal/cityops/worker@sha256:9f2c…a41d",
  "git_sha": "c41d88f",
  "config_version": "compose.prod v17",
  "model_identity": "cityops-classifier v3.2.1 / onnx int8",
  "deployed_by": "cd-pipeline (merge c41d88f by dev)"
}

Digest-pinned images, the exact git SHA, the config version, the model artifact identity — everything needed to reproduce or roll back the running system, in one file. At 7:42, Tom's triage question "is the fix actually live?" is answered by one cat, not a Slack thread. This record also feeds the audit trail: every sensitive action's run_id can be joined to the deployment that served it.

Monitoring, secrets, RBAC, runbook — the rest of the loop

The remaining pieces close the loop from the earlier posts:

  • Monitoring: Prometheus scrapes /metrics (request latency, error rate, ingestion lag — the 311 feed's records-per-minute, which is what would have caught the 2:14 AM flatline). Alert rule: ingestion lag > 15 minutes pages Tom. Dashboards show Maria the counts she demos.
  • Secrets: database passwords and signing keys live in the secrets manager, mounted as files, rotated every 90 days without redeploys. Nothing secret appears in the compose file, the repo, or the logs — the CI PII scan's stricter sibling.
  • RBAC: the access matrix from the security-review post is enforced in the API — operator, viewer, admin, and the insert-only service:ingest role. The 311 worker physically cannot read PII back, which bounds the incident's blast radius below.
  • Runbook: docs/runbook.md in the repo — symptoms, triage queries, who to page, how to pause ingestion, how to backfill. The team follows it at 6:30 AM instead of improvising.

The night-before ritual: Tom's go-live checklist

Cutovers fail on the boring stuff — an expired certificate, a full disk, a secret that was rotated in staging but not prod. Tom runs a written checklist the night before, and every item is a command with an expected output, not a vibe:

CheckCommand / whereExpected
All services healthydocker compose -f docker-compose.prod.yml ps3/3 healthy
Readiness, not just livenesscurl https://cityops.internal/ready{"status":"ready"} + model identity
Secrets mounted, none in envdocker compose config | grep -i passwordNo output
Backup ran todayBackup job logToday's timestamp, restore test < 30 days old
Dashboards currentMaria's dashboardLatest report < 10 min old
On-call roster correctPager scheduleTom primary, Dev secondary
Runbook printed (yes, printed)docs/runbook.mdOn Tom's desk — incidents don't wait for Wi-Fi

The checklist passed completely at 10 PM. The incident still happened at 2:14 AM — which is the point worth internalizing: go-live checklists verify the system you built, not the world around it. The upstream feed wasn't on the checklist because nobody owned it. After the incident, "feed contract test green" becomes line one of Tom's nightly check. Checklists grow; that's their job.

graph TB subgraph Internet U[Mayor's office demo
+ operators] end subgraph Production host R[Reverse proxy
TLS termination] A[cityops/api
FastAPI] W[cityops/worker
311 ingest] D[(Postgres 16
pgdata volume)] S[Secrets
mounted files] end F[City 311 feed
external] M[Prometheus + dashboards] L[Immutable audit log] U --> R --> A --> D W -->|poll every 5 min| F W --> D S -.-> A S -.-> W S -.-> D A -.->|/metrics| M W -.->|/metrics| M A -->|audit JSONL| L style D fill:#3b1d1d,stroke:#f87171 style S fill:#1d2a45,stroke:#22D3EE

By 10 PM the night before launch, this is all green. The cutover is done, the dashboard shows yesterday's reports, Tom goes to bed. Then the 311 feed changes its schema.

Break it: the 2:14 AM schema change

What actually happened is mundane, which is the point. The city's 311 feed renamed a field — category became service_code — in a routine upstream update, with no notice. The intake worker's parser did a direct key lookup, raised KeyError, crashed, and the poll loop's supervisor restarted it every five minutes into the same crash. Zero records ingested from 2:14 AM. The dashboard at 6 AM shows a flat line.

This is the "trust no input" lesson arriving at 2 AM with interest: the team validated citizen input rigorously and trusted the upstream feed implicitly. External data is input too.

Triage: the runbook, not improvisation

At 6:30 the team opens the runbook instead of guessing. The triage sequence is deliberately ordered — scope first, cause second, fix third:

  1. Scope: Is the API down, or just ingestion? /ready is green, the dashboard loads, yesterday's data is intact. Blast radius: new 311 reports only. The mayor's demo shows stale-but-real data, not an error page — important for what Maria says at 9.
  2. Cause: Worker logs show the crash loop. The last successful poll was 2:09 AM; the first crash 2:14 AM. Diffing a 2:09 payload against a 2:14 payload shows the rename: category → service_code. Root cause identified in eleven minutes — because the worker logs the raw payload shape on failure, a decision made in the Stage 4 tracing post.
  3. Data safety: Did we lose reports? The 311 feed is pull-based with cursor pagination — the missed window is re-fetchable. Nothing is lost; it's late. This distinction determines the whole response: backfill, not apologies for lost data.

The actual triage commands, because runbooks are specific or they're useless:

# 1. Scope: is the API serving? (readiness, not just liveness)
$ curl -s https://cityops.internal/ready | head -c 120
{"status":"ready","model":"cityops-classifier v3.2.1 / onnx int8"}

# 2. Cause: what did the worker die on?
$ docker compose -f docker-compose.prod.yml logs worker --since 6h | grep -m2 -A3 Traceback
Traceback (most recent call last):
  File "ingest/poll_311.py", line 41, in poll
    batch = [parse_report(r) for r in fetch()]
KeyError: 'category'

# 3. When did it start, and what changed? (payload shape on failure — the
#    logging decision from Stage 4 that paid for itself at 6:41)
$ grep "payload_keys" /var/log/cityops/worker.log | tail -2
02:09 payload_keys=[id,category,text,ts] ok=47
02:14 payload_keys=[id,service_code,text,ts] KeyError='category'

# 4. Data safety: what's the re-fetchable window?
$ curl -s "https://311.city.example/api?since=2026-10-04T02:14" | jq '.count'
312

Four commands, eleven minutes, root cause confirmed. The third one is the lesson: the worker logs the incoming payload's key set on every failure — a one-line logging decision from the Stage 4 tracing post that turned "something changed upstream" into "the field was renamed from category to service_code at 02:14." Triage speed is a function of what you chose to log months ago.

The old parser's behavior, reproduced exactly:

OLD parser at 02:14 AM: KeyError: 'category' -> pipeline crashes, 0 records ingested

The fix: tolerate, quarantine, backfill

Dev writes the fix at 7:05 — not a rename of one field, but the general principle the Pydantic lesson taught: validate at the boundary, quarantine what you can't map, never crash the pipeline on a shape you didn't expect.

def parse_report_tolerant(raw: dict) -> dict:
    # Accept known aliases; quarantine what you can't map.
    category = raw.get("category", raw.get("service_code"))
    if category is None:
        return {"quarantined": True,
                "reason": "no mappable category field",
                "fields_seen": sorted(raw.keys())}
    return {"id": raw["id"],
            "category": category,
            "text": raw.get("text", "")}
NEW parser: {"id": 9917, "category": "POTHOLE", "text": "Elm St crater"}
NEW parser, unknown shape: {"quarantined": true, "reason": "no mappable category field",
                             "fields_seen": ["id", "zzz"]}

The quarantine path matters more than the alias: the next schema change won't be a rename the team anticipates. Unknown shapes land in the quarantine table — visible on Tom's dashboard, countable, alertable — instead of crashing the worker. And the customer owns the quarantine policy (the Pydantic review's durable rule): Maria decides that quarantined reports get human review within one business day, and the team implements the queue.

By 7:40 the fix is through CI (tests, PII scan, build) and deployed by the pipeline — the same pipeline from the cutover, which is the point of having one. By 7:55 the backfill completes: the worker re-pulls the 2:14–7:40 window, 312 reports land, the dashboard is current. Total data delay: under six hours, zero records lost.

gantt title Launch-day incident timeline dateFormat HH:mm axisFormat %H:%M section Feed Schema change ships :02:14, 02:15 Missed ingestion window :02:14, 07:40 section Team Tom paged (ingestion lag) :06:12, 06:13 War room + runbook triage :06:30, 06:55 Fix written + CI :07:05, 07:40 Deploy + backfill (312) :07:40, 07:55 Dashboard verified current :07:55, 08:00 section Customer Mayor's office demo :09:00, 09:30 Postmortem delivered :12:00, 12:01

Must know

  • Triage order: scope → cause → data safety. Knowing "nothing is lost, it's late" changes the entire response.
  • Log the raw payload shape on ingestion failure. The eleven-minute root cause came from one logging decision.
  • The fix is general, not specific: tolerate known aliases, quarantine unknown shapes, never crash on unexpected input. External feeds are untrusted input.
Myth: "A crash loop is the safe failure mode — at least it pages someone." The crash loop paged Tom at 6:12, four hours after the break, via a lag alert — not via the crash. A quarantining parser would have kept 311 of 312 reports flowing and surfaced one quarantined record at 2:15. Crashing is the loudest failure mode, not the safest.

What Maria told the mayor's office at 9

The fix was technical; the 9 AM demo was communicational. At 8:15, Dev gives Maria a three-sentence brief she can use verbatim — written for a non-technical audience, honest about the morning, forward-looking:

Dev to Maria: "Say this: 'Overnight, our data supplier changed the format of one field without telling us. Our system flagged it instead of guessing — that's why this morning's reports arrived late rather than wrong. Everything is current now, and we've added an automatic check so format changes get caught in testing instead of in production.'"

Maria: "Late rather than wrong. I like that. Is it true?"

Dev: "Every word. The quarantine queue proves it — 312 reports delayed, zero corrupted."

Three things make this brief work. It names the external cause without blaming ("changed the format without telling us" — a fact, not an accusation). It reframes the delay as a safety property ("flagged instead of guessing") — which is true, because the alternative to the crash was silently misparsed data. And it ends with the mechanism, not a promise: the contract test, which Maria can mention by name if asked. The demo goes ahead on current data, and nobody in the room needs the postmortem yet — but when it arrives at noon, it confirms everything she said.

Productionize: what changes permanently

The fix restored service. Productionizing means the same failure gets cheaper next time — ideally free. The team makes four permanent changes, each mapped to the process gap the incident exposed:

Gap the incident exposedPermanent changeOwner
Upstream schema changed without noticeContract test on the feed: a CI job pulls the feed's sample payload hourly and asserts the fields the parser needs still exist. A rename breaks the contract test at 3 AM, not the pipeline at 2 AM.Dev
Unknown shapes crashed the workerQuarantine queue + dashboard: unmappable records land in quarantine, counted and alertable. Maria's policy: human review within one business day.Dev + Maria
Four hours from break to pageAlert tuning: ingestion-lag threshold 15 min → 10 min, plus a new alert on worker restart rate. The crash loop itself becomes a signal.Tom
No feed version anywhereUpstream version recorded: every ingested batch records the feed's reported schema version (or payload hash when unversioned) in the audit trail — the Stage 4 "version the system" rule extended to inputs.Dev

Notice the shape of these fixes: none of them is "be more careful." Each is a mechanism — a test, a queue, an alert, a recorded version — that makes the failure mode structurally cheaper. That's what productionize means in this curriculum: convert the lesson of the incident into a system property, the same way Stage 4 converted "don't trust the model" into guardrails.

The weekly ops review: keeping it live

Launch isn't a moment; it's a cadence. Tom institutes a 30-minute weekly ops review — the same four people, the same four questions, every Friday:

  1. What paged us? Every alert from the week, with the action taken. An alert that fired and needed nothing is a tuning candidate — noisy alerts get ignored, and ignored alerts miss the real 2 AM.
  2. What's in quarantine? The count, the oldest item, whether Maria's one-business-day review SLA held. A growing quarantine queue is the early warning for the next schema drift.
  3. What changed upstream? Feed contract test results, dependency updates, any vendor notice. The 2:14 AM incident becomes a standing agenda item: "did any supplier change shape this week?"
  4. Are the artifacts fresh? Lisa's rule from the security-review post — no artifact without a date — applied weekly: runbook, access matrix, evidence pack. Stale documentation is a future incident wearing a trench coat.

The review's output is a one-paragraph note Maria can read in sixty seconds. Four weeks of those notes become the operational history that the next security questionnaire asks for. The evidence pack from post 7 doesn't stay fresh by itself — this meeting is what keeps it fresh.

Later, not now

  • Multi-region or hot-standby deployment — the current single-host Compose stack with volume backups meets the customer's stated RTO/RPO; geographic redundancy is a later scaling decision, not a launch requirement.
  • A formal SLA with error budgets — worth defining once there's a quarter of production data to calibrate it against.

Communicate: the postmortem

By noon, Maria has the postmortem on her desk — and forwards it to the mayor's office unedited. That forwardability is the quality bar. Read it as the deliverable it is: blameless (no names attached to mistakes, only to actions), client-facing (no jargon without translation), and honest about what "fixed" means.

# Postmortem: CityOps 311-ingestion delay — 2026-10-04
Status: resolved. Document owner: Dev. Review date: 2026-10-11.

## Summary
On 2026-10-04 at 02:14, the city's 311 feed changed the name of one data
field without notice. Our intake worker could not read the new field name
and stopped ingesting new reports. The gap was detected at 06:12, fixed and
deployed by 07:40, and all 312 delayed reports were backfilled by 07:55.
No data was lost. The public dashboard showed data up to 02:09 during the
gap; it is current as of this writing.

## Impact
- 312 citizen reports delayed by up to 5h41m. Zero reports lost.
- The 09:00 demonstration used current data; no demo impact.
- No citizen-facing outage: the dashboard, API, and operator tools worked
  throughout. Only newly arriving 311 reports were delayed.

## Timeline (all times local)
- 02:14 — 311 feed ships renamed field (category -> service_code); worker
  begins crash loop. No alert fires (lag threshold not yet reached).
- 06:12 — Ingestion-lag monitor pages on-call. War room opened 06:30.
- 06:41 — Root cause identified: field rename, confirmed by diffing payloads.
- 07:05 — Fix written: parser accepts both field names; unknown shapes go
  to a quarantine queue instead of crashing the worker.
- 07:40 — Fix deployed through the standard pipeline (tests + scans green).
- 07:55 — Backfill of the 02:14–07:40 window complete; dashboard verified.

## Root cause
Our ingestion trusted the upstream feed's shape the way we used to trust
user input: implicitly. The parser assumed a field name that was never
contractually guaranteed. When the assumption broke, the failure mode was
"stop everything" instead of "set aside what you can't read."

Note what this is NOT: nobody made an error. The feed owner changed a
field, as feed owners do. Our system was brittle to a normal event.

## What went well
- Triage followed the runbook; root cause in 11 minutes.
- Pull-based feed meant zero data loss — backfill, not recovery.
- Standard pipeline deployed the fix; no emergency procedures needed.
## Corrective actions 1. Hourly contract test against the live feed schema (Dev, by 10/11). 2. Quarantine queue for unmappable records + dashboard count (Dev, done). 3. Ingestion-lag alert 15min -> 10min; new alert on worker restart rate (Tom, done). 4. Record upstream feed version/hash per batch in the audit trail (Dev, by 10/11). 5. Quarterly review of this postmortem's action items with Maria (recurring). ## Appendix: for the technically curious The old parser did raw["category"] — a direct lookup that throws when the key is absent. The new parser tries known names, then quarantines. The full diff is 14 lines, reviewed and tested like any other change.

Study the language choices. "Our system was brittle to a normal event" — the cause is a property of the system, not a person. "Note what this is NOT" — the document explicitly forecloses blame, because blame is what makes the next incident go unreported. "No demo impact" sits in Impact, not buried — the client's actual worry, answered first. And the appendix translates the jargon without condescending: one sentence on what changed, one on how big the change was.

Maria's feedback, verbatim: "I forwarded it as-is. They replied asking when the next department can onboard." The postmortem didn't just restore trust — it demonstrated the operational maturity that procurement questionnaires (Stage 5, post 7) ask about. The postmortem is the product isn't a slogan; it's what happened.

Don't memorize

  • Postmortem templates from big tech — the five questions (what happened, impact, root cause, learnings, changes) matter; the exact format doesn't.
  • The incident's specific times — they're evidence for this incident, not a standard. Your timeline will differ.
Myth: "Blameless means nobody is accountable." The postmortem names owners for every corrective action — Dev, Tom, Maria, Lisa — with dates. Blameless means accountability attaches to fixing the system, not to the person nearest the failure. The team that confuses the two gets silence after the next incident instead of a timeline.

CityOps: live

Step back and name what the milestone delivered. CityOps is now a production system in the full sense:

  • Deployed by pipeline, not by hand — every change flows through CI (tests, PII log-scan, dependency audit), builds an immutable image, and rolls forward only past green health checks, with rollback as the default.
  • Observable — liveness vs. readiness correctly separated, metrics on dashboards Maria demos, ingestion-lag alerts that page Tom, and an audit trail recording every sensitive action with per-request correlation.
  • Hardened — secrets in the manager, RBAC on every endpoint, the database unreachable from outside its network, PII boundaries drawn and inventoried, and a security-review evidence pack that answers Lisa's questionnaire with artifacts.
  • Operated — a runbook the team actually used at 6:30 AM, a quarantine policy Maria owns, corrective actions with owners and dates, and a postmortem practice that turns incidents into sales assets.
  • Resilient to its suppliers — the 311 feed is now treated as untrusted input: contract-tested, version-recorded, tolerated and quarantined rather than trusted.

And the handover is explicit. Tom owns the infrastructure and the alerts. Dev owns the code and the pipeline. Maria owns the data policies — retention, quarantine review, the evidence pack. Lisa owns the questionnaire and the quarterly access reviews. Nobody owns "everything," which is why everything is owned.

Acceptance: what "done" looked like

Milestones need acceptance criteria, or they drift. Maria signed off on these six lines — the definition of done for Stage 5:

#Acceptance criterionVerified how
1Production serves the mayor's-office demo URLDemo ran 09:00 on current data
2Every deploy goes through CI/CD; no hand-deploysDeploy record exists for every running image
3Monitoring pages a human within 10 min of ingestion stallAlert fired 06:12 on the real incident
4Audit trail records every sensitive action, denials includedSampled allow + deny records, Lisa reviewed
5Security questionnaire answered with dated artifactsEvidence pack accepted by procurement
6Incident produced a blameless postmortem within 6 hoursDelivered 12:00, forwarded unedited

Criterion 6 is the one other teams skip, and it's the one that matters most. A milestone that ends with "we launched" is a deployment. A milestone that ends with "we launched, it broke, here's the postmortem" is a production operation. Maria signed line 6 with the comment: "This is why we'll survive the next one." Stage 5 is complete.

Stage 5's arc, in one line: deployment is not the moment you go live; it's the system that keeps you live. The Compose file got CityOps into production. The pipeline, the monitors, the audit trail, the runbook, and the postmortem are what keep it there — including, especially, on the mornings when the city's feed renames a field at 2 AM.

That's the whole curriculum's bet, stated one last time: the FDE's job was never to write the cleverest code. It's to build the system that survives the questionnaire, the schema change, and the 9 AM demo — and to leave behind the artifacts that prove it. CityOps is live. The evidence says so.

Field check

  1. The incident's root cause says "nobody made an error." A stakeholder asks: "Then why did it take four hours to detect?" Is that a blame question or a systems question? How does the postmortem answer it?
  2. Your feed contract test breaks at 3 AM on a Saturday — the upstream renamed another field. The pipeline is green and production is healthy. What do you do before Monday, and why is this outcome better than the 2:14 AM version?
  3. Maria wants the postmortem "softened" before forwarding — less mention of the crash loop. What do you tell her, and what principle is at stake?
  4. Map each Stage 5 post (1–7, by title) to one component of the final production system. Which component would hurt most to remove, and why?
  5. The mayor's office asks: "Prove this won't happen again." What's the honest answer — and which corrective action comes closest to a real guarantee?
Field check answers

1. It's a systems question wearing blame's clothes — and the postmortem answers it as one: detection took four hours because the lag threshold was 15 minutes and the crash loop wasn't itself alerted on. The fix is a mechanism (tighter threshold + restart-rate alert), not a person to watch harder. Blameless doesn't mean questionless; it means every question terminates in a system change.

2. You update the parser's alias list (or confirm quarantine caught it), run it through the normal pipeline, and go back to sleep — no war room, because nothing is broken in production. This is better because the failure was converted from "silent pipeline death discovered at dawn" into "a test failure on a Saturday." Same upstream event, two orders of magnitude cheaper.

3. You tell her no — gently, and with the reason: the postmortem's value is its candor. A softened version that hides the crash loop teaches the mayor's office that you manage impressions; the unedited version teaches them you manage systems. The principle: the postmortem is the product, and products don't ship with the defects edited out of the brochure.

4. Roughly: Compose → the stack; CI/CD → the pipeline; health checks → the readiness gates; monitoring → the lag alerts and dashboards; secrets → the mounted credentials; RBAC → the endpoint authorization; security review → the audit trail and evidence pack. Hardest to remove: the pipeline — without it, every deploy (including incident fixes) becomes a hand-operated ceremony, and the 7:40 fix would have been a 10:40 fix.

5. The honest answer: you can't prove a negative — upstream feeds will change again. What you can show is that the failure mode changed: the next rename hits the contract test or the quarantine queue, not the crash loop. The quarantine queue comes closest to a guarantee, because it bounds the worst case ("some records wait for human review") instead of promising the best case ("nothing ever breaks").