Tom (ops, 6:12 AM, launch day): "The intake pipeline broke overnight. Zero records since 2:14 AM. The city's 311 feed changed something — I don't know what yet."
Maria (customer): "The mayor's office demo is at 9:00. They open the dashboard, they expect to see this morning's reports. What do I tell them?"
Dev: "Give us until 8. We'll have it fixed, verified, and written up. And Maria — the write-up is part of the deliverable, not an apology note."
Milestone 5 is the production cutover: CityOps goes live on real infrastructure, deployed by pipeline, watched by monitors, guarded by secrets and roles — and then, on launch morning, it breaks. This is the capstone because it integrates everything: the Compose stack, the CI/CD pipeline, the health checks, the audit trail, the evidence pack. And it proves the hardest lesson of the stage: the incident is part of the product.
The customer problem
Two problems arrive at once, which is exactly how production works. The planned problem: cut CityOps over from the pilot environment to production — real Compose stack, real CI/CD, real secrets, real monitoring, real runbook. The unplanned problem: at 2:14 AM on launch day, the city's 311 feed ships a schema change, the intake pipeline crashes on it, and by 6 AM the dashboard is showing a flat line where this morning's reports should be.
Maria's 9 AM demo is the forcing function. She doesn't need a root-cause analysis by 9 — she needs a working dashboard and an honest story. Dev's team has until 8 AM to do three things: triage (what broke, what's the blast radius), fix (restore ingestion without losing data), and communicate (a blameless, client-facing postmortem on Maria's desk by noon).
This is why the milestone is the right end to Stage 5. "Deploy Like a Pro" was never about YAML — it was about what happens when the YAML meets reality. The launch-day break is the exam.
Clarify the ask
In the war-room call at 6:30, the team separates the two asks before touching anything — because conflating "fix the incident" with "finish the cutover" is how you make both worse.
Ask 1 — the cutover (planned): CityOps runs in production behind the Compose stack built across this stage: containers for the API, Postgres, and the intake worker; CI/CD deploying on merge; health checks gating traffic; secrets from the manager, not env files; RBAC on every endpoint; dashboards on metrics; the audit trail from the security-review post recording every sensitive action. Done means Maria can point the mayor's office at a URL and Tom can go back to sleep.
Ask 2 — the incident (unplanned): restore the 311 feed ingestion, backfill the missed records, and deliver the postmortem by noon. Done means the dashboard is current and Maria has a document she can forward to the mayor's office without editing it.
Dev states the principle that governs the whole milestone: a milestone is integration, not invention. Nothing in this post is new technology. The Compose file, the pipeline, the health checks, the audit middleware, the evidence pack — all built in posts 1–7. The milestone's job is to compose them into a system that survives contact with a real Tuesday morning. If you find yourself inventing a new mechanism during a milestone, stop: you're papering over a gap in the earlier posts.
Must know
- A milestone is integration, not invention. Compose what the stage built; don't invent new machinery under pressure.
- Separate the planned work (cutover) from the unplanned work (incident) before you touch anything. Different goals, different clocks.
- The postmortem is a deliverable, not an apology. Maria should be able to forward it unedited.
The minimum concept
The cutover needs one concept, and the incident needs one concept. Both are small.
Cutover concept: the deployment is a closed loop. Code moves through CI (build, test, scan) into images; images move through CD into the Compose stack; the stack reports health; health gates traffic; traffic generates metrics and audit records; metrics feed monitors; monitors page Tom. Every arrow in that loop was built in an earlier post. The milestone just closes it and proves it runs end to end.
Incident concept: the postmortem is the product. The fix restores service; the postmortem restores trust. A blameless postmortem answers five questions — what happened, what was the impact, what was the root cause (a process gap, never a person), what did we learn, what changes now — in language a non-engineer can forward. Maria's demo audience doesn't need to know what a KeyError is. They need to know it won't silently eat their data again.
The integration map: where posts 1–7 live in the final system
"Integration, not invention" is checkable. Every component of the production system below was built in an earlier Stage 5 post — if you can't point at the post, you're inventing:
| Stage 5 post (by title) | What it built | Where it runs at launch |
|---|---|---|
| Containers & Compose | Images, Compose services, networks | The three-service prod stack |
| CI/CD pipelines | Test → scan → build → deploy loop | Deploys the 7:40 incident fix |
| Health checks & monitoring | Liveness/readiness, metrics, alerts | Gates traffic; pages Tom at 6:12 |
| Secrets management | Manager-backed, file-mounted secrets | DB passwords, signing keys |
| RBAC & access control | Roles, least-privilege service accounts | Every API endpoint; the insert-only worker |
| Security review evidence | Audit trail, PII inventory, access matrix | Every sensitive action logged; Lisa's pack |
| This milestone | Composition + the runbook + the postmortem — the only new artifacts | |
The milestone's genuinely new content is small on purpose: the runbook (an operations document) and the postmortem (a communication document). Everything else is wiring. When a milestone's "new" column is longer than its "integrated" column, the stage plan had a hole.
+ image build] B --> C[CD: deploy Compose stack] C --> D{Health checks
liveness + readiness} D -->|pass| E[Traffic to new containers] D -->|fail| F[Automatic rollback
previous images stay] E --> G[Metrics → dashboards
Audit → immutable log] G --> H[Monitors → page Tom] H -.->|incident| I[Postmortem
by noon]
Build it: the production cutover
The cutover is assembly, not invention. Each piece below comes from its Stage 5 post (referenced by title); what's new is that they now run together, against production data, behind one URL.
The Compose stack
Three services, one network, no exposed database port. The API serves traffic, the worker ingests the 311 feed, Postgres holds the reports. Secrets arrive as files from the secrets manager (Stage 5's secrets post), never as environment variables in the compose file:
# docker-compose.prod.yml — the production cutover
services:
api:
image: registry.internal/cityops/api:${IMAGE_TAG:?IMAGE_TAG required}
depends_on:
db:
condition: service_healthy
secrets:
- db_password
- api_signing_key
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/ready"]
interval: 15s
timeout: 5s
retries: 3
start_period: 30s
deploy:
resources:
limits: {cpus: "2", memory: 1G}
networks: [cityops]
worker:
image: registry.internal/cityops/worker:${IMAGE_TAG:?IMAGE_TAG required}
depends_on:
db:
condition: service_healthy
secrets:
- db_password
command: ["python", "-m", "ingest.poll_311", "--interval", "300"]
networks: [cityops]
db:
image: postgres:16
volumes:
- pgdata:/var/lib/postgresql/data
secrets:
- db_password
environment:
POSTGRES_PASSWORD_FILE: /run/secrets/db_password
healthcheck:
test: ["CMD-SHELL", "pg_isready -U cityops"]
interval: 10s
retries: 5
networks: [cityops]
# no ports: published — the database is not reachable from outside
# this Compose network. That absence is a security control.
volumes:
pgdata:
secrets:
db_password:
external: true
api_signing_key:
external: true
networks:
cityops:
Read the security properties out of the YAML: the database has no published port (least privilege at the network layer), secrets are external (rotation doesn't require redeploys), every service has a health check, and the image tag is mandatory — no latest, because "pin the exact tag" may not be enough for models (Stage 4), but for containers an immutable digest-pinned tag is exactly the deployment identity you want.
Health checks: liveness vs. readiness, correctly
The Stage 4 model-operations standard applies verbatim here: liveness proves the process responds; readiness proves it can serve. A health GET that returns 200 only proves the server is up. The /ready endpoint the Compose file hits must prove the dependencies are actually usable:
@app.get("/live")
def liveness():
return {"status": "ok"} # the process is alive. That's all this claims.
@app.get("/ready")
def readiness(db=Depends(get_db)):
# Readiness proves the API can actually serve requests:
try:
db.execute("SELECT 1") # database reachable?
model_registry.ping() # required model artifact loadable?
except Exception as e:
raise HTTPException(503, f"not ready: {type(e).__name__}")
return {"status": "ready",
"model": model_registry.identity()} # immutable artifact id
This distinction is load-bearing in the incident below: during the outage, /live stayed green while /ready would have caught a sick dependency. Monitors watch readiness; pagers fire on readiness. Liveness is for the orchestrator's restart decisions only.
The CI/CD pipeline
Every merge to main runs the same loop — test, scan, build, deploy — so the cutover is a non-event rather than a ceremony:
# .github/workflows/deploy-prod.yml
name: deploy-prod
on:
push:
branches: [main]
jobs:
ci:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: pytest -q # unit + contract tests
- run: python scripts/pii_log_scan.py # fail on PII in test logs
- run: pip-audit -r requirements.txt # dependency vulnerabilities
- run: docker build -t registry.internal/cityops/api:${{ github.sha }} ./api
- run: docker push registry.internal/cityops/api:${{ github.sha }}
cd:
needs: ci
runs-on: [self-hosted, prod]
steps:
- run: |
IMAGE_TAG=${{ github.sha }} docker compose \
-f docker-compose.prod.yml up -d --no-deps api worker
# Compose healthchecks gate traffic: unhealthy containers
# never join the network. Previous images stay until the
# new ones pass /ready — that IS the rollback plan.
Note the rollback story, because Maria will ask: there isn't a separate rollback procedure. The previous containers keep running until the new ones pass their health checks; if the new ones never pass, traffic never moves. Rollback is the absence of rollforward — the safest kind, because it has no new code paths to break at 2 AM.
Deployment identity: knowing exactly what is running
The Stage 4 standard — record an immutable deployment identity for every deployment — applies to containers too. "We deployed the fix" is not an identity. The team writes a deploy record on every CD run, and Tom checks it during triage instead of asking Dev "did the fix go out?":
# /var/lib/cityops/deploys/2026-10-04-0740.json — written by the CD job
{
"deployed_at": "2026-10-04T07:40:12-07:00",
"api_image": "registry.internal/cityops/api@sha256:9f2c…a41d",
"worker_image": "registry.internal/cityops/worker@sha256:9f2c…a41d",
"git_sha": "c41d88f",
"config_version": "compose.prod v17",
"model_identity": "cityops-classifier v3.2.1 / onnx int8",
"deployed_by": "cd-pipeline (merge c41d88f by dev)"
}
Digest-pinned images, the exact git SHA, the config version, the model artifact identity — everything needed to reproduce or roll back the running system, in one file. At 7:42, Tom's triage question "is the fix actually live?" is answered by one cat, not a Slack thread. This record also feeds the audit trail: every sensitive action's run_id can be joined to the deployment that served it.
Monitoring, secrets, RBAC, runbook — the rest of the loop
The remaining pieces close the loop from the earlier posts:
- Monitoring: Prometheus scrapes
/metrics(request latency, error rate, ingestion lag — the 311 feed's records-per-minute, which is what would have caught the 2:14 AM flatline). Alert rule: ingestion lag > 15 minutes pages Tom. Dashboards show Maria the counts she demos. - Secrets: database passwords and signing keys live in the secrets manager, mounted as files, rotated every 90 days without redeploys. Nothing secret appears in the compose file, the repo, or the logs — the CI PII scan's stricter sibling.
- RBAC: the access matrix from the security-review post is enforced in the API — operator, viewer, admin, and the insert-only
service:ingestrole. The 311 worker physically cannot read PII back, which bounds the incident's blast radius below. - Runbook:
docs/runbook.mdin the repo — symptoms, triage queries, who to page, how to pause ingestion, how to backfill. The team follows it at 6:30 AM instead of improvising.
The night-before ritual: Tom's go-live checklist
Cutovers fail on the boring stuff — an expired certificate, a full disk, a secret that was rotated in staging but not prod. Tom runs a written checklist the night before, and every item is a command with an expected output, not a vibe:
| Check | Command / where | Expected |
|---|---|---|
| All services healthy | docker compose -f docker-compose.prod.yml ps | 3/3 healthy |
| Readiness, not just liveness | curl https://cityops.internal/ready | {"status":"ready"} + model identity |
| Secrets mounted, none in env | docker compose config | grep -i password | No output |
| Backup ran today | Backup job log | Today's timestamp, restore test < 30 days old |
| Dashboards current | Maria's dashboard | Latest report < 10 min old |
| On-call roster correct | Pager schedule | Tom primary, Dev secondary |
| Runbook printed (yes, printed) | docs/runbook.md | On Tom's desk — incidents don't wait for Wi-Fi |
The checklist passed completely at 10 PM. The incident still happened at 2:14 AM — which is the point worth internalizing: go-live checklists verify the system you built, not the world around it. The upstream feed wasn't on the checklist because nobody owned it. After the incident, "feed contract test green" becomes line one of Tom's nightly check. Checklists grow; that's their job.
+ operators] end subgraph Production host R[Reverse proxy
TLS termination] A[cityops/api
FastAPI] W[cityops/worker
311 ingest] D[(Postgres 16
pgdata volume)] S[Secrets
mounted files] end F[City 311 feed
external] M[Prometheus + dashboards] L[Immutable audit log] U --> R --> A --> D W -->|poll every 5 min| F W --> D S -.-> A S -.-> W S -.-> D A -.->|/metrics| M W -.->|/metrics| M A -->|audit JSONL| L style D fill:#3b1d1d,stroke:#f87171 style S fill:#1d2a45,stroke:#22D3EE
By 10 PM the night before launch, this is all green. The cutover is done, the dashboard shows yesterday's reports, Tom goes to bed. Then the 311 feed changes its schema.
Break it: the 2:14 AM schema change
What actually happened is mundane, which is the point. The city's 311 feed renamed a field — category became service_code — in a routine upstream update, with no notice. The intake worker's parser did a direct key lookup, raised KeyError, crashed, and the poll loop's supervisor restarted it every five minutes into the same crash. Zero records ingested from 2:14 AM. The dashboard at 6 AM shows a flat line.
This is the "trust no input" lesson arriving at 2 AM with interest: the team validated citizen input rigorously and trusted the upstream feed implicitly. External data is input too.
Triage: the runbook, not improvisation
At 6:30 the team opens the runbook instead of guessing. The triage sequence is deliberately ordered — scope first, cause second, fix third:
- Scope: Is the API down, or just ingestion?
/readyis green, the dashboard loads, yesterday's data is intact. Blast radius: new 311 reports only. The mayor's demo shows stale-but-real data, not an error page — important for what Maria says at 9. - Cause: Worker logs show the crash loop. The last successful poll was 2:09 AM; the first crash 2:14 AM. Diffing a 2:09 payload against a 2:14 payload shows the rename:
category→service_code. Root cause identified in eleven minutes — because the worker logs the raw payload shape on failure, a decision made in the Stage 4 tracing post. - Data safety: Did we lose reports? The 311 feed is pull-based with cursor pagination — the missed window is re-fetchable. Nothing is lost; it's late. This distinction determines the whole response: backfill, not apologies for lost data.
The actual triage commands, because runbooks are specific or they're useless:
# 1. Scope: is the API serving? (readiness, not just liveness)
$ curl -s https://cityops.internal/ready | head -c 120
{"status":"ready","model":"cityops-classifier v3.2.1 / onnx int8"}
# 2. Cause: what did the worker die on?
$ docker compose -f docker-compose.prod.yml logs worker --since 6h | grep -m2 -A3 Traceback
Traceback (most recent call last):
File "ingest/poll_311.py", line 41, in poll
batch = [parse_report(r) for r in fetch()]
KeyError: 'category'
# 3. When did it start, and what changed? (payload shape on failure — the
# logging decision from Stage 4 that paid for itself at 6:41)
$ grep "payload_keys" /var/log/cityops/worker.log | tail -2
02:09 payload_keys=[id,category,text,ts] ok=47
02:14 payload_keys=[id,service_code,text,ts] KeyError='category'
# 4. Data safety: what's the re-fetchable window?
$ curl -s "https://311.city.example/api?since=2026-10-04T02:14" | jq '.count'
312
Four commands, eleven minutes, root cause confirmed. The third one is the lesson: the worker logs the incoming payload's key set on every failure — a one-line logging decision from the Stage 4 tracing post that turned "something changed upstream" into "the field was renamed from category to service_code at 02:14." Triage speed is a function of what you chose to log months ago.
The old parser's behavior, reproduced exactly:
OLD parser at 02:14 AM: KeyError: 'category' -> pipeline crashes, 0 records ingested
The fix: tolerate, quarantine, backfill
Dev writes the fix at 7:05 — not a rename of one field, but the general principle the Pydantic lesson taught: validate at the boundary, quarantine what you can't map, never crash the pipeline on a shape you didn't expect.
def parse_report_tolerant(raw: dict) -> dict:
# Accept known aliases; quarantine what you can't map.
category = raw.get("category", raw.get("service_code"))
if category is None:
return {"quarantined": True,
"reason": "no mappable category field",
"fields_seen": sorted(raw.keys())}
return {"id": raw["id"],
"category": category,
"text": raw.get("text", "")}
NEW parser: {"id": 9917, "category": "POTHOLE", "text": "Elm St crater"}
NEW parser, unknown shape: {"quarantined": true, "reason": "no mappable category field",
"fields_seen": ["id", "zzz"]}
The quarantine path matters more than the alias: the next schema change won't be a rename the team anticipates. Unknown shapes land in the quarantine table — visible on Tom's dashboard, countable, alertable — instead of crashing the worker. And the customer owns the quarantine policy (the Pydantic review's durable rule): Maria decides that quarantined reports get human review within one business day, and the team implements the queue.
By 7:40 the fix is through CI (tests, PII scan, build) and deployed by the pipeline — the same pipeline from the cutover, which is the point of having one. By 7:55 the backfill completes: the worker re-pulls the 2:14–7:40 window, 312 reports land, the dashboard is current. Total data delay: under six hours, zero records lost.
Must know
- Triage order: scope → cause → data safety. Knowing "nothing is lost, it's late" changes the entire response.
- Log the raw payload shape on ingestion failure. The eleven-minute root cause came from one logging decision.
- The fix is general, not specific: tolerate known aliases, quarantine unknown shapes, never crash on unexpected input. External feeds are untrusted input.
What Maria told the mayor's office at 9
The fix was technical; the 9 AM demo was communicational. At 8:15, Dev gives Maria a three-sentence brief she can use verbatim — written for a non-technical audience, honest about the morning, forward-looking:
Dev to Maria: "Say this: 'Overnight, our data supplier changed the format of one field without telling us. Our system flagged it instead of guessing — that's why this morning's reports arrived late rather than wrong. Everything is current now, and we've added an automatic check so format changes get caught in testing instead of in production.'"
Maria: "Late rather than wrong. I like that. Is it true?"
Dev: "Every word. The quarantine queue proves it — 312 reports delayed, zero corrupted."
Three things make this brief work. It names the external cause without blaming ("changed the format without telling us" — a fact, not an accusation). It reframes the delay as a safety property ("flagged instead of guessing") — which is true, because the alternative to the crash was silently misparsed data. And it ends with the mechanism, not a promise: the contract test, which Maria can mention by name if asked. The demo goes ahead on current data, and nobody in the room needs the postmortem yet — but when it arrives at noon, it confirms everything she said.
Productionize: what changes permanently
The fix restored service. Productionizing means the same failure gets cheaper next time — ideally free. The team makes four permanent changes, each mapped to the process gap the incident exposed:
| Gap the incident exposed | Permanent change | Owner |
|---|---|---|
| Upstream schema changed without notice | Contract test on the feed: a CI job pulls the feed's sample payload hourly and asserts the fields the parser needs still exist. A rename breaks the contract test at 3 AM, not the pipeline at 2 AM. | Dev |
| Unknown shapes crashed the worker | Quarantine queue + dashboard: unmappable records land in quarantine, counted and alertable. Maria's policy: human review within one business day. | Dev + Maria |
| Four hours from break to page | Alert tuning: ingestion-lag threshold 15 min → 10 min, plus a new alert on worker restart rate. The crash loop itself becomes a signal. | Tom |
| No feed version anywhere | Upstream version recorded: every ingested batch records the feed's reported schema version (or payload hash when unversioned) in the audit trail — the Stage 4 "version the system" rule extended to inputs. | Dev |
Notice the shape of these fixes: none of them is "be more careful." Each is a mechanism — a test, a queue, an alert, a recorded version — that makes the failure mode structurally cheaper. That's what productionize means in this curriculum: convert the lesson of the incident into a system property, the same way Stage 4 converted "don't trust the model" into guardrails.
The weekly ops review: keeping it live
Launch isn't a moment; it's a cadence. Tom institutes a 30-minute weekly ops review — the same four people, the same four questions, every Friday:
- What paged us? Every alert from the week, with the action taken. An alert that fired and needed nothing is a tuning candidate — noisy alerts get ignored, and ignored alerts miss the real 2 AM.
- What's in quarantine? The count, the oldest item, whether Maria's one-business-day review SLA held. A growing quarantine queue is the early warning for the next schema drift.
- What changed upstream? Feed contract test results, dependency updates, any vendor notice. The 2:14 AM incident becomes a standing agenda item: "did any supplier change shape this week?"
- Are the artifacts fresh? Lisa's rule from the security-review post — no artifact without a date — applied weekly: runbook, access matrix, evidence pack. Stale documentation is a future incident wearing a trench coat.
The review's output is a one-paragraph note Maria can read in sixty seconds. Four weeks of those notes become the operational history that the next security questionnaire asks for. The evidence pack from post 7 doesn't stay fresh by itself — this meeting is what keeps it fresh.
Later, not now
- Multi-region or hot-standby deployment — the current single-host Compose stack with volume backups meets the customer's stated RTO/RPO; geographic redundancy is a later scaling decision, not a launch requirement.
- A formal SLA with error budgets — worth defining once there's a quarter of production data to calibrate it against.
Communicate: the postmortem
By noon, Maria has the postmortem on her desk — and forwards it to the mayor's office unedited. That forwardability is the quality bar. Read it as the deliverable it is: blameless (no names attached to mistakes, only to actions), client-facing (no jargon without translation), and honest about what "fixed" means.
# Postmortem: CityOps 311-ingestion delay — 2026-10-04
Status: resolved. Document owner: Dev. Review date: 2026-10-11.
## Summary
On 2026-10-04 at 02:14, the city's 311 feed changed the name of one data
field without notice. Our intake worker could not read the new field name
and stopped ingesting new reports. The gap was detected at 06:12, fixed and
deployed by 07:40, and all 312 delayed reports were backfilled by 07:55.
No data was lost. The public dashboard showed data up to 02:09 during the
gap; it is current as of this writing.
## Impact
- 312 citizen reports delayed by up to 5h41m. Zero reports lost.
- The 09:00 demonstration used current data; no demo impact.
- No citizen-facing outage: the dashboard, API, and operator tools worked
throughout. Only newly arriving 311 reports were delayed.
## Timeline (all times local)
- 02:14 — 311 feed ships renamed field (category -> service_code); worker
begins crash loop. No alert fires (lag threshold not yet reached).
- 06:12 — Ingestion-lag monitor pages on-call. War room opened 06:30.
- 06:41 — Root cause identified: field rename, confirmed by diffing payloads.
- 07:05 — Fix written: parser accepts both field names; unknown shapes go
to a quarantine queue instead of crashing the worker.
- 07:40 — Fix deployed through the standard pipeline (tests + scans green).
- 07:55 — Backfill of the 02:14–07:40 window complete; dashboard verified.
## Root cause
Our ingestion trusted the upstream feed's shape the way we used to trust
user input: implicitly. The parser assumed a field name that was never
contractually guaranteed. When the assumption broke, the failure mode was
"stop everything" instead of "set aside what you can't read."
Note what this is NOT: nobody made an error. The feed owner changed a
field, as feed owners do. Our system was brittle to a normal event.
## What went well
- Triage followed the runbook; root cause in 11 minutes.
- Pull-based feed meant zero data loss — backfill, not recovery.
- Standard pipeline deployed the fix; no emergency procedures needed.
## Corrective actions
1. Hourly contract test against the live feed schema (Dev, by 10/11).
2. Quarantine queue for unmappable records + dashboard count (Dev, done).
3. Ingestion-lag alert 15min -> 10min; new alert on worker restart rate (Tom, done).
4. Record upstream feed version/hash per batch in the audit trail (Dev, by 10/11).
5. Quarterly review of this postmortem's action items with Maria (recurring).
## Appendix: for the technically curious
The old parser did raw["category"] — a direct lookup that throws when the
key is absent. The new parser tries known names, then quarantines. The full
diff is 14 lines, reviewed and tested like any other change.
Study the language choices. "Our system was brittle to a normal event" — the cause is a property of the system, not a person. "Note what this is NOT" — the document explicitly forecloses blame, because blame is what makes the next incident go unreported. "No demo impact" sits in Impact, not buried — the client's actual worry, answered first. And the appendix translates the jargon without condescending: one sentence on what changed, one on how big the change was.
Maria's feedback, verbatim: "I forwarded it as-is. They replied asking when the next department can onboard." The postmortem didn't just restore trust — it demonstrated the operational maturity that procurement questionnaires (Stage 5, post 7) ask about. The postmortem is the product isn't a slogan; it's what happened.
Don't memorize
- Postmortem templates from big tech — the five questions (what happened, impact, root cause, learnings, changes) matter; the exact format doesn't.
- The incident's specific times — they're evidence for this incident, not a standard. Your timeline will differ.
CityOps: live
Step back and name what the milestone delivered. CityOps is now a production system in the full sense:
- Deployed by pipeline, not by hand — every change flows through CI (tests, PII log-scan, dependency audit), builds an immutable image, and rolls forward only past green health checks, with rollback as the default.
- Observable — liveness vs. readiness correctly separated, metrics on dashboards Maria demos, ingestion-lag alerts that page Tom, and an audit trail recording every sensitive action with per-request correlation.
- Hardened — secrets in the manager, RBAC on every endpoint, the database unreachable from outside its network, PII boundaries drawn and inventoried, and a security-review evidence pack that answers Lisa's questionnaire with artifacts.
- Operated — a runbook the team actually used at 6:30 AM, a quarantine policy Maria owns, corrective actions with owners and dates, and a postmortem practice that turns incidents into sales assets.
- Resilient to its suppliers — the 311 feed is now treated as untrusted input: contract-tested, version-recorded, tolerated and quarantined rather than trusted.
And the handover is explicit. Tom owns the infrastructure and the alerts. Dev owns the code and the pipeline. Maria owns the data policies — retention, quarantine review, the evidence pack. Lisa owns the questionnaire and the quarterly access reviews. Nobody owns "everything," which is why everything is owned.
Acceptance: what "done" looked like
Milestones need acceptance criteria, or they drift. Maria signed off on these six lines — the definition of done for Stage 5:
| # | Acceptance criterion | Verified how |
|---|---|---|
| 1 | Production serves the mayor's-office demo URL | Demo ran 09:00 on current data |
| 2 | Every deploy goes through CI/CD; no hand-deploys | Deploy record exists for every running image |
| 3 | Monitoring pages a human within 10 min of ingestion stall | Alert fired 06:12 on the real incident |
| 4 | Audit trail records every sensitive action, denials included | Sampled allow + deny records, Lisa reviewed |
| 5 | Security questionnaire answered with dated artifacts | Evidence pack accepted by procurement |
| 6 | Incident produced a blameless postmortem within 6 hours | Delivered 12:00, forwarded unedited |
Criterion 6 is the one other teams skip, and it's the one that matters most. A milestone that ends with "we launched" is a deployment. A milestone that ends with "we launched, it broke, here's the postmortem" is a production operation. Maria signed line 6 with the comment: "This is why we'll survive the next one." Stage 5 is complete.
Stage 5's arc, in one line: deployment is not the moment you go live; it's the system that keeps you live. The Compose file got CityOps into production. The pipeline, the monitors, the audit trail, the runbook, and the postmortem are what keep it there — including, especially, on the mornings when the city's feed renames a field at 2 AM.
That's the whole curriculum's bet, stated one last time: the FDE's job was never to write the cleverest code. It's to build the system that survives the questionnaire, the schema change, and the 9 AM demo — and to leave behind the artifacts that prove it. CityOps is live. The evidence says so.
Field check
- The incident's root cause says "nobody made an error." A stakeholder asks: "Then why did it take four hours to detect?" Is that a blame question or a systems question? How does the postmortem answer it?
- Your feed contract test breaks at 3 AM on a Saturday — the upstream renamed another field. The pipeline is green and production is healthy. What do you do before Monday, and why is this outcome better than the 2:14 AM version?
- Maria wants the postmortem "softened" before forwarding — less mention of the crash loop. What do you tell her, and what principle is at stake?
- Map each Stage 5 post (1–7, by title) to one component of the final production system. Which component would hurt most to remove, and why?
- The mayor's office asks: "Prove this won't happen again." What's the honest answer — and which corrective action comes closest to a real guarantee?
Field check answers
1. It's a systems question wearing blame's clothes — and the postmortem answers it as one: detection took four hours because the lag threshold was 15 minutes and the crash loop wasn't itself alerted on. The fix is a mechanism (tighter threshold + restart-rate alert), not a person to watch harder. Blameless doesn't mean questionless; it means every question terminates in a system change.
2. You update the parser's alias list (or confirm quarantine caught it), run it through the normal pipeline, and go back to sleep — no war room, because nothing is broken in production. This is better because the failure was converted from "silent pipeline death discovered at dawn" into "a test failure on a Saturday." Same upstream event, two orders of magnitude cheaper.
3. You tell her no — gently, and with the reason: the postmortem's value is its candor. A softened version that hides the crash loop teaches the mayor's office that you manage impressions; the unedited version teaches them you manage systems. The principle: the postmortem is the product, and products don't ship with the defects edited out of the brochure.
4. Roughly: Compose → the stack; CI/CD → the pipeline; health checks → the readiness gates; monitoring → the lag alerts and dashboards; secrets → the mounted credentials; RBAC → the endpoint authorization; security review → the audit trail and evidence pack. Hardest to remove: the pipeline — without it, every deploy (including incident fixes) becomes a hand-operated ceremony, and the 7:40 fix would have been a 10:40 fix.
5. The honest answer: you can't prove a negative — upstream feeds will change again. What you can show is that the failure mode changed: the next rename hits the contract test or the quarantine queue, not the crash loop. The quarantine queue comes closest to a guarantee, because it bounds the worst case ("some records wait for human review") instead of promising the best case ("nothing ever breaks").