Your deployment is a promise: CityOps will answer when the city calls. This is the lesson about keeping that promise — the health checks, logs, metrics, and alerts that let you discover you're broken before Maria's inbox does.
Monday, 8:04 AM. Maria forwards an email with no comment, which from Maria is the angriest possible message. It's from the city's operations desk: "CityOps has been returning errors since about 6 AM. Residents are calling us instead. Please advise."
Dev pulls up the service. "It's responding fine for me," he says — the health endpoint returns 200. Tom checks his phone: no pages, no alerts, because there are no alerts to receive. The deploy went out at 5:40 AM with a bad database setting, and every request touching the database has failed since. The health check never touched the database, so as far as the infrastructure was concerned, everything was fine.
Maria's reply is one line: "Why did our customer tell us our system was broken before we knew?"
The customer problem
The bad config wasn't the real failure. Configs go bad; that's what deploys are for. The real failure was blindness: the team had no signal that distinguished "the server answered my ping" from "the server can serve a customer." Dev's health check proved the process was alive. It said nothing about whether the process could do its job — and the job is all the customer cares about.
Three gaps made a two-hour outage out of a five-minute fix:
- No honest health signal. One
/healthendpoint, returning 200 if the Python process was breathing. It checked nothing the service actually depends on. - No usable logs. The app printed ad-hoc strings to stdout —
print("got request")in some places, nothing in others. When Tom tried to reconstruct the morning, there was no request ID, no timestamp discipline, no way to ask "show me every failed request from City Hall's IP between 6 and 8." - No metrics, no alerts. Nobody measured the error rate, so nobody could alert on it. The first monitoring system the team had was Maria's forwarded email.
Maria doesn't want a postmortem full of jargon. She wants an answer to her one-line question — and she wants it to be "never again."
Clarify the ask
Dev replays Maria's question back to her as three concrete asks, because "better monitoring" is a wish, not a work order:
- "Tell me it's broken before the customer does." — a health signal the infrastructure can act on: restart what's dead, stop sending traffic to what can't serve.
- "When it breaks, tell me what happened." — logs and metrics that reconstruct the incident: what failed, when it started, how many customers it touched.
- "Tell the right person, not everyone." — alerts routed by severity: Tom gets paged for customer-impacting failures; Maria gets a status note, not a 3 AM wake-up.
Out of scope, deliberately: picking a vendor. Datadog, Grafana Cloud, and friends are fine products, but tooling is not the lesson. The lesson is instrumentation — the signals your app emits. A well-instrumented app can move between vendors in an afternoon; a black box stays a black box no matter whose dashboard it's on.
That framing gives us the FDE principle for this lesson:
One principle
"If you don't know it's broken, the customer will tell you." Monitoring isn't a dashboard — it's the difference between finding out from your own systems and finding out from an angry email. Every signal below exists so the team hears it first.
The minimum concept: four signals, two questions
Production observability is four signals. Each answers a different question, and confusing them is exactly how Monday happened:
| Signal | Question it answers | Who consumes it |
|---|---|---|
| Liveness | Is the process alive? | The orchestrator — restart it if not |
| Readiness | Can it serve traffic right now? | The orchestrator — hold traffic if not |
| Logs | What happened, request by request? | Humans debugging an incident |
| Metrics | How many, how long, how wrong — over time? | Alerting rules and dashboards |
Liveness and readiness deserve their own paragraph, because teams merge them into one endpoint and then wonder why the orchestrator can't help. Liveness asks about the process. Is it breathing? If not, restart it — that's the only fix a dead process needs. A liveness check must be cheap and must not touch the network: if your liveness probe depends on the database, a database blip becomes a restart cascade, which is how a hiccup becomes an outage. Readiness asks about the workload. Can this instance serve a customer request right now — database reachable, model loaded, migrations applied? If not, hold its traffic and let the healthy instances carry the load. Readiness failing is not a reason to restart; it's a reason to wait.
The myth that caused Monday
"Our health check returns 200, so the service is healthy." A health GET that returns 200 proves exactly one thing: the server responded to a GET. It says nothing about the database, the model, or the last deploy's config. Dev's endpoint was telling the truth — it was just answering a question nobody asked.
Logs are the narrative: what happened, in order, with enough context (which request, which tenant, how long it took) to reconstruct an incident. Metrics are the numbers: counts and durations aggregated over time, which is what makes "error rate over the last 5 minutes" a computable thing. Alerts are the contract on top: when this symptom appears, tell this person, and here's what they do. Metrics detect; logs explain. That split matters — you'll see the same "detect vs. explain" discipline from the evals lesson, and it's the same instinct.
Build: health endpoints that tell the truth
Two endpoints, two questions, no sharing. The liveness endpoint touches nothing but the process itself. The readiness endpoint proves the instance can do its actual job — including the checks the model-operations lesson demands: the required model artifact is loaded, and it's the digest this deploy pinned.
# cityops/health.py — drop-in health endpoints for the CityOps API
from __future__ import annotations
import time
from fastapi import APIRouter
from fastapi.responses import JSONResponse
router = APIRouter()
STARTED_AT = time.time()
# --- dependency probes: small, fast, honest --------------------------------
# Each probe answers one question and finishes in well under a second.
# Probe timeouts come from the latency SLO; the 1.0s below is an EXAMPLE,
# not a default. Your timeout bounds YOUR wait — it says nothing about the
# remote operation, which may still be running after you give up on it.
def check_database(timeout_s: float = 1.0) -> tuple[bool, str]:
"""Prove we can actually run a query — not just that the driver imports."""
try:
with psycopg.connect(DB_DSN, connect_timeout=timeout_s) as conn:
with conn.cursor() as cur:
cur.execute("SELECT 1")
cur.fetchone()
return True, "SELECT 1 ok"
except Exception as exc:
return False, f"{type(exc).__name__}: {exc}"
def check_model() -> tuple[bool, str]:
"""Readiness proves the required model can serve — not just that the
server booted. _model_digest is set once at startup (see below)."""
if _model_digest is None:
return False, "model not loaded"
if _model_digest != EXPECTED_MODEL_DIGEST:
return False, "digest mismatch: wrong artifact for this deploy"
return True, f"digest {_model_digest[:12]}..."
@router.get("/health/live")
def liveness():
# ONE question: is this process alive? Cheap, local, no network.
return {"status": "ok", "uptime_s": round(time.time() - STARTED_AT, 1)}
@router.get("/health/ready")
def readiness():
# A DIFFERENT question: can this instance serve traffic NOW?
checks = {
"database": check_database(),
"model": check_model(),
}
ok = all(passed for passed, _ in checks.values())
body = {
"status": "ready" if ok else "not_ready",
"checks": {
name: {"ok": passed, "detail": detail}
for name, (passed, detail) in checks.items()
},
}
return JSONResponse(status_code=200 if ok else 503, content=body)
Two supporting pieces make this real rather than decorative. First, EXPECTED_MODEL_DIGEST is pinned per deploy (a config value, not a secret) and _model_digest is set once at startup — if the digest doesn't match, the process refuses to start at all. A model that loaded the wrong artifact is not "degraded"; it's wrong, and wrong shouldn't serve. Second, DB_DSN comes from config with the password supplied by the secret store (next lesson's territory — for now, note the separation: connection string shape in config, credentials in the store).
Now wire the two endpoints to the orchestrator, each with its own job:
# kubernetes/deployment.yaml (excerpt) — probes with DIFFERENT jobs
livenessProbe:
httpGet: {path: /health/live, port: 8000}
periodSeconds: 15
timeoutSeconds: 2 # EXAMPLE — derive from your latency SLO
failureThreshold: 3 # dead 45s before we restart it
readinessProbe:
httpGet: {path: /health/ready, port: 8000}
periodSeconds: 10
timeoutSeconds: 3 # EXAMPLE — readiness may wait slightly longer
failureThreshold: 2 # 20s of not-ready before traffic is held
Notice the asymmetry, because it encodes the concept: liveness failure means restart, readiness failure means wait. And liveness never checks the database — if it did, every database blip would restart every pod, which turns a slow dependency into a self-inflicted outage.
Build: logs that answer questions
When Tom tried to reconstruct Monday, the logs were print("got request") in some handlers and silence in others. A log line is only useful if a tired human (or a log aggregator) can filter it, sort it, and correlate it. That means three properties: structured (JSON, one object per line — machines parse it without regex archaeology), timestamped (UTC, ISO format, always), and correlated (every line from one request carries the same request ID).
# cityops/logging_config.py
import json
import logging
import sys
from datetime import datetime, timezone
# Deliberately small allowlist. If a field isn't here, it doesn't get logged —
# which is how "never log secrets" becomes a property of the code, not a wish.
SAFE_CONTEXT_KEYS = ("run_id", "route", "tenant", "latency_ms", "status")
class JsonFormatter(logging.Formatter):
def format(self, record: logging.LogRecord) -> str:
payload = {
"ts": datetime.now(timezone.utc).isoformat(),
"level": record.levelname.lower(),
"logger": record.name,
"msg": record.getMessage(),
}
for key in SAFE_CONTEXT_KEYS:
if hasattr(record, key):
payload[key] = getattr(record, key)
return json.dumps(payload)
def configure_logging(level: str = "INFO") -> None:
handler = logging.StreamHandler(sys.stdout)
handler.setFormatter(JsonFormatter())
root = logging.getLogger()
root.handlers.clear()
root.addHandler(handler)
root.setLevel(level)
# Uvicorn's per-request access logs are noise in prod; ours carry more.
logging.getLogger("uvicorn.access").setLevel(logging.WARNING)
One middleware then gives every request its correlation ID and its timing — the two things Tom needed on Monday and didn't have:
# cityops/request_logging.py
import logging
import time
import uuid
from fastapi import Request
log = logging.getLogger("cityops.http")
async def log_requests(request: Request, call_next):
run_id = request.headers.get("x-run-id") or uuid.uuid4().hex[:12]
start = time.perf_counter()
status = 500
try:
response = await call_next(request)
status = response.status_code
finally:
latency_ms = round((time.perf_counter() - start) * 1000, 1)
log.info(
"request completed",
extra={
"run_id": run_id,
"route": request.url.path,
"latency_ms": latency_ms,
"status": status,
},
)
response.headers["x-run-id"] = run_id
return response
Two deliberate choices here. First, run_id is the same correlation ID the Stage 4 audit trail uses — one identifier joins the HTTP log to the model-call audit, so an incident can be traced end to end. Second, the extra fields are drawn from the allowlist in the formatter: a developer can't accidentally log a password by stuffing it into extra, because the formatter drops anything not on the list. Minimization before redaction — the same instinct the guardrails lesson taught for prompts.
Don't memorize levels — use this rule
DEBUG in development, INFO for "a request completed normally," WARNING for "something odd but handled" (a retried request, a slow dependency), ERROR for "a request failed and a human may need to know." If everything is ERROR, nothing is.
Build: metrics worth alerting on
Logs tell you what happened to one request. Metrics tell you what happened to all of them — which is what makes "the error rate is climbing" a sentence you can compute instead of a feeling. The prometheus_client library exposes the standard exposition format; anything that speaks Prometheus can scrape it.
# cityops/metrics.py
import time
from fastapi import APIRouter, Request, Response
from prometheus_client import (
Counter, Histogram, Gauge,
generate_latest, CONTENT_TYPE_LATEST,
)
router = APIRouter()
REQUESTS = Counter(
"cityops_requests_total",
"Requests handled by the CityOps API",
["route", "method", "status"],
)
LATENCY = Histogram(
"cityops_request_seconds",
"Request latency in seconds",
["route"],
# EXAMPLE buckets — tune to your latency SLO, not to this blog post.
buckets=(0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0),
)
IN_FLIGHT = Gauge(
"cityops_requests_in_flight",
"Requests currently being handled",
)
@router.get("/metrics")
def metrics():
return Response(generate_latest(), media_type=CONTENT_TYPE_LATEST)
async def observe_requests(request: Request, call_next):
# Route TEMPLATE, not raw path: "/reports/12345" as a label value would
# create one metric series per report — a cardinality explosion.
route = request.scope.get("route")
route_name = getattr(route, "path", "unknown") if route else "unknown"
IN_FLIGHT.inc()
start = time.perf_counter()
status = 500
try:
response = await call_next(request)
status = response.status_code
return response
finally:
elapsed = time.perf_counter() - start
REQUESTS.labels(
route=route_name, method=request.method, status=status
).inc()
LATENCY.labels(route=route_name).observe(elapsed)
IN_FLIGHT.dec()
The label discipline in that middleware is the whole lesson in one function: route gets the template (/reports/{report_id}), never the raw path. Metric labels are for dimensions you group by — route, method, status — not for identifiers. Put a report ID or a user email in a label and every unique value becomes a new time series; that's how a metrics bill (or a crashed Prometheus) happens. Tenant ID as a label is a judgment call: fine for a handful of cities, dangerous for thousands — know which one you are.
With those three series, the questions Maria asked become queries:
- "When did it start?" —
sum(rate(cityops_requests_total{status=~"5.."}[5m]))plotted over the morning shows the exact minute the 5xx rate lifted. - "How many customers?" — the same rate, integrated over the window, counts failed requests. (It's a count of failures, not of humans — don't claim the code computes what it doesn't.)
- "Is it still broken?" —
cityops_requests_in_flightstuck high plus climbing latency means the pool is still exhausted.
Break: three failures, demonstrated
Dev builds the new instrumentation, then deliberately breaks it three ways — because each of these has bitten a real team, and "we'd never do that" is what every team says first.
Failure 1: readiness missing, liveness lying by omission. Dev deploys with only /health/live wired to both probes. The database password rotates (more on that next lesson); every readiness-worthy check would fail, but there is no readiness check. Liveness returns 200, the orchestrator is content, and customers get 500s for as long as it takes Maria to forward another email. The fix isn't a better liveness probe — it's the second probe.
Failure 2: liveness checks the database. Trying to be thorough, Dev points the liveness probe at /health/ready. The database slows down under load; readiness goes 503; the orchestrator reads that as dead and restarts every pod at once. The restart stampede hammers the recovering database, which slows further, which fails more probes. A dependency hiccup becomes a full outage, manufactured entirely by the monitoring. Liveness tests the process, not the neighborhood.
Failure 3: instrumentation nobody reads. The metrics endpoint works, the JSON logs flow — and nobody looks at either, because there's no alert rule and no dashboard anyone opens. During the next incident Tom greps logs without a run_id to join on (one handler was added without the middleware) and stares at a metrics graph with no thresholds. Instrumentation without alerting is decoration. Logs without correlation IDs are confetti.
Productionize: probes, retention, and alerts that respect humans
Getting the signals right is half the job. The other half is the operating discipline around them — the part that decides whether the 3 AM page is worth waking Tom for.
Probe discipline. Timeouts come from the latency SLO and measured performance — never from a blog post's example values. A readiness probe that takes longer than the orchestrator's patience flaps: not-ready, ready, not-ready, churning traffic between instances. Keep readiness checks fast (a SELECT 1, a digest comparison — milliseconds, not model inference), and keep the expensive verifications at startup, where refusing to boot is the correct behavior.
Log discipline. JSON everywhere, including background workers — one format, or the aggregator can't parse it. Retention is a policy, not an accident: hot searchable storage for days-to-weeks (an example starting point — size it to your incident-review cadence and your contract), colder archive after. And the allowlist in the formatter is load-bearing: it's what lets you tell Lisa, in writing, that secrets and raw PII can't reach the logs by construction rather than by developer discipline.
Alert discipline — the part teams get wrong. Alert on symptoms (error rate, latency, readiness flapping), not causes ("disk 80% full" pages nobody at 3 AM; "requests failing" does). Every alert names one owner and links one runbook. An example starting set, tuned to your SLOs:
# alerts/cityops.yaml — EXAMPLE thresholds; yours come from your SLOs
groups:
- name: cityops
rules:
- alert: CityOpsHighErrorRate
expr: |
sum(rate(cityops_requests_total{status=~"5.."}[5m]))
/
sum(rate(cityops_requests_total[5m])) > 0.05
for: 5m
labels: { severity: page, owner: tom }
annotations:
summary: "CityOps 5xx rate above 5% for 5 minutes"
runbook: "https://wiki.internal/runbooks/cityops-5xx"
- alert: CityOpsReadinessFlapping
expr: changes(cityops_ready_status[10m]) > 4
for: 0m
labels: { severity: ticket, owner: dev }
annotations:
summary: "Readiness flapping — check probe timeouts and DB pool"
Two routing rules keep the humans functional: Tom gets paged for customer-impacting symptoms; Maria gets a status note, not a page — she needs to answer the city, not to debug the pool. And every alert must be actionable: if the runbook's first step is "acknowledge and wait," it's not an alert, it's a notification wearing a costume. Delete it or demote it.
Deploy discipline. Readiness is what makes rolling deploys safe: the new pod doesn't receive traffic until /health/ready passes, so a bad config fails the deploy instead of failing the customers. Monday's 5:40 AM deploy, under this discipline, would have been a pod that never became ready, an alert to Dev, and a rollback — with Maria's inbox untouched.
Communicate: the status update Maria actually wants
Maria doesn't read dashboards. She reads the two-paragraph note that lets her answer the city with confidence. After the new instrumentation ships, Dev writes the template once, and it's the same shape every time:
The status note template
- What happened, in one sentence: "CityOps returned errors for 11 minutes starting 06:02; our monitoring caught it at 06:03."
- Customer impact: "Roughly 340 requests failed; no data was lost; the service is healthy now."
- What we did: "Traffic was held from the bad instance automatically; we rolled back the config at 06:13."
- What changes: "The deploy check that missed this is now a readiness gate, tested in staging."
Notice the first line does quiet, enormous work: "our monitoring caught it at 06:03" — one minute after it started, and before any customer report. That sentence is the entire return on this lesson. It's also why the alert routing matters: Maria hears about incidents from Dev's note, proactively, instead of from the city's operations desk. The team that pages itself is the team the customer trusts.
Internally, Tom keeps a runbook per alert — what the alert means, the first three commands to run, and the exact condition for escalating to Maria ("customer impact confirmed" or "not resolved in 30 minutes"). A runbook is how the 3 AM version of Tom benefits from the 3 PM version's thinking.
Where this lands in CityOps
The Ask CityOps pilot is the first CityOps component real residents depend on, which makes it the first component that gets the full treatment:
- Readiness gates the model. The pilot's readiness check verifies the pinned model digest and a fresh vector index — per the model-operations rule, readiness proves the required model can actually serve requests, not just that the server booted. A deploy with the wrong artifact never takes traffic.
- One
run_idacross the stack. The HTTP middleware'srun_idis the same identifier the Stage 4 audit trail records, so a failed answer can be traced from the resident's request through the model call and back. Logs detect nothing; they explain — and now they can. - Metrics per route, alerts per symptom.
/metricsexposes request rate, error rate, and latency for the API built in Milestone 3: The CityOps API; the alert rules page Tom on customer-impacting symptoms and file tickets for everything else.
Maria's one-line question finally has its one-line answer: "We know before the customer does — our systems page us, and here's the note we send you." Next lesson, Lisa asks her own one-line question — "who can see whose data?" — and the answer is locks, keys, and tenant boundaries.
Field check
- Your liveness probe returns 200 but every customer request fails. Which probe is misconfigured — and is the fix to improve liveness or to add something else?
- A developer points the liveness probe at a check that queries the database. The database slows down for 60 seconds. Walk through what the orchestrator does, and why this turns a hiccup into an outage.
- You're adding a
user_emaillabel to thecityops_requests_totalcounter "for debugging." What breaks, and what's the correct way to get per-user debugging context? - Write the readiness check for a worker that needs (a) the database, (b) a Redis cache, and (c) a 4 GB model file. Which of the three belongs in the probe, which belongs at startup, and why?
- Draft the alert rule for "CityOps latency is degrading" — symptom-based, with owner, threshold source, and runbook link. Then explain why "disk 80% full" is not a page-worthy alert.
Answers
1. Neither probe is misconfigured — readiness is missing. Liveness is correctly reporting "the process is alive," which was never the question. The fix is the second probe: a readiness endpoint that checks the database (and the model), wired to the orchestrator's readiness gate so traffic is held from instances that can't serve.
2. The database slowdown makes the liveness check time out; the orchestrator concludes the process is dead and restarts it. Restarting doesn't fix a slow database, so the new process fails its probe too — and meanwhile every other pod's probe is failing for the same reason, so the orchestrator restarts all of them, stampeding the recovering database. A dependency problem became a full outage because liveness tested the neighborhood instead of the process.
3. Cardinality explosion: every unique email becomes a new metric series, which can exhaust the metrics backend's memory and makes the series useless for aggregation. The correct debugging context is the run_id in the structured logs — look up the failing request's log lines by run_id, which carry the user context without polluting the metrics.
4. Database and Redis belong in the probe — they're fast SELECT 1 / PING checks that answer "can I serve right now?" The 4 GB model file belongs at startup: verifying it takes seconds-to-minutes, which would make the probe flap. Load and verify the model once at boot (refusing to start on digest mismatch); the probe then only checks the cheap in-memory flag that says it's loaded.
5. Something like: alert when the 95th-percentile latency from the cityops_request_seconds histogram exceeds the latency SLO for 10 minutes, owned by Tom, runbook linked. The threshold comes from the SLO — the contract with the customer — not from a guess. "Disk 80% full" isn't page-worthy because it's a cause, not a symptom: disks can sit at 80% for months without customer impact, so paging on it trains the team to ignore pages. Alert on the symptom (requests failing or slowing); let the runbook mention disk as a possible cause.