Nearly every integration you will ever debug is one of three failures wearing a disguise: a name that won't resolve, a conversation the other side rejects, or a sealed envelope that won't open. Learn to tell them apart in thirty minutes and you will never again say "the API is down" when it isn't.
Thursday, 9:03 AM. Maria's message lands in Slack: "The dashboard can't reach the 311 data. Nobody knows why. The borough meeting is tomorrow."
Dev chimes in: "The API is fine. I just hit it from my laptop."
Lisa: "Did the certificate change? Nobody touched the certificate, right?"
Tom: "I don't care what it is. Just make it work."
Three theories, one deadline, zero evidence. This is the most common emergency in deployed engineering — and the fix starts not with a patch, but with a mental model.
The client problem
The CityOps dashboard pulls live service-request data from New York City's open 311 API (data.cityofnewyork.us). This morning the dashboard shows a blank panel and a timeout error. Maria needs it working before tomorrow's borough meeting. Dev is confident the API itself is healthy. Lisa is worried about security configuration. Tom wants his numbers.
Notice what nobody has: a fact about which part of the request is actually failing. "Can't reach the API" is not a diagnosis — it is a symptom that could live in any of three places. Your job is to find out which one, with evidence, in minutes.
Why it happened, explained simply
Every request your dashboard makes is a short trip with three legs. Picture asking a question at a secure government building:
Leg 1 — Find the building. You know the name, data.cityofnewyork.us, but couriers need a street address. DNS is the address book that turns the name into a numeric address. If the address book is wrong, stale, or unreachable, the trip never starts.
Leg 2 — Get through the secure entrance. Before any question is asked, your machine and the server open a private, tamper-proof channel. TLS is that sealed envelope: it proves you are talking to the real server (not an impostor) and scrambles the conversation so nobody in between can read or change it. If the seal fails, the conversation never happens.
Leg 3 — Ask the question and get an answer. Your machine sends a request ("give me one 311 record") and the server sends back a response with a status code and the data. HTTP is the language of this conversation. If the question is malformed, the server says so — precisely, in a status code most people ignore.
Nearly every "the dashboard can't reach the API" emergency is exactly one of these three legs failing — and each one fails differently, which means you can tell them apart if you know what to look at.
The minimal concept: one trip, in order
Here is what happens, step by step, when the dashboard asks the 311 API for a single record. Memorize the order, not the details:
1. DNS: the dashboard asks a resolver, "what address is data.cityofnewyork.us?" and gets back a numeric IP address.
2. Connect: your machine opens a raw connection to that address on port 443 (the standard door for encrypted web traffic).
3. TLS handshake: the two sides agree on encryption, and the server presents a certificate proving its identity. Your machine verifies it.
4. HTTP request: inside the now-sealed channel, the dashboard sends its question: GET /resource/erm2-nwe9.json?$limit=1.
5. HTTP response: the server answers with a status code (200 means the question was understood and answered), headers describing the answer, and the data itself.
Steps 1–3 are the setup. Steps 4–5 are the conversation. Debugging is simply figuring out which step never completed — and the tools below let you watch each step happen.
Fix it together: read the wire
Open a terminal. You are going to reproduce the dashboard's request by hand, leg by leg, against the real 311 API. Everything below runs as-is — no accounts, no keys.
Leg 1: check the address book
Before anything else, ask whether the name resolves. dig (or nslookup on Windows) interrogates the address book directly. A resolver is simply the server your machine asks to do lookups — usually run by your network provider or your company's IT team:
$ dig data.cityofnewyork.us +short
# annotated output — yours may differ, and that's fine:
52.86.200.11 # the numeric address the name currently points to
52.86.222.94 # large services usually return several addresses
What you are reading: each line is an address the name currently points to. If instead you see a header line like this, the address book is telling you the name doesn't exist at all:
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 12345
# NXDOMAIN = "no such domain" — the lookup itself failed
And if the command just hangs with no answer, the resolver itself is unreachable — a network problem one level below the API. Either way, the trip never got past step 1, and no amount of restarting the dashboard will help.
One habit that saves hours: run this from the machine that is failing, not your laptop. Dev's laptop resolving fine tells you nothing about the dashboard server's DNS. "Works on my machine" has never once settled a DNS argument.
Legs 2 and 3: watch the whole conversation
curl -v is the debugger's microscope — it narrates every step of the trip. Run the dashboard's actual question:
$ curl -v 'https://data.cityofnewyork.us/resource/erm2-nwe9.json?$limit=1'
* Trying 52.86.200.11:443... # step 2: connecting to the address
* Connected to data.cityofnewyork.us port 443
* TLSv1.3 (OUT), TLS handshake, Client hello
* TLSv1.3 (IN), TLS handshake, Server hello
* SSL certificate verify ok. # step 3: identity confirmed, channel sealed
> GET /resource/erm2-nwe9.json?$limit=1 HTTP/2 # step 4: the question
< HTTP/2 200 # step 5: the answer's status
< content-type: application/json; charset=utf-8
# ...then the actual data — one real 311 record, trimmed for readability:
[{"unique_key":"70591261","created_date":"2026-10-01T02:05:23.000",
"agency":"NYPD","agency_name":"New York City Police Department",
"complaint_type":"Noise - Street/Sidewalk","descriptor":"Loud Music/Party",
"borough":"BRONX","incident_zip":"10452","status":"In Progress",
"latitude":"40.831669762919574","longitude":"-73.9287544719845"}]
Read it like a story. The * lines are the setup: connection, handshake, certificate verification. The > line is your question. The < lines are the server's answer. If the API is healthy, you see exactly this shape: setup succeeds, a 200 comes back, and JSON data follows.
Read the status code before you read the body
The status code is the server's one-word verdict on your request. Learn the families, not the full registry:
| Family | Meaning | The debugger's question |
|---|---|---|
| 2xx | Success — the server understood and answered | Is the body actually what I expected? A 200 with an empty result is a different bug than a failure. |
| 3xx | Redirect — ask over there instead | Am I following redirects? In scripts, decide deliberately whether redirects are allowed. |
| 4xx | Client error — the request was wrong | What did I get wrong? 400: bad parameters. 401/403: credentials or permissions. 404: wrong path. 429: too many requests, slow down. |
| 5xx | Server error — their side is broken | Is it transient? Retry with backoff. Is it persistent? That's their incident, not your bug. |
Myth: "A 200 status means it worked."
A 200 means the conversation worked — not that the answer was right. An API can return 200 with an empty dataset, a cached stale response, or even an error message dressed up as JSON. The status code tells you the request was understood. Only the body tells you whether the answer is useful.
Back to Thursday morning: run that same curl -v from the dashboard server. Handshake plus a 200 means the API is fine and the bug is in the dashboard; anything earlier tells you exactly which leg broke — and so does everyone on the call.
Break it: what each failure looks like
A model you cannot break on purpose is a model you don't trust yet. Here is what each leg's failure looks like in the terminal, so you recognize it in production.
Breaking leg 1 (DNS): ask for a name that doesn't exist.
$ curl -v 'https://data.cityofnewyork.us.example/resource/erm2-nwe9.json?$limit=1'
* Could not resolve host: data.cityofnewyork.us.example
* Closing connection
curl: (6) Could not resolve host
The signature: it dies before any connection attempt. No TLS lines, no HTTP lines. The trip never started. In the real incident, this is what you'd see if the dashboard server's DNS resolver were misconfigured or a firewall were blocking DNS queries.
Breaking leg 2 (TLS): ask for a server whose certificate is invalid. The site expired.badssl.com is a real public test server whose entire job is to present an expired certificate — it exists so engineers can practice exactly this failure:
$ curl -v "https://expired.badssl.com/"
* Connected to expired.badssl.com port 443
* TLSv1.3 (OUT), TLS handshake, Client hello
* TLSv1.3 (IN), TLS handshake, Server hello
* SSL certificate problem: certificate has expired
* Closing connection
curl: (60) SSL certificate problem: certificate has expired
The signature: the connection succeeds, the handshake starts, then verification fails and curl refuses to continue. Notice that curl stops rather than sending your request over an untrusted channel — that refusal is a security feature, not a bug. When Lisa asks "did the certificate change?", this output is the evidence that answers her.
Breaking leg 3 (HTTP): ask a valid server a bad question — a dataset path that doesn't exist.
$ curl -s -o /dev/null -w "%{http_code}\n" \
'https://data.cityofnewyork.us/resource/does-not-exist.json?$limit=1'
404
The signature: setup completes perfectly and the server answers — with a rejection. The -w "%{http_code}" flag prints just the verdict, which is ideal for scripts. A 404 here means your path is wrong; in a 311-API context it usually means a typo in the dataset ID.
Three failures, three unmistakable signatures. That is the entire diagnostic skill: match the symptom to the leg, then fix the leg.
Make it production-safe
Debugging by hand is for incidents. Production code needs the same three legs checked automatically. Four rules:
1. Time out everything. A request with no timeout can hang forever, and one hung request can quietly take down a whole dashboard. Set a connection timeout and a total timeout on every outbound call you write from here on.
$ curl --connect-timeout 5 --max-time 20 \
'https://data.cityofnewyork.us/resource/erm2-nwe9.json?$limit=1' \
-s -o /dev/null -w "%{http_code}\n"
200
2. Check the status code, not just "did it run." A script that only checks whether curl exited cleanly will happily accept a 500 as success. Read the code, branch on the family: 2xx → parse, 429 → back off and retry, 5xx → retry a few times then alert, 4xx → alert immediately because retrying a bad request never helps.
3. Never disable certificate verification in production. You will see curl -k suggested as a "quick fix" for TLS errors. It is not a fix — it is the removal of the seal on the envelope. It tells your code to trust any server, including an impostor. Fine for a lab experiment; a security incident waiting to happen anywhere real. (Python's equivalent is verify=False in the requests library — the same warning applies, and Lisa will ask about it in Stage 2.) If TLS verification fails in production, the certificate is the problem — fix the certificate.
4. Retry with backoff, and respect 429s. When a server says "too many requests," hammering it faster is how a small bug becomes a blocked IP address. Backoff just means waiting longer between each retry — a second, then two, then four — and only retrying failures that are the server's fault (5xx), never bad requests (4xx). If the response includes a Retry-After header, honor it — the server is telling you exactly when to come back.
Must know
- DNS turns names into addresses — and it is the first leg to check when "can't reach" appears. Debug it on the failing machine, not yours.
- TLS seals the channel and proves identity — a verification failure is a security event, not an annoyance. Never ship with verification disabled.
- Read the status-code family before the body — 2xx/3xx/4xx/5xx tells you whose fault it is and what to do next.
curl -vshows the whole trip — connection, handshake, request, response. When in doubt, watch it happen.- 200 means "understood," not "correct" — always validate the body against what you expected.
Useful later
- Certificate chains and intermediate authorities — why a cert can be valid yet untrusted on one machine
- HTTP/2 and HTTP/3 — what changes under the hood (Post 6 revisits this with load balancers and CDNs)
- DNS caching and TTL — why a fixed DNS record can take time to be seen everywhere
- Reading a certificate by hand with
openssl s_client— for the day curl's error message isn't enough
Don't memorize this
- The exact names of TLS handshake messages — know the shape (agree, prove identity, seal), look up the labels
- Every field in
digoutput — the answer section and the status line carry the signal - The full status-code registry — the four families plus 401/403/404/429 cover the vast majority of real incidents
Explain it to the customer
Tomorrow's borough meeting needs an explanation, not a lecture. Here is how you say it to Maria and Tom — two paragraphs, no jargon:
"When the dashboard shows 311 numbers, it asks the city's data service a question in three steps. First it looks up the service's address, like finding a building in a directory. Then it opens a private, sealed line to that address so nothing in between can eavesdrop or tamper with the numbers. Then it asks the question and the service answers. Yesterday morning the first step failed on our dashboard's server — the address lookup wasn't working there — so the question never even left the building. The city's service was healthy the whole time, which is why it looked fine from our laptops."
"We've fixed the lookup on the server, verified the dashboard is pulling live numbers again, and added an automatic check that tests all three steps every five minutes. If any step fails in the future, we'll know which one within minutes instead of discovering it the morning of a meeting."
Notice the structure: what it is, what broke, what we did, what changes going forward. That shape works for nearly every incident explanation you will ever write.
Add it to CityOps
This is where the model becomes infrastructure. Stage 1's milestone asks you to ship a deployed CityOps status page — and a status page that can't tell you why something is red is decoration. Build the three-leg check into it.
Add a small health script to the repo — call it scripts/check_311.sh — that performs the same three checks you just ran by hand, in order, and reports which leg failed:
#!/bin/bash
# CityOps 311 connectivity check: DNS, then TLS, then HTTP — in order.
HOST="data.cityofnewyork.us"
URL="https://data.cityofnewyork.us/resource/erm2-nwe9.json?\$limit=1"
# Leg 1: DNS
dig +short "$HOST" > /dev/null || { echo "FAIL: DNS — cannot resolve $HOST"; exit 1; }
# Legs 2+3: TLS + HTTP, with timeouts, verifying the certificate
CODE=$(curl --connect-timeout 5 --max-time 20 -s -o /dev/null -w "%{http_code}" "$URL")
case "$CODE" in
2*) echo "OK: 311 API reachable (HTTP $CODE)" ;;
429) echo "WARN: rate limited (HTTP 429) — backing off" ; exit 2 ;;
4*) echo "FAIL: HTTP $CODE — our request is wrong, check the URL" ; exit 1 ;;
5*) echo "FAIL: HTTP $CODE — server-side problem, retry then escalate" ; exit 1 ;;
*) echo "FAIL: no usable response (TLS or connection failure)" ; exit 1 ;;
esac
Run it on a schedule and surface the result on the status page. When it's healthy, the output is reassuringly boring:
$ ./scripts/check_311.sh
OK: 311 API reachable (HTTP 200)
Then write a one-paragraph runbook entry — a runbook is just a short playbook of what to do when something breaks, written for the person who gets paged at 2 AM:
If the check reports a DNS failure, look at the dashboard server's resolver configuration first. If it reports a TLS failure, check whether the certificate expired and whether the server's clock is correct (certificates are time-sensitive — a wrong clock breaks verification). If HTTP 4xx, our request changed — check the URL and dataset ID. If HTTP 5xx, check the provider's status page, wait out the backoff, then escalate.
You have just turned this post's mental model into a production artifact — and into the first entry of your incident runbook, which Stage 5 will grow into a full discipline.
This post was a Discover-and-Learn loop in miniature: observe the failure, scope which leg broke, build the check, and feed what you learned back into the system.
Field check
- The dashboard log says
Could not resolve host: data.cityofnewyork.us, but the API loads fine on your laptop. Which leg failed, and what is your very first check? - Dev points at the dashboard and says, "The API returns 200, so the data must be right." What's wrong with that reasoning, and where do you look next?
- Lisa asks why the pipeline uses encrypted connections instead of plain HTTP. Answer her in two sentences, in plain language, with no technical terms she would have to look up.
What a good answer looks like
1. Leg 1 — DNS — failed on the dashboard server specifically. Your laptop is irrelevant evidence; the first check is to run the name resolution from the failing machine itself (dig or nslookup there), then look at that machine's resolver configuration and whether anything between it and the resolver is blocking the query. 2. A 200 means the server understood the request, not that the answer is correct — the next place to look is the response body: is it an empty array, a stale cached payload, or an error object wearing a 200 status? Then verify the query parameters against the data contract. 3. Something like: "Plain connections send every request and every number readable across each network in between, so anyone along the way could see or quietly change the data. The encrypted connection scrambles everything in transit and proves we're talking to the real city server, not an impostor."