"The site is down" is not a diagnosis. It is a symptom with at least six possible locations. This chapter teaches you to read the request path — client, DNS, CDN, load balancer, firewall, server — and point at the right layer with evidence, not guesses.

Wednesday, 8:47 AM. An email from Maria, the operations director:

"The CityOps status page is down. Nobody in the office can load it. This is the page you said would prove the system is healthy. Please advise."

You check from your phone: it loads fine. Dev, the data engineer, writes back: "Server logs are clean, uptime is 100%." Lisa, security, asks: "Did anything change on the network?" Tom wants to know whether to tell the borough meeting the system is broken.

The server is healthy. The client is unhappy. Both statements are true — because "down" depends on where you stand.

Forward Deployed Engineer Solutions Architect Customer Engineer Sales Engineer

Why it happened, simply

The answer you will prove below: the office's egress proxy was blocking the status page's domain. City IT keeps an allowlist of approved domains for the office network. The status page was deployed to a new subdomain, nobody added it to the allowlist, and every office request died at the proxy before reaching the internet.

This is the shape of many FDE "the site is down" incidents: the application is healthy, but something in the path between the user and the application is not. Your job is to trace the request's journey, find exactly where it dies, and prove it with evidence the client can act on.

Myth: "The site is down" describes one thing.
"Down" is a symptom reported by a person in a place, not a diagnosis. Office: unreachable. Your phone: fine. Server monitoring: green. All three reports were accurate — of three different vantage points. Treat "down" as the start of an investigation, never the answer.

The minimal concept: the request path

When someone loads a page, the request does not teleport from their laptop to your server. It walks a path, and every step can fail:

Client → DNS → CDN → Load balancer → Firewall → Server

TCP/IP, just enough. Data travels in packets, each stamped with a destination address. IP is the addressing system — every machine gets an address, like a street address. TCP is the agreement on top: your machine opens a connection with a quick handshake, and both sides acknowledge receipt so nothing silently goes missing. A port is the door number — 443 means "the HTTPS door." If the handshake never completes, nothing else can happen.

DNS (post 4's phone book) translates the name you type into an IP address. If DNS fails, you never learn the address, and nothing after it can work.

CDN (Content Delivery Network) is a worldwide fleet of edge servers that cache copies of your content close to users — local warehouses instead of one distant factory. If the CDN is misconfigured or its cache is stale, users see errors or old content while the origin server sits there perfectly healthy.

Load balancer spreads incoming requests across your servers. One name, many machines behind it. If its health checks (periodic "are you alive?" probes) decide all backends are sick, it stops forwarding traffic — and the site looks dead even though the servers might be fine.

Firewall / proxy is the gatekeeper. A firewall blocks or allows traffic by rule (which addresses, which ports). A proxy makes requests on your behalf — in corporate offices, outbound web traffic usually goes through an egress proxy that can block domains not on an allowlist. This was the layer that failed in our story.

Server is the machine that runs your application. The last stop, not the first suspect.

Myth: "Ping works, so the site is fine."
Ping tests whether a machine answers a knock on a side door. It says nothing about port 443, TLS, the CDN's content, or the application — and many networks ignore pings entirely. A successful ping proves almost nothing about a website; a failed one proves even less. Use the right probe for the layer you are testing.

The symptom table: which layer each failure implicates

Memorize this table's shape, not its exact words:

SymptomLikely layerWhat to run
Works from home, fails from the officeOffice firewall / proxy (your story)Test from a second network; compare
curl fails before any timing data — "Could not resolve host"DNSdig the name, try a public resolver
curl timing shows DNS fine, then hangs on connectFirewall / proxy / network pathCheck proxy env vars, try --noproxy
curl timing shows connect + TLS fine, but first byte takes secondsServer or overloaded backendCompare against a static asset on the same host
HTTP 5xx while origin server logs are cleanCDN or load balancer (something in the middle is erroring)Read response headers: via, cache headers
Users see old content; a hard refresh shows the new pageCDN cacheCheck age / cache headers; purge the cache
Browser shows a certificate warning for the right domainTLS / CDN edge configopenssl s_client -connect, check cert dates
Nobody can reach it from anywhere; monitoring agreesServer or the load balancer's health checksCheck server logs + LB health-check status

The most important row is the first. When the failure depends on which network you're on, the network you're on is the suspect — not the site.

Fix it together: the debugging runbook

Here is the exact sequence, the way you would run it on the call with Maria. Every command below is real and runnable — the outputs shown are from this story's scenario (the hostnames are fictional), so run the same commands against your own URLs to see your live results.

Step 1 — Confirm it is not you: test from two networks

From the office Wi-Fi:

$ curl -sS -o /dev/null -w "code=%{http_code} time=%{time_total}s\n" \
    https://status.cityops.example/ --max-time 10
curl: (28) Connection timed out after 10001 milliseconds
code=000 time=10.001s

Then tether to your phone's data connection — same laptop, same command, different network:

$ curl -sS -o /dev/null -w "code=%{http_code} time=%{time_total}s\n" \
    https://status.cityops.example/ --max-time 10
code=200 time=0.318s

Read: same machine, same URL, two networks — one fails, one returns 200 in a third of a second. The site is alive; the office network path is killing the request. Server, CDN, load balancer, application: all eliminated in under a minute.

Step 2 — Find which stage dies: time each phase

curl can time each phase of the request separately — the single most useful networking command an FDE knows:

$ curl -sS -o /dev/null -w \
  "dns=%{time_namelookup}s connect=%{time_connect}s tls=%{time_appconnect}s firstbyte=%{time_starttransfer}s total=%{time_total}s code=%{http_code}\n" \
  https://status.cityops.example/ --max-time 10

From the phone (working):

dns=0.021s connect=0.048s tls=0.102s firstbyte=0.214s total=0.318s code=200

From the office (failing):

dns=0.019s connect=0.000s tls=0.000s firstbyte=0.000s total=10.001s code=000

Read: DNS resolved in 19ms — the name is fine. Then connect=0.000s: the TCP handshake never completed. The request died between your machine and the first hop — it never reached the open internet. That is the signature of a local firewall or egress proxy blocking the destination, not the CDN, the load balancer, or the server.

Step 3 — Check for the office proxy

Corporate networks usually route outbound traffic through a proxy. Check whether one is configured:

$ env | grep -i proxy
HTTPS_PROXY=http://proxy.cityhall.internal:8080
HTTP_PROXY=http://proxy.cityhall.internal:8080

There it is. Now ask the proxy directly what it thinks of the domain, using -I to fetch just headers and -v to watch the conversation:

$ curl -sSI https://status.cityops.example/ --max-time 10
HTTP/1.1 403 Forbidden
X-Proxy-Block: domain-not-on-allowlist
Content-Type: text/html

<html><body>Access to this site is blocked by policy.
Contact IT Security to request an exception.</body></html>

Read: the proxy answered 403 Forbidden with its own block page — domain-not-on-allowlist. Case closed, with a receipt to forward to IT Security.

Step 4 — Read the headers like an FDE

From the working network, look at the full response headers once. They tell you who actually answered:

$ curl -sSI https://status.cityops.example/
HTTP/2 200
server: nginx
via: 1.1 varnish, 1.1 cdn-edge-ewr12
x-cache: HIT
age: 142

Read: via shows two intermediaries — a caching layer and a CDN edge node (ewr12 = Newark). x-cache: HIT and age: 142 mean this response was served from the CDN's cache, 142 seconds old — you never touched the origin server.

$ dig +short status.cityops.example
status-cityops.cdnprovider.example.
192.0.2.44

The name resolves to a CNAME pointing at the CDN, not at a server directly — confirming that requests normally travel through the CDN edge before reaching your infrastructure.

One more tool for the kit, for the day the failure is past the proxy: mtr --report --tcp --port 443 maps each hop on the path. Hops answering fine, then everything beyond a point going dark, means packets die at that boundary. (One caution: many networks ignore probe packets, so rows of ??? can also mean "this router doesn't answer probes." A hint, not a verdict.)

Break it: other failure shapes at each layer

Now that you can walk the path forward, walk it with a hammer. The detective logic: the geography of the failure tells you the layer. Everyone everywhere fails the same way → look at shared infrastructure (DNS, CDN edge config, load balancer). Only one network fails → look at that network's gatekeepers (firewall, proxy). Only some users fail intermittently → look for the layer that splits traffic (load balancer, CDN cache state).

LayerFailure shapeWhat the user sees
DNSDomain expired, or the CNAME to the CDN was deleted"Could not resolve host" everywhere, from every network
CDNStale cache after a deploy; misconfigured cache rulesHalf the users see yesterday's page; hard refresh fixes it
CDNTLS certificate on the edge expiredBrowser certificate warnings for everyone at once
Load balancerHealth checks misconfigured — all backends marked sick503s while every server reports healthy
Load balancerOne bad backend not drained; sticky sessions pin users to itOnly some users error; others fine
FirewallRule change blocks the CDN's new edge IP rangeSudden outage for one office, site fine elsewhere
ProxyYour story: new subdomain missing from the allowlistBlock page or timeout from the corporate network only
ServerApp exhausts database connections after a traffic spikeSlow first byte, then 500s

Make it production-safe: know before Maria does

Debugging well is good. Finding out before the client emails you is better. Three practices, all cheap:

1. Monitor from more than one vantage point. An uptime check in a Virginia datacenter will never see the office proxy block the domain. Run a second lightweight probe from inside the client's network — a small agent on an office machine, or a scheduled check from their VPN. The Virginia probe says "the site is up"; the office probe says "the site is up for the client."

2. Alert on the layer, not just the outcome. Record the timed phases — DNS, connect, TLS, first byte — not just up/down. Then the alert itself carries the diagnosis: "office probe: connect phase failing, DNS fine."

3. Keep the runbook you just ran. Write down the two-network test, timed curl, proxy check, header read, and mtr; link them from the status page. And add the lesson to the launch checklist: new subdomains get allowlist-requested before launch, with an office-network smoke test.

Must know

  • The request path: client → DNS → CDN → load balancer → firewall/proxy → server — and one plain sentence for what each hop does
  • "Works on one network, fails on another" implicates the failing network's gatekeepers, not the site
  • curl -w timing phases: time_namelookup (DNS), time_connect (TCP), time_appconnect (TLS), time_starttransfer (first byte) — each slow or zero phase names its layer
  • Reading response headers to spot intermediaries: via, cache headers like x-cache/age

Useful later

  • Reading mtr/traceroute output; anycast and how CDNs route you to the nearest edge
  • The TLS handshake in more detail (certificates, SNI) — post 4's model is enough today
  • WAFs (Web Application Firewalls): the firewall variant that inspects requests for attacks — and sometimes blocks legitimate traffic

Don't memorize this

  • TCP header fields or CIDR subnet math — look them up when you need them
  • Any CDN vendor's console clicks — the concepts (edge, cache HIT/MISS, purge) transfer; the buttons do not
  • Port lists beyond 80 (HTTP) and 443 (HTTPS) — the config file is the source of truth

Explain it to the customer: the email you actually send

No jargon, no blame. Here is the email that closes Maria's incident:

Subject: Status page unreachable from the office — found and fixed

Maria —

Found it. The status page itself was healthy the whole time; the issue was that the office network was blocking the page's new address. City IT keeps a list of approved sites for the office network, and the new status page address wasn't on it yet, so office computers couldn't reach it. Computers outside the office were never affected.

What I've done: I've submitted the request to IT Security to add the address to the approved list, and I've confirmed the page loads correctly from outside the office. I'll confirm with you once IT approves it — typically within a day — and I'll verify from an office computer myself.

So this doesn't repeat: before we launch anything with a new address, I'll send the address to IT Security for approval first, and I'll test it from an office computer as part of the launch checklist.

No data was affected, and the underlying CityOps system was never down.

— You

Notice the structure: what happened in plain words, what is fixed, what is still pending and who owns it, what changes going forward, and an explicit statement about data. Lisa gets a one-line version: "New subdomain needs egress allowlist approval — request submitted, details below." Tom gets: "The system was never down; it was an office network setting. Resolving today."

Add it to CityOps

Turn the runbook into a living check. This small Python probe hits the status page with phase-aware error handling, names the failing layer, and prints it — the seed of the dual-vantage monitoring above:

import requests

URLS = {
    "status page": "https://status.cityops.example/",
    "311 API": "https://data.cityofnewyork.us/resource/erm2-nwe9.json?$limit=1",
}

def check(name, url):
    try:
        r = requests.get(url, timeout=10)
        detail = f"first byte {r.elapsed.total_seconds():.3f}s, HTTP {r.status_code}"
        layer = "healthy" if r.status_code < 400 else "application"
    except requests.exceptions.ConnectTimeout:
        layer, detail = "network/firewall/proxy", "connection never established"
    except requests.exceptions.ConnectionError:
        layer, detail = "network/firewall/proxy", "connection refused or reset"
    except requests.exceptions.Timeout:
        layer, detail = "server/load balancer", "connected, but response too slow"
    print(f"{name}: {layer} | {detail}")

for name, url in URLS.items():
    check(name, url)

Sample output from the office during the incident:

status page: network/firewall/proxy | connection never established
311 API: healthy | first byte 0.412s, HTTP 200

Read: the status page dies at the connection phase, but the public 311 API — a different domain, already on the allowlist — works fine from the same machine. Same machine, two domains: the failing layer is confirmed as the path to that domain. That contrast is exactly the evidence you attach to the IT Security ticket.

Run this probe from two places — a cloud VM and a small agent inside the client's network — and alert when the office vantage fails while the cloud vantage passes. This probe joins the CityOps repo in the Stage 1 milestone; Stage 5 promotes it into the full monitoring stack.

Discover→Scope→Build→Deploy→Adopt→Learn

This chapter lives at Deploy → Learn: the request path is the terrain every deployment walks, and every incident teaches the next launch's checklist.

One real record, to keep it concrete

The 311 API the probe checks above serves millions of real service requests — the actual data your CityOps platform will be built on. A typical record:

{
  "unique_key": "70602054",
  "created_date": "2026-10-01T02:01:04.000",
  "agency": "NYPD",
  "complaint_type": "Noise - Residential",
  "descriptor": "Loud Music/Party",
  "borough": "BROOKLYN",
  "status": "In Progress"
}

Every hop you learned today exists to deliver records like this one — from a Socrata server, through a CDN, to a browser in a city office. When it breaks, you now know how to find where.

Field check

  1. The status page loads over your phone's data connection but not from the client's office Wi-Fi. The server's own monitoring shows 200 OK. Which layer is the most likely suspect, and what is the first command you run from an office machine?
  2. curl -w reports time_namelookup=0.02s, time_connect=0.05s, time_appconnect=0.10s, time_starttransfer=8.40s. Which layer does this implicate — the network, the CDN, the load balancer, or the application server — and why?
  3. Lisa asks: "Could the firewall be blocking the CDN's new edge IP range after their rotation?" What would you check to test that hypothesis, and what do you tell Maria while you investigate?
What a good answer looks like

1. The office network's gatekeepers — the egress firewall or proxy — not the site. First: the two-network comparison (timed curl from the office), then check proxy settings (env | grep -i proxy) and ask the proxy directly what it thinks of the domain. 2. The application server (or an overloaded backend behind the load balancer). DNS, TCP, and TLS all completed quickly — the network path is fine — but the first byte took 8.4 seconds, so the server accepted the request and was slow to answer. Compare with a static asset on the same host: if the asset is fast, it is the application, not the infrastructure. 3. Test the hypothesis: run the timed curl from the office and from outside, and check the firewall's recent rule changes against the CDN's published IP ranges. Tell Maria: "The site is healthy from outside the office; we're checking whether a recent network change on the office side is blocking part of the path. No data is affected. I'll update you within the hour."