Every API lesson so far has been about calling out. This week the direction reverses: the city can now push 311 status changes to CityOps the moment they happen, instead of us polling for them. A webhook is just an HTTP request — except this time, you're the server, the internet is the client, and anyone on the internet can knock. The lesson: never trust a knock you didn't verify.

Assumes: Post 6 (fetching APIs) and the HMAC signing from the auth lesson earlier in this stage. Everything runnable below is stdlib-only on localhost — no accounts, no keys, no network.

Wednesday, 2:15 PM. Dev: "Good news — the city's 311 system can push status-change events to us now. No more polling every 15 minutes. I just need to expose an endpoint and trust whatever arrives."

You: "Trust whatever arrives. From the open internet. On an endpoint that updates Maria's dashboard."

Dev: "Well — when you say it like that."

Maria (overhearing): "I want the live updates. Yesterday a council member asked about a complaint we'd already resolved and I found out from him."

Lisa (also overhearing): "And I want proof that nobody except the city can trigger those updates. A forged 'complaint resolved' event is a lie on Maria's dashboard with your name on it."

Before you code: clarify the ask

You: "Maria — what does 'live' mean? Seconds, minutes?"

Maria: "Minutes. If a status change shows up within five minutes, that's live to me."

You: "Lisa — what does 'proof' mean? What's the attack you're worried about?"

Lisa: "Someone forging an event, or capturing a real one and replaying it later. Either way, your dashboard updates and nobody can tell it wasn't the city."

You: "Dev — what's the retry story? If our endpoint is down for a minute, does the city resend?"

Dev: "Yes — they'll retry for up to a day. So we might get the same event twice."

Input: HTTP POSTs from the city's 311 system, arriving at our endpoint at unpredictable times
Output: status updates Maria can trust, with proof of authenticity Lisa can audit
Deadline: before the council briefing next week

Notice what's new: every API you've built or called so far had you as the client. Now you're the server, which means the threat model flips. When you call an API, you decide what to trust. When the internet calls you, every request is guilty until proven innocent.

The minimal concept

A webhook is a reverse API call: instead of you polling "anything new?", the vendor POSTs "here's something new" to a URL you give them. Three threats, three defenses — and they compose:

ThreatDefenseHow
Forgery — an attacker invents an eventSignature verificationHMAC-SHA256 over the raw bytes, shared secret from env
Capture-and-replay — an attacker resends a real event laterTimestamp freshnessReject anything older than a few minutes
Honest redelivery — the vendor retries and you get it twiceReplay protectionRemember seen event IDs; reject duplicates

The key insight: no single defense covers everything. The signature proves the event came from someone holding the secret — but a captured real event still has a valid signature, so freshness and dedupe close the replay window. The freshness window (five minutes here) is what makes a finite dedupe store safe: anything older than the window is rejected by the timestamp check anyway.

Build it: the receiver

The receiver verifies in a deliberate order — cheapest check first, and note that json.loads only happens after verification. You don't parse attacker-controlled bytes before you've established they're from the city:

import hashlib, hmac, json, os, time
from http.server import BaseHTTPRequestHandler, HTTPServer

SECRET = os.environ["WEBHOOK_SECRET"].encode()   # env, never code (Post 8)
MAX_SKEW = 300  # seconds
seen_ids = set()

def sign(timestamp: str, body: bytes) -> str:
    msg = timestamp.encode() + b"." + body
    return hmac.new(SECRET, msg, hashlib.sha256).hexdigest()

class Handler(BaseHTTPRequestHandler):
    def do_POST(self):
        length = int(self.headers.get("Content-Length", 0))
        body = self.rfile.read(length)          # raw bytes — verify these
        ts = self.headers.get("X-Webhook-Timestamp", "")
        event_id = self.headers.get("X-Webhook-Id", "")
        sig = self.headers.get("X-Webhook-Signature", "")

        try:
            age = abs(time.time() - int(ts))
        except (TypeError, ValueError):
            age = None
        if age is None or age > MAX_SKEW:
            return self._reject(401, event_id, "stale timestamp")
        if not event_id or event_id in seen_ids:
            return self._reject(409, event_id, "duplicate event id")
        if not hmac.compare_digest(sign(ts, body), sig):
            return self._reject(401, event_id, "bad signature")
        seen_ids.add(event_id)
        event = json.loads(body)               # parsed only after verification
        self._accept(event_id, event)

The sender signs the same way — HMAC(secret, timestamp + "." + raw_body). Timestamp in the signed payload is what binds the signature to a moment: a captured signature can't be transplanted onto a newer timestamp because the signature won't match.

Run it: seven scenarios, all real

Sender and receiver on localhost, every scenario executed — the sender's view:

valid            -> 200 {"ok": true, "received": "evt-0001"}
tampered body    -> 401 {"ok": false, "error": "bad signature"}
reserialized     -> 401 {"ok": false, "error": "bad signature"}
forged signature -> 401 {"ok": false, "error": "bad signature"}
replay, 1st send -> 200 {"ok": true, "received": "evt-0005"}
replay, 2nd send -> 409 {"ok": false, "error": "duplicate event id"}
stale timestamp  -> 401 {"ok": false, "error": "stale timestamp"}

And the receiver's decision log — this is the audit trail Lisa asked for:

07:59:54 accept id=evt-0001 status=IN PROGRESS
07:59:54 reject id=evt-0002 bad signature
07:59:54 reject id=evt-0003 bad signature
07:59:54 reject id=evt-0004 bad signature
07:59:54 accept id=evt-0005 status=IN PROGRESS
07:59:54 reject id=evt-0005 replay or missing event id
07:59:54 reject id=evt-0006 stale timestamp (age=3600.32)

Seven scenarios, seven correct verdicts. The three attacks all fail closed; the honest redelivery is absorbed without double-processing.

Break it, three ways

1. The endpoint that trusts the internet. Dev's first proposal was an endpoint with no verification at all — parse the JSON, update the dashboard. We never shipped it, but the forged-signature run above is the proof of what it would have allowed: that request carried a made-up signature and a plausible payload. Against the naive endpoint it's a 200 and a fake "complaint resolved" on Maria's dashboard. Against ours it's a 401 and a log line. The naive version wasn't wrong about the feature — it was wrong about the threat model. When you're the server, the threat model is everyone.

2. The reserialization trap. Look at the third scenario: same content, different whitespace — the sender signed the pretty-printed JSON but transmitted the compact form. Nothing was tampered with, nothing was forged, and it still failed. This is the classic webhook bug: the signature is over bytes, and JSON re-serialization changes bytes. Sign the exact bytes you transmit; verify the exact bytes you receive. The moment you parse-then-resign, you've changed the evidence.

3. The replay that isn't an attack. The vendor retries for up to a day — the duplicate evt-0005 is expected traffic, not malice. Without the seen-ids store, the retry double-applies the status change: double-counted in reports, double-triggered downstream alerts. The 409 isn't punishment — it's the receiver saying "already handled, you're good." Note the deliberate choice: we check replay before the signature, so a redelivered event is recognized even if the clocks have drifted a few seconds between sends.

Productionize: the store, the rotation, the posture

The in-memory seen_ids set dies on restart. The production version is a tiny table with a TTL — and this run shows the honest tradeoff:

import sqlite3, time
db = sqlite3.connect("webhooks.db")
db.execute("CREATE TABLE IF NOT EXISTS seen(event_id TEXT PRIMARY KEY, first_seen REAL)")

def note(event_id):
    try:
        db.execute("INSERT INTO seen VALUES (?, ?)", (event_id, time.time()))
        return "new"
    except sqlite3.IntegrityError:
        return "duplicate"   # PRIMARY KEY does the dedupe — Post 3's trick

def prune(older_than_seconds):
    cutoff = time.time() - older_than_seconds
    return db.execute("DELETE FROM seen WHERE first_seen < ?", (cutoff,)).rowcount
new
duplicate
pruned: 1
new

Read that last line carefully: after the TTL expired, the old event id was accepted as "new" again. That's safe because the freshness check bounds the replay window — any replayed request must carry its original timestamp, and a timestamp older than 300 seconds is rejected before the store is even consulted. The checks compose: freshness limits the window, the store covers the window.

Secrets rotate, so verification must accept two of them during the rollover:

def verify(ts, body, sig):
    msg = ts.encode() + b"." + body
    return any(hmac.compare_digest(
        hmac.new(k, msg, hashlib.sha256).hexdigest(), sig)
        for k in (PRIMARY, SECONDARY))   # new secret, then old
old-secret sig accepted: True
forged sig accepted: False

The rollout: deploy the dual-secret receiver, switch the sender to the new secret, retire the old after the freshness window plus a margin. No downtime, no dropped events — and the secret lives in env (Post 8), never in the code. And one posture rule for the whole endpoint: every rejection is a controlled 4xx with a log line. Attacker input must never produce a 500 — a traceback on malicious input is both a log-spam vector and an information leak.

Explain it to the customer

"Lisa — every inbound event is now verified three ways before it touches the dashboard: the HMAC signature proves it came from someone holding the shared secret, the timestamp proves it isn't a captured replay, and the event-id store proves we haven't already processed it. Every decision — accept or reject, with the reason — lands in the webhook log, so the audit trail is the log file, not my memory. Maria — status changes land within minutes of the city sending them, and redeliveries can't double-apply. Dev — the receiver is one endpoint plus a tiny SQLite table; the sender needs three headers, and there's a test script that exercises all seven scenarios against localhost."

The pattern, again: the guarantees, then the evidence, then the operator's view. Lisa got proof, Maria got latency, Dev got a runbook.

Must know

  • A webhook is the vendor calling YOU — every inbound request is untrusted until verified
  • Verify three things, in order: freshness (timestamp), replay (event id), signature (HMAC over raw bytes)
  • hmac.compare_digest for signature comparison — never == on secrets
  • Sign the exact bytes transmitted; verify the exact bytes received — re-serialization breaks signatures
  • Parse attacker-controlled bytes only after verification

Useful later

  • mTLS for webhook receivers, when the vendor supports it — signatures plus transport identity
  • Queue-then-ack: return 200 fast, process asynchronously, so slow handlers don't trigger vendor retries
  • Per-event-type handlers and payload schema versioning
  • Alerting on rejection spikes — the same instinct as the DLQ spike in Milestone 2

Don't memorize this

  • Header names — every vendor names them differently; read their docs
  • The exact HMAC construction — remember shared secret + raw bytes + compare_digest, look up the rest
  • Which status code for replays — we used 409; some contracts want 200; match the vendor's contract

Where this lands in CityOps

This lesson is the receiver half of Milestone 3's webhook alerts: when a critical 311 request changes status, the city pushes an event and CityOps fans out the alert. The sender half — CityOps emitting webhooks to downstream subscribers — is the mirror image, with the same three-header contract. And the verification order here (freshness, then replay, then signature) is the same defense-in-depth instinct as the intake pipeline's contract gate: cheap checks first, expensive trust last. Post 5's boundary lesson shows up too — the 300-second freshness window is a time boundary, and like all boundaries it rejects loudly instead of absorbing quietly.

The signature line for this one: never trust a knock you didn't verify. Post 6 taught you to call APIs like a careful client. This is the other side of the table — receiving like a careful server.

Field check

  1. A webhook arrives with a valid signature but a timestamp from yesterday. Accept or reject? Why?
  2. The receiver gets the same event id twice, both with valid signatures and fresh timestamps. What happens — and what does that tell you about the sender?
  3. You change the sender to sign json.dumps(payload) after parsing instead of the raw bytes. Everything breaks. Why?
  4. The seen-ids store is wiped on a restart. What's the worst that happens — and which of the other checks limits the damage?
  5. Lisa asks: "Rotate the webhook secret without dropping events." What do you do?
What good answers look like

1. Reject — 401, stale timestamp. A valid signature on an old message is exactly what capture-and-replay looks like; freshness is what bounds the replay window, and "valid signature" alone proves nothing about recency. 2. First is a 200, second a 409. It tells you the sender retried — normal at-least-once delivery, not an attack. The duplicate is absorbed; the 409 is the receiver saying "already handled." 3. Because the signature is computed over bytes, and re-serialization changes bytes (whitespace, key order, indent). The receiver verifies the raw bytes it received; if you signed different bytes, the signature dies even though the content is identical. Sign what you transmit. 4. A replayed event could be processed twice after the restart. The damage is limited by the freshness check: only events younger than 300 seconds can pass verification anyway, so the double-processing window is five minutes. The checks compose — each covers the others' gaps. 5. Dual-secret verification: the receiver accepts signatures from both the old and new secret during the rollover; switch the sender to the new secret; retire the old after the freshness window plus a margin. The secret itself lives in env (Post 8), never in the code — so rotation is a config change, not a deploy.