Nobody hires an FDE to draw boxes. They hire one to stand in front of a whiteboard, ask the questions nobody asked yet, and turn "we need it to scale" into numbers — before a single box gets drawn. This post is one worked scenario, end to end: clarify, compute, then architect.

Assumes: Posts 1 (REST) and 3 (pagination, rate limits, retries). No installs — the arithmetic runs in your head and in one short Python script.

Wednesday, 9:00 AM. A new customer — a regional logistics company — puts you in a room with their CTO and says: "We're launching same-day delivery tracking for 40,000 drivers. Customers refresh the tracking page constantly. We need it to scale. Can you design it?"

The trap: nodding, opening a diagram tool, and drawing boxes labeled "API", "Database", "Cache". Every box you draw before you have numbers is a guess wearing a rectangle costume.

What you do instead: ask questions for twenty minutes, do arithmetic for ten, and then draw — three boxes, each one justified by a number.

Step 1: Clarify the ask (the 20 minutes that matter most)

System design interviews — and real customer rooms — reward the same move: refuse the vague ask. "Scale" is not a requirement; it's a mood. Here's the question set, and what the CTO actually says:

QuestionWhy you askAnswer
How many drivers, how often do they report location?Write load40,000 drivers, every 30 seconds
How many customers watch tracking?Read load~200,000 daily; each refreshes ~12 times per delivery
How fresh must the location be?Consistency vs. caching"Within a minute is fine" — not real-time
How long do we keep the data?Storage sizingLive for 7 days; trip history 1 year for disputes
What breaks if tracking is down for 10 minutes?Availability target"Support tickets, not lawsuits" — important, not critical
Budget?Every design is a budget"Don't embarrass me" — i.e., boring and cheap wins

Notice what happened: "we need it to scale" became six numbers. The freshness answer ("within a minute") just killed every real-time architecture you might have drawn. The availability answer just killed multi-region. Clarifying is designing — most of the architecture is already decided by the time you stop asking questions.

Clarify before you compute

  • Users, reads/sec, writes/sec — get all three as numbers, not adjectives
  • Freshness requirement decides caching; downtime cost decides availability spend
  • Data retention decides storage; budget decides "boring wins"
  • Write the numbers down — they're the spec your arithmetic checks against

Step 2: Do the arithmetic (the 10 minutes that justify the boxes)

Back-of-the-envelope math, all of it checkable. Round aggressively; the point is order of magnitude, not precision:

# --- writes: driver location pings ---
drivers = 40_000
ping_interval_s = 30
writes_per_sec = drivers / ping_interval_s          # ~1,333/s
print(f"writes/sec: {writes_per_sec:,.0f}")        # 1,333

# --- reads: customers refreshing tracking ---
daily_watchers = 200_000
refreshes_per_delivery = 12
# deliveries spread over ~12 active hours
reads_per_sec = daily_watchers * refreshes_per_delivery / (12 * 3600)
print(f"reads/sec: {reads_per_sec:,.0f}")          # ~56

# --- storage: live window ---
bytes_per_ping = 200          # driver_id, lat, lon, timestamp, status
live_days = 7
live_bytes = writes_per_sec * bytes_per_ping * 86400 * live_days
print(f"live storage: {live_bytes / 1e9:.1f} GB")  # ~161 GB

# --- storage: trip history (1 year, 1 ping/min kept per trip) ---
history_pings_per_day = drivers * 1440 / 30        # 1.92M pings/day at full rate
# keep 1-in-30 for history (one point per 30 min is plenty for disputes)
history_bytes = history_pings_per_day / 30 * bytes_per_ping * 365
print(f"history storage: {history_bytes / 1e9:.1f} GB")  # ~4.7 GB

# --- peak factor: lunch rush doubles the refresh rate ---
peak_reads = reads_per_sec * 2
print(f"peak reads/sec: {peak_reads:,.0f}")        # ~111

Read those numbers again, because they just designed the system:

  • ~1,333 writes/sec, ~56 reads/sec (111 at peak). Write-heavy, but neither number is scary — a single Postgres handles this with room to spare. No sharding. No exotic store.
  • ~161 GB live, ~5 GB history. Fits on one disk. The "big data" framing evaporates.
  • Freshness: 1 minute. Cache reads for 30–60 seconds and the write load never touches the database on the read path.

This is the post's principle: do the arithmetic before you draw the boxes. Every number above eliminated an architecture. The remaining design is almost forced — which is exactly what you want. A design that falls out of the numbers is defensible; a design that fell out of a diagram tool is a preference.

Step 3: Draw the boxes (three of them, each with a receipt)

Only now, the diagram — and every component carries the number that justifies it:

Driver app ──POST /pings──▶ API ──▶ Postgres (locations table)
                                    │
Customer ◀──GET /tracking── API ◀── Redis (30s TTL cache)
                                    │
                          nightly job ──▶ object storage (trip history)
BoxJustified byDeliberately not
Single Postgres1,333 writes/sec — one node handles 10x thisSharding, Cassandra — the numbers don't require it
Redis, 30s TTL56→111 reads/sec; 1-minute freshness SLAReal-time push — customer said a minute is fine
Nightly history job161 GB live vs 5 GB history — different access, different storeKeeping everything hot — paying for disk nobody reads

Three boxes. Each one has a receipt — a number from the arithmetic that says "this, and not the fancier thing." When the CTO asks "why not Kafka?", the answer isn't an opinion: "1,333 writes/sec doesn't need a distributed log; here's the math."

Break it: what the numbers don't cover

The arithmetic justifies the boxes for the numbers you were given. The design review is where you attack your own assumptions:

The lunch rush lies. Peak reads at 2x is a guess. What if a viral delay — a snowstorm, a holiday — makes 10x the customers refresh 50x? The cache TTL is the shock absorber: at 30s TTL, even a 10x read spike costs the database nothing extra. The failure mode isn't reads; it's cache invalidation stampedes if every key expires at once — jitter the TTLs.

The write path has no shock absorber. 1,333 writes/sec is the average; drivers reconnecting after a tunnel outage arrive in bursts. Postgres absorbs bursts with connection pooling and batched inserts — but the API needs Post 3's polite-client treatment in reverse: a bounded ingest queue with backpressure, so a burst slows drivers' apps instead of killing the database.

The "boring" database is a single point of failure. "Support tickets, not lawsuits" justified skipping multi-region — but it doesn't justify skipping backups. One managed Postgres with point-in-time recovery: the boring answer to durability, and the memo says so explicitly.

Myth: "System design is about knowing which technology to pick."

Reality: Technology selection is the last 10%. The first 90% is turning adjectives into numbers — then the technology is usually obvious, and "boring" is usually the right answer. Senior engineers aren't distinguished by knowing more systems; they're distinguished by asking the questions that make the system obvious.

Productionize: the design memo

You don't hand the CTO a diagram. You hand them the numbers, the boxes, and the tradeoffs — one page:

DESIGN MEMO: same-day delivery tracking
Load:      ~1,333 writes/s, ~56 reads/s (111 peak)
Storage:   ~161 GB live (7d), ~5 GB history (1y)
Freshness: 60s  →  Redis 30s TTL on read path
Avail:     single region + managed Postgres w/ PITR ("tickets, not lawsuits")
NOT built: sharding, multi-region, real-time push, Kafka
Triggers:  re-shard review at 10k writes/s sustained;
           multi-region when downtime cost exceeds 2x infra cost

Note the last line: every "not built" carries a trigger — the number at which you'd revisit the decision. That's what makes the memo a living document instead of a snapshot. "We didn't shard" is a guess; "we'll revisit sharding at 10k writes/sec sustained" is engineering.

Explain it to the customer

You, Wednesday: "Twenty minutes of questions, ten of arithmetic. Your load is 1,333 location writes a second and about a hundred reads — a single database handles that with headroom, so we're not sharding anything. Customers get tracking data up to a minute old, which is what you asked for, so we cache reads for 30 seconds. Trip history goes to cheap storage nightly. Three boxes, each one justified by a number — and the memo lists the exact numbers that would make us redesign."

CTO: "And when we 10x?"

You: "The memo says: sharding review at 10k writes a second sustained. You'll know it's coming before it arrives."

Useful later

  • Load testing (k6, Locust) — verify the arithmetic against reality before launch
  • Cache stampede protection — jittered TTLs, request coalescing
  • Backpressure on ingest — bounded queues, Post 3's rate limiting in reverse
  • Multi-region — when the downtime math changes, not before

Don't memorize this

  • Postgres max-connections tuning — remember pool and batch the writes, look up the flags
  • Redis eviction policies — remember TTL per key, jittered, look up the policy names
  • Exact bytes-per-row — remember round aggressively, check the order of magnitude

Post 3's principle was be a polite client. This post's: do the arithmetic before you draw the boxes.

Field check

  1. The CTO says "actually, make it real-time — locations must be under 5 seconds old." Which box changes, which numbers change, and what gets more expensive?
  2. Driver count 10x's to 400,000. Recompute writes/sec. Does the single Postgres survive? What breaks first?
  3. A teammate proposes Cassandra "for scale" on day one. Using this post's numbers, write the two-sentence rebuttal.
  4. The memo's sharding trigger is "10k writes/s sustained." Why is a trigger better than just sharding now "to be safe"?
What good answers look like

1. The Redis box changes: 30s TTL becomes ~2–3s, or polling becomes server-push (SSE/websockets). Freshness math changes — cache hit ratio collapses, so reads hit the database at ~full rate. What gets expensive: database read capacity and connection churn; the "boring" design stops being boring. 2. 400,000 / 30 ≈ 13,333 writes/sec — past the 10k trigger. Single Postgres doesn't survive comfortably: WAL throughput, connection count, and index maintenance on the hot locations table break first. That's exactly when the sharding review fires. 3. "Our measured load is 1,333 writes/sec and 161 GB live — one Postgres handles 10x that. Cassandra buys us operational complexity we'd pay for every deploy; the numbers don't require it." 4. Sharding now pays complexity tax on every query, migration, and hire from day one, for a load you don't have. A trigger converts the decision into a monitored threshold — you get simplicity now and a planned response later, instead of complexity now and regrets throughout.