Prerequisites: Parts of this bank assume you've worked through the Path to FDE curriculum's spine projects — the Milestone 2 intake pipeline, the Stage 3 CityOps API integrations, the Stage 4 RAG system, and the Stage 5 production deployment. Each group below lists the lessons to study first. All candidates, companies, and interviewers named here are illustrative composites drawn from common patterns, not real people or companies.

System-design interviews for FDE roles are not architecture-beauty contests. They are simulations of the job: someone hands you a messy client, an unreliable legacy system, and a deadline — and asks how you'd build something that survives contact with reality.

A composite hiring manager at a composite logistics firm runs the whiteboard round. "Walk me through how you'd design an ingestion pipeline for our shipment manifests," she says. The candidate draws three perfect boxes — S3, Lambda, Postgres — and stops. Silence.

"Great start," she says. "Now: the carrier sends us CSVs with a new column every quarter, their SFTP server goes down for days at a time, and the warehouse team doesn't trust your numbers because last quarter a dashboard double-counted 40,000 containers. Redesign it."

That second question — not the first — is the FDE interview. The boxes were never the test. The failure modes are the test. Every question in this bank works the same way: a reasonable-looking design prompt, and behind it, the messy reality a Forward Deployed Engineer has to absorb.

The FDE system-design twist

In a backend-engineering interview, "design a URL shortener" rewards clean components and big-O thinking. In an FDE interview, you're designing for a specific client's reality, and the interviewer keeps moving the cheese: the schema changes, the third-party API has a 10-request-per-minute quota, the data owner won't give you prod credentials, the approval workflow needs a human signature before anything executes.

That gives you the single biggest scoring lever in this bank:

The principle

Design for the client's reality, not the whiteboard's.

A design that names three failure modes and two customer-facing tradeoffs beats a pristine architecture that assumes the inputs behave. Say out loud: "Here's what breaks, here's what I'd measure, here's what the customer gives up if we choose this path." That's the sound of an FDE thinking.

How to use this bank

There are 32 questions in four groups: data pipelines, API integration, AI systems, and deployment & operations. Each group opens with the published lessons to study first, then numbered questions with a what good looks like hint — one or two sentences describing what a strong answer demonstrates, so you can practice and self-score.

A practice loop that works for these:

  1. Answer out loud, whiteboard-style, for 5 minutes — don't just read the hint. The muscle being built is structured, spoken reasoning under mild pressure.
  2. Then read the hint and grade yourself on three axes: did you name failure modes? did you name tradeoffs? did you say what you'd measure?
  3. Re-answer once, incorporating what you missed. The second pass is where the learning happens.
  4. End by explaining your answer to a non-technical human — a partner, a friend, even your own notes in plain language. FDEs sell designs to operators and executives, not just to engineers.

What this bank will not do

  • No trick questions, no trivia ("name five HTTP methods") — every question is a design conversation with a messy-real-world twist.
  • No claims that a single answer is the correct design. System design has tradeoffs, not solutions; the hints describe what a strong answer demonstrates, not a transcript to memorize.
  • No invented statistics ("87% of FDE interviews ask about Kafka"). Everything below is grounded in the actual curriculum you can study on this site.
flowchart TD A[The design prompt] --> B[Name the failure modes] B --> C[Name the customer-facing tradeoffs] C --> D[Say what you would measure] D --> E[A design the client can actually operate]

Group 1 — Data pipelines

Q1.1 — Design an ingestion pipeline for weekly CSV drops from a third-party carrier

A client receives CSV manifests from a carrier every Monday at 6 AM via SFTP. Sketch the pipeline from drop to queryable data. Then: the carrier's server is occasionally unreachable for days, and their files arrive with different encodings across weeks. What changes?

What good looks like: a pipeline that separates landing the file from parsing it (raw zone vs. validated zone), retries and checkpointing on the fetch step, and encoding sniffing or normalization before validation — plus an explicit answer for what happens to Monday's reporting when Friday's file finally arrives three days late.

Q1.2 — The schema drifts

Your pipeline has been running for six months when the carrier adds a new column, renames two existing ones, and occasionally ships rows with extra fields. Your downstream dashboards break silently — no errors, just wrong numbers. How do you design the pipeline so schema changes are detected, not silently absorbed?

What good looks like: strict validation at the boundary (a Pydantic-style schema that rejects unknown or renamed fields by default), a schema-version log per batch, alerts on new or missing columns before data reaches dashboards, and a quarantine path for non-conforming batches instead of best-effort guessing.

Q1.3 — Reconciliation: proving the numbers are right

The client's warehouse team doesn't trust your pipeline's totals because an old script once double-counted 40,000 containers. Design a reconciliation mechanism that lets an operator verify, on any given day, that what entered the pipeline equals what the dashboards show. What do you store, and where?

What good looks like: row counts and checksums (or per-batch digests) captured at each pipeline stage, a reconciliation report comparing landed vs. validated vs. loaded counts, and a clear story for where the records diverged when the numbers disagree — the mechanism matters more than the exact hash chosen.

Q1.4 — Dead-letter queues: what breaks goes here

Design the dead-letter path for your intake pipeline: rows that fail validation, files that fail to parse, batches that can't reconcile. What information does each dead-letter record carry, how does an operator investigate and reprocess it, and — critically — what stops a single malformed row from blocking the entire Monday batch?

What good looks like: poison records isolated at row granularity (not batch granularity), each carrying the original payload, the validation errors, and the batch lineage; an operator workflow to fix-and-replay; and a rule for when a dead-letter record gets escalated to the data owner instead of retried forever.

Q1.5 — Scheduling: cron, queues, or event-driven?

You have three ingestion patterns in one engagement: (a) a weekly carrier CSV drop, (b) a carrier API that must be polled every 15 minutes, (c) manifests that arrive at unpredictable times via an upload portal. Compare how you'd schedule each, and defend the choice. Then the twist: the client asks "what happens if the scheduler itself fails silently for a weekend?"

What good looks like: different triggers for different patterns (cron for the weekly drop, a polling loop or lightweight scheduler for the 15-minute API, event-driven for unpredictable arrivals), each with idempotent processing so reruns are safe — plus heartbeat monitoring or a watchdog that pages someone when a scheduled run never happens.

Q1.6 — Late-arriving and out-of-order data

Your pipeline computes "shipments per week" for the client's KPI dashboard. Carrier files routinely arrive late — a manifest for week 12 lands during week 14. Design the pipeline and the warehouse tables so that late data corrects history without rewriting the past in a way nobody can audit. How do you explain this to the warehouse manager?

What good looks like: separating event time from processing time, an explicit restatement or correction policy (append corrections rather than silently updating rows), and a human-readable explanation: "your week-12 number changed because late manifests arrived; here's the audit trail of exactly which rows changed and why."

Q1.7 — Exactly-once, at-least-once, and the lie in between

An interviewer says: "Design this pipeline for exactly-once processing." The pipeline writes shipment records into Postgres and emits events to a downstream queue. Where can duplicates creep in, and what mechanisms get you closest to exactly-once — or convince you that at-least-once plus idempotency is the honest answer?

What good looks like: naming the duplicate sources (retried fetches, replayed batches, crash between write and checkpoint), then either idempotent consumers with stable record keys, or transactional outbox-style coordination — and the judgment to say "true exactly-once across two systems isn't achievable here, so here's how we make duplicates harmless," rather than pretending otherwise.

Q1.8 — The backfill: rebuilding six months of history

The client signed off on a new, stricter schema. They now want the last six months of raw carrier files reprocessed through the new pipeline. Design the backfill: how do you run it without disrupting the live weekly pipeline, how do you verify the backfilled data, and what do you do when the backfill produces different totals than the numbers the client already reported to their board?

What good looks like: backfills running against the raw archive (proof the raw zone paid off), isolated from the live path with its own outputs for comparison, reconciliation between old and new totals, and a communication plan for discrepancies — the FDE move is treating the differing board numbers as a customer conversation, not just a data bug.

Group 2 — API integration

Study these lessons first

This group is grounded in the real Stage 3 lessons — the messy reality of integrating with systems you don't control:

Q2.1 — Integrating with a legacy system that has no API

The client's warehouse management system is a twenty-year-old on-prem application: no API, nightly CSV exports to a shared folder, and a "database" the vendor warns you not to query directly. Design the integration that feeds this data into your modern pipeline. What's the strangest constraint you plan for?

What good looks like: treating the CSV export (or a read replica / CDC-style capture) as the integration seam instead of fighting the vendor; checksums and row counts to detect truncated exports; planning for the export format to change without notice; and an honest list of what you cannot do (real-time reads, writes back) stated to the client up front.

Q2.2 — Pagination at scale

You need to pull 2 million work-order records from a vendor API that paginates 100 records per page with cursor-based pagination. The interviewer adds: the dataset mutates while you paginate, and the vendor's cursors expire after 10 minutes. Design the extraction, and explain what correctness means here.

What good looks like: cursor handling with resume/checkpoint state, dealing with mid-extraction mutations (snapshot semantics vs. accepting drift, dedupe by stable record ID), parallel page fetching only within the vendor's documented limits, and defining correctness as "a consistent, auditable extraction" rather than "a perfect point-in-time snapshot" when the API can't offer one.

Q2.3 — Rate limits: 10 requests per minute

A carrier's tracking API allows 10 requests per minute per API key, and you need fresh status for 5,000 active shipments. Design the polling strategy. Then the twist: the client asks for "real-time" tracking on their customer portal. What do you tell them?

What good looks like: request budgeting (prioritize shipments by status volatility), caching with explicit staleness, backoff and queueing when the limit is hit, and a frank customer conversation — translating 10 req/min into "your portal refreshes every N minutes" and proposing webhooks or event-driven alternatives as the real fix for "real-time."

Q2.4 — Retries that don't make things worse

Your integration POSTs delivery confirmations to a client's legacy endpoint that times out intermittently and occasionally processes a request twice. Design the retry strategy. What's the difference between retrying this endpoint and retrying an idempotent GET — and how do you make the POST safe to retry?

What good looks like: exponential backoff with jitter, a bounded retry budget, idempotency keys minted by your application so the receiver can dedupe, and a clear line: retries on the GET are harmless, retries on the POST need the idempotency key plus verification (read-after-write) before assuming success — with a dead-letter path when retries are exhausted.

Q2.5 — Designing a webhook receiver

The carrier now offers webhooks for shipment status changes. Design your receiver: verification, ordering, duplicates, and downtime. Then the twist: during your demo, a burst of 10,000 webhooks arrives in one minute and your handler falls over. What broke, and how do you fix it without losing events?

What good looks like: signature verification on every event, idempotent processing by event ID, accepting-then-queueing (return 200 fast, process asynchronously) instead of doing work in the request handler, out-of-order handling by event timestamp, and a replay/reconciliation mechanism for the downtime window — the fix for the burst is architectural (queue + workers), not a bigger server.

Q2.6 — Auth in the wild

You're integrating three client systems: one uses OAuth 2.0 client credentials, one uses long-lived API keys rotated quarterly by a human, and one uses mutual TLS with certificates that expire yearly. Design the credential management for your deployment: where secrets live, how rotation happens, and what breaks at 2 AM when a certificate expires. Then: the client's security team asks how you prove no credential ever appeared in a log.

What good looks like: secrets in a proper secrets store (never env files in the repo, never logs), rotation procedures per mechanism with calendar or automated triggers, monitoring for expiry with alerts weeks ahead, and log redaction plus a reviewable policy for credential handling — the 2 AM answer is "we get paged before it expires, and the rotation runbook is tested."

Q2.7 — Designing under constraints: the 48-hour integration

The client's biggest customer demands EDI-style shipment data in their portal by end of week. You have 48 hours, no access to the client's production network (VPN approval takes two weeks), and the only documentation is a 200-page PDF from 2019. Sketch your plan. What do you deliberately leave out, and how do you say so?

What good looks like: ruthless scoping — a minimal viable integration (one shipment type, one direction of data flow, manual steps documented where automation isn't feasible), a local mock or recorded fixtures to develop against without prod access, explicit "not in scope" list communicated early, and a hardening plan for week two. The FDE skill being tested is saying no to scope, kindly and early.

Q2.8 — Debugging someone else's API

Your integration worked in testing but fails in the client's staging environment with a generic 400 and no useful error message. The vendor's support says "works on our end." Walk through your debugging method: what do you capture, what do you compare, and how do you prove where the fault lies without access to the vendor's servers?

What good looks like: capturing full request/response pairs (headers, body, timing), diffing the exact bytes between the working and failing environments, reproducing with a minimal request outside your application code, and building an evidence packet (timestamps, request IDs, correlation) that lets the vendor reproduce it — methodical elimination rather than guessing, and never blaming the vendor without evidence.

Group 3 — AI systems

Q3.1 — Design a RAG system for 50,000 internal operations documents

The client wants operators to ask natural-language questions over 50,000 internal SOPs, incident reports, and policy PDFs. Sketch the architecture: ingestion, chunking, retrieval, generation. Then the twist: half the PDFs are scanned images, and the ops team updates SOPs weekly. What breaks first?

What good looks like: an ingestion pipeline with OCR for scans, chunking choices tied to document structure (not one fixed size for everything), embeddings with a pinned model version, retrieval with exact filters for exact facts, and an update strategy (re-index deltas, versioned documents) — plus naming the first failure: stale SOPs served as current answers, fixed by document versioning and freshness metadata in retrieval.

Q3.2 — Retrieval quality before prompt cleverness

Your RAG system's answers are wrong in a specific way: fluent, confident, and citing documents that don't support the claims. The team wants to "improve the prompt." You suspect retrieval. How do you diagnose whether the problem is retrieval, generation, or the eval — and what do you change first?

What good looks like: instrumenting the pipeline to inspect retrieved chunks per query (separating citation presence from citation support), testing retrieval in isolation with known question-to-document pairs, and fixing retrieval quality (chunking, filters, ranking) before touching the prompt — the principle from the curriculum: retrieval quality comes before prompt cleverness.

Q3.3 — Evals: proving the AI got better

The client asks: "How do we know the new version of the assistant is actually better than the old one?" Design the evaluation approach: what dataset, what scoring, what process. Then the twist: the product team changed the system prompt AND the eval set in the same release. Why is that a problem, and how do you fix the process going forward?

What good looks like: a versioned eval set of representative real questions, a rubric defining correctness with a scorer that operationalizes it, and the discipline to version the system, the eval set, and the scoring logic independently — because if two change at once, you can't tell whether the product changed or the test did.

Q3.4 — Guardrails: the AI must not go rogue

Your assistant can look up shipment records and draft customer-facing status emails. The client worries it might disclose one customer's data to another, or send an email with a fabricated delivery date. Design the guardrail architecture: what sits around the model, where, and what happens when a guardrail fires?

What good looks like: controls around the model, not inside the prompt — tenant-scoped retrieval so the model never sees another customer's data, output validation against the retrieved evidence (citations checked against what was actually retrieved), PII handling on the way in and out, and a defined failure behavior (block, redact, or escalate) with audit logging of every guardrail decision.

Q3.5 — Human approval for agentic actions

The client wants the assistant to not just answer questions but take actions: reschedule a delivery, issue a credit, update a shipment record. Design the approval workflow. What's the difference between authorization ("may this caller act?") and approval ("did a human approve this exact action?") — and how do you implement both?

What good looks like: separating the two concepts — authorization checks the caller's permissions on the resource, approval requires a human to sign off on an immutable snapshot of the exact action (parameters frozen, shown verbatim) before execution; approval records stored where the agent cannot rewrite them; idempotent execution so a double-approved action can't double-fire; and a dry-run preview so the approver sees consequences, not just parameters.

Q3.6 — Cost and latency tradeoffs

Your RAG assistant costs roughly $0.08 per query and takes 6 seconds end-to-end. The client wants it under $0.02 and under 2 seconds for a 10x rollout to every operator. Walk through the levers you'd pull, what each one trades away, and how you'd decide — without guessing at numbers you haven't measured.

What good looks like: naming real levers (smaller/faster model for the draft with a larger model for verification, caching frequent queries, trimming retrieved context, streaming for perceived latency, cheaper embeddings, local inference where privacy allows) with each lever's cost in quality or complexity stated honestly — and refusing to promise the target until you've measured the quality impact on the client's own eval set.

Q3.7 — Structured output you can actually trust

The assistant must extract structured fields from incident reports (incident ID, severity, affected systems, timeline) into the client's ticketing system. Free-text extraction hallucinates fields. Design the extraction so the output is machine-consumable: schema, validation, and what happens when the model can't fill a required field.

What good looks like: schema-constrained generation (the model must conform to a declared schema), server-side validation of every output (Pydantic-style — schema proves shape, not truth), explicit handling of missing fields (null with a reason, not invented values), and human review routing for low-confidence extractions — validation and truth are separate problems, and the design treats them that way.

Q3.8 — The model changed under you

Your assistant's answer quality drops on a Tuesday morning. Nothing in your code changed — but the API provider shipped a new default model version overnight. Design the deployment identity and change-management practices that would have caught this before customers noticed. What do you pin, what do you monitor, and what's your rollback?

What good looks like: pinning the exact model version (not "latest"), recording a full deployment identity (model artifact/version, serving config, prompt version, eval set version), monitoring answer quality on a live eval slice (not just uptime), canary or staged rollout of model changes, and a rollback that restores the previous pinned configuration in minutes — the lesson: "pin the exact tag" is the beginning, not the end, of deployment identity.

Group 4 — Deployment & operations

Q4.1 — Docker Compose for the whole CityOps stack

Design the local-and-staging deployment for a stack with a FastAPI service, a Postgres database, a Redis queue, a RAG worker, and an operator console frontend. One docker compose up should bring it all to life. Then the twist: the client's security team says no container may run as root, and secrets must not live in the compose file. What changes?

What good looks like: a compose file with explicit service dependencies and healthcheck-gated startup ordering, named volumes for data, non-root users in each image, secrets injected from a secrets store or env files excluded from version control — and the judgment to say compose is for dev/staging while production gets the hardened orchestrator, not a bigger compose file.

Q4.2 — Design the CI/CD pipeline

Sketch the CI/CD pipeline for the CityOps services: what runs on every pull request, what runs on merge to main, and what gates the deploy to staging and production. Then the twist: a bad migration shipped to staging last month and took the staging database down for a day. What does the pipeline do differently now?

What good looks like: fast feedback on PRs (lint, unit tests), heavier checks on merge (integration tests, image builds, migration dry-runs), deployment gates with manual approval for production, and — after the incident — migration safety checks (backwards-compatible migrations, backup-before-migrate, tested rollback) built into the pipeline, not left to memory.

Q4.3 — Monitoring: know before the customer does

Design the monitoring for the production CityOps deployment: health checks, logs, metrics, and alerts. The interviewer adds a constraint: the on-call engineer is a client employee who has never operated this system, and the runbook must fit on one page. What do you monitor, what pages them, and what deliberately does NOT page them?

What good looks like: liveness vs. readiness checks separated (the server responding is not the same as the model being servable), structured logs with correlation IDs, a small set of paging alerts on user-facing symptoms (not on every internal metric), dashboards for the rest, and a one-page runbook covering the three most likely failures — alert fatigue is treated as a design failure, not an ops problem.

Q4.4 — Secrets, RBAC, and tenant isolation in production

The CityOps deployment serves three client business units whose data must never mix. Design the secrets management, role-based access, and tenant isolation: where credentials live, how access is granted and revoked, and how you prove to an auditor that unit A's operator cannot see unit B's shipments. What happens when an employee leaves?

What good looks like: secrets in a managed store with rotation and audit logging, RBAC with least-privilege roles per business unit, tenant isolation enforced at the data layer (tenant ID on every query, not just the UI), offboarding as a tested procedure (key revocation, session invalidation), and evidence an auditor can actually inspect — the design assumes the audit will happen.

Q4.5 — Choose a release strategy

You need to ship a new version of the RAG answer service. Compare blue-green, canary, and feature flags for this specific service: which do you choose, and why? Then the twist: the new version answers differently (not wrongly, differently) and the client's ops team notices within an hour. How does your strategy make that safe?

What good looks like: a real comparison tied to the service's properties (canary for gradual exposure with quality metrics, blue-green for fast rollback, flags for decoupling deploy from release), choosing one with reasons rather than reciting definitions — and the twist answered by the strategy's observability: traffic shifting plus answer-quality monitoring on a live eval slice, with instant rollback when behavior drifts.

Q4.6 — Backups and disaster recovery

Design the backup and disaster recovery plan for CityOps: the Postgres database, the vector index, uploaded documents, and configuration. Define RPO and RTO in terms the client's operations director understands. Then the twist: "We've been backing up nightly for a year." What's the one question that determines whether those backups are worth anything?

What good looks like: per-component backup strategy (database dumps plus point-in-time recovery, vector index rebuildable from source documents, config in version control), RPO/RTO stated as business impact ("we can lose at most N hours of data; we're back in M hours"), and the twist answered directly: "when did you last restore from backup?" — an untested backup is a hope, proven by a scheduled restore drill, not by the backup job's green checkmark.

Q4.7 — Deploying behind enterprise proxies and firewalls

The client's production environment has no direct internet access: all traffic goes through an allowlisted forward proxy, inbound connections are forbidden, and your container registry is outside their network. Design the deployment: how do images, model weights, and Python packages get in? What breaks the first time a library tries to phone home?

What good looks like: an internal registry mirror or vendored artifacts, proxy configuration for every layer (container runtime, package managers, the application), offline-capable images with all dependencies baked in, and a pre-flight checklist that catches "phone home" behavior (telemetry, license checks, model downloads) before it fails silently in production — the design treats the network boundary as a first-class constraint, not an afterthought.

Q4.8 — The 2 AM incident

It's 2 AM. The client's status page shows the operator console is down, the on-call client engineer is paging you, and the last deploy was six hours ago. Walk through your first 30 minutes: what do you check, in what order, and what do you communicate to the client — and what do you deliberately NOT do?

What good looks like: a calm, ordered triage (is it the app, the platform, or the network? check health endpoints and recent changes before touching anything), communicating early with what you know and what you're checking next, resisting the urge to redeploy or restart everything blindly, and capturing a timeline for the post-incident review — the FDE move is managing the customer's confidence while you manage the system.

Putting it together: the full-loop answer pattern

If you take one answering structure into every system-design round, make it this four-beat loop — it maps directly onto the principle at the top of this post:

BeatWhat you sayWhat it proves
Scope"Here's what I'm building, for whom, and what's explicitly out of scope."You clarify before you architect — the FDE discovery habit.
Happy path"The straightforward design looks like this…"You can design clean systems, not just critique them.
Failure modes"Here's what breaks: schema drift, expired credentials, late data…"You design for the client's reality, not the whiteboard's.
Tradeoffs & measurement"We trade X for Y; here's what I'd measure to know it's working."You think in customer-facing consequences, not just components.

A note on "correct" answers

None of these 32 questions has a single correct design. Two strong candidates can choose different architectures — one picks canary, the other picks blue-green — and both can pass, because what interviewers score is the reasoning: did you surface the constraints, name the failure modes, weigh the tradeoffs, and say how you'd know it's working? Practice the reasoning, not the answers.

Field check — grade yourself before the interview

  1. Pick one question from each group. Answer each out loud for five minutes, then check your answer against the hint: did you name at least two failure modes and one customer-facing tradeoff?
  2. Explain your Q3.5 (approvals) answer to a non-technical friend in under two minutes. If they can't repeat back the difference between authorization and approval, your explanation needs work — and so does your interview answer.
  3. Write down the three monitoring alerts from your Q4.3 answer that would page the on-call engineer. For each, write the runbook step. If any alert has no runbook step, it's noise, not monitoring.
How to know you're ready

You're ready when you can take any question in this bank, spend the first minute scoping it ("who's the customer, what's out of scope?"), the next three on the happy path plus failure modes, and close with tradeoffs and measurement — without notes, without rambling, and without inventing numbers you can't defend. The whiteboard is just a prop; the thinking is the product.

Next in the bank

Part C: The Behavioral Bank — telling FDE stories that prove you ship in messy reality: the STAR-format answers, the failure stories, and the customer conversations that turn "tell me about a time" into an offer.