Prerequisite: Stages 1–5 of this curriculum — you've built the CityOps intake pipeline, the API, the RAG pilot, and shipped it to production. This post assumes you know how the work is done and asks a different question: how do you think when a human is angry, confused, or at risk in front of you?

The scene: Forty minutes into your final-round interview, the hiring manager stops taking notes, leans forward, and says:

"Forget the whiteboard. Here's what happened last month. Our biggest customer's security team blocked our rollout forty-eight hours before the demo their CEO was attending. The engineer on site froze. What do you do?"

There is no test case to pass. There is no correct answer in the back of the book. There is only you, thinking out loud, while someone watches how you think.

FDE interviews don't test whether you can sort an array. They test whether a customer would trust you at 2 a.m. with their production system and their career on the line.

The problem: interviews can't watch you work, so they watch you think

A hiring manager for a Forward Deployed Engineering role has a specific, expensive problem: the person they hire will sit across a table from a paying customer within their first weeks, and that customer will bring problems that are ambiguous, emotional, and time-pressured — the exact opposite of a LeetCode prompt. No coding test can tell the hiring manager whether a candidate will clarify before building, communicate before things break, or hold a line when pressured to ship something unsafe.

So instead of testing your knowledge, scenario interviews test your judgment under pressure. The interviewer invents a customer situation — often a composite of real incidents from their own deployments — and asks you to think through it live. What they are actually grading is a small set of behaviors:

  • Do you clarify before solving? Weak candidates hear a scenario and immediately propose a fix. Strong candidates spend the first minutes making sure they understand what the real problem is, who is affected, and what constraints exist.
  • Do you see the humans, not just the systems? The angry stakeholder is not a bug to be fixed; they're a person whose deadline, budget, or reputation is on the line. Candidates who acknowledge that explicitly tend to keep trust. Candidates who jump to "here's the technical fix" tend to lose it.
  • Do you name tradeoffs instead of pretending they don't exist? Every real customer decision costs something. Interviewers watch for whether you say "we could do X, but it would cost us Y" — and whether you involve the customer in choosing.
  • Do you communicate upward and outward? The engineer who quietly fixes things without telling anyone creates surprises. FDE interviewers want to hear you say who you'd tell, what you'd tell them, and when.
  • Do you hold an ethical line? Some scenarios are designed so that the easy answer is the wrong one — shipping something unsafe, hiding a problem, massaging numbers. The interview is checking whether you'll notice the trap.

One more thing to know before the prompts: every scenario below is an illustrative composite — fictionalized patterns drawn from common FDE interview practice, not accounts of any real customer, company, or interviewer. The curriculum's CityOps characters (Maria, Tom, Lisa, Dev) appear as composites too, used here because you've met them across Stages 1–5 and can pattern-match their concerns quickly.

The minimum concept: clarify → propose → tradeoffs → communicate

You don't need a twenty-step framework. You need one that survives adrenaline. Every scenario in this post can be answered with four moves, in this order:

flowchart TD A["1. CLARIFY — what is actually broken,
for whom, how badly?"] --> B["2. PROPOSE — a concrete next step,
not a perfect solution"] B --> C["3. TRADEOFFS — what does the fix cost,
and who chooses?"] C --> D["4. COMMUNICATE — who do you tell,
what, and by when?"]

The principle

In scenario interviews, the process is the answer. Nobody remembers whether your plan for the broken pipeline was the optimal one. They remember whether you asked what "broken" meant before designing it, whether you named what it would cost, and whether you told the right people before it got worse. Walk the four steps out loud and you will outperform most candidates, even if your specific proposal isn't the one the interviewer would have picked.

Here's why each step exists, and what skipping it sounds like:

StepWhat it looks likeWhat skipping it sounds like
Clarify"Before I do anything — what does 'worse' mean here? Slower, more errors, or both? When did it start?""I'd rewrite the form in React and add caching." (You don't know what's wrong yet.)
Propose"Here's what I'd do in the next hour, the next day, and the next week.""We should probably look into it." (No plan, no ownership.)
Tradeoffs"The fast fix gets you running today but keeps the flaky dependency. The real fix takes two weeks and needs your sign-off.""I'll just fix it." (Hides the cost; the customer discovers it later.)
Communicate"I'd tell the ops lead today, give you a status by 5 p.m., and write up what happened so the next team doesn't repeat it."Silence. (The most common way good engineers fail scenario interviews.)

This same arc — clarify the ask, build the minimum, break it, harden it, explain it to a human — has run through every lesson in this curriculum. Scenario interviews are simply that arc played at conversation speed, with a person watching your face while you do it.

How a weak answer sounds (and how to fix it)

Take a simple prompt: "The customer's ops lead says your new intake tool is slower than the spreadsheet it replaced. Walk me through what you do."

A weak answer jumps straight to step 2 and skips everything else:

Weak: "I'd profile the tool, find the slow query, add an index, and redeploy. Maybe cache the dropdowns."

Technically, that's not wrong — it's just incomplete in every way the interviewer is grading. Nothing about what "slower" means. Nothing about whether the claim is even true. Nothing about the ops lead's actual workflow, the cost of the fix, or who gets told what. The candidate demonstrated debugging skill and zero customer judgment.

Now the same answer, walked through the four steps out loud:

Strong (same scenario)

"First I'd clarify what 'slower' means — is the page loading slowly, or does the whole task take longer than the spreadsheet? Because those are different problems. If it's the whole task, I might have automated the wrong part of their workflow, which is a design problem, not a performance problem. Then I'd shadow them for one real task to see where the time goes. My proposal would depend on that — maybe it's an index, maybe it's a form redesign, maybe it's both. Either way, I'd be honest with the ops lead about what I found, including if the tool is genuinely worse at something the spreadsheet did well, and agree on what 'fixed' looks like before I change anything."

The difference is not technical depth. It's that the strong answer treats the customer's complaint as a hypothesis to investigate rather than a ticket to close. Every prompt below is graded on exactly that instinct.

The 20 prompts

How to use these: for each scenario, answer out loud for three to five minutes as if you're in the room. Then open the grading notes and check yourself against all three bullets. The notes are written from the interviewer's side of the table — what they're listening for, what sinks candidates, and the follow-up they'd use to probe deeper.

Group A — The angry stakeholder

These scenarios put you across the table from someone who is upset, and part of the test is whether you can stay regulated and curious instead of defensive. The grading pattern across this group is consistent: acknowledge the human cost first, investigate second, promise third. Candidates who reverse that order — promising a fix before understanding the damage — read as people who will say anything to end an uncomfortable conversation.

1. The missed deadline

Maria, an operations director, opens the call with: "You committed to the dashboard by Friday. It is Tuesday, and I just told my VP it would be ready. What do you want me to tell her?" You know the delay is real — an upstream data feed changed format over the weekend and half your team spent Monday fixing it.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: owning the miss without excuses ("we missed it, here's why" — the upstream feed change is context, not a defense), separating what Maria needs right now (a message she can take to her VP) from the engineering work (the feed fix), and giving her a concrete, checkable next update time. The best candidates ask what her VP actually needs — the full dashboard, or one number for a meeting?

The common failure mode: defending the timeline ("we couldn't have known about the feed change") or promising a new date on the spot without knowing whether it's achievable. Both tell the interviewer this candidate will keep a customer calm for one meeting and create a second angry meeting next week.

The follow-up the interviewer should ask: "Maria says that's not good enough — her VP needs something by end of day. What do you offer, and what do you refuse to offer?"

2. "Your tool made things worse"

Dev, the ops lead at a CityOps depot, tells you in front of his team: "This new intake form takes my people longer than the spreadsheet did. We liked the spreadsheet." His team is nodding. You spent six weeks building the form, and the spreadsheet was genuinely error-prone — but right now, in this room, that doesn't matter.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: not arguing with the room. The candidate validates the complaint ("if it's slower, we built the wrong thing — help me see where"), asks to watch one real task end-to-end, and treats "we liked the spreadsheet" as data about the workflow, not ignorance. Bonus: acknowledging that adoption is part of the deliverable — a tool nobody uses is a failed tool, no matter how correct its validation logic is.

The common failure mode: explaining why the spreadsheet was bad ("but it had no validation!"). Correcting the customer in front of their team is the fastest way to lose the room, and interviewers watch for the instinct to do it.

The follow-up the interviewer should ask: "You shadow Dev and find the form really is slower for one common task, but catches errors the spreadsheet missed. How do you decide what to change?"

3. Security blocks your launch, 48 hours before the demo

Lisa from the customer's security team emails: your service can't go live because its data flow was never reviewed — the review queue is three weeks. The demo for the customer's executive team is in 48 hours, and your sponsor has already sent the invitations. Lisa is not being difficult; she is doing her job.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: treating Lisa as an ally, not an obstacle. The candidate separates the demo (a presentation, which can often run on anonymized or synthetic data with no real data flow) from the launch (the real thing, which waits for the review). They ask Lisa what a review-safe demo looks like, escalate the review request through the proper channel, and — the tell — they note the process failure: the review should have been requested weeks ago, and they'd change the launch checklist so this can't happen again.

The common failure mode: pressuring Lisa to "make an exception" or going around her to the sponsor. Interviewers read this as a candidate who will trade the customer's compliance posture for their own demo — exactly the instinct the enterprise security review exists to filter out.

The follow-up the interviewer should ask: "Your sponsor says the demo must use real customer data or it's meaningless. What's your answer?"

4. The escalation email

You arrive Monday morning to an email from Tom, a department head, cc'd to your manager and his VP: "The vendor team doesn't understand our process. We've explained the requirements three times. I need someone competent on this account." You believe the requirements genuinely were contradictory across his team — but that's not what the email says.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: not replying-all in anger, and not being defensive in the reply. The candidate acknowledges Tom's frustration, takes the contradiction problem seriously (it's real and it's theirs to solve), and proposes a reset: a short working session to reconcile the conflicting requirements with all parties in the room, with a written summary afterward so "explained three times" can't happen a fourth time. Strong candidates also tell their own manager proactively before the manager asks.

The common failure mode: litigating the email ("we did understand, his team contradicted each other") — either to Tom or to the manager. Interviewers are checking whether the candidate can absorb an unfair public criticism and still move the work forward.

The follow-up the interviewer should ask: "In the reset meeting, two of Tom's people give you directly contradictory requirements again, in front of him. What do you do in the room?"

5. Your fix broke their workaround

You shipped an automation that replaced a manual reconciliation step the customer's team ran every morning. It worked in testing. But it turns out the team had quietly built their own checks into that manual step — checks nobody documented — and your automation skipped them. Two bad records went to a regulator-facing report before anyone noticed. The customer is furious, and technically, the bug is yours.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate owns it immediately and completely — no "well, they didn't document it." They lay out the response in order: contain (stop the automation, assess the blast radius of the two bad records, help draft the correction), then learn (the undocumented checks existed because the team didn't trust the old process — that's a discovery failure on the FDE's part), then prevent (shadow-mode runs and sign-off gates before automating human steps in the future). The regulator angle must be named explicitly, not minimized.

The common failure mode: blaming the customer for the undocumented workaround, or treating it as purely a technical bug ("I'll add the checks to the script"). Interviewers are grading ownership and whether the candidate updates their process, not just their code.

The follow-up the interviewer should ask: "The customer now wants every future automation change approved by their team in writing before it ships. Reasonable? How do you respond?"

Group B — The vague requirement

These scenarios test whether you can turn fog into a plan without either freezing or inventing certainty. The pattern the interviewer wants: make the ambiguity visible, propose a small concrete step that reduces it, and get agreement on what "done" means before building. This is the discovery and scoping muscle from the early curriculum — Discovery and Scoping & Success Metrics — compressed into a conversation.

6. "Make reporting better"

A stakeholder's entire brief is: "Our reporting is bad. Make it better." No examples of bad reports, no definition of better, no users named. The interviewer adds: "You have thirty minutes with this person next Tuesday. What do you do before, during, and after?"

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate refuses to start building from a four-word brief. Before: gather whatever exists (sample reports, who consumes them, what decisions they drive). During: ask for one concrete recent pain ("tell me about the last time a report failed you"), identify the actual users and their decisions, and push for a measurable definition of "better" (faster? fewer errors? one fewer manual step?). After: a one-page written summary of what was agreed — problem, users, success metric — sent back for confirmation. The written artifact is the tell; it proves the candidate knows alignment evaporates without it.

The common failure mode: jumping to a solution ("I'd build a dashboard with drill-downs") or asking questions with no structure — twenty random questions that leave the stakeholder more confused than before.

The follow-up the interviewer should ask: "The stakeholder can't give you a single concrete example of a bad report. Now what?"

7. "You're the expert"

Every question you ask, the stakeholder deflects: "You're the expert — you tell me." "Whatever you think is best." "I don't have time for the details, just make it work." You suspect they're disengaged, but you can't build on shrugs.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate recognizes this as a risk, not a compliment. "You're the expert" means every later decision can be disowned. Strong candidates change tactics: stop asking open questions and start bringing concrete options ("Option A gives you X by Friday, Option B gives you Y next month — which matters more?"), make the stakeholder choose in writing, and escalate the disengagement itself as a project risk to their own manager early. They also check the hypothesis: maybe the stakeholder is the wrong person and the real users are two levels down.

The common failure mode: accepting the flattery and building on assumptions, or complaining about the stakeholder. The interviewer is checking whether the candidate protects themselves and the project with written decisions instead of hoping it works out.

The follow-up the interviewer should ask: "You bring two options and the stakeholder picks one, then blames you when it doesn't work. What did you miss?"

8. Two stakeholders, opposite asks

Maria wants fewer alerts — her team is drowning in noise. Tom wants every event logged and surfaced — his team missed an incident last quarter and he's still answering for it. Both are your stakeholders. Both are right about their own problem. What do you do?

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate refuses to pick a side or average the two requests. They reframe: Maria has a signal problem (too much noise), Tom has a coverage problem (fear of missing something). Those can both be solved — severity tiers, routing rules, digest vs. real-time channels — but only if the candidate gets both stakeholders to agree on what each alert costs (attention) and what missing one costs (incidents). The tell: proposing a joint 30-minute session with both of them rather than shuttling between them.

The common failure mode: building a compromise nobody asked for (a settings page with fifty toggles) or escalating immediately ("let my manager decide"). Interviewers want to see the candidate hold the tension themselves.

The follow-up the interviewer should ask: "In the joint session, Tom says any alert Maria's team ignores is a liability for him. How do you get to an agreement?"

9. "We'll know it when we see it"

The customer can't define a success metric for the pilot: "We'll know it when we see it. Just build something and we'll react." You know from the curriculum that unevaluated work can't be shown to improve — but you also can't force a customer to write a rubric.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate meets the customer where they are while smuggling in measurement. They propose a tiny, time-boxed slice ("give me two weeks and ten real cases"), define the metric themselves as a starting draft ("here's how I'll measure whether this helped — tell me what's wrong with it"), and make the customer's "we'll know it" concrete by asking them to judge the ten cases blind. Strong candidates name the failure mode honestly: without a metric, every demo becomes an argument about taste.

The common failure mode: either lecturing the customer about evals or happily building with no success criteria. The interviewer is checking whether the candidate can introduce rigor without making the customer feel managed.

The follow-up the interviewer should ask: "The ten cases come back split — five look great, five are arguable, and the customer disagrees with your scoring on three. What now?"

10. The growing scope

Your sponsor adds one more table to the dashboard in every standup. It's always small — "just one more column," "just one more filter" — and each time, saying no feels petty. Three weeks in, the dashboard is late and the original goal is buried under additions nobody prioritized.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate spots the pattern early and addresses it as a process problem, not a series of individual requests. They propose a visible backlog with costs attached ("each addition moves the date by roughly X — here's the list, you rank them"), tie every addition back to the original agreed goal, and — the key move — have the scope conversation with the sponsor outside the standup, where "just one more column" can't hide behind the rhythm of the meeting. They also own their part: they let it happen three weeks without flagging it.

The common failure mode: saying yes to everything and then missing the date, or saying no bluntly in the standup and damaging the relationship. Interviewers watch for whether the candidate makes the tradeoff visible to the person who owns it.

The follow-up the interviewer should ask: "The sponsor says all ten additions are must-haves and the date can't move. What's your response?"

Group C — The launch at risk

Something is broken, the clock is running, and people are watching. These scenarios grade triage discipline: can you separate "stop the bleeding" from "fix the disease," communicate a realistic status instead of a comforting one, and learn the right lesson afterward? Several echo real CityOps interrupts from the curriculum — the broken pipeline, the pilot demo, the rate-limited API — because interviewers at FDE-style companies routinely probe exactly the incidents their teams have lived through.

11. The schema change that broke the pipeline overnight

Sunday night, an upstream vendor renamed three fields in the feed your CityOps intake pipeline depends on. Monday is go-live. Your monitoring caught it — the pipeline is failing loudly — but the vendor's support doesn't open until 9 a.m. and the launch meeting is at 10.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: a two-track response. Track one, immediate: can the pipeline be made resilient to the rename (a field-mapping layer, a fallback to the old names) before 10 a.m., and if not, what's the honest status for the launch meeting? Track two, structural: a renamed field shouldn't be able to kill a launch — the candidate proposes schema validation at ingestion with explicit contracts, so the next change fails safe instead of failing loud at the worst moment. The tell: the candidate tells the sponsor the real status before the 10 a.m. meeting, not during it.

The common failure mode: either hacking a silent fix with no validation ("just map the fields and ship it") or waiting for the vendor with no plan B. Interviewers are grading whether the candidate protects the launch without hiding the fragility.

The follow-up the interviewer should ask: "You get it working by 9:30, but you're not confident in the data. Do you launch?"

12. The demo crashes on the stakeholder's laptop

Your demo worked perfectly on your machine. On the stakeholder's laptop, in front of six people, it crashes on the second screen. The stakeholder smiles thinly and says, "Well. This is awkward." You have the room for another twenty minutes.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: composure and a pivot, in that order. The candidate doesn't debug live for twenty minutes while the room watches — they acknowledge it plainly ("that's not supposed to happen, and I owe you an explanation of why"), switch to a fallback (screenshots, a recorded walkthrough, the staging environment on their own machine), and use the remaining time to talk through what the demo would have shown. Afterward: a real root-cause note to the stakeholder, not silence. The best candidates add the prevention: demos run on the customer's hardware in rehearsal, or they don't run at all.

The common failure mode: live-debugging in front of the audience, blaming the laptop ("it worked on mine"), or pretending the crash didn't matter. All three tell the interviewer the candidate has never carried a demo that mattered.

The follow-up the interviewer should ask: "The stakeholder later asks: 'Can we trust this in production if it can't survive a demo?' How do you answer?"

13. The rate limit collapses mid-sprint

Halfway through the sprint, your integration starts returning 429s — the third-party API you depend on has a lower rate limit than anyone verified during design, and your usage pattern blows through it by noon every day. The feature can't work like this. The vendor's sales team says a higher tier "can be arranged" on an enterprise contract your customer hasn't signed.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate treats this as a design assumption that failed, not just an ops problem. Immediate: reduce demand (batching, caching, backoff with jitter — the standard rate-limit toolkit from the APIs stage) to buy time. Structural: the limit should have been verified during design, so the candidate proposes a spike to measure the real ceiling and reframes the feature around it. And the commercial angle must be named: the "higher tier" is a procurement decision for the customer, so the candidate lays out the options and costs honestly instead of quietly building something that depends on a contract that doesn't exist.

The common failure mode: either hammering the API harder (retries without backoff — the classic way to turn a limit into a ban) or assuming the contract will materialize. Interviewers watch for whether the candidate separates the technical fix from the commercial dependency.

The follow-up the interviewer should ask: "Caching would fix the volume, but the customer needs real-time data. Now what?"

14. The model was wrong in front of the customer

During the Ask CityOps pilot demo — the RAG system from Milestone 4 — the assistant confidently states a permit status that is wrong. The stakeholder notices. The room goes quiet. This is the failure mode the evals stage warned about: a confident wrong answer in front of the person whose trust you're trying to earn.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate doesn't minimize it ("it's just a demo") and doesn't over-apologize into a spiral. They name what happened accurately — the system produced an unsupported claim, which is the known failure mode of retrieval-backed generation — show how they'd diagnose it (was the right document retrieved? did the model ignore it? was the source itself stale?), and reframe the pilot's purpose: this is why you run a pilot with evals before production. Strong candidates propose the concrete guardrail honestly: citation-backed answers, confidence thresholds, and human review on high-stakes outputs. They also tell the stakeholder what changes before the next demo.

The common failure mode: blaming the model ("the LLM hallucinated, nothing we can do") or promising it won't happen again. Both dodge the real question, which is whether the candidate has a system for catching wrong answers, not just hope.

The follow-up the interviewer should ask: "The stakeholder asks whether they should trust any AI answer at all now. What's your honest answer?"

15. Dirty data, hard deadline

You're migrating 40,000 customer records before a cutover date that can't move. This week you discover the data is dirty — duplicates, missing required fields, conflicting addresses — and there's no clear owner for cleanup decisions. The customer says "just migrate it, we'll clean it later." You know "later" means never.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate refuses the false choice between "migrate garbage" and "miss the date." They propose triage: quarantine the bad records into a visible, countable bucket (with reasons attached to each), migrate the clean majority on schedule, and give the customer a concrete cleanup backlog with an owner and a deadline — making "later" a scheduled thing instead of a wish. The tell: the candidate quantifies the dirt ("about 8% of records fail validation, here's the breakdown") because unmeasured data problems get dismissed and measured ones get staffed.

The common failure mode: either migrating everything silently (the customer's "just do it" becomes your production incident) or blocking the whole migration on perfect data. Interviewers are grading whether the candidate can hold quality and schedule at the same time by making the tradeoff explicit.

The follow-up the interviewer should ask: "The customer won't assign anyone to own the quarantined records. Who owns them?"

Group D — Ethics & judgment

These are the scenarios where the easy answer is the wrong one. Interviewers use them to check whether you'll notice the trap — and whether you'll hold a line when holding it is uncomfortable. The grading pattern here is the simplest in the post: the candidate must name the ethical problem out loud, refuse the specific harmful action, and propose a legitimate alternative. Partial credit for noticing; full credit for acting.

16. "Just ship it — we'll fix the safety issue later"

Your customer wants to auto-dismiss fraud flags on new accounts to speed up onboarding. You've seen the data: the flags catch real fraud, and auto-dismissing them removes the only check before money moves. Your sponsor says, "Ship it now, we'll add the review back next quarter." Next quarter is doing a lot of work in that sentence.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate names the stakes plainly — this isn't a UX tradeoff, it's removing a fraud control, and "next quarter" is not a control. They refuse the specific ask (not the customer, not the relationship — the ask), and they come with alternatives: a risk-tiered flow where low-risk accounts move fast and high-risk ones still get reviewed, or a time-boxed experiment with explicit fraud-rate monitoring and a kill switch. The tell: the candidate puts the refusal and the reasoning in writing, because verbal "we'll fix it later" agreements are how accountability evaporates.

The common failure mode: shipping it with a vague concern noted somewhere, or framing the refusal as personal discomfort ("I'm not comfortable with that") rather than a professional judgment about the customer's risk. Interviewers want to hear the candidate protect the customer from the customer's own short-term pressure.

The follow-up the interviewer should ask: "The sponsor says you're being difficult and asks your manager to replace you on the account. What do you do?"

17. PII in the chat thread

A contractor drops a CSV into your team's chat thread "for a quick look." You open it and realize it contains social security numbers and home addresses — real customer PII — sitting in a chat log that half the company can read. The contractor says, "Oh, I didn't realize. Can you just delete the file?"

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate knows "just delete the file" doesn't undo exposure — chat systems keep history, caches, and backups. They treat it as an incident: contain (get the file removed through the proper channel, confirm who could have accessed it), assess (was the data real? whose? how widely was it visible?), and report (tell the customer's security contact and their own lead — this is the customer's data and the customer's incident process governs). They also fix the cause: the contractor needs a safe way to share data samples, like redacted fixtures or a governed sandbox. Deleting quietly and moving on is the failure the scenario is designed to catch.

The common failure mode: deleting the file and considering the matter closed, or panicking and blasting the discovery to a wider audience. Interviewers are checking for incident discipline: contain, assess, report through the right channel, fix the process.

The follow-up the interviewer should ask: "Your lead says reporting this will cause a huge fuss and suggests handling it quietly between you two. Now what?"

18. The feature that violates their own policy

The customer asks for a raw data export feature — every tenant's records, unfiltered, downloadable by any admin. You happen to know their own data policy prohibits cross-tenant visibility; you read it during onboarding. When you mention this, the requester says, "That policy is outdated. Just build it."

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate doesn't get to declare the customer's policy outdated — that's the customer's governance decision, not the vendor engineer's. They pause the build, route the request to whoever owns the policy (with the requester, not behind their back), and propose what they can build meanwhile: scoped exports per tenant, which solves the likely real need without violating the policy. The tell: the candidate distinguishes "the customer asked for it" from "the customer is authorized to ask for it" — a distinction the guardrails stage of this curriculum exists to teach.

The common failure mode: building it because the customer insisted, or refusing with a lecture about compliance. Interviewers want the middle path: don't build the violating thing, do route the decision to its rightful owner, and keep delivering value on the non-violating version.

The follow-up the interviewer should ask: "The policy owner updates the policy to allow the export, in writing. Do you build it now?"

19. "Just exclude the bad slices"

Your eval numbers for the pilot are mixed — overall decent, but two slices of the eval set fail badly. Your manager, under pressure from sales, suggests: "Just exclude those slices from the report. They're edge cases." You know those "edge cases" are the customer's highest-volume query types.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: a flat refusal to alter the measurement, stated calmly and professionally. The candidate explains why: the eval set and the scoring are the measuring stick — changing them to flatter the result destroys the one thing that lets the team know whether the product actually improved, and it misleads the customer into a launch decision based on fiction. They redirect: report the numbers honestly, lead with the failing slices as the roadmap ("here's exactly what we'd fix before launch"), and let the customer decide with true information. Strong candidates also note they'd document the request, because being asked to massage numbers is itself something a paper trail should exist for.

The common failure mode: complying, or refusing theatrically ("I could never do that!") without explaining the professional reasoning. Interviewers are grading whether the candidate understands why measurement integrity matters — it's what makes every future decision possible — not just whether they know the rule.

The follow-up the interviewer should ask: "The manager says it's just internal reporting, the customer will never see it. Does that change your answer?"

20. The fabricated demo

You discover that a teammate's celebrated demo — the one that won the pilot extension — used hand-written outputs presented as live system results. Nobody else seems to know. The teammate is well-liked, the extension is signed, and raising it will cause real damage to someone you work with every day.

Grading notes — what the interviewer is listening for

A strong answer demonstrates: the candidate doesn't treat this as a loyalty test — it's a trust test, and the trust at stake belongs to the customer, who just signed an extension based on fiction. Strong candidates start privately (talk to the teammate first, give them the chance to correct it themselves — there may be context you don't have), but they don't end privately: if the teammate won't correct it, the candidate escalates through the proper channel. They name the downstream harm clearly: every future commitment the team makes will be priced against results that never existed. The tell: the candidate is uncomfortable and does it anyway — interviewers watch for whether discomfort changes the answer.

The common failure mode: staying silent for the sake of team harmony, or going straight to public accusation without giving the teammate a chance to explain. The interviewer is checking for the hardest combination in professional life: loyalty to people and honesty to the customer, in the right order.

The follow-up the interviewer should ask: "You talk to the teammate and they say the outputs were 'representative of what the system will do' and it's standard practice. Your response?"

How to rehearse

Reading twenty scenarios is not preparation. Answering them out loud, badly, several times, is. Here's a rehearsal method that actually transfers to the interview room:

  1. Pick five, not twenty. Choose one scenario from each group plus one wild card. Depth beats coverage — the interviewer is grading your thinking pattern, which transfers across scenarios, not your memorized answers, which don't.
  2. Answer out loud with a timer. Three minutes per scenario, no notes. The time pressure is the point: it forces you to reach for the four-step structure instead of wandering. Record yourself on your phone.
  3. Grade yourself with the notes — harshly. After each attempt, open the grading notes and check all three bullets. Most people discover the same gap: they proposed a fix before clarifying, or they never said who they'd tell. That's the gap to drill, not the scenario.
  4. Do one with a partner. Have a friend read the scenario and ask the follow-up question no matter how good your answer was. The follow-up is where interviews are actually decided — anyone can survive the first question; the second one tests whether your thinking holds up under pressure.
  5. Build your two-minute stories. For each group, prepare one real (or realistic composite) story from your own experience that demonstrates the same muscle: a time you clarified a vague ask, a launch you triaged, a line you held. Scenario answers show how you think; stories show you've done it. Interviewers ask for both.

What not to rehearse

  • Memorized scripts for specific scenarios. Interviewers change the details deliberately; a memorized answer to a scenario you half-recognize sounds worse than thinking fresh.
  • Clever technical solutions. Nobody ever got hired from a scenario interview because their rate-limit strategy was elegant. They get hired because the customer in the story would have trusted them.
  • The "right" answer. There isn't one. There's a defensible process, honestly walked through, with tradeoffs named and people told. That's the whole game.

Avoid

  • Promising dates, fixes, or outcomes before you understand the problem — in the interview or in real life.
  • Blaming the customer, the vendor, or your teammate in your answer. Describe what happened; assign the fix, not the fault.
  • Treating ethics scenarios as hypothetical. Answer them as commitments: "I would not ship that" is a statement about who you are, not a debate position.

Putting it together

Every stage of this curriculum has been building toward the moment these interviews simulate. Discovery taught you to turn fog into a problem statement — that's the Clarify step. Scoping taught you to make tradeoffs visible and get agreement in writing — that's Propose and Tradeoffs. Guardrails taught you that trust is a system property and some lines don't move — that's Group D. Milestone 5 taught you what a launch actually costs — that's Group C, at 2 a.m., with the customer watching.

The scenario interview is not a separate skill bolted onto the engineering. It is the engineering, with the customer in the room. Walk in with clarify → propose → tradeoffs → communicate, hold the lines that matter, and let the process be the answer.

Field check

Before you call yourself ready for scenario interviews, prove it — out loud, timed, with no notes:

  1. Take scenario 3 (the security block, 48 hours before the demo). Give your full answer in under three minutes, then write down the one sentence you'd actually send to Lisa. Does it treat her as an ally?
  2. Take scenario 9 ("we'll know it when we see it"). Explain your measurement plan to a non-technical friend. If they can't repeat it back, your plan isn't clear enough for a customer either.
  3. Take scenario 16 (the unsafe ship). Say your refusal out loud, including the alternative you offer. If the refusal sounds like an apology, rehearse it until it sounds like a professional judgment.
  4. Record yourself answering scenario 12 (the crashed demo). Watch it back and count: did you clarify anything before proposing? Did you name who you'd tell afterward?
  5. Write your one-paragraph version of the principle — "the process is the answer" — in your own words, and test it against a scenario from your own experience. If it doesn't fit a real situation you've lived, it's a slogan, not a tool.
What "ready" looks like

You're ready when: you can walk any of the twenty scenarios through clarify → propose → tradeoffs → communicate without notes, in under four minutes, while naming at least one tradeoff and at least one person you'd communicate with. You can state a refusal (scenarios 16–20) plainly, without hedging or apologizing, and offer a legitimate alternative in the same breath. And you have two real stories from your own experience — one where you clarified a vague ask, one where you managed a difficult human situation — that you can tell in under two minutes each. If any of that feels shaky, rehearse that piece specifically; don't re-read the post.