Prerequisites: Stages 1–5 of this curriculum, especially Discovery, Legacy Systems, RAG End-to-End, Evals & Tracing, and the Deploy stage. All names, companies, and people in this brief are fictional and illustrative. This is a drafting-and-execution brief — a worked reference solution will appear later as an upcoming end-to-end chapter; attempt it yourself first.

Maya has been a hiring manager for forward deployed teams for six years. Two portfolios sit on her desk.

Candidate A lists five completed tutorials: a to-do app, a weather dashboard, a chatbot demo built on a sample dataset, a Kaggle notebook, and a clone of a popular site. Every repo has a README that says "built to learn X." Nothing is finished in the way a business would need it to be finished: no data contract, no error handling, no runbook, no evidence anyone ever used it.

Candidate B has one project: a six-week engagement brief for a fictional retailer, executed end to end. There is a discovery memo. A data contract with a quarantine policy. A migration pipeline whose reconciliation report proves the numbers tie out. A connector for a nasty legacy API with auth, pagination, retries, and webhook handling. A live dashboard. An AI assistant over the client's SOP docs — with an eval set and measured results, not vibes. A runbook an on-call engineer could actually follow at 2 a.m.

Maya interviews Candidate B. This post is how you become Candidate B.

A project brief is the closest thing to a real FDE engagement you can build without a client. This one gives you the whole thing — the messy data, the creaky API, the stakeholders, the deadline — plus the rubric it will be graded against. The brief is the job; the rubric is the contract.

The principle

"The brief is the job; the rubric is the contract." In a real engagement, the client's definition of done is the only definition that matters. Here, the rubric plays that role: it grades outcomes — reconciled data, a robust connector, measured copilot quality — not effort, not completion theater. Read the rubric before you write a line of code, the way you'd read a contract before signing it.

Why this brief exists

Stage 6 is the Get Hired stage, and its currency is evidence. A hiring manager for a forward deployed role is not buying certificates; they are buying a prediction about how you behave when a client's data is a mess, their API fights back, and their executives want a demo on Friday. Tutorials cannot provide that evidence. A fresh-build engagement brief can — because it forces you to make the same trade-offs a working FDE makes: what to clarify, what to measure, what to harden, what to document, and what to refuse.

This brief is deliberately fresh build, not a guided walkthrough. You are given the engagement the way a real one arrives: incomplete information, competing stakeholders, hard constraints, and a deadline. What you produce is a portfolio artifact that reads like field work — discovery notes, a data contract, a reconciliation report, a connector, a dashboard, a copilot with evals, a runbook. Each artifact answers a question a hiring manager actually asks: can this person turn a drowning client into a live system without breaking the business?

Two rules govern everything below. First, grade outcomes, not completion — the curriculum's milestone-scorecard rule. A pipeline that runs but silently drops 3% of orders is worse than an honest report that says "we can migrate 97% today and here is exactly what the other 3% needs." Second, separate what you observed from what you decided: profile the data before you clean it, measure the API before you trust it, evaluate the copilot before you demo it.

Useful background, if you need a refresher before starting: Discovery: Turning "We're Drowning" Into a Problem Statement, Scoping & Success Metrics: What You'll Refuse to Build, and Designing Systems Under Constraints.

1. The engagement: Acme Corp

Everything in this section is fictional and illustrative — a composite engagement assembled from common patterns, so you can execute it without a real client. Treat it as real while you work: make decisions, write them down, and be ready to defend them.

Who you are

You are a Forward Deployed Engineer at a software company that sells Nimbus, a fictional business-operations platform: dashboards, data pipelines, integrations, and AI assistants, all tenant-isolated per customer. A new customer has signed. You are the engineer sent to get them live.

Who the client is

Acme Corp (fictional) is a mid-size retailer: a few dozen stores plus an online shop, operating for roughly a decade. They are successful in the way that creates the mess — growth outpaced systems. Their ask, in the sales rep's one-liner: "We need one view of our business and an AI assistant that can answer questions from our SOP documents." Your job in week 1 is to turn that sentence into a problem statement you can build against.

What they bring you

Three assets, each a classic FDE headache:

  1. A decade of orders trapped in Excel and CSV exports. Not a database — exports, produced by hand, by different people, over ten years. Expect the full catalog of real-world damage: inconsistent date formats (03/04/2019 next to 2019-04-03 next to Apr 3 19), duplicate customers ("Acme Retail Ltd" vs "ACME RETAIL LTD." vs "acme-retail"), missing SKUs on a meaningful share of rows, and shifting column names — the same field called OrderID, order_id, Order No, and id depending on the year and who exported it. Some files are UTF-8, some are not. Some have 12 columns, some have 31.
  2. A legacy inventory system behind a creaky REST API. It is the system of record for stock levels and it is not going anywhere — Acme's warehouse runs on it. The API has authentication (an aging token scheme), pagination (inconsistent page sizes, occasional duplicate pages), undocumented rate limits, intermittent 500s, and webhook support that fires unreliably. There is partial documentation, and you should trust it the way you'd trust a stranger's directions: verify everything. Background reading: Legacy Systems, SDK Design & Debugging Someone Else's API, Pagination, Rate Limits & Retries, Auth in the Wild, Webhooks.
  3. Stakeholders who need two things live: a dashboard and an AI assistant over their SOP docs. The dashboard must show the business truthfully — orders, revenue, inventory — on data you migrated and integrated yourself. The assistant must answer staff questions from their standard operating procedures: returns, refunds, shipping cutoffs, discount rules. It must cite its sources, and it must not invent policy.

The goal and the deadline

Get Acme live on Nimbus in six weeks without breaking their business. "Live" means: their order history migrated and reconciled, inventory flowing from the legacy API, a dashboard their CFO trusts, an SOP assistant their staff can use, and a handoff package their team (or yours, after you leave) can operate. "Without breaking their business" means the legacy inventory system keeps running untouched throughout — you integrate read-only, cut over in stages, and can roll back.

Stakeholder notes (illustrative)

The following notes are fictional and illustrative — composites of the kinds of things stakeholders actually say. Read them the way you'd read discovery interview notes: every line hides a requirement, a risk, or a trap.

Priya, Operations Lead (fictional): "I just need to know what's actually in the warehouses before I promise a customer anything. Right now I check three spreadsheets and call Marcus. If your dashboard is wrong even once, my team will go back to the spreadsheets and never come back."

Marcus, Warehouse Manager (fictional): "The inventory system is old but it works. Nobody touches it during the holiday rush — if your integration slows it down or locks a table, that is on you. Also the webhooks sometimes fire twice. Or not at all."

Dana, CFO (fictional): "The board deck goes out on the first Monday of every month. The revenue number in your dashboard has to match the number in our accounting exports to the cent, or I can't use it. And I need to know which orders you couldn't migrate and why."

Sam, Store Associate (fictional): "Customers ask me about the return policy and I give three different answers depending on my mood. If the assistant can just tell me the real policy with a link to the doc, that would save my life. But if it makes up a policy, I'll get fired before you get debugged."

Read these notes like an FDE

  • Priya's note is a trust requirement: one wrong number kills adoption. Your reconciliation report is the answer to her.
  • Marcus's note is a constraint: read-only integration, no load on the legacy system during peak, idempotent webhook handling (fire twice = process once).
  • Dana's note is a success metric: revenue reconciles to the cent; unmigrated orders are enumerated, not hidden.
  • Sam's note is a safety requirement: citations against the retrieved evidence set, and a measured hallucination rate — not "the model seems fine."

For the discovery technique behind turning notes like these into a buildable problem statement, see Discovery. For deciding what you will not build in six weeks, see Scoping & Success Metrics.

2. The six-week plan

This is the milestone spine. Each milestone has exit criteria — things that must be true, not things you must have attempted. You may compress or reorder weeks, but you may not skip an exit criterion. If a criterion is not met, the milestone is not done, and you say so in your status notes: honesty about a red milestone is itself graded (see the rubric).

flowchart LR A[Discovery<br/>notes & problem statement] --> B[Intake<br/>profiling & data contract] B --> C[Migration pipeline<br/>validation + reconciliation] C --> D[Legacy API integration<br/>auth, pagination, retries, webhooks] D --> E[Dashboard & SOP copilot<br/>trustworthy surfaces] E --> F[Deploy & harden<br/>security, monitoring, rollback] F --> G[Demo & handoff<br/>runbook, sign-off]
WeekMilestoneExit criteria (must be true)
1Discovery notesA written problem statement exists and each stakeholder's core need is quoted back to them in their own words. Explicit out-of-scope list (what you will refuse to build in six weeks). Named risks with owners (e.g., "legacy API undocumented rate limits — owner: you, mitigation: measured backoff"). A one-page engagement plan with the milestones below.
1–2Data profilingA profiling report over the actual exports: every distinct date format found, column-name variants mapped per year, measured duplicate-customer rate, share of rows with missing SKUs, encoding/row-length anomalies. Nothing cleaned yet — this milestone is observation, and its output is the evidence the data contract is built on.
2–3Data contractA written contract: canonical schema for orders, customers, products; normalization rules (dates → ISO, customer dedup keys, SKU handling); a quarantine policy for dirty rows that the customer approves (you propose, they decide — never the reverse); validation rules with severities (reject vs. quarantine vs. warn). Stored as a versioned document, not tribal knowledge.
3–4Migration pipelineA repeatable pipeline (code, not a notebook you ran once) that ingests the exports, applies the contract, and emits two reports: a validation report (rows rejected/quarantined/warned, by rule) and a reconciliation report (source row counts vs. migrated counts, revenue totals vs. the accounting exports — to the cent, per Dana). Re-running the pipeline on the same inputs produces identical outputs (idempotent, deterministic).
4–5Legacy API integrationA connector that authenticates, pages through inventory without assuming page sizes, retries with backoff on the measured failure modes, deduplicates unreliable webhooks by idempotency key, and never writes to the legacy system. Documented rate-limit behavior you measured, not what the docs claim. A load test or reasoned argument that it will not disturb the warehouse during peak.
5Dashboard + SOP copilotDashboard: the numbers a stakeholder can cross-check (revenue, orders, inventory) match the reconciliation report; every figure traces to a source. Copilot: answers grounded in the SOP docs with citations; an eval set of real staff questions with measured pass rates; guardrails (no invented policy, PII handling, audit trail). Both tenant-isolated on Nimbus.
6Deploy, harden, demo, hand offDeployed behind the project's security posture (secrets managed, RBAC, tenant isolation), with health checks, logging, and monitoring that would page you before Priya notices. A stakeholder demo against the agreed success metrics. A runbook: how to operate it, how to roll back, what to do when each known failure mode fires, and who owns what after you leave. Written sign-off from the (fictional) stakeholders.

3. Deliverables checklist

When the six weeks are over, this is the package on the table. Each item is an artifact a real engagement would leave behind — and each one is something you can show in an interview without violating anyone's confidentiality, because the client is fictional.

#DeliverableWhat it proves
1Discovery memo — problem statement, stakeholder needs in their own words, out-of-scope list, risks with ownersYou clarify before you build.
2Data profiling report — measured mess: date formats, column variants, duplicate-customer rate, missing-SKU share, encoding anomaliesYou observe before you decide.
3Data contract (versioned) — canonical schemas, normalization rules, customer-approved quarantine policy, validation severitiesYou make data agreements explicit and durable.
4Migration pipeline (code) — repeatable, deterministic, idempotentYou build things that can be re-run, not notebooks you ran once.
5Validation report — per-rule reject/quarantine/warn countsYou measure what your pipeline did to the data.
6Reconciliation report — source vs. migrated counts and revenue totals, to the centYou can prove the numbers tie out. (Dana's board deck.)
7Legacy API connector — auth, pagination, retries with backoff, idempotent webhook handling, read-onlyYou integrate with systems you don't control without breaking them.
8Connector test notes — measured rate limits, failure modes found, peak-load reasoningYou verify the stranger's directions. (Marcus's warehouse.)
9Live dashboard on Nimbus — figures traceable to sourcesYou ship a surface stakeholders trust. (Priya's team.)
10SOP copilot — cited answers over the SOP docs, guardrails, audit trailYou put AI in front of staff responsibly. (Sam's job.)
11Copilot eval set + results — real staff questions, pass rates, failure analysisYou measure AI quality instead of asserting it.
12Security posture notes — secrets handling, RBAC, tenant isolation, what the security review askedYou treat trust as a system property.
13Monitoring & alerting — health checks, logs, alerts that fire before the customer noticesYou operate what you ship.
14Demo script + stakeholder sign-off — demo mapped to the agreed success metricsYou communicate outcomes, not activity.
15Runbook — operations, rollback plan, known failure modes with responses, ownership after handoffYou leave the client (or the next engineer) able to run it at 2 a.m.

What "done" is not

A demo that works on your laptop is not done. A pipeline you ran once is not done. A copilot that "seems good" is not done. Done is: reconciled, measured, documented, deployable by someone else, and signed off against the success metrics. The rubric below is the instrument that checks this — read it now, before you plan your weeks.

Anatomy of three artifacts (worked excerpts)

Abstract checklists don't teach; artifacts do. Below are condensed excerpts of what three of the highest-weight deliverables actually look like. They are illustrative examples in the fictional Acme context — yours will differ in the details, but should match in rigor.

Excerpt 1 — data contract (versioned YAML). Note what it does not do: it never guesses. Ambiguous dates go to quarantine with a rule name, not to a silent default.

# data-contract v0.3 — Acme Corp orders (fictional example)
# Approved by: Dana (CFO), Acme — 2026-10-20 (illustrative)
entities:
  order:
    fields:
      order_id:      {type: string, required: true}
      order_date:    {type: date, required: true, format: ISO-8601}
      customer_key:  {type: string, required: true}   # normalized name + postcode
      sku:           {type: string, required: false}  # missing -> quarantine, not drop
      line_total:    {type: money, required: true, currency: USD}
normalization:
  - dates: all observed formats mapped -> ISO-8601; unparseable -> quarantine(rule=date_unparseable)
  - customers: lowercase, strip punctuation/suffixes -> match key; collisions -> quarantine(rule=customer_collision)
  - columns: per-year alias map (OrderID/order_id/Order No/id -> order_id)
validation:
  reject:    [order_id missing, line_total <= 0, order_date in future]
  quarantine: [date_unparseable, customer_collision, sku missing]
  warn:      [line_total rounding > $0.01 vs source]
quarantine_policy:
  owner: Acme ops (Priya) reviews weekly; unreviewed rows stay quarantined,
         never auto-promoted. Reject-rate tolerance set by Acme, not by us.

Excerpt 2 — reconciliation report (tail). This is the page Dana reads. Every source row is accounted for in exactly one bucket, and revenue ties to the cent.

RECONCILIATION — orders migration, run 2026-11-02 (fictional example)
Source exports scanned ............ 1,284,306 rows (212 files, 2016-2026)
Migrated clean .................... 1,251,882 rows  (97.47%)
Quarantined .......................    28,914 rows  (2.25%)  -> review table, reasons attached
Rejected ..........................     3,510 rows  (0.27%)  -> enumerated in validation report
Accounted for ..................... 1,284,306 rows (100.00%)  -- no silent drops
Revenue, source exports ........... $48,213,904.17
Revenue, migrated ................. $47,092,311.55
Revenue, quarantined (est.) ....... $ 1,118,082.40
Revenue, rejected ................. $     3,510.22
Totals reconcile to the cent ..... YES
Determinism check (re-run diff) .. IDENTICAL

Excerpt 3 — weekly status note (week 4). Notice the honest red milestone, the plan attached to it, and the absence of adjectives doing the work of evidence.

To: Priya, Marcus, Dana (fictional) — Week 4 status

Green: Migration pipeline re-runs deterministically; reconciliation ties to the cent on the full history. Dashboard prototype shows orders and revenue matching the report.

Yellow: Legacy API webhooks fire twice on ~4% of deliveries in our tests; connector deduplicates by idempotency key, but we have not yet observed a full peak-day cycle.

Red: SOP copilot evals: 88% overall, but abstention tests failing (12 of 20) — the model answers from general knowledge when the SOPs are silent. We are not demoing the copilot next week as planned. Instead: narrowing scope to returns/refunds/shipping (the documented sections), adding a strict abstention path, re-running evals. Demo moves to week 6; dashboard demo stays on schedule.

Needs from you (Dana): sign-off on the quarantine review cadence — 28,914 rows need an owner before go-live.

That status note is doing three jobs at once: it reports truthfully (rubric: stakeholder communication), it protects the deadline by re-scoping instead of hoping (constraints), and it asks the customer to own the decision that is theirs (quarantine policy). If you can write that note every Friday for six weeks, you can do this job.

4. Constraints: don't break the business

Constraints are what separate an engagement from a side project. These are non-negotiable; violating one fails the brief no matter how shiny the dashboard is.

Constraint 1: The legacy system is untouchable

The inventory API is the system of record for a working warehouse. Your integration is read-only: no writes, no schema changes, no "quick fixes" to their side. You may not add load that risks the warehouse during peak — Marcus's warning about the holiday rush is a hard boundary. Design the connector accordingly: cache aggressively, sync on a schedule the business approves, back off the moment the API shows stress, and have a documented kill switch. If the API goes down, Acme's warehouse keeps working and your dashboard shows stale-but-labeled data, never silently wrong data.

Constraint 2: The data is guilty until proven innocent

Ten years of hand-made exports will contain contradictions no one remembers creating. Your job is not to "clean the data" as a heroic solo act — it is to make the mess visible and get decisions from the customer. The quarantine policy belongs to Acme: you propose thresholds and consequences, they approve. Every row you reject or quarantine is enumerated in the validation report with a reason. Silent drops are a trust violation — Priya's team will find the one wrong number, and that will be the number they remember.

A concrete trap to plan for: 03/04/2019 is March 4th or April 3rd depending on who typed it. You cannot resolve this by guessing. The data contract must say how ambiguous dates are handled (quarantine? customer ruling? per-file locale evidence?), and the reconciliation report must show the decision was applied consistently. Date handling background: Dates, Timezones & Money.

Constraint 3: The copilot may not invent policy

Sam's fear is the correct fear. An assistant that answers "what is the return policy?" with a confident, plausible, wrong answer is worse than no assistant. So: citations against the retrieved evidence set on every answer, an explicit "I don't find this in the SOPs" behavior, evals that measure groundedness (not just fluency), and guardrails around PII and audit. If the eval numbers are bad, you report them and narrow the scope — you do not ship and hope. Background: Guardrails, Approvals & Audit and Evals & Tracing.

Constraint 4: Six weeks means saying no

You cannot rebuild their warehouse system, unify a decade of master data perfectly, and ship a general-purpose AI platform in six weeks — and pretending otherwise is how engagements die. The out-of-scope list in your discovery memo is a deliverable, not an apology. A good one names the tempting things you refused (real-time sync? full customer-master dedup across all history? multi-language SOPs?) and says what the client gets instead.

5. The rubric: graded on outcomes, not completion

This is the contract. Score yourself — or have a peer score you — against outcomes that can be checked, not effort you put in. Weights sum to 100. For each criterion, "excellent" means the outcome is demonstrated with evidence; "acceptable" means it works but the evidence is thin; "needs work" means a hiring manager would spot the gap in an interview.

Criterion (weight)ExcellentAcceptableNeeds work
Data integrity & reconciliation (25%)Reconciliation report ties source → migrated counts and revenue to the cent against the accounting exports. Every rejected/quarantined row is enumerated with a rule and a reason. Pipeline is deterministic: re-runs produce identical outputs. Ambiguous cases (dates, dup customers) resolved by a documented, customer-approved rule.Counts reconcile but revenue is off by small amounts with a hand-waved explanation, or quarantine reasons are grouped so coarsely ("bad rows") that no one could act on them. Re-runs mostly reproduce.Numbers don't tie out and the report doesn't say so. Silent drops. Date/customer ambiguities "resolved" by guessing, undocumented. Re-running the pipeline gives different results.
API integration robustness (20%)Connector handles auth, pagination (no assumed page sizes), retries with backoff tuned to measured failure modes, and idempotent webhook processing (duplicate deliveries processed once). Read-only by construction. Documented evidence it won't disturb the legacy system at peak, plus a kill switch. Measured rate-limit behavior, not docs quotes.Works against the happy path; retries exist but with unmeasured defaults; webhooks deduplicated "probably." No evidence about peak-load impact. Docs trusted where they should have been verified.Breaks on page-size changes, duplicate webhooks double-count inventory, or — worst case — writes to the legacy system. No retry strategy; one 500 kills the sync. Assumes the docs are accurate.
Copilot quality, with evals (20%)Answers carry citations to the retrieved SOP evidence. A labeled eval set of realistic staff questions with measured pass rates and a failure analysis (what failed, why, what changed). Explicit abstention behavior ("not in the SOPs"). Guardrails: no invented policy, PII handling, audit trail. Scope was narrowed honestly if numbers were weak.Citations present but eval set is small or ad hoc; pass rate reported without failure analysis. Abstention exists but untested. Guardrails mentioned, not demonstrated.No evals — "it looks good in the demo." Answers without citations. Hallucinated policy in the demo. Claims about quality the code doesn't measure.
Deployment & hardening (15%)Secrets managed (never in code), RBAC and tenant isolation in place, health checks + logging + monitoring that alert before the customer notices. Rollback plan exists and is plausible; backups/DR considered for the migrated data. Security-review-style self-critique included.Deployed and reachable, but secrets handling is sloppy, monitoring is "I'll check the logs," rollback is "redeploy the old version (untested)."Runs on a laptop. Credentials in the repo. No monitoring, no rollback story, no tenant isolation. "We'll harden it later."
Stakeholder communication (10%)Discovery memo quotes stakeholders' needs back accurately; status notes report red milestones honestly with a plan; demo maps to the agreed success metrics; Dana's revenue question and Priya's trust question are answered with artifacts, not adjectives.Status is green-yellow-red but vague; demo shows features rather than success metrics; stakeholder needs paraphrased loosely.No written communication artifacts. Demo is a feature tour. Bad news hidden until asked. Promises the brief didn't make.
Runbook & handoff (10%)An on-call engineer who's never seen the system can operate it: startup/shutdown, normal ops, each known failure mode with symptoms → diagnosis → fix, rollback steps, escalation contacts, and ownership after you leave. Tested by having someone else follow it.Runbook exists but reads like architecture notes; failure modes listed without responses; rollback steps untested.No runbook, or a README that says "contact me with questions." Knowledge leaves when you leave.

How to use this rubric

  • Before you start: read it as a contract. Every "excellent" cell is a requirement you are signing up for.
  • During: when a week goes sideways, the rubric tells you what to protect. A red milestone reported honestly with a recovery plan beats a green milestone that's fiction — "stakeholder communication" grades the honesty.
  • After: score yourself harshly, write down the score with evidence links (report files, eval outputs, runbook sections), and put the top-line result in your portfolio README. A self-score of 78/100 with evidence is more credible than a claimed 100.
  • In interviews: walk through one rubric row and its evidence. "Here's my reconciliation report — revenue ties to the cent, and here are the 2,140 quarantined rows with reasons" ends more interviews well than any certificate.

How engagements typically fail each row

Read the rubric as a list of the six ways this engagement dies in practice. Each failure below is a composite of mistakes working engineers actually make — check your plan against them:

  • Data integrity: the "looks right" migration — spot-checked ten rows, never reconciled totals, and the 2% of dropped rows are discovered by the CFO, not by you. The fix was always the same: a reconciliation report run before anyone asks for it.
  • API integration: the happy-path connector — built against the docs, never against the API's actual behavior, so the first duplicate webhook page double-counts inventory in production. The fix: probe first, assume nothing, deduplicate everything.
  • Copilot quality: the demo-driven copilot — tuned on the five questions in the demo script, never evaluated on the questions staff actually ask, and the first abstention failure happens in front of a user. The fix: an eval set built from real question distributions before the demo.
  • Deployment: the laptop deployment — everything works until the engineer goes on vacation, because secrets, monitoring, and rollback lived in their head. The fix: the runbook test — hand it to a stranger.
  • Communication: the green-status fiction — weeks of "on track" followed by a week-6 surprise, which teaches the client that your status notes are marketing. The fix: report the red milestone the week it turns red, with the recovery plan attached.
  • Handoff: the "call me" handoff — no runbook, no ownership transfer, and six months later nobody knows why the sync job has a cron entry nobody understands. The fix: write the runbook as if you will never be reachable again.

Myth: "Completion is what gets graded." A finished dashboard over unreconciled data scores worse than an honest, incomplete migration with a precise account of what's missing and why. The rubric has no row for "tried hard." It has rows for outcomes a business can rely on. That is the entire point of grading outcomes, not completion.

6. Hints (expand when stuck)

Stuck is part of the brief — a real engagement doesn't come with a solutions manual either. But getting permanently blocked helps no one, so here are nudges. Try for at least an hour before opening one.

Hint: Where do I even start in week 1?

Start with the stakeholder notes, not the data. Write the problem statement first: one paragraph per stakeholder, in their words, then one paragraph of what "live in six weeks" concretely means. Then write the out-of-scope list — it forces the scoping conversation early. Only then open the first CSV. Technique refresher: Discovery.

Hint: How do I profile ten years of messy exports without drowning?

Don't read everything — sample and measure. Write a small profiler that reports, per file: column names, row counts, detected date formats per date-like column, null share per column, and duplicate-key rates. The output of week 2 is this report, not cleaned data. Resist cleaning during profiling; observation first, decisions later. See Files & the Messy Real World and Data Wrangling with pandas.

Hint: What does a data contract actually look like?

A short versioned document: (1) canonical schema per entity (field name, type, required/optional), (2) normalization rules ("all dates → ISO 8601; customer match key = normalized name + postcode"), (3) validation rules with severities — reject (never load), quarantine (load to a review table), warn (load, flag), and (4) the quarantine policy signed off by the customer: thresholds, who reviews the quarantine table, and what happens to unreviewed rows. Encode the rules in code (e.g., Pydantic models — see Trust No Input), not just prose.

Hint: Validation vs. reconciliation — aren't they the same thing?

No, and the rubric grades both. Validation checks each row against the contract as it flows through ("does this row satisfy the rules?"). Reconciliation checks the whole migration against the source after the fact ("did every source row land somewhere — migrated, quarantined, or rejected — and do the totals match?"). Validation without reconciliation can silently lose rows; reconciliation without validation can't tell you why rows failed. Dana's board deck needs the reconciliation report.

Hint: The legacy API has partial docs. How do I integrate safely?

Treat the docs as hypotheses. Probe: page through and check for duplicate/overlapping pages and changing page sizes; measure when 429s/500s actually appear under load you control; test webhook delivery with a receiver that logs everything, then verify idempotency by replaying deliveries. Wrap it all in a connector with auth, pagination, retries with backoff, and idempotency keys — and keep it read-only. Background: SDK Design & Debugging Someone Else's API, Pagination, Rate Limits & Retries, Webhooks. Debugging method: Debugging Like a Detective.

Hint: How many eval questions does the copilot need?

Enough to cover the real question distribution, not a magic number. Start with 30–50 questions written from Sam's perspective (returns, refunds, shipping cutoffs, discount rules), including trick questions whose answers are not in the SOPs (to test abstention) and paraphrases of the same question (to test robustness). Label expected answers with the source section. Report pass rate per category plus a failure analysis — "failed 4/12 abstention tests because the model answered from general knowledge" is the kind of sentence that scores. Method: Evals & Tracing; ingestion: Feeding the Machine.

Hint: What goes in the runbook, concretely?

Write it for a stranger at 2 a.m.: how to start/stop each component, what "healthy" looks like (dashboards, log lines), the five most likely failures with symptoms → diagnosis → fix (stale inventory sync, webhook backlog, pipeline validation spike, copilot eval regression, expired API token), the rollback procedure step by step, backup/restore for migrated data, and who owns what after handoff. Then hand it to someone and watch them follow it — every place they get stuck is a runbook bug. Ops background: Health Checks, Logs & Monitoring, Backups & Disaster Recovery.

7. The reference solution unlocks after the attempt

A fully worked reference solution for Project Acme will appear later as an upcoming end-to-end chapter — the same brief, executed with all artifacts, so you can compare your decisions against a second opinion. It unlocks after your attempt for a reason: reading the solution first converts the engagement into a tutorial, and tutorials are exactly what this brief is the antidote to.

The honest order of operations:

  1. Attempt the brief yourself — all six weeks (or a compressed version you declare up front), scored against the rubric above.
  2. Write down your score with evidence before looking at anything else.
  3. Then read the reference chapter and diff your decisions against it: where did you scope differently, what did the reference measure that you didn't, which of your calls survived contact with a second opinion?

The diff in step 3 is where the learning lives — and it's a superb interview story: "here's what I built, here's how I scored it, and here's what I'd change after seeing a second approach."

Field check

  1. Why does the rubric grade reconciliation (25%) separately from validation, instead of folding both into "the pipeline works"?
  2. Name two concrete signals, visible in production, that your data contract is being violated by new data.
  3. Marcus's warehouse runs on the legacy system. Your sync job starts failing with 500s at 2 p.m. on a peak day. What do you do — and what do you not do?
  4. Your copilot eval shows 90% pass rate overall but fails most abstention tests (it answers questions the SOPs don't cover). Do you ship it? What do you change?
  5. A peer says "I'll just re-run the notebook if the data changes." Which rubric row does that violate, and what artifact would fix it?
Check your answers

1. Because they catch different failures: validation checks rows against rules in flight; reconciliation checks the whole migration against the source after the fact (counts, revenue to the cent). A pipeline can validate every row it processes while silently never processing some rows. Dana's board deck needs reconciliation, not just validation.

2. Any two of: a spike in quarantine/reject rates in the validation report; reconciliation totals drifting from source totals; new column names or date formats appearing in intake logs; dashboard figures diverging from accounting exports; alerts on schema-validation failures in the pipeline's monitoring.

3. Do: back off immediately (don't hammer a struggling legacy system), serve stale-but-labeled inventory data, alert per the runbook, and investigate off-peak. Do not: retry aggressively, "fix" anything on their side, or silently serve fresh-looking but wrong numbers. The constraint is read-only and do-no-harm — the warehouse keeps running no matter what.

4. Do not ship as-is: a 90% headline with failed abstention means it invents policy — Sam's firing scenario. Narrow the scope (fewer question types, stricter retrieval thresholds), add an explicit abstention path, re-eval, and report the numbers honestly. A smaller, trustworthy copilot beats a broad, confident one.

5. Data integrity & reconciliation (25%): a notebook re-run is neither deterministic nor reviewable — no versioned contract enforcement, no emitted validation/reconciliation reports, no identical-output guarantee. The fix is the migration pipeline as code (deliverable #4) emitting the validation and reconciliation reports (#5, #6).

Project Acme is the second portfolio brief in the Get Hired stage: a fresh build, executed by you, graded by outcomes. The artifacts it leaves behind — discovery memo, data contract, reconciliation report, connector, evals, runbook — are the evidence a hiring manager is actually buying. Build them like the business depends on them, because in the story of this brief, it does.

When you're ready to talk about this work the way an interviewer will probe it, the Get Hired page is your home base for the stage.