Prerequisites: Comfort writing and testing a small Python project, the testing with pytest workflow, and the scoping mindset from Scoping & Success Metrics. Everything else is taught here.

Thursday, 6:40 p.m. A recruiter's email lands in your inbox: "We'd like you to do a take-home. Build a small service that ingests a CSV of support tickets and exposes a summary endpoint. Please budget about four hours. Send us the repo link."

You have a strong instinct to open your editor immediately and start coding. Four hours feels generous — you'll write something impressive. Maybe a database, a dashboard, caching, rate limiting, Docker. By Sunday night you'll have worked nine hours and built a beautiful system that does far more than the brief asked.

Monday, the hiring team's verdict arrives in one line: "Didn't quite demonstrate what we were looking for." No working code complaint — your service ran. What exactly didn't demonstrate?

This post answers that question. (All candidates and hiring managers in this post are illustrative composites based on common patterns, not any specific person or company.)

A take-home assignment is a work sample graded by strangers. The grader never meets you, never watches you code, and spends limited time on your submission. What they grade is not how much you typed — it's the evidence of judgment your submission leaves behind.

The principle

The take-home grades your judgment, not your typing speed. Every artifact — the code, the tests, the README, the commit history — is a piece of evidence. Weak candidates ship volume; strong candidates ship decisions, documented and defended.

Minimum concept: what the grader can actually see

Strip away the interview mystique and the take-home is brutally simple. A person you have never met opens a repository you built alone, with somewhere between twenty minutes and an hour of attention, and tries to answer one question: would I want to work with the person who built this?

They cannot see your thought process. They cannot ask you why you chose SQLite over Postgres, or why you skipped the CLI, or what the vague sentence in the brief meant to you. They only see artifacts — files. And files are evidence of judgment: what you prioritized under a time budget, what you considered an edge case, how you communicate to a stranger.

This is the minimum concept that reorganizes everything else in this post:

  • Working code is the price of entry, not the differentiator. Most submissions run. The ones that advance are the ones where the grader can reconstruct why each decision was made.
  • Ambiguity is the test. Take-home briefs are deliberately (or accidentally) underspecified. The grader isn't checking whether you guessed their hidden preference — they're checking what you did with the gap: did you state an assumption, or did you silently guess?
  • Time is a constraint to be managed, visibly. Four hours is part of the brief. A submission that nails the core in four hours beats a sprawling one that clearly took twelve, because the job being hired for has deadlines too.

Everything below — the hour-by-hour plan, the README template, the ambiguity playbook, the repo structure — is one system for manufacturing legible evidence of judgment. Let's build it.

Part 1: The 4-hour execution plan

The most common failure mode of take-homes isn't bad code. It's bad budgeting: two hours lost to tooling setup and an over-engineered data model, forty frantic minutes of README at the end, zero tests, and a commit history consisting of one commit titled final final v2.

A four-hour take-home is a tiny project, and tiny projects still need a plan. Here is the plan, as a table you can tape to your monitor. It is a composite pattern distilled from what strong candidates tend to do — not a survey of any company's process.

TimePhaseWhat "done" looks like
0:00–0:30Read, clarify, designYou have written down (in a scratch note or the README draft) what the brief asks for, what it doesn't say, your assumptions, and a rough design: the two or three core behaviors, the data flow, what you will NOT build. No production code yet.
0:30–2:30Build the coreThe smallest version that satisfies the brief end-to-end works: ingest the CSV, compute the summary, serve it over the endpoint. Happy path first, wired together, runnable. You can demo it to yourself.
2:30–3:15Harden: edges, tests, errorsEdge cases handled and documented: empty file, malformed rows, duplicate tickets, huge input. A handful of pytest tests covering the core logic and at least one unhappy path. Errors are handled, not swallowed — and user-facing errors are safe to show.
3:15–3:45README + polishThe README is complete: how to run, what you assumed, known limitations, what you'd do next. Repo structure is clean. No secrets, no junk files, no 40MB CSV committed by accident.
3:45–4:00Final review as the graderFresh clone (or at least a fresh terminal) and follow your own README exactly. Fix anything that breaks. Read the diff of every commit. Submit.

The mermaid view: the 4-hour plan as a flow

flowchart TD A["0:00 — Read the brief twice"] --> B["List asks, gaps & assumptions
in writing"] B --> C["Rough design: core behaviors
+ explicit NOT building list"] C --> D["0:30 — Build the core
happy path, end to end"] D --> E["Demo to yourself:
does it run?"] E -->|"No"| D E -->|"Yes"| F["2:30 — Harden:
edge cases + pytest"] F --> G["Errors handled,
nothing swallowed"] G --> H["3:15 — README:
run, assumptions,
limitations, next steps"] H --> I["Clean repo:
no secrets, no junk"] I --> J["3:45 — Fresh-clone review:
follow your own README"] J -->|"Broken"| K["Fix, re-verify"] K --> J J -->|"Clean"| L["Submit"]

What to cut when time runs short

You will run short of time. Plan for it now, so the cuts are strategic instead of panicked. The ordering below is deliberate: cut from the bottom.

PriorityIf you're behind, ...Why it's safe to cut
Never cutCore behavior working end-to-end; a README that tells the grader how to run itWithout these, nothing else can be graded. A broken core with beautiful docs still fails the work-sample test.
Cut lastA few tests on the core logic and one unhappy pathTests are cheap evidence of judgment. Keep at least a skeleton suite — even three focused tests change the signal.
Cut nextEdge-case hardening beyond the obvious (empty input, one malformed row)Document the unhandled edges in the README's "Known limitations" instead of silently ignoring them. Acknowledged gaps score; silent gaps don't.
Cut firstExtra features: dashboards, caching layers, auth, Docker, a second interfaceEvery extra feature is evidence you didn't budget time. Graders read scope discipline as a proxy for on-the-job judgment.

Don't memorize — internalize

  • The exact minute boundaries matter less than the shape: design → core → harden → document → review-as-stranger.
  • If you finish the core early, the extra time goes to tests and README, not to features. Features are the trap.
  • The 15-minute final review as the grader catches more submission-killing mistakes than any other single activity. Protect it.

Part 2: READMEs that score

The README is the first file the grader opens and often the only prose they read in full. Think of it as a cover letter for your code: it frames everything they are about to see. A weak README forces the grader to reverse-engineer your intent from the code. A strong README tells them what to look for — and they will look for it kindly, because you've made their job easy.

Strong candidates' READMEs tend to share four sections, in this order:

  1. What this is — one paragraph: the problem, the approach, the scope. Written for a stranger.
  2. How to run — copy-pasteable commands that work from a fresh clone. This is the section that fails most often, and it's the one the grader hits first.
  3. What I assumed — the ambiguities you found and the calls you made. This is where judgment becomes visible.
  4. Known limitations + what I'd do next — the honest gap list and the roadmap. This turns "didn't do X" from a weakness into a demonstration of prioritization.

The template

Copy this skeleton for every take-home. Fill it in as you go — write the assumptions section during the 0:00–0:30 design phase, not at 3:40 in a panic.

# Ticket Summary Service

One-paragraph summary: what the service does, who it's for, and the
approach in a sentence. (2-3 lines. If you can't say it in 3 lines,
your scope is too big.)

## How to run

Prerequisites: Python 3.11+, nothing else.

    python -m venv .venv
    source .venv/bin/activate
    pip install -r requirements.txt
    python -m tickets summary data/tickets.csv --serve --port 8000

Then open http://localhost:8000/summary — you should see a JSON payload
with per-category counts for the sample data.

Run the tests:

    pytest -q

## Design notes

- The ingestion step streams the CSV row by row instead of loading it
  into memory, so a 1M-row file won't blow up RAM.
- Summaries are computed in a single pass and cached in memory;
  re-ingesting replaces the cache. (Chose this over a database:
  the brief asks for a summary view, not persistence.)

## Assumptions

1. The brief doesn't say how to treat rows with a missing category.
   I count them under "uncategorized" and report the count, rather
   than dropping them silently. Rationale: dropping data the user
   supplied is a decision the operator should make, not the tool.
2. "Summary" is interpreted as per-category counts plus the 5 oldest
   open tickets. If you meant something richer (trends over time),
   the aggregation layer is isolated in `tickets/aggregate.py` and
   easy to extend.
3. Timestamps are assumed to be UTC ISO-8601. Anything unparseable
   is logged to stderr and skipped, with a skipped-row count in
   the summary output.

## Known limitations

- No authentication on the endpoint (out of scope for a 4-hour build;
  in production this would sit behind the org's gateway).
- In-memory cache only — a restart loses the ingested state.
- Single-process; no concurrency handling beyond the stdlib server.

## What I'd do next (with more time)

1. Persist ingestion state (SQLite) so restarts are safe.
2. Add a `--since` filter for incremental summaries.
3. Structured logging + a `/health` endpoint, per the deployment
   checklist in our monitoring lesson.

## Repo layout

    tickets/          # the package: ingest.py, aggregate.py, api.py
    tests/            # pytest suite for ingestion + aggregation
    data/             # small sample CSV (2KB), not real customer data
    README.md         # you are here

Before / after: the same candidate, two READMEs

The examples below are composites. The "before" is what a capable engineer writes when they treat the README as an afterthought. The "after" is the same project, documented with the template. Same code — different grade, because the evidence of judgment is different.

Before — the afterthought README:

# tickets

ticket summary thing

run: python main.py

What's missing, from the grader's chair: Which Python version? What arguments does main.py take — a file path? a port? What does "run" even produce? What did the candidate decide about the malformed rows that are definitely in the test data? The grader must now reverse-engineer all of this from the code, on a clock, and every minute spent guessing is a minute not spent appreciating the good parts. Silence reads as carelessness, even when the code is fine.

After — the same project, scored

  • How to run is copy-pasteable from a fresh clone — the grader's first five minutes are frictionless.
  • Assumptions are numbered and reasoned. The grader can disagree with assumption #1 (maybe they'd drop the rows) — but they can see the reasoning, which is exactly what the job requires: defending a call to a customer who disagrees.
  • Known limitations pre-empt the "but there's no auth" objection by showing it was a scope decision, not an oversight.
  • "What I'd do next" is ordered by value, which reads as a mini-roadmap — the grader sees product thinking, not just code.

Notice what the strong README never does: it never apologizes ("sorry this is rough"), never blames the brief ("the instructions were unclear"), and never invents completeness ("fully production-ready"). Confidence without absolutes — the same editorial standard as the rest of this curriculum.

Part 3: Handling ambiguity

Here the curriculum's DNA pays off directly. From Stage 1, this course has drilled one reflex: clarify the ask before you build (Scoping & Success Metrics; Discovery). The take-home is the same skill under harsher conditions: you cannot ask the customer a follow-up question, so you must do the next best thing — state your assumptions explicitly, document the tradeoffs, and never silently guess.

Why "never silently guess"? Because a silent guess has two failure modes and no upside. If you guess right, the grader can't tell you reasoned it out — it looks like luck. If you guess wrong, it looks like you didn't notice the ambiguity at all. An explicit assumption, by contrast, is gradable evidence: the grader sees the gap, sees your reasoning, and can evaluate the quality of your judgment even when they'd have chosen differently.

The working method, compressed:

  1. Spot the gap. During the 0:00–0:30 read, mark every sentence in the brief that could mean two things. If a sentence makes you pause, that's a gap.
  2. Pick the most defensible reading — the one a reasonable operator would expect — and note the alternative you rejected and why.
  3. Write it in the README's Assumptions section, numbered, with a one-line rationale each.
  4. Isolate it in the code. Put the ambiguous decision behind a small function or a named constant, so a different reading is a one-line change. Graders notice this; it reads as "designed for the customer to overrule me," which is exactly the FDE posture.

Three composite ambiguity scenarios

Each scenario below is a composite based on common patterns in take-home briefs. For each: the vague brief line, what a weak response looks like, and what a strong candidate does.

Scenario 1: "Handle large files efficiently"

The gap: "Large" is undefined — 10MB? 10GB? — and "efficiently" could mean memory, speed, or both. The brief gives no numbers and no environment.

Weak response: Loads the whole CSV into a list of dicts, adds a comment saying # efficient enough, and moves on. Or the opposite failure: builds a chunked parallel processing pipeline with worker pools for a problem that might involve a 2MB file — impressive machinery, unjustified by anything in the brief.

Strong response: Streams the file row by row (constant memory, one pass), states the assumption — "Assumed 'large' means bigger than RAM is comfortable with; streaming handles anything from KB to GB with the same code. If 'efficiently' meant latency under a deadline, I'd add chunked parallelism — noted in 'What I'd do next.'" — and isolates the ingestion in one function so the strategy is swappable. The grader sees a measured response to an unmeasured requirement.

Scenario 2: "The API should be intuitive"

The gap: "Intuitive" is in the eye of the beholder. REST purists, RPC pragmatists, and CLI-first operators will all read this differently.

Weak response: Picks whatever style the candidate used last, with no explanation. The grader, who prefers a different style, reads it as a coin flip.

Strong response: Chooses one convention and documents the choice: "Assumption: 'intuitive' means a resource-oriented endpoint (GET /summary) returning JSON with stable field names, since the consumer is likely another script or a dashboard. Rejected a CLI-only design because the brief says 'exposes an endpoint.'" Field names are documented with an example response in the README, so "intuitive" is verifiable, not asserted. Bonus signal: the example response in the README matches the actual output byte-for-byte — graders do check.

Scenario 3: "Include appropriate error handling"

The gap: "Appropriate" for whom? The end user, the operator, the developer integrating the API? Every audience wants different errors.

Weak response: Wraps everything in try/except Exception: pass, or lets raw tracebacks leak to the API consumer. Both are evidence the candidate didn't think about who reads the error — a lesson this curriculum covers in Debugging Like a Detective.

Strong response: Separates audiences explicitly: malformed input rows are logged to stderr with line numbers and counted in the summary (operator audience); the API returns a clean 422 with a short, safe message (consumer audience) — never a traceback, never an internal path. The README's assumptions section names the policy in one line: "Errors are classified by audience: operators get details in logs, API consumers get safe summaries." That one line is worth more than a page of exception handlers.

Later — when you're on the job

This is the exact muscle FDEs use with customers: the brief is always vague, the customer always means something specific, and the cost of silently guessing is a demo that misses. The take-home is a rehearsal for the discovery conversations in Discovery — except here your "customer" is a grader who can only read what you wrote down.

Part 4: Repo structure template

Graders form an impression of your engineering habits within seconds of seeing the file tree. The template below is deliberately boring — and boring is the point. It mirrors the packaging conventions from Packaging, Envs, Logging & CLIs: a real package, a real test suite, pinned dependencies, and nothing committed that shouldn't be.

take-home-tickets/
├── README.md              # the graded cover letter (see Part 2)
├── requirements.txt       # pinned: requests==2.32.3, pytest==8.3.4
├── .gitignore             # .venv/, __pycache__/, *.pyc, .env, data/*.csv
├── tickets/               # the actual package
│   ├── __init__.py
│   ├── ingest.py          # CSV reading: streaming, row validation
│   ├── aggregate.py       # summary computation: pure functions, easy to test
│   └── api.py             # thin HTTP layer: routes call aggregate, nothing else
├── tests/
│   ├── __init__.py
│   ├── test_ingest.py     # happy path + malformed rows + empty file
│   └── test_aggregate.py  # counts, edge cases (no tickets, ties)
├── data/
│   └── sample_tickets.csv # tiny sample (a few KB) so the grader can run it
└── NOTES.md               # optional: your 0:00-0:30 scratch design notes,
                           # left in deliberately — shows your thinking

Three rules that separate strong submissions from sloppy ones:

  1. No secrets, ever. If your code needed an API key during development, it comes from an environment variable, and .env is in .gitignore. Committing a key — even a "test" key — is the fastest way to fail a take-home at a security-conscious company. The Secrets & Security Basics lesson exists for exactly this reason.
  2. Small, honest commits. git log --oneline should read like a work diary: scaffold package layout, ingest CSV with streaming reader, add summary endpoint, handle malformed rows + tests, write README. One giant initial commit at 11:58 p.m. tells the grader nothing about your process — and the process is part of what's being graded.
  3. The sample data is tiny and obviously fake. A 40MB CSV in the repo, or real customer-looking data, raises eyebrows. A 2KB sample with names like "Test User" signals you thought about what you're shipping.

NOTES.md deserves a word: leaving your scratch design notes in the repo is a deliberate, slightly unusual move — and that's why it works. It shows the grader your 0:00–0:30 thinking verbatim: the gaps you spotted, the options you weighed, the things you chose not to build. It's the closest a take-home gets to letting the grader watch you think.

Part 5: What graders actually score

A note on framing, because this curriculum doesn't invent data: there is no published survey of "how companies grade take-homes" cited here, and you should be skeptical of anyone who claims there is. What follows is a composite pattern — what strong candidates' submissions tend to demonstrate, mapped to the signal each artifact sends. Use it as a checklist for your own submission, not as a claim about any specific company's rubric.

ArtifactSignal it sendsWhat "strong" looks like
Tests (tests/)Judgment about what mattersA few focused tests on core logic + at least one unhappy path. Test names read like specifications. (The pytest lesson is the playbook.)
READMECommunication to strangersCopy-pasteable run instructions, numbered assumptions with rationales, honest limitations, ordered next steps.
Git historyProcess under a time budgetSmall, message-meaningful commits in a sensible order — evidence you planned, built, hardened, documented.
Repo structureEngineering habitsBoring, conventional layout; pinned dependencies; .gitignore present; no secrets, no junk, tiny sample data.
Edge-case handlingCare for the real worldEmpty input, malformed rows, and duplicates handled or explicitly listed as known limitations — never silently ignored.
Scope disciplinePrioritizationCore done well; extras absent or explicitly deferred. No second interface, no dashboard, no auth system in a 4-hour brief.
Error messagesEmpathy for the next readerErrors classified by audience; safe messages for consumers, details in logs; no raw tracebacks leaking.

Myth: "The best take-home is the most impressive one."

Impressive is a gamble: it bets the grader shares your taste in impressive. Legible is a strategy: it bets the grader can recognize good judgment when it's written down. Strong candidates optimize for legibility — every decision labeled, every gap acknowledged, every tradeoff priced. The grader may disagree with a call and still advance you, because disagreeing with a well-reasoned call is a conversation, and conversations are what interviews are for.

Productionize: your reusable take-home checklist

The curriculum's "productionize" step means turning a one-off effort into a repeatable system. Here is yours — a checklist to run before every take-home submission, in order:

  1. Fresh-clone test: clone into a new directory (or a new terminal with a clean env) and follow your README verbatim. If any step needs something not in the README, the README is wrong — fix the README.
  2. Secret scan: git log -p | grep -iE "api[_-]?key|secret|token|password" and eyeball the hits. Check .env is ignored and untracked.
  3. Assumption audit: re-read the brief once more and check every "hmm" moment has a numbered assumption in the README. Unwritten assumptions are silent guesses.
  4. Commit hygiene: git log --oneline reads as a sensible story; no WIP, no fix without context, no 200-file "cleanup" commit.
  5. Scope check: for every file and feature, ask "did the brief ask for this, or did I want to show off?" Cut or document.
  6. The stranger test: read your README as if you've never seen the project. Would you know what it does, how to run it, and what the author decided — in five minutes? If not, rewrite.

Bridge to the rest of Stage 6

The take-home gets you to the interview loop; the loop itself is a different performance. Pair this playbook with the Get Hired page for the full arc — and when the take-home involves a live system to operate rather than just code to write, the deployment instincts from Health Checks, Logs & Monitoring are what keep your demo alive under a grader's click.

Field check: explain it to a human

  1. You receive a take-home brief that says "build a dashboard for our metrics." It doesn't say which metrics, what "dashboard" means (web page? CLI table? PDF?), or who the audience is. You have four hours and cannot ask questions. Walk through exactly what you would do in the first 30 minutes, and write the three README assumptions you would commit to.
  2. Your take-home is done with 45 minutes to spare. List two things you would spend the time on and two things you would explicitly NOT add, with a one-line justification for each.
  3. A friend's take-home submission runs perfectly but has a one-line README ("run: python main.py") and a single git commit. Name three specific, checkable improvements from this post's playbook and explain what signal each one sends to the grader.
How a strong candidate answers

1. First 30 minutes: read the brief twice, marking every ambiguous phrase; decide the most defensible reading of "dashboard" for an unknown audience (a local web page with a table + a JSON endpoint, since both a human and a script can consume it); write down what I will NOT build (auth, persistence, real-time updates); sketch the data flow in NOTES.md. Three committed assumptions: (a) "Metrics" means the columns present in the provided sample data — if the sample has timestamp/value/category, those are the metrics, and I'll say so; (b) "Dashboard" means a single-page summary view plus a JSON API, chosen because the audience is unknown and this serves both humans and scripts; (c) "Our metrics" are assumed to be small enough to aggregate in memory, with streaming ingestion so the code doesn't care either way. Each gets a one-line rationale and the ambiguous choice is isolated behind one function.

2. Spend the time on: (a) the fresh-clone README test — the highest-leverage 15 minutes in the whole playbook; (b) two or three more tests on the unhappy paths, since tests are cheap judgment-evidence. Explicitly NOT: (a) a second interface (CLI and web) — scope discipline is the signal, and a second surface doubles the grader's confusion; (b) Docker/packaging ceremony the brief didn't ask for — impressive-looking, unrequested, and a classic "didn't budget time" tell.

3. (a) Expand the README with copy-pasteable run instructions — signal: communication to strangers, and it removes the grader's first friction point. (b) Split the work into small, message-meaningful commits — signal: process under a time budget; checkable via git log --oneline. (c) Add a "Known limitations + what I'd do next" section — signal: honest prioritization; turns every missing feature from a silent gap into a documented scope decision.