Prerequisites: The Terminal and Git, Demystified (you'll run scripts from the command line) and Data Formats and Databases in One Sitting (CSV basics). Everything else is taught here.
Every FDE spends a large part of every deployment staring at something that broke. Debugging isn't a talent — it's a procedure. Here's the procedure, practiced on a real pipeline.
The scene. Tuesday, 9:40 AM. You're embedded with the city's 311 team — the same CityOps scaffold you built in Milestone 1. Maria from the mayor's office calls: "The dashboard is broken. It was working Friday. The numbers just… stopped."
You open the intake pipeline script — the one that loads the daily 311 request CSV and feeds the dashboard. You run it. It crashes.
Your job in the next hour: find out why, fix the actual cause (not the symptom), and prove it's fixed. That is this entire lesson.
Clarify the ask: "it's broken" is not a bug report
Before touching code, pin down the symptom. "It's broken" could mean ten different things, and each one sends you down a different path. Ask:
- What exactly fails? The script crashes? The dashboard shows old numbers? The page won't load?
- When did it start? "Friday it worked" tells you to look at what changed since Friday — new data file, a deploy, a config edit.
- Can you make it happen again? A bug you can reproduce on demand is a bug you can fix. An intermittent one is a ghost story.
- What changed? New data, new code, new machine, new user input. Something changed; your job is to find what.
This is the same discipline as Discovery: separate the complaint from the problem. Maria's complaint is "the dashboard is broken." The problem, as you'll discover, is one bad row in a CSV.
The minimum concept: the debugging loop
Debugging is a loop with six steps. Memorize it, because you will run it hundreds of times:
The FDE principle for this lesson:
One principle
"The error message is a witness, not the verdict." It tells you where the program gave up and what it was choking on — not why. The why is always one step behind the message. Your job is to interrogate the witness, not to obey it.
Build: reproduce the crash
Here is the intake script, close to what Milestone 1 produced. It reads the daily 311 CSV and computes the average days a request stays open:
import csv
def load_requests(path):
rows = []
with open(path) as f:
reader = csv.DictReader(f)
for row in reader:
rows.append({
"id": row["request_id"],
"borough": row["borough"],
"days_open": int(row["days_open"]),
})
return rows
requests = load_requests("requests.csv")
print(f"Loaded {len(requests)} requests")
avg = sum(r["days_open"] for r in requests) / len(requests)
print(f"Average days open: {avg:.1f}")
And here is requests.csv — the file the city uploaded this morning:
request_id,borough,days_open,status
SR-10421,Queens,3,open
SR-10422,Brooklyn,7,open
SR-10423,Manhattan,N/A,open
SR-10424,Bronx,12,closed
SR-10425,Queens,5,open
Step one of the loop: reproduce. Run it:
$ python3 intake.py
Traceback (most recent call last):
File "intake.py", line 15, in <module>
requests = load_requests("requests.csv")
File "intake.py", line 11, in load_requests
"days_open": int(row["days_open"]),
ValueError: invalid literal for int() with base 10: 'N/A'
Good — it crashes on demand. Now step two: read the witness statement.
Read the traceback like a detective
Start with the exception message at the bottom, then walk backward through the traceback frames to understand how execution reached it. The last line is the punchline; the lines above it are the trail:
- Last line first:
ValueError: invalid literal for int() with base 10: 'N/A'— the program tried to turn the text'N/A'into an integer and gave up. That's what it choked on. - The frames above: line 15 called
load_requests("requests.csv"), and inside that, line 11 ranint(row["days_open"]). That's where. - What's missing: which row. The traceback doesn't say. The witness saw the crime but not the criminal's face. That's normal — the next step is yours.
Don't memorize
Don't start by memorizing every exception type. Learn to read the traceback first; over time you'll naturally recognize the common categories. You need the habit: read the last line, then walk up the frames. Those two places are usually your best starting point.
Hypothesize, then test small
Form one hypothesis: "some row has a non-numeric days_open." Test it with the smallest possible check — don't rewrite the loader yet, just interrogate the data:
import csv
with open("requests.csv", newline="", encoding="utf-8-sig") as f:
reader = csv.DictReader(f)
for n, row in enumerate(reader, start=2): # header is line 1
try:
int(row["days_open"])
except ValueError:
print(f"Line {n}: request_id={row['request_id']!r}, days_open={row['days_open']!r}")
$ python3 inspect_row.py
Line 4: request_id='SR-10423', days_open='N/A'
Hypothesis confirmed. Line 4 of the CSV — request SR-10423 from Manhattan — has N/A where a number should be. The dashboard "stopped" because the loader crashes on the first bad row and never finishes.
Evidence, not stories
We have identified the immediate data condition that triggered the failure: SR-10423 contains days_open=N/A. We have not proven why the source system emitted N/A — it could be manual entry, an export transformation, or an upstream default. Don't convert a plausible story into a root cause. The failure mechanism (bad value crashes the loader) is established; the upstream origin is still unknown, and that's fine to say out loud.
Notice what you did not do: you didn't guess, you didn't rewrite the whole script, and you didn't "fix" it by deleting the row from the CSV by hand (that would just hide this morning's symptom and guarantee a repeat tomorrow).
Break: the error message can point at the symptom, not the cause
That was the easy version — the witness pointed near the crime. Often it doesn't. Here are misdirections you'll meet in the field:
| Symptom you see | What it might actually be |
|---|---|
FileNotFoundError: 'requests.csv' | The script is fine — you're running it from the wrong directory. The file is next door, not here. |
ModuleNotFoundError | Wrong virtualenv, or you installed the package for a different Python. The code didn't change; the environment did. |
KeyError: 'borough' | The CSV header has a hidden BOM character or a trailing space — the column is secretly named '\ufeffborough'. |
| Dashboard shows yesterday's numbers | Don't assume it's the browser cache. The pipeline, the database, the API/cache layer, or the browser could each be stale — identify which layer is behind before you "fix" anything. |
| Works on your laptop, crashes on the server | Different Python version, missing file, different working directory. "Works on my machine" is a clue about environments, not code. |
Watch the first one happen. Same script, same files — just run from the wrong folder:
$ cd subfolder
$ python3 ../intake.py
Traceback (most recent call last):
File "../intake.py", line 15, in <module>
requests = load_requests("requests.csv")
File "../intake.py", line 5, in load_requests
with open(path) as f:
FileNotFoundError: [Errno 2] No such file or directory: 'requests.csv'
The witness says "no such file." The verdict: nothing is wrong with the file or the code — you are standing in the wrong room. When an error makes no sense, widen the interrogation: what directory am I in, which Python is this, what changed since it last worked? (This is why knowing your terminal matters — half of debugging is orienting yourself.)
Productionize: fix the cause, log the evidence
Now the fix. The root cause isn't "line 4 is bad" — it's "the loader assumes every row is clean, and real-world data never is." (You met messy real-world files in Data Formats and Databases; this is what you do about them.)
But hold on — before writing code, there's a decision here that isn't yours alone to make. Quarantining the bad row is a business policy, not just a technical fix. So you call Dev from the 311 team:
You: "One malformed row currently blocks the whole daily load. Should one invalid days_open stop publication, or should we quarantine that request, publish the remaining valid rows, and flag the incomplete input?"
Dev: "Quarantine it and publish — but make the incomplete count visible. If the dashboard pretends the data is complete, we'll get burned."
Don't invent failure policy
Don't silently invent failure policy. When bad input arrives, clarify with the stakeholder whether it should fail the batch, quarantine the record, or be repaired — before you code the handler. The technical fix is easy; the correct technical fix depends on a business answer.
With Dev's approval, fix the assumption, not the row:
import csv
import logging
logging.basicConfig(level=logging.INFO, format="%(levelname)s: %(message)s")
def load_requests(path):
rows, quarantined = [], []
with open(path, newline="", encoding="utf-8-sig") as f:
reader = csv.DictReader(f)
for n, row in enumerate(reader, start=2):
try:
days = int(row["days_open"])
except ValueError:
logging.warning("line %d: quarantined %r (days_open=%r)",
n, row["request_id"], row["days_open"])
quarantined.append(row)
continue
rows.append({"id": row["request_id"],
"borough": row["borough"],
"days_open": days})
return rows, quarantined
requests, quarantined = load_requests("requests.csv")
if not requests:
raise RuntimeError("no valid requests loaded")
print(f"Loaded {len(requests)} requests, quarantined {len(quarantined)}")
avg = sum(r["days_open"] for r in requests) / len(requests)
print(f"Average days open: {avg:.1f} across {len(requests)} valid requests; "
f"{len(quarantined)} request quarantined")
$ python3 intake_fixed.py
WARNING: line 4: quarantined 'SR-10423' (days_open='N/A')
Loaded 4 requests, quarantined 1
Average days open: 6.8 across 4 valid requests; 1 request quarantined
The newline="" and encoding="utf-8-sig" are one-line insurance: they handle Windows line endings and strip a BOM if the city's export tool adds one — the same BOM that can cause a phantom KeyError: 'borough'.
Three decisions in that fix, all deliberate:
- Quarantine, don't silently drop. The bad row is logged and kept aside — Maria can see exactly which request needs a real value. Silently skipping it would make the numbers wrong without anyone knowing.
- Log, don't print-debug-and-delete.
logging.warningstays in the code and shows up in server logs at 2 AM. Aprint("HERE!!!")you delete tomorrow helps nobody. A useful production log answers three questions: what failed, on which input or entity, and what the system did about it — our warning names the line, the request ID, the bad value, and the quarantine decision. - Catch only what you can handle. The
try/exceptwraps exactly one conversion, and the handler does something useful: records the bad row and moves on. Contrast that with the classic sin:
# Don't silently swallow broad exceptions.
try:
process(row)
except Exception:
pass # the bug is now invisible AND still there
The sin isn't the broad except itself — at a top-level worker or process boundary, catching broadly to log, report, and shut down cleanly is legitimate. The sin is the silent pass: no log, no record, no signal. If you catch broadly, you must do something visible with it.
Quarantine is not a blank check
Quarantine lets isolated bad records avoid taking down the batch. It does not mean "accept unlimited bad data." Production pipelines need an agreed threshold at which the batch fails or alerts instead of publishing. Conceptually: 1 quarantined out of 5 → publish with a warning (what we just did). 4,000 quarantined out of 5,000 → fail the batch and page someone, because the source feed is probably broken. The threshold itself is another business conversation with Dev — the engineering point is that the threshold must exist.
And one more assumption to check: what if every row is bad? The loader would quarantine everything, then crash dividing by zero:
$ python3 intake_fixed.py # with a file where all rows have N/A
WARNING: line 2: quarantined 'SR-10001' (days_open='N/A')
WARNING: line 3: quarantined 'SR-10002' (days_open='N/A')
Traceback (most recent call last):
File "intake_fixed.py", line 25, in <module>
raise RuntimeError("no valid requests loaded")
RuntimeError: no valid requests loaded
That's the if not requests: raise RuntimeError(...) guard doing its job — failing loudly instead of crashing mysteriously on ZeroDivisionError. Recovery from one failure mode can expose the next assumption; check for it deliberately.
Deliberately scoped
We're validating only the failure we've reproduced — a non-numeric days_open. The loader still trusts request_id and borough blindly. Later ingestion lessons expand this into an explicit row contract: required columns, allowed values, and reconciliation checks. "Productionize" here means this failure is handled properly, not that every failure is.
The exception rule
Only catch an exception when you can add useful context or recover from it. Otherwise, let Python show the original error. A bare except: pass doesn't fix bugs — it hides them, and hidden bugs bill you later with interest. When you do catch, add context so the next reader knows which row, which file, what value:
try:
days = int(row["days_open"])
except ValueError as e:
raise ValueError(f"bad days_open in row {row['request_id']!r}") from e
$ python3 context_demo.py
SR-10421 3
SR-10422 7
Traceback (most recent call last):
File "context_demo.py", line 8, in process
days = int(row["days_open"])
ValueError: invalid literal for int() with base 10: 'N/A'
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "context_demo.py", line 13, in <module>
for rid, days in process("requests.csv"):
File "context_demo.py", line 10, in process
raise ValueError(f"bad days_open in row {row['request_id']!r}") from e
ValueError: bad days_open in row 'SR-10423'
The original error is preserved (Python chains it with from e), and the new message names the row. That's catching with purpose.
Final loop step: verify. "It ran without crashing" is not verification. Check the specific expectations:
- The original failing input now completes end to end.
- Exactly the 4 expected valid rows loaded.
- Exactly the 1 expected row (SR-10423) quarantined, named in the log.
- The metric matches the hand calculation: 6.8 across 4 valid requests, 1 quarantined — a number with its context, not a naked average.
- The dashboard consumes the new output and shows fresh numbers.
A fix you haven't re-tested against expectations is a hypothesis, not a fix.
Communicate: the bug report Maria actually needs
You're the FDE on site — Maria needs to tell the mayor's office what happened, and your team needs to know what changed. Write it up in four lines:
- Symptom: dashboard froze; intake script crashed with
ValueErroron this morning's CSV. - Root cause: row SR-10423 contained
days_open=N/A; the loader assumed everydays_openvalue was numeric. (Source cause — why the upstream system emittedN/A— still being investigated. The failure mechanism is confirmed; the upstream origin is not.) - Fix: loader now quarantines bad rows with a logged warning instead of crashing; SR-10423 flagged for a real value. (Quarantine-vs-fail policy confirmed with Dev.)
- Verified: re-ran on today's file — 4 valid rows loaded, 1 quarantined; average 6.8 days across 4 valid requests with 1 quarantined; dashboard updating.
Notice the shape: symptom, cause, fix, proof. That's the same shape as a good problem statement — and it's what makes people trust you with the next outage.
CityOps: put it together
Back to Tuesday morning. Here's the full play in order:
- Reproduce: ran
intake.py— crashed on demand. Not intermittent, not the dashboard, not the network. - Read: bottom-up traceback —
ValueErroronint('N/A')at line 11. Witness: bad value. Missing: which row. - Hypothesize & test: one bad row — confirmed with a 10-line probe: line 4, SR-10423.
- Fix the cause: loader quarantines bad rows with logging instead of assuming clean data.
- Verify: re-ran against expectations — 4 valid rows loaded, SR-10423 quarantined and logged, metric 6.8 across 4 valid with 1 quarantined, dashboard live. Called Maria with the four-line report.
The important part wasn't speed; it was that every step reduced uncertainty. No guessing, no rewriting the world, no 2 AM heroics. That's what "experience the job before you get the job" feels like — the job is mostly this loop, run calmly.
Later in the curriculum
Stage 2's Testing with pytest turns today's manual verification into automated regression tests, so this exact bug can never silently return. Stage 3's SDK Design & Debugging Someone Else's API applies the same loop to other people's broken APIs.
Field check
- A script crashes with
KeyError: 'borough'on a CSV that looks fine. Name two things you'd check before changing the code.Reveal
The header row for hidden characters (BOM, trailing spaces) — print
repr()of the fieldnames — and whether you're reading the file you think you're reading (right directory, right filename). - Why is
except Exception: passworse than letting the program crash?Reveal
A crash tells you something is wrong, where, and with what value. A swallowed exception tells you nothing while the bad data (or bad logic) keeps flowing downstream, corrupting everything quietly.
- Your fix works, but how do you prove it to a skeptical teammate?
Reveal
Re-run the exact reproduction on the original failing input, show the before/after output, and point at the log line naming the quarantined row. "It works on my machine" is a claim; a re-run on the failing input is evidence.
- The loader receives 10,000 rows and quarantines 9,500. The script completes successfully. Did your fix work?
Reveal
Technically the loader survived — but operationally the batch is almost certainly unhealthy. Quarantine without an agreed threshold just converts a crash into silent data loss at scale. The fix needs a fail/alert threshold and a data-completeness policy, agreed with the stakeholder, before you'd call this "working."