You don't need to master Python this weekend. You need to read a client's data, answer a stakeholder's question, and write a script someone else can run on Monday morning. This is the 20% of Python that does 80% of field work.

Prerequisite: you've seen variables, loops, and functions in at least one programming language. If this is your first time coding, start with our Python Primer (coming soon) — this crash course moves fast on purpose.

Thursday, 4:47 PM. Tom emails you a file called 311_export_october.csv — 84,000 rows of service requests. "Can you tell me which complaint types are up in Queens this month? I need it for tomorrow's borough meeting." Maria is CC'd. Dev, the data engineer who would normally do this, is on vacation.

You have a terminal, a CSV file, and tonight. This post is the Python you need.

Why Python is the FDE's field language

You'll meet engineers who argue about languages the way people argue about sports teams. In the field, the question is simpler: which language gets you unstuck fastest inside a client's environment? More often than any other, the answer is Python.

Most popular APIs have Python SDKs — and even when they don't, Python makes raw HTTP requests easy. Every CSV, JSON file, and database has a Python story. Python runs across Windows, macOS, and Linux, it's common in data and engineering environments, and the analyst or junior dev you hand your script to can probably read it. It isn't the fastest language or the most elegant — it's the one that turns "Tom needs an answer by tomorrow" into a solved problem by tonight.

This stage treats Python as a working field language: scripting, data, APIs, tests, and code someone else can deploy. Not computer science — operations.

Setup: four commands, then move on

Check what you have, then isolate your work in a virtual environment so you never break a client's system Python:

python3 --version        # want 3.10 or newer
python3 -m venv .venv
# macOS / Linux:
source .venv/bin/activate
# Windows PowerShell:
.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip

That's enough for this post. Post 8 covers packaging, environments, and logging properly. The one judgment rule that matters now: never install packages into system Python on a machine you don't own. A virtual environment is a ten-second habit that prevents a very awkward call with the client's IT team.

You've seen code — here's the translation

If you've written JavaScript, Java, or any C-family language, Python mostly asks you to unlearn punctuation. The logic is the same; the ceremony is gone:

You know thisPython's version
// comment# comment
{ } blocksIndentation — 4 spaces, always
null / undefinedNone
console.log(x)print(x)
===== (and is means identity — don't worry about it yet)
true / falseTrue / False — capitalized

The indentation point deserves emphasis because it bites everyone once: indentation isn't style in Python, it's syntax. The block under an if or a for is defined by being indented four spaces. Mix tabs and spaces and Python will refuse to run your program — which is actually a kindness, because it fails loudly instead of silently doing the wrong thing.

The 20% that does 80%

1. Variables and types — Python figures it out, you stay responsible

borough = "Queens"        # str
open_cases = 12408        # int
avg_days = 6.4            # float
escalated = False         # bool

Python is dynamically typed: you don't declare types, and a variable can hold anything. The practical rule is simple — Python won't stop you from putting the wrong thing in, so check what's actually in your data. Half of field debugging is discovering that a column you assumed was numbers is actually strings. type(x) tells you what something really is.

2. f-strings — your favorite feature, starting today

borough = "Queens"
count = 12408
print(f"{borough}: {count:,} open cases")   # Queens: 12,408 open cases

An f before the quotes lets you drop expressions straight into strings. The :, formats numbers with commas — the kind of small touch that makes output Maria can read without squinting. You'll use f-strings for messages, filenames, log lines, and API URLs for the rest of your career.

3. Lists and dicts — the only two data structures you need this week

top_types = ["Noise", "Parking", "Sanitation"]   # a list: ordered
top_types.append("Heat/Hot Water")
print(top_types[0])        # Noise — indexing starts at 0

row = {"borough": "Queens", "complaint_type": "Noise", "days_open": 6}
print(row["borough"])      # Queens — lookup by key

A list is an ordered bag of things. A dict is a labeled bag — keys map to values, exactly like a JSON object, which is why dicts are everywhere once you start touching APIs. If you only internalize two data structures this weekend, make it these two.

4. Loops and conditionals

for ctype in top_types:
    if ctype == "Noise":
        print(f"Escalating: {ctype}")
    elif ctype == "Parking":
        print(f"Watching: {ctype}")
    else:
        print(f"Normal: {ctype}")

for x in things iterates directly over items — no index juggling. The in operator also tests membership: "Queens" in boroughs. Read it like English and you'll rarely be wrong.

5. Functions — name the idea, reuse it

def summarize(borough, count):
    """Return a one-line summary a stakeholder can read."""
    return f"{borough}: {count:,} open cases"

print(summarize("Queens", 12408))

The triple-quoted string under def is a docstring — a one-line habit that pays off the first time Dev opens your script six months later and needs to know what a function does without reading its body.

6. Comprehensions — read them now, write them soon

# the loop version
upper = []
for ctype in top_types:
    upper.append(ctype.upper())

# the comprehension version — same thing, one line
upper = [ctype.upper() for ctype in top_types]

You don't need to write these fluently yet. You need to read them, because every Python codebase you'll inherit is full of them.

7. Reading files — the bridge to everything after this

import csv

with open("311_export_october.csv", encoding="utf-8") as f:
    for row in csv.DictReader(f):
        print(row["borough"], row["complaint_type"])

Three things to notice. with open(...) closes the file for you, even if something crashes mid-loop. csv.DictReader turns each row into a dict keyed by the header row — so row["borough"] reads like the data, not like column positions. And encoding="utf-8" is stated explicitly, because the day Tom sends you a file exported from a 2007 spreadsheet, the default guess will be wrong. (Post 2 goes deep on exactly that disaster.)

Before you code: clarify the ask

Read Tom's email once more: "which complaint types are up in Queens this month?" Up compared with what? Last month? Last October? The average of the previous six months?

Never let an ambiguous word silently become a technical assumption. "Up" is doing a lot of work in that sentence, and if you guess wrong you'll deliver a confident answer to a question nobody asked. So before writing a line of code, you message Tom:

You: "Quick check — when you say 'up,' do you want October compared with September, or with the same month last year?"

Tom: "For tomorrow, just give me the top complaint types in Queens. We'll do the trend comparison next."

Now the requirement is precise:

Input: the October 311 export
Filter: Queens
Output: top 10 complaint types by count
Deadline: tomorrow morning

FDE skill #1 isn't Python. It's knowing which question you're actually answering. The code comes second.

Fix it together: Tom's question, answered

With the clarified requirement in hand, let's build it. The script reads all 84,000 rows, counts complaint types in Queens, prints the top 10, and saves a summary file Tom can forward to the borough meeting.

import csv
from collections import Counter

KNOWN_BOROUGHS = {"MANHATTAN", "BROOKLYN", "QUEENS", "BRONX", "STATEN ISLAND"}

def count_by_type(path, borough):
    """Count complaint types for one borough in a 311 CSV export."""
    counts = Counter()
    total = 0
    bad_borough = 0
    with open(path, encoding="utf-8") as f:
        for row in csv.DictReader(f):
            total += 1
            cleaned = row["borough"].strip().upper()
            if cleaned not in KNOWN_BOROUGHS:
                bad_borough += 1   # blank, misspelled, or garbage — measure it, don't guess
                continue
            if cleaned == borough.upper():
                complaint = row["complaint_type"].strip()
                counts[complaint] += 1
    return total, bad_borough, counts

def write_summary(path, borough, total, bad_borough, counts):
    with open(path, "w", encoding="utf-8") as f:
        f.write(f"311 summary — {borough}, October\n")
        f.write(f"{total:,} total rows, {sum(counts.values()):,} in {borough}\n")
        f.write(f"{bad_borough:,} rows ({bad_borough/total:.1%}) had a blank or unrecognized "
                f"borough and were excluded\n\n")
        for ctype, n in counts.most_common(10):
            f.write(f"{n:>6,}  {ctype}\n")

# Only run this when we execute queens.py directly.
# We'll explain why this matters below.
if __name__ == "__main__":
    total, bad, queens = count_by_type("311_export_october.csv", "Queens")
    write_summary("queens_summary.txt", "Queens", total, bad, queens)
    print(f"Loaded {total:,} rows, {sum(queens.values()):,} in Queens.")
    print(f"Skipped {bad:,} rows with blank/unrecognized borough ({bad/total:.1%}).")
    print("Wrote queens_summary.txt")

Counter is a dict specialized for counting — counts[x] += 1 works even the first time it sees x. The .strip().upper() normalization handles trailing spaces and mixed case ("queens " becomes "QUEENS"), and anything that still doesn't match a known borough — blank cells, "Qeens", garbage — gets counted as bad data instead of silently vanishing. That count becomes the data-quality caveat in your summary. Run it:

$ python queens.py
Loaded 84,213 rows, 12,408 in Queens.
Skipped 2,526 rows with blank/unrecognized borough (3.0%).
Wrote queens_summary.txt

And queens_summary.txt now holds the top 10 — an artifact Tom can attach to an email without ever seeing your code. Notice what you send Tom: two sentences and the top three complaint types, not the script. The script is how you got the answer; the answer is what he asked for.

Break it: reading tracebacks like a local

Your script will crash. Everyone's does. Python's error messages — tracebacks — are unusually helpful once you learn the ritual: read from the bottom up. The last line names the actual complaint; the lines above show the path your code took to get there.

ErrorWhat it really meansThe usual fix
IndentationErrorA block isn't indented consistently — often mixed tabs and spacesConvert to 4 spaces; configure your editor to insert spaces for Tab
NameError: name 'borogh' is not definedA typo, or the variable was never assignedCheck the spelling; check it was assigned before this line ran
TypeError: can only concatenate str to strYou mixed text and numbers, e.g. "total: " + 12408Use an f-string instead of +
KeyError: 'Borough'The dict has no such key — the CSV header is probably borough, lowercase, or has a trailing spacePrint row.keys() once and look at what's actually there
FileNotFoundErrorPython looked for the file relative to where you ran the script, not where the script livescd to the right directory, or pass the full path

The KeyError row is the most FDE-flavored bug on the list. The data is never quite what the header promises — trailing spaces, renamed columns, a BOM character at the start of the file. When a lookup fails, don't argue with the data. Print the keys and look.

Production-safe: the habits that separate a script from a liability

Tom's script works. Here's what keeps a useful script from becoming tomorrow's production problem:

Wrap the work in functions. When Dev returns from vacation and wants your counting logic inside the real pipeline, she should be able to import your function — not copy-paste from a script that also prints things and writes files on import. The if __name__ == "__main__": guard at the bottom means "only run this part when the file is executed directly, not when it's imported."

Don't hardcode paths. "311_export_october.csv" sitting in the middle of your logic is fine for tonight; next week it becomes an argument or a constant at the top of the file, so nobody has to hunt through code to point the script at November's export.

Fail with a message a human can act on. When file handling can fail — and it can — catch it at the boundary and say what happened:

try:
    total, bad, queens = count_by_type("311_export_october.csv", "Queens")
except FileNotFoundError:
    print("Could not find 311_export_october.csv — run this from the folder holding the export.")
    raise

Note the raise at the end: you explained the problem in human terms, then still let the program exit loudly. Swallowing errors silently — the bare except: pass — is how pipelines produce wrong numbers for three weeks before anyone notices.

The deeper rule: only catch an exception when you can add useful context or recover from it. Otherwise, let Python show the original error. Catching everything "just in case" teaches exactly the wrong instinct — a noisy traceback you can read beats a silent wrong answer every time.

Print what you're doing. Loaded 84,213 rows, 12,408 in Queens is the cheapest observability there is. When this script runs on a schedule next month, those lines are how you'll know it did its job — or didn't.

Explain it to the customer

The lesson template asks for this every time, so practice it now. Tom doesn't want a code review. He wants to trust the number he's carrying into a meeting:

"Tom — I went through all 84,213 rows in the October export and pulled out the 12,408 Queens requests. The top three complaint types were Noise (2,314), Illegal Parking (1,876), and Sanitation (1,402). The full top-10 list is in the attached summary. One caveat: 2,526 rows (3.0%) had a blank or unrecognized borough value, so I excluded them from the Queens count rather than guessing — happy to dig into those if you want."

That last sentence is doing real work. You disclosed a data-quality caveat before anyone could discover it, stated what you did about it, and offered a next step. That's what makes a stakeholder trust your numbers twice.

Must know

  • f-strings, lists, dicts, for/if, functions — the working core
  • Reading a CSV with csv.DictReader and encoding="utf-8" stated explicitly
  • Reading tracebacks bottom-up; the last line is the real complaint
  • Counter for counting things — reach for it before hand-rolled dict logic

Useful later

  • Comprehension fluency — read them everywhere before writing them
  • pathlib for file paths that work on every OS
  • defaultdict, enumerate, zip — the stdlib's greatest hits

Don't memorize this

  • Every string and list method — look them up; that's what docs are for
  • Decorator syntax, metaclasses, the GIL debate — not field work
  • Which Python version added which syntax — write for 3.10+ and move on

Where this lands in CityOps

This script is the seed of Milestone 2. Right now the flow is "Tom emails a CSV, you run a script." Over the next posts, that flow grows up: Post 2 teaches you to survive the encodings and broken files clients actually send; Post 3 moves the data into SQL; and by the milestone, "Tom emails a CSV" has become a scheduled pipeline that fetches, cleans, validates, stores, and reconciles the 311 feed — with every record accounted for.

Save queens.py. You'll recognize its grandchildren.

Field check

  1. Modify the script to count complaints for a different borough, passed as a command-line argument (sys.argv). What breaks if the borough name has different capitalization in the data?
  2. Tom's next export renames complaint_type to ComplaintType and your script fails with a KeyError. Would you (a) automatically normalize all headers, (b) explicitly support a small list of known aliases, or (c) fail loudly? What are the tradeoffs?
  3. Why is the if __name__ == "__main__": guard worth keeping even in a 30-line throwaway script?
  4. Tom originally asked which complaint types are "up." Why couldn't the script as written answer that question? What additional data or clarification would you need?
What good answers look like

1. Nothing breaks — the .strip().upper() normalization on both sides handles "queens", "QUEENS", and "Queens " alike. That's the point of normalizing at the boundary: the mess gets cleaned once, in one place. 2. Start with (c) fail loudly — a crash is an alert, and a pipeline that silently maps an unfamiliar column to the wrong meaning is more dangerous than one that stops. Graduate to (b) explicit aliases once you've seen the same rename twice: it's predictable, reviewable, and documents the drift you've actually observed. Avoid (a) blind auto-normalization — "resilient" parsing that guesses wrong produces confident wrong numbers. And tell Tom the format drifted, so the next export doesn't surprise either of you. 3. Because throwaway scripts have a way of becoming permanent. The guard costs one line and means Dev can import your function into the real pipeline later without the script executing on import. Write every script as if someone will reuse it — someone will. 4. "Up" is a comparison, and the script only ever sees one month. You'd need a comparison period — September's export, last October's, or a rolling average — plus agreement on what "up" means: raw count, share of total volume, or percentage change? The clarification question to Tom was the real fix; no amount of code resolves ambiguity the stakeholder hasn't settled.