Six posts in, your pipeline does real work: it ingests messy files, reconciles counts, parses money to the cent, and calls live APIs with honest retries. Now Dev wants to put it on the city's scheduler — and asks the question that should make you nervous: "How do we know it still works after the next change?" This week: tests. Not testing theory — a suite that proves your pipeline's promises, in seconds.

Assumes: Posts 1–6. One install: pip install pytest.

Monday, 10:15 AM. Dev: "We're moving your pipeline to the city's scheduler on Friday. Nightly run, no human watching."

You: "Okay — what's the rollback plan if a change breaks it?"

Dev: "That's what I'm asking you. Last month a 'harmless' edit to a parsing helper silently broke a downstream report for three days. Nobody noticed until Finance called. So: how do we know your pipeline still works after the next change?"

The pipeline in question is everything from Posts 2 through 6: the quarantine loader, the money and timestamp parsers, the retrying API client. It works today. Dev wants proof it works every day — proof he can run himself, in seconds, before Friday.

Before you code: clarify the ask

You: "What does 'works' mean here? The scheduler runs it and it exits zero — is that the bar?"

Dev: "The bar is the numbers it writes are numbers we'd defend. If a change breaks that, I want to know in seconds, not on Thursday when Finance calls."

You: "So which behaviors are load-bearing? Money parsing exactness, timestamp rejection, retry behavior, quarantine counts?"

Dev: "All four. If any of those silently change, the report lies."

You: "And the proof has to be fast enough that you'll actually run it?"

Dev: "If it takes longer than making coffee, I won't."

Input: the pipeline helpers — parse_money, parse_ts (Post 5), fetch_json (Post 6), the quarantine loader (Post 2's instinct)
Output: a test suite Dev can run in seconds, plus a one-line confidence report
Deadline: Friday's scheduler cutover

Notice what just happened: Dev defined "works" as four specific behaviors, not a vibe. A test suite is that definition made executable — and executable definitions can be checked every day, for free.

The minimal concept

Three ideas, and the third one is the whole post:

A test is a function that asserts something about your code. def test_parse_money_exact(): followed by assert parse_money("$4,250,000.00") == Decimal("4250000.00"). If the assertion holds, the test passes; if not, pytest shows you exactly what differed. assert is the entire assertion API for now — no frameworks-within-frameworks.

pytest finds and runs them. Any file named test_*.py, any function named test_* — pytest collects them all and runs them with one command. No test runner to configure, no boilerplate class to inherit.

A green suite is a set of promises your code keeps in public. Each passing test is a promise: "money parsing is exact," "naive timestamps are rejected," "the quarantine path isolates bad rows." When someone edits a helper next month, the suite re-checks every promise in seconds. That's the answer to Dev's question — not "I think it works," but "here are the promises, and here's them holding, right now."

Build it: the first promises

Layout first — two files side by side, convention over configuration:

cityops/
├── cityops.py        # the pipeline helpers (Posts 2, 5, 6)
└── test_cityops.py   # the promises
$ pip install pytest
$ python -m pytest test_cityops.py -q

Start with the money parser, because Post 5's exactness lesson deserves a promise that never silently weakens. One test, four cases — parametrize keeps it that way instead of four copy-pasted functions:

from decimal import Decimal
import pytest
from cityops import parse_money

@pytest.mark.parametrize("raw,expected", [
    ("$4,250,000.00", Decimal("4250000.00")),
    ("472222.22",     Decimal("472222.22")),
    ("$0.10",         Decimal("0.10")),
    ("  $1,000 ",     Decimal("1000")),
])
def test_parse_money_exact(raw, expected):
    assert parse_money(raw) == expected

def test_parse_money_never_touches_float():
    # the Post 5 trap, pinned down: string in, exact out
    assert parse_money("0.1") + parse_money("0.2") == Decimal("0.3")

Then the timestamp boundary — including the failure path, because a promise that only covers the happy path is Post 6's warning all over again:

from datetime import timezone
from cityops import parse_ts

def test_parse_ts_rejects_naive():
    with pytest.raises(ValueError, match="missing offset"):
        parse_ts("2026-11-01 01:30:00")

def test_parse_ts_normalizes_to_utc():
    ts = parse_ts("2026-11-01T00:30:00-04:00")
    assert ts.tzinfo == timezone.utc
    assert (ts.hour, ts.minute) == (4, 30)

pytest.raises is how you promise about errors: "this input must fail, and fail this way." A boundary that silently accepts bad input is worse than no boundary — the test makes the loudness load-bearing.

And the quarantine path, with a fixture — because tests need sample data without littering the repo with CSVs. tmp_path is pytest's built-in fixture: a fresh temporary directory per test, cleaned up afterward:

def load_rows(path):
    """Read a CSV of raw rows into dicts. No cleaning here."""
    with open(path, newline="") as f:
        return list(csv.DictReader(f))

def quarantine_rows(rows):
    """Split rows into (good, quarantined). Bad money or bad
    timestamps go to quarantine: kept, counted, never trusted."""
    good, bad = [], []
    for r in rows:
        try:
            amount = parse_money(r["amount"])
            ts = parse_ts(r["created_at"])
        except (KeyError, ValueError):
            bad.append(r)
        else:
            good.append({**r, "amount": amount, "created_at": ts})
    return good, bad

@pytest.fixture
def complaints_csv(tmp_path):
    p = tmp_path / "complaints.csv"
    p.write_text(
        "id,amount,created_at\n"
        "1,$12.50,2026-09-20T08:00:00-04:00\n"
        "2,$9.99,2026-09-20 08:00:00\n"   # naive timestamp: quarantined
    )
    return p

def test_quarantine_end_to_end_from_csv(complaints_csv):
    good, bad = quarantine_rows(load_rows(complaints_csv))
    assert len(good) == 1 and len(bad) == 1
    assert len(good) + len(bad) == 2   # every record accounted for

That last assertion is Post 2's reconciliation instinct, executable: good + quarantined must equal everything that came in. A row that vanishes silently is the failure mode this whole pipeline was built to prevent.

Break it, two ways

Tests can lie, too. Here are the two lies you'll meet first — I wrote both on purpose, ran them, and kept the receipts.

1. The test that costs more than the code. The obvious way to test the weather call is to actually call it. I wrote that test with a 5-second stand-in for the network to show you the price:

import time

def test_with_network():
    time.sleep(5)   # stand-in for one real API call
    assert True
1 passed in 5.01s

It passes. It also takes five seconds — per test, per run, plus flakiness whenever the API hiccups, plus you're now rude to a free API on every save. A suite Dev won't run is a suite that doesn't exist. Network tests don't belong in the fast suite; the next section shows what goes in their place.

2. The test that passes for the wrong reason. Read this one carefully:

def test_totals_look_right():
    totals = {"2026-09-20": 8839}
    # ...forgot the assert. Spot the bug.
1 passed in 0.01s

Green. And it proves nothing — there's no assertion, so there's no promise. pytest can only check what you assert; a test without an assert is a promise you never made. This is the cheapest bug in testing and one of the most common: the fix is one line, assert totals["2026-09-20"] == 8839, and the discipline is reading your tests the way you'd read a contract — what, exactly, is being promised here?

Productionize: mock the network, trust the suite

The replacement for the network test is a scripted fake: an object with the same shape as requests.Session that plays pre-written responses instead of touching the network. The test controls the weather — including bad weather:

class FakeResponse:
    def __init__(self, status, headers=None, body=None):
        self.status_code = status
        self.headers = headers or {}
        self._body = body
    def raise_for_status(self):
        if 400 <= self.status_code < 600:
            raise requests.HTTPError(f"{self.status_code} error")
    def json(self):
        return self._body

class FakeSession:
    """A stand-in for requests.Session. Each get() plays the next
    response in the script, so the test controls the network."""
    def __init__(self, script):
        self.script = list(script)
        self.calls = 0
    def get(self, url, params=None, timeout=None):
        self.calls += 1
        item = self.script.pop(0)
        if isinstance(item, Exception):
            raise item
        return item

def test_fetch_json_recovers_after_503s():
    session = FakeSession([
        FakeResponse(503, {"Retry-After": "1"}),
        FakeResponse(503, {"Retry-After": "1"}),
        FakeResponse(200, {}, {"rain_mm": 10.9}),
    ])
    assert fetch_json(session, "https://x.test/data") == {"rain_mm": 10.9}
    assert session.calls == 3

def test_fetch_json_gives_up_after_tries():
    session = FakeSession([requests.ConnectionError("no route")] * 5)
    with pytest.raises(RuntimeError, match="failed after 2 tries"):
        fetch_json(session, "https://x.test/data", tries=2)
    assert session.calls == 2

def test_fetch_json_does_not_retry_a_400():
    session = FakeSession([FakeResponse(400, {}, "bad params")])
    with pytest.raises(requests.HTTPError):
        fetch_json(session, "https://x.test/data")
    assert session.calls == 1   # your bug, not the network's: no retry

These are Post 6's retry promises, now executable: recovery works, giving up works, and a 400 is never retried. The network is no longer a dependency of the suite — it's a script the test directs.

And now the part I didn't plan. When I wired these helpers into the suite, it failed — on the very first run:

E       decimal.InvalidOperation: [<class 'decimal.ConversionSyntax'>]
cityops.py:25: InvalidOperation
=========================== short test summary info ============================
FAILED test_cityops.py::test_quarantine_splits_good_and_bad - decimal.Invalid...
1 failed, 11 passed in 2.84s

parse_money("not money") raised decimal.InvalidOperation — which is not a ValueError, so quarantine_rows didn't catch it. In production, that malformed row would have crashed the pipeline instead of being quarantined. The suite found a real bug in under three seconds, before Friday, before Finance called. That's the entire post in one traceback.

The fix follows the exception rule — catch only to add context:

from decimal import Decimal, InvalidOperation

def parse_money(s):
    """Money string in -> exact Decimal out. Floats never touch this path."""
    try:
        return Decimal(s.strip().replace("$", "").replace(",", ""))
    except InvalidOperation:
        # translate the low-level error into the boundary's language,
        # with the offending input attached
        raise ValueError(f"not a money amount: {s!r}")

Then the whole suite, for real:

............                                                            [100%]
12 passed in 2.89s

Twelve promises, all kept, faster than making coffee. That's what Dev runs on Friday.

Explain it to the customer

"Dev — the pipeline's promises are now executable. Twelve tests, all passing, under three seconds: money parsing exactness, naive-timestamp rejection, retry recovery and surrender, the 400-is-not-retried rule, and the quarantine path with full reconciliation. Run pytest -q any time — if it's green, the four load-bearing behaviors hold. And it already earned its keep: the suite caught a real crash-on-bad-row bug while I was writing it, before the scheduler ever saw it."

The pattern by now: the verdict, then what it covers, then the invitation to verify it yourself. "Run it any time" is doing quiet work — confidence you can re-check beats confidence you have to take on faith.

Must know

  • pytest discovers test_*.py files and test_* functions — assert is the assertion API
  • @pytest.mark.parametrize: many cases, one test, no copy-paste
  • Fixtures (tmp_path and your own): test data without repo litter
  • pytest.raises: promise about failure paths, not just happy paths
  • Mock the network with scripted fakes — real APIs don't belong in the fast suite
  • A test with no assert proves nothing

Useful later

  • Coverage: measuring how much of your code the suite actually touches
  • Markers (-m "not slow"): keeping a slow integration tier separate from the fast suite
  • Running the suite on every push — the day your tests guard the main branch, not just your laptop
  • Property-based testing (Hypothesis): the machine invents the edge cases for you

Don't memorize this

  • Every pytest CLI flag — remember -q and -v, look up the rest
  • Fixture scopes — the default (fresh per test) is right until slow setup proves otherwise
  • Mocking libraries — a hand-written fake teaches the idea; reach for unittest.mock when fakes get tedious

Where this lands in CityOps

The suite joins the pipeline as its proof. In Milestone 2, pytest -q runs before every scheduled pull — green means the promises hold, red means the pipeline stops before it writes numbers nobody would defend. The quarantine test's reconciliation assertion becomes the nightly "every record accounted for" check, and the retry tests guard the enrichment step Post 6 built. Post 4's principle was if you can't rerun it, you didn't clean it — the suite is what makes "rerun" mean something.

And the signature line for this one: a test is a promise your code keeps in public. Post 6 said don't trust the happy path — the suite is how you stop trusting it and start verifying it, every day, in under three seconds.

Field check

  1. Your suite hits the real weather API and takes 40 seconds. Name two changes you'd make.
  2. A test passes but contains no assert. What does it prove, and what's the one-line fix?
  3. The quarantine test asserts len(good) + len(bad) == 2 instead of only checking the good rows. Why?
  4. The retry tests use a FakeSession instead of the real network. What exactly is being tested — and what isn't?
  5. Dev asks, "Can I just run the suite before every deploy?" What's your one-line answer — and what would make you nervous about saying yes?
What good answers look like

1. First, replace the network call with a scripted fake (like FakeSession) so the test is fast and deterministic. Second, if you still want an occasional real-API check, mark it slow (@pytest.mark.slow) and keep it out of the default run with -m "not slow" — the fast suite stays fast. 2. It proves nothing — pytest can only check what you assert. The fix is adding the assertion (and the discipline is re-reading each test as a contract: what, exactly, is promised here?). 3. It's the reconciliation promise: good + quarantined must equal everything that came in. Checking only the good rows would let a silently-dropped row pass the suite — the exact failure mode the quarantine path exists to prevent. 4. What's tested: fetch_json's logic — retry counting, backoff, giving up, the 400 rule — under fully controlled conditions. What isn't: the real network, the real API's actual responses, TLS, DNS, or whether the endpoint still exists. The fake tests your code's decisions; only a (slow, separate) integration check tests the outside world. 5. "Yes — that's exactly what it's for." What should make you nervous: saying yes unconditionally. The suite proves the twelve promises it contains, nothing more. If Dev deploys a change to behavior the suite doesn't cover — a new column, a changed schedule, a different API endpoint — green means nothing about that change. The honest answer names the coverage boundary.