This is the post where the talking stops and the system starts. By the end, you will own a real repository, a real page on the real internet with real HTTPS, real city data in your terminal — and the two documents every FDE engagement lives or dies by.
Friday, 4:40 PM. You are sharing your screen with Maria, the operations director. On the call: Dev, the skeptical data engineer, and Tom, the borough manager who thinks in spreadsheets.
Your screen shows a plain, dark status page at https://cityops.yourname.github.io (or your own domain): a headline number, a short table of the newest service requests, and a timestamp that says when the numbers were pulled. Nothing fancy. It loads fast. The little lock icon sits next to the address.
Maria leans in. "Is that real data?"
"Pulled live from the city's 311 feed this morning," you say. "Same data your team stares at in spreadsheets — but this refreshes itself."
Dev squints. "From that API? The one with the garbage dates?"
"We will get to the dates," you say. "I wrote down exactly what we can and cannot trust about that feed. It is in the contract."
Maria smiles for the first time in six weeks. "Okay," she says. "What now?"
Rewind three weeks. All you had was a vague email, a messy public API, and a team that did not agree on what they needed. This milestone is the story of turning that into the page Maria is looking at. Let us build it.
What you are building (and why each piece matters)
A milestone is a checkpoint with a working system and written proof, not a pile of exercises. Here is your build list, and — because FDEs always know the "why" — what each piece proves to a customer:
| You build | What it proves |
|---|---|
| A clean GitHub repo | You can organize work other humans can read, clone, and trust |
| Dev / staging / prod environment discipline | You will never again experiment on the thing the customer sees |
| A tiny status page served over HTTPS | You can put something on the real internet, securely, today |
| A live 311 API pull from your terminal | You can touch real customer-adjacent data and say what it means |
| A discovery note and a data contract v0 | You write things down — the difference between a hacker and an engineer |
On the FDE loop, this milestone walks the first four links:
Milestone 1: Discover and Scope (the two documents), then Build and Deploy (the repo and the page). Adopt and Learn come later.
Step 1 — Create the repo with a clean structure
A repository is just a folder whose history Git tracks. A clean structure means a stranger can open it and guess where everything lives in ten seconds. Here is yours:
$ mkdir cityops && cd cityops
$ git init -b main
$ gh repo create cityops --public --source=. --push
Now lay out the structure. Every folder earns its place by answering one question:
cityops/
├── README.md # what this is, how to run it, who it is for
├── docs/
│ ├── discovery-note.md # Milestone deliverable 1
│ └── data-contract-v0.md # Milestone deliverable 2
├── site/
│ └── index.html # the status page
├── scripts/
│ └── pull_311.sh # the exact API pull, saved so it is repeatable
├── .env.example # placeholder for secrets (never real values)
└── .gitignore # keeps secrets and junk out of git
Commit it like you mean it — a customer may read this history:
$ git add .
$ git commit -m "Milestone 1: scaffold repo, status page, API pull scripts"
$ git push -u origin main
Step 2 — Set up dev, staging, and prod
An environment is simply a place your code runs, with its own settings. You need three, because the cardinal sin of client work is experimenting on the thing the customer sees:
| Environment | Where it lives | Who sees it | Used for |
|---|---|---|---|
| Dev | Your laptop (python3 -m http.server in site/) | Only you | Trying things, breaking things |
| Staging | The staging branch, deployed to a second Pages URL | You + Dev | Testing exactly what prod will serve |
| Prod | The main branch, deployed to the public URL | Maria, Tom, everyone | The truth the customer relies on |
Two rules make this real instead of theater. Rule one: secrets never enter git. A secret is any credential — API keys, passwords, tokens. You do not need one for this milestone (the 311 API is public), but the habit starts now: put a template in .env.example, the real values in .env, and .env in .gitignore.
$ cat .gitignore
.env
__pycache__/
$ cat .env.example
# copy to .env and fill in real values; .env is never committed
EXAMPLE_API_KEY=your-key-here
Rule two: nothing reaches prod except through a merge. You work on a branch, push it, look at it on the staging URL, and only then merge to main:
$ git checkout -b staging
$ git push -u origin staging
# ... review on the staging URL, fix what looks wrong ...
$ git checkout main
$ git merge staging
$ git push
Step 3 — Deploy the status page with HTTPS
HTTPS is HTTP with encryption plus identity: the "S" means the connection is scrambled so nobody in the middle can read it, and a certificate — a small cryptographic document issued by a trusted authority — proves the server is who it claims to be. Browsers show the lock icon only when both hold.
The fastest honest path to HTTPS for a static page is GitHub Pages: push an index.html, flip one switch, and GitHub issues and renews the certificate for you. Keep the page tiny and honest — a headline, a table, a timestamp:
<!DOCTYPE html>
<html lang="en">
<head><meta charset="utf-8"><title>CityOps Status</title></head>
<body>
<h1>CityOps — 311 Operations Status</h1>
<p>Newest service requests, pulled live from NYC Open Data.</p>
<p>Last updated: 2026-10-02 09:15 ET (refresh: daily)</p>
<table>
<tr><th>ID</th><th>Type</th><th>Borough</th><th>Status</th></tr>
<tr><td>70591261</td><td>Noise - Street/Sidewalk</td><td>Bronx</td><td>In Progress</td></tr>
</table>
</body>
</html>
Deploy it:
$ git add site/index.html
$ git commit -m "Add CityOps status page v0"
$ git push origin main
# then: repo Settings → Pages → Deploy from branch → main → Save
Verify like an engineer, not a hope-and-pray clicker. curl (the command-line tool from Post 5 for making raw web requests) shows you the truth:
$ curl -sI https://yourname.github.io/cityops/ | head -3
HTTP/2 200
content-type: text/html; charset=utf-8
strict-transport-security: max-age=31536000
200 means success. The https:// in the URL plus that response means your page is encrypted in transit. Maria can open it on her phone in a taxi. That is what "deployed" means at this stage.
Step 4 — Your first live 311 API pull
The 311 dataset lives behind a public API — a URL you can ask for data instead of a human. The pattern: the base address names the dataset, and ?$limit=5 is a query parameter (an instruction tacked onto the URL) saying "give me just 5 records." Run it:
$ curl -s 'https://data.cityofnewyork.us/resource/erm2-nwe9.json?$limit=5' \
| python3 -c "import json,sys; [print(r['unique_key'],'|',r['complaint_type'],'|',r['status'],'|',r['created_date']) for r in json.load(sys.stdin)]"
70591261 | Noise - Street/Sidewalk | In Progress | 2026-10-01T02:05:23.000
70598315 | Noise - Street/Sidewalk | In Progress | 2026-10-01T02:04:49.000
70592880 | Noise - Residential | In Progress | 2026-10-01T02:04:13.000
...
That is real city data, pulled seconds ago. Save the exact command so it is repeatable — repeatability is the whole game — in scripts/pull_311.sh:
#!/bin/bash
# First live pull: 5 newest 311 service requests (public API, no key needed)
curl -s 'https://data.cityofnewyork.us/resource/erm2-nwe9.json?$limit=5' -o data/sample_311.json
echo "Saved $(python3 -c "import json; print(len(json.load(open('data/sample_311.json'))))") records to data/sample_311.json"
Now look at one full record — this is the raw material everything in CityOps will be built from. Read it slowly; every field is a future decision:
{
"unique_key": "70591261",
"created_date": "2026-10-01T02:05:23.000",
"agency": "NYPD",
"agency_name": "New York City Police Department",
"complaint_type": "Noise - Street/Sidewalk",
"descriptor": "Loud Music/Party",
"incident_zip": "10452",
"incident_address": "963 WOODYCREST AVENUE",
"city": "BRONX",
"borough": "BRONX",
"status": "In Progress",
"latitude": "40.831669762919574",
"longitude": "-73.9287544719845",
"open_data_channel_type": "PHONE"
}
A real complaint: loud music on Woodcrest Avenue in the Bronx, phoned in at 2:05 AM, still in progress. Notice three things already. One: latitude is a string, not a number — you will have to convert it before doing math. Two: the field names are decent but the vocabulary is the city's, not yours ("Noise - Street/Sidewalk" is free text, not a fixed category). Three: Dev is right to be suspicious of created_date — we will deal with that below. Writing these observations down is exactly what the data contract is for.
Deliverable 1 — The discovery note (filled example)
A discovery note is a one-page write-up of what you learned before writing code: who the stakeholders are, what they actually asked for (in their words), what the real problem is underneath, and what you will and will not build. Post 2 taught the method — stakeholder interviews, the 5 Whys, separating symptoms from problems. Here is that template, filled in for CityOps. This is a real artifact: commit it as docs/discovery-note.md.
Discovery Note — CityOps Milestone 1
Engagement lead: you. Date: 2026-09-25. Attendees: Maria (Director of Operations), Dev (Data Engineer), Lisa (Security), Tom (Borough Manager).
In their words. Maria's opening email: "We are drowning in unresolved 311 requests. I think AI could help us get ahead of them. Can you build something?" Tom, same meeting: "I just need my borough's numbers in Excel every Friday." Dev: "Before anyone builds anything, that 311 feed has garbage in it. The dates do not always make sense." Lisa: "Whatever you build, complaint text does not leave our building."
Problem statement. Operations staff cannot see today's 311 queue in one trusted place, so assignments lag and requests age. The 5 Whys: crews are assigned late because nobody sees the queue building until Friday's spreadsheet; reporting is manual because nobody trusts the feed enough to automate it. Milestone 1 delivers a single live status view they can check every morning.
Users. 6 triage operators across 3 borough offices (daily users); Maria, Director of Operations (success owner). Tom, borough manager, consumes the weekly Excel summary.
Current workflow. 311 feed → weekly raw export → Tom's spreadsheet → operators sort by received date → oldest requests handled first → overdue items discovered by complaint, not by process.
Constraints. Complaint text never leaves the building (Lisa); created_date is unreliable — quarantined, never used for aging or ordering (Dev); Tom's weekly Excel habit is non-negotiable; six-week timeline, deadline will not move.
Explicitly out of scope. No AI triage, no predictions, no Excel export builder, no per-borough drill-downs, no complaint-text processing (Lisa's constraint stands until the security review in Stage 5).
Success metric. Maria opens the status page at least 3 mornings in the first week and can state the current queue size without opening a spreadsheet.
Open questions. (1) How fresh must "live" be — daily or hourly? (2) What counts as "unresolved" — everything not Closed, or only Open + In Progress? (3) Who owns the page after handoff?
Risks. The 311 feed's dates are suspect (Dev's flag — quarantined in the data contract). Tom will ask for Excel before we are ready (interrupt planned for below).
Notice what the note does: it turns "build something with AI" into a concrete, small, measurable first step — and it writes down the refusals. The refusals are the most valuable sentences on the page.
Deliverable 2 — Data contract v0
A data contract is a written promise about the data you will rely on: which fields exist, what type each one is, how fresh the data is, and exactly how it can fail. "v0" means first draft — honest, incomplete, and already useful. Commit it as docs/data-contract-v0.md.
What the 311 API promises (verified 2026-10-02)
| Field | Type | Example | Our promise |
|---|---|---|---|
unique_key | String of digits | "70591261" | Present on every record; treat as the ID |
created_date | ISO 8601 string | "2026-10-01T02:05:23.000" | FLAGGED — see open issue 1. Do not use for aging or ordering yet |
complaint_type | String | "Noise - Street/Sidewalk" | Always present in our sample; free-form city vocabulary, not a fixed list |
status | String | "In Progress" | Values seen: Open, In Progress, Closed. Assume others exist |
borough | String | "BRONX" | Sometimes blank — counts must handle missing |
latitude / longitude | String (not number!) | "40.831669762919574" | Cast to float before any math; blank on some records |
incident_zip | String | "10452" | Keep as string (leading zeros matter) |
Freshness. New records appear throughout the day; the city refreshes the open dataset at least daily. Our promise to Maria: numbers on the page are never older than 24 hours, and the page always shows exactly when they were pulled.
Failure modes (what can break, and what we do).
| Failure | How we would see it | What we do |
|---|---|---|
| API returns HTTP 429 (rate limit — "slow down") | curl exits with an error body instead of JSON | Wait 60 seconds, retry once, then serve yesterday's numbers with a "stale" banner |
| API is down entirely | curl times out | Keep the last good page live; timestamp tells the truth |
| Schema change (a column disappears or renames) | Our pull script errors on a missing key | Script fails loudly (never silently); fix the field list; note the change in the contract |
Records missing borough or created_date | Counts do not add up | Exclude from that count, log how many, show "N records excluded" on the page |
Open issue 1 — created_date reliability (Dev's flag). Dev noticed timestamps that do not line up with when records appear in the feed. Until we investigate in Stage 2's pipeline work, the contract quarantines this field: we display it, we never compute ages or orderings from it. A contract that admits what it does not know is more trustworthy than one that pretends.
The demo checkpoint — your 10 minutes with Maria
A stakeholder demo is not a tour of your code; it is a 10-minute story that ends with the customer knowing what changed and what happens next. Rehearse it out loud. Here is the script:
Minutes 0–2: the hook. Open the live page. Say: "This is this morning's 311 queue, live. Before today, getting this view took Tom's Friday spreadsheet. Now it is a bookmark." Let the headline number land.
Minutes 2–6: the tour (three clicks, no more). Scroll the newest-requests table. Point at one real row: "Loud music, Woodcrest Avenue, Bronx, phoned in at 2 AM, still in progress — that is a real request from this morning." Then the timestamp: "And here is when we pulled it, so you always know how fresh this is."
Minutes 6–8: the trust part. Open the data contract. Say: "Here is what I trust about the feed and what I do not — yet. Dev flagged the dates, so I quarantined them instead of building on them. You will always see the caveats in writing." Watch Dev nod. That nod is worth more than any feature.
Minutes 8–10: the close. One slide, three lines: what shipped (repo, page, live pull), what is next (the intake pipeline in Stage 2), one ask: "Maria, will you open this three mornings next week and tell me if the number matches your gut? That is our success metric."
Then stop talking. Demos die in minute 11.
Client interrupt 1 — Tom wants borough numbers "by Friday"
It is Wednesday. This email lands:
Subject: quick add to the status page?
Hi — the page looks great. Could you add a section with my borough's numbers broken out, plus a download button for Excel? I need it for the Friday leadership meeting. Should be quick, right? — Tom
The interrupt. A reasonable-sounding request, a hard deadline, and a hidden iceberg: "broken out" means per-borough aggregation logic, and "download button for Excel" means a file-generation feature with formatting decisions — neither is quick, and both expand the milestone's scope two days before the demo.
Your decision. Ship the smallest useful slice: a simple per-borough counts table on the page (five rows, computed by hand from the API with a $select query — no new infrastructure). Cut the Excel download to Stage 2's pipeline milestone, where exports belong. Renegotiate, don't refuse: give Tom something real for Friday while protecting the milestone.
The customer communication — written out, because writing it is the skill:
Subject: Re: quick add to the status page?
Hi Tom — glad the page is useful. Here is what I can do for Friday:
I will add a borough breakdown table to the status page by Thursday evening — one row per borough with open request counts, pulled from the same live feed. You can screenshot it or print the page straight from the browser for the leadership meeting.
The Excel download button is a bigger build than it looks (file formatting, what "current" means when data refreshes, who gets which rows), so I am moving it into the next milestone where we build the data pipeline properly — that way the export is generated from clean, reconciled data instead of a one-off hack. I will confirm the date with you once Stage 2 is scoped.
If the borough table needs different columns for Friday, tell me by Thursday noon and I will adjust. — You
What this email does: says yes to the outcome (numbers for Friday), no to the implementation (the download button), explains why in one sentence a non-engineer can follow, and moves the cut item to a named future instead of a vague "later." Tom gets his meeting; the milestone stays intact.
Client interrupt 2 — Dev says created_date is unreliable
Thursday morning, in the team channel:
Dev: heads up — I pulled 500 records and checked the dates. A bunch of them have closed_date BEFORE created_date — the request supposedly closed before it was even created. I would not build anything on those columns.
The interrupt. Your most skeptical stakeholder just handed you a data-quality finding that undermines any "oldest requests" or "average age" feature you might have imagined. He is also testing you: will you argue, or will you listen?
Your decision. Believe him — he did the work. Quarantine created_date in the data contract (done above: displayed, never used for aging or ordering). Do not try to fix the city's data; that is not your job and not your milestone. Promise a proper investigation in Stage 2, when the pipeline milestone adds reconciliation and can quantify exactly how bad it is. The cost of this decision: no "oldest unresolved" sorting in Milestone 1. The benefit: every number you show Maria is defensible.
The customer communication — two messages, because Dev and Maria need different things:
To Dev (same channel, within the hour): "Good catch — thank you for checking. I have quarantined created_date in the data contract: we display it but nothing is computed from it. In the next milestone I will run the full reconciliation and quantify how many records are affected, and we will decide together whether to repair, replace, or drop the field. Flagged as open issue 1 so it does not get lost."
To Maria (demo, 30 seconds): "One thing Dev found: the city's timestamps are not reliable enough to rank requests by age, so I left any 'oldest first' sorting out of this version rather than show you a ranking I cannot defend. The counts and the newest-requests table are solid. We will dig into the dates properly in the next phase."
What these messages do: Dev hears that his finding changed the plan (skeptics become allies when their expertise visibly matters). Maria hears a tradeoff explained in plain language, with the cut named and the reason given. Nobody hears "the data is bad, trust me anyway."
The milestone scorecard
Milestones are graded on outcomes, not effort. Score yourself honestly — this is the same scorecard every CityOps milestone uses:
| Dimension | What "good" looks like for Milestone 1 |
|---|---|
| Technical correctness | Repo clones cleanly; page loads over HTTPS; the curl pull runs and returns real records |
| Data reconciliation | Counts on the page match a manual API check; excluded records are counted and labeled, not silently dropped |
| Reliability | The page survives an API outage (last good numbers stay up with an honest timestamp) |
| Security & compliance | No secrets in git; no complaint text processed; Lisa's constraint documented in the discovery note |
| Time-to-value | Maria sees a live number in week one, not a big-bang reveal in week six |
| Customer outcome | Maria checks the page instead of asking Tom for the spreadsheet — the success metric from the discovery note |
| Scope discipline | Something was explicitly cut or refused with rationale (the Excel button; created_date aging) |
| Communication | Tom's email answered with a yes/no/why he can forward; Maria's demo ends with one clear ask |
| Operability | Someone besides you can re-pull the data and redeploy the page from the README alone |
| Adoption | Evidence Maria opened it three mornings (ask her; write down the answer) |
Notice the deliberate lesson in the grading: the fanciest status page does not win. The page that Maria actually opens, whose numbers Dev cannot poke holes in, and whose caveats are written down — that is the A.
Must know
- Environment discipline: dev is your laptop, staging is the rehearsal, prod is the customer's truth — and code only moves forward through merges
- HTTPS in one sentence: encrypted connection plus a certificate proving the server's identity — the lock icon is a promise you can verify with curl
- Repeatability: if a command mattered, it lives in a script in the repo; "I ran something once" is not engineering
- The two documents: the discovery note (what we learned, what we refused) and the data contract (what the data promises, what it does not) — write them before you need them
- Scope negotiation: yes to the outcome, no to the implementation, with the cut item moved to a named future — in writing
Useful later
- GitHub Actions (automating the daily pull so no human runs the script — Stage 5 makes this real)
- Custom domains on Pages (the
CNAMEfile and DNS record from Post 4, applied) - Socrata's
$whereand$selectfor server-side filtering and aggregation
Don't memorize this
- The exact GitHub Pages settings clicks — they move; the principle (branch → deploy → HTTPS) does not
- The full 311 field list — the contract holds what matters; the principle (branch → deploy → HTTPS) does not
- Socrata query syntax details —
$limittoday, the docs tomorrow
Field check
- What could fail? The 311 API starts returning HTTP 429 (rate limiting) on the morning of Maria's demo. Walk through exactly what breaks and what the page shows.
- How would you detect it? You cannot stare at the page all day. What is the cheapest check that tells you the data pipeline is unhealthy before Maria notices?
- What would you tell the customer? Maria asks why the numbers did not change today. Write the two-sentence explanation you would give her — no jargon.
What a good answer looks like
1. The pull script gets an error instead of JSON, so there is nothing new to publish — but the page itself still loads, because the last good deployment is untouched. What Maria sees: yesterday's numbers with yesterday's timestamp. Nothing breaks visibly; the page just goes stale. 2. The cheapest detection is the timestamp: a tiny check (even a manual one, or a scheduled script in a later milestone) that compares "data pulled at" against "now" and alerts if it is older than 24 hours. You detect staleness, not the API error itself. 3. "The city's data feed was temporarily limiting our requests this morning, so today's numbers are yesterday's. The page shows the date the data was pulled, and I will refresh it as soon as the feed responds normally — no action needed from your team."
What comes next
You now have the scaffold: a repo with discipline, a live page, real data in your terminal, and — most importantly — written proof of what you learned and what the data promises. A junior engineer built a page. An FDE built a page plus the trust around the page.
But look honestly at what you have: the pull is manual, the numbers are hand-placed, nothing is cleaned, nothing is reconciled, and Dev's date problem is quarantined, not solved. That is exactly where Stage 2: Python & Data begins. You will turn the one-off curl into a real intake pipeline — data pulled on a schedule, cleaned with Python, stored in a database, reconciled so every record is accounted for, with the data contract upgraded from a written promise to code that rejects bad data automatically.
The scaffold holds. Now we build the machine that lives on it.