Your pipeline runs. It pulls the data, reconciles the counts, logs everything. Then Lisa from security reads the deployment notes and asks the question nobody asked yet: "Who exactly is 'cityops-pull-user' — and where is its password?" This week: the auth zoo. API keys, OAuth2 for machines, and what "enterprise login" actually means when a customer says it.

Assumes: Post 6 (fetching APIs), Post 8 (secrets live in env). Everything runnable below is stdlib-only and runs on localhost — no accounts, no keys, no network.

Tuesday, 9:40 AM. Lisa: "I audited the vendor integration. Your pipeline authenticates as cityops-pull-user with a shared password. Where is the password stored?"

You: "In the deployment notes. The file Tom has."

Lisa: "So: in a document, readable by everyone who deploys. When someone leaves the team, how do we revoke it?"

Dev (arriving with worse news): "Also, the vendor is deprecating API keys on the 1st. Everything moves to OAuth2. Can we just put the new password in the script?"

You: "Nobody is awake at 6 AM to type a password into anything. Whatever we build has to log itself in."

Before you code: clarify the ask

You: "Lisa — when you say 'enterprise login,' what has to be true for you to sign off?"

Lisa: "Three things. Every action maps to an identity I can audit. Every credential can be revoked without a code deploy. And no human's password ever does a machine's job."

You: "Dev — the vendor's OAuth2: does the pipeline get its own identity, or does it borrow someone's?"

Dev: "Its own. They call it a service account: cityops-pipeline, with a client ID and secret."

You: "And Maria and Tom — when the operator console exists, they log in how?"

Lisa: "With their company accounts. The same login they use for email. That's the part 'enterprise login' means."

Input: a vendor API moving to OAuth2 on the 1st, and Lisa's audit finding on the shared password
Output: the pipeline authenticates as its own service account with short-lived tokens; humans use the company login; every secret lives in env and can be rotated
Deadline: the vendor cutover

Notice the shape of the ask: it's two different problems wearing one word ("auth"). The pipeline needs a machine identity that works at 6 AM with no human present. Maria and Tom need human identities tied to the company directory. Confuse the two and you get exactly what Lisa found — a human-shaped credential (a password) doing a machine's job.

The minimal concept

Three ideas. The third one is the post.

Authentication is "who are you"; authorization is "what are you allowed to do." A password proves identity. A token usually carries permission. OAuth2 has "auth" in the name but it's mostly an authorization framework — a standard way to hand out scoped, expiring permission slips without ever sharing the password. That confusion causes half the bad designs you'll see.

The zoo has three animals, and each has a habitat.

MethodWho it's forThe shape of it
API keysSimple scripts and integrationsA long random string in a header; the server looks it up
OAuth2 client credentialsService-to-service — the pipeline's caseClient ID + secret buy a short-lived token; the token does the work
SSO / OIDCHumans"Log in with your company account"; the app never sees the password

Machines get service accounts and short-lived tokens; humans get SSO. A password sitting in a script is a human credential doing a machine's job — it can't be scoped tightly, it never expires on its own, and revoking it means editing code. The fix isn't a better hiding place for the password. The fix is a credential designed for machines in the first place.

runnable conceptual — sections below are tagged. Runnable ones ran on this machine; their outputs are pasted from real runs. Conceptual ones are the parts no laptop can demo (a company's identity provider), taught with tables and diagrams instead.

Build it, part 1: API keys without the leak conceptual

An API key is the simplest credential in the zoo: a long random string, sent in a header, that the server maps to your account. The mechanism is trivial. The discipline is everything.

The right shape, all of it from Post 8's rule — secrets live in env, never in code:

import os

api_key = os.environ["VENDOR_API_KEY"]  # KeyError on a misconfigured box: loud, not silent
headers = {"X-API-Key": api_key}
$ VENDOR_API_KEY=demo-key-123 python3 keydemo.py
header set, key length: 12
$ python3 -c "import os; os.environ['VENDOR_API_KEY']"
Traceback (most recent call last):
  File "<string>", line 1, in <module>
KeyError: 'VENDOR_API_KEY'

The KeyError is the point: a missing secret fails at startup, loudly, instead of sending an unauthenticated request halfway through the night's run. Same instinct as Post 5's boundaries — reject bad input at the edge, with noise.

How API keys actually die in practice — Lisa's greatest-hits list, from real audits:

The leakWhy it happens
Committed to git"Just this once, I'll rotate it later." Git history never forgets; later never comes.
Pasted in a chat or ticketDebugging at midnight; the ticket system keeps it forever.
In the URL query stringURLs land in server logs, browser history, and proxies. Headers are the only home.
Screenshared in a demoThe terminal with the key is visible for four seconds. Four seconds is enough.
One key for everythingNo scopes, no per-environment keys — so revoking it is an outage, and nobody revokes it.

The honest summary: API keys are fine for what they're for — simple integrations, low blast radius — and the vendor is deprecating them anyway. The pipeline is moving to the next animal.

Build it, part 2: the token dance runnable

OAuth2 client credentials is the machine-to-machine flow. The choreography, exactly as the vendor's docs will describe it:

6 AM: pipeline wakes up, no human present
  → POST /oauth/token  "I'm cityops-pipeline, here's my secret"
  ← {"access_token": "eyJ...", "token_type": "Bearer", "expires_in": 300}
  → GET /api/readings  with  Authorization: Bearer eyJ...
  ← 200, the data
5 minutes later the token is worthless. That's the feature.

The secret is used once per token, over the wire, to buy a short-lived token — and then the token does the work. If a token leaks, it dies on its own in minutes. If the secret leaks, you revoke it at the vendor and issue a new one: no code deploy, which is exactly what Lisa asked for.

To make this runnable without any vendor account, I built the vendor: a tiny token server on localhost (client-credentials grant, HTTP Basic for the client secret, 3-second tokens so expiry is observable), and the client the pipeline would actually run — token cached, refreshed automatically on 401:

class TokenClient:
    """Machine-to-machine client: token cached, refreshed on expiry/401."""
    def __init__(self, base_url, client_id, client_secret):
        self.base_url = base_url
        self.client_id = client_id
        self.client_secret = client_secret
        self._token = None
        self._expires_at = 0.0

    def _fetch_token(self):
        data = urllib.parse.urlencode(
            {"grant_type": "client_credentials"}).encode()
        creds = base64.b64encode(
            f"{self.client_id}:{self.client_secret}".encode()).decode()
        req = urllib.request.Request(
            self.base_url + "/oauth/token", data=data, method="POST",
            headers={"Authorization": "Basic " + creds})
        try:
            with urllib.request.urlopen(req, timeout=5) as r:
                body = json.load(r)
        except urllib.error.HTTPError as e:
            # add context the caller can't get from the bare status code
            detail = json.load(e)
            raise RuntimeError(
                f"token request failed: {detail.get('error')}") from e
        self._token = body["access_token"]
        # renew a little early: never use a token in its last breath
        self._expires_at = time.time() + body["expires_in"] - 0.5
        return self._token

    def get_token(self, force=False):
        if force or time.time() >= self._expires_at:
            return self._fetch_token()
        return self._token

    def get(self, path):
        def _call(tok):
            req = urllib.request.Request(
                self.base_url + path,
                headers={"Authorization": "Bearer " + tok})
            with urllib.request.urlopen(req, timeout=5) as r:
                return json.load(r)
        try:
            return _call(self.get_token())
        except urllib.error.HTTPError as e:
            if e.code != 401:                       # not an auth problem: let it fly
                raise
            # 401: token may have died server-side. Refresh once, retry once.
            return _call(self.get_token(force=True))

Two things to notice before the output. First, the except blocks follow the exception rule: the token fetch catches only to translate a bare status into a message with the vendor's error attached; the 401 handler catches only to recover — refresh and retry once. A non-401 error re-raises untouched. Second, the client renews tokens half a second early, because a token that expires mid-request is a 401 you scheduled yourself.

The real run — server and client both on this machine, tokens living 3 seconds so you can watch them die:

$ python3 auth_demo.py
1) token fetched: jINgF5Ko... (type=Bearer, expires_in=3s)
   server has issued 1 token(s)
2) GET /api/readings -> 200, 2 readings
3) second call reuses the cached token; server has still issued 1 token(s)
4) waiting 3.5s for the token to die...
   401 hit, token refreshed, GET -> 200, 2 readings; server has now issued 2 token(s)
5) wrong client secret -> token request failed: invalid_client
6) HMAC: valid=True tampered_body=False replayed_1h_later=False

Read the story in the numbers: one token fetched, reused for the second call (the server's issue counter stays at 1 — caching works), then the token dies, the 401 fires, the client refreshes on its own, and the call succeeds. Line 5 is the impostor: a wrong secret gets invalid_client, translated from a bare 401 into a sentence. (Line 6 is the next section's output — keep reading.)

One loud caveat, because the demo would be dishonest without it: this runs over plain HTTP on localhost. Tokens on the wire are credentials. In production this entire dance happens over HTTPS or it doesn't happen — a bearer token sent over HTTP is a password shouted across the room.

Build it, part 3: proving a request wasn't tampered with runnable

Tokens prove who's calling. Sometimes you also need to prove nobody rewrote the message — that's HMAC request signing, and it's the mechanism behind signed webhooks (a later lesson in this stage leans on it, so learn it here). Both sides share a secret; the sender hashes the request with it and attaches the signature; the receiver recomputes and compares:

import hashlib, hmac, time

def sign_request(secret: bytes, method: str, path: str, body: str, ts: int) -> str:
    msg = "\n".join([str(ts), method.upper(), path, body]).encode()
    return hmac.new(secret, msg, hashlib.sha256).hexdigest()

def verify_request(secret, method, path, body, ts, sig, *, max_skew=300):
    if abs(time.time() - ts) > max_skew:     # stale: possible replay
        return False
    return hmac.compare_digest(sign_request(secret, method, path, body, ts), sig)

The timestamp is part of the signed message, which kills two attacks at once: change the body and the signature won't match; capture a valid request and replay it an hour later and the skew check rejects it. hmac.compare_digest instead of == — a timing-safe comparison, so an attacker can't guess the signature one byte at a time by measuring response times.

Real run, from the same demo script (this was line 6 of the output above):

6) HMAC: valid=True tampered_body=False replayed_1h_later=False

The untampered request verifies. Flip one character of the body — rejected. Replay it an hour later — rejected. Signing doesn't encrypt anything (the body is still readable); it makes forgery evident. Different job from the token, complementary to it.

Build it, part 4: what "enterprise login" actually means conceptual

No laptop can demo a company's identity provider, so this part is conceptual — but it's the part Lisa cares about most. When a customer says "enterprise login," they mean this flow:

Maria clicks "Log in with company account" in CityOps
  → CityOps redirects her to the company's identity provider
  → she authenticates THERE (password, MFA, whatever IT mandates)
  → the provider redirects back with an ID token: "this is maria@city.gov"
  → CityOps verifies the token's signature, starts its own session
CityOps never sees Maria's password. Ever.

OIDC (OpenID Connect) is the standard vocabulary for that conversation; SSO (single sign-on) is the experience. Two tokens, two jobs — the distinction that trips people up:

TokenAnswersExample
ID tokenWho is this human?"maria@city.gov, from the city's identity provider"
Access tokenWhat may they do?"read the dashboard, approve escalations, until 5 PM"

And the table Lisa would tape to the wall — users vs service accounts:

Human user (Maria)Service account (cityops-pipeline)
Proves identity byCompany login (SSO/OIDC), MFAClient ID + secret, from env
Credential lifetimeSession hours; password rotated by IT policyTokens live minutes; secret rotated on schedule
Revoke whenShe leaves — IT disables one directory accountOne API call at the vendor; no code deploy
Must neverShare the login, or use it from a scriptBe a human's password in a config file

That last row is the whole lesson in one line. Lisa's audit finding wasn't "the password was in the wrong file" — it was a category error: a human-shaped credential doing a machine's job, in a place no human's credential should ever be.

Break it, three ways

1. The secret in the repo. The failure Lisa actually found: the vendor password sitting in the deployment notes, readable by everyone, revokable by nobody without a code change. The demo's line 5 shows the server side of the fix — a wrong secret doesn't get a polite error page, it gets invalid_client, and the client translates the bare 401 into a sentence instead of crashing with a traceback at 6 AM. But the deeper fix is structural: with client credentials, the secret lives in env on the server (Post 8), and rotation is an ops task, not a code change. Lisa's test is always the same: "someone leaves the team — show me revocation." If the answer involves opening an editor, you failed.

2. The token that died at 6:00:03. Here's what happens to a client without the refresh logic — the naive version that fetches a token once and trusts it forever. Real run, token TTL 3 seconds:

$ python3 naive_client.py
naive client holds token txxQbPvT...
GET with dead token -> 401 Unauthorized

That's the 6 AM page: the pipeline fetched its token, did some work, the token expired mid-run, and every subsequent call 401s. The resilient client's answer (built above): cache with an expiry margin, and on 401 refresh once and retry once. Notice what's not in the resilient client: retrying forever, or catching the 401 and continuing without data. One refresh, one retry, then the error surfaces loudly.

3. The forged webhook. Back to line 6 of the demo output: tampered_body=False. Someone — or some proxy, or some "helpful" middleware — rewrites the payload in transit, and without signing the receiver happily processes the forged version. With HMAC, the signature stops matching and the request is rejected. And replayed_1h_later=False: even a byte-perfect copy of a legitimate request is useless after the skew window. Signing is cheap; the class of bugs it eliminates is not.

Productionize: the auth runbook

Six habits that turn the demo into something Lisa signs off on:

Secrets from env, always. Post 8's rule, now with a security reason: the client ID may appear in logs and configs (it identifies, it doesn't authenticate), but the client secret appears nowhere except the environment of the machine that needs it. Different secrets per environment — the staging secret must not open production.

Short-lived tokens, scoped narrowly. Minutes, not days. And ask the vendor for the smallest scope that works (readings:read, not admin): a leaked token with a narrow scope is a contained incident; a leaked token with admin scope is a breach. This is the same least-privilege instinct as Foundations Post 3's "what you'll refuse to build" — decide the boundary before you're asked to cross it.

Rotate on a schedule, not on a scare. Quarterly secret rotation, practiced until it's boring. The drill matters more than the interval: the team that has rotated twice can do it at 2 PM on a Tuesday; the team that never has will do it at 2 AM during an incident.

Never log the token. Log about auth, never the credentials: "token refreshed, expires in 300s" is observability; the token string in the log file is a credential store you didn't mean to build. Maria's Monday-morning log reading (Post 8) must never include secrets.

Audit the failures. Every invalid_client and every 401 spike goes to the log with a timestamp. A wrong secret at 3 AM is either a misconfigured deploy or someone else's script trying your credentials — either way you want to know. The DLQ instinct from Post 2 applies to auth events too: rejected attempts are evidence, kept and counted.

HTTPS or nothing. The demo ran on localhost HTTP because there's no network to snoop. The vendor integration runs over TLS, with the token endpoint's certificate verified — a bearer token over plain HTTP is the password shouted across the room, and no expiry window fixes that.

Explain it to the customer

"Lisa — the shared password is gone. The pipeline now authenticates as its own service account, cityops-pipeline: the client secret lives in the server's environment, never in code or docs. It buys 5-minute tokens, scoped to read-only, and refreshes them itself — if a token leaks, it dies on its own in minutes. Revocation is one call at the vendor, no deploy. Every failed auth attempt lands in the log with a timestamp for your audits. And the humans: when the operator console ships, Maria and Tom log in with their company accounts — the app never sees their passwords. Your three conditions: every action maps to an identity (service account for the machine, directory accounts for the humans), every credential revokes without a code change, and no human password does a machine's job anymore."

"Dev — the vendor cutover is covered. Same client handles it; the only config change is the token endpoint URL and the new client ID and secret, both in env."

The pattern, as always: her three conditions, answered in order, then Dev's one-line reassurance. Lisa doesn't need to understand the token dance — she needs to hear her audit criteria met, in her vocabulary.

Must know

  • Authentication ("who are you") vs authorization ("what may you do") — OAuth2 is mostly the second
  • API keys: simple, in a header, in env — and every way they leak (git, chat, URLs, screenshots)
  • OAuth2 client credentials: the machine flow — secret buys a short-lived token, token does the work, refresh on 401
  • Service accounts for machines, SSO/OIDC for humans — a password in a script is a category error
  • HMAC signing: prove a message wasn't tampered with or replayed; timestamp in the signed payload
  • Never log credentials; scope tokens narrowly; HTTPS or nothing

Useful later

  • Other OAuth2 grants — authorization code (the human login dance behind SSO), device flow (TVs and CLIs)
  • JWT structure — what's actually inside those dot-separated tokens, and why you verify the signature
  • mTLS — when both sides prove identity with certificates instead of shared secrets
  • Secret managers (Vault, cloud KMS) — env vars are the floor, not the ceiling

Don't memorize this

  • OAuth2 endpoint paths and parameter names — remember secret buys token, token does work, look up the rest per vendor
  • HMAC construction details — remember sign the request, check the clock, look up the code
  • OIDC claim names — remember ID token says who, access token says what, look up the fields

Where this lands in CityOps

This lesson is the identity foundation Milestone 3 builds on. The intake pipeline keeps its service account and its token dance — nothing about the nightly pull changes except that it now meets Lisa's audit. But the next milestone puts a FastAPI service in front of human operators, and that's where the two halves of this lesson meet: the service authenticates to vendors as a machine, while Maria and Tom authenticate to the service as humans via the company login, with RBAC-lite deciding what each of them may do. The users-vs-service-accounts table becomes an architecture diagram.

And the signature line: identity is part of architecture. Not a login screen bolted on at the end — a design decision, made up front, about which kind of identity every actor gets. Post 6 taught you to fetch from APIs; this lesson taught you to prove who you are when you knock.

Field check

  1. Lisa asks: "A contractor's laptop had the vendor client secret on it. Walk me through revocation." What's the answer — and what would it have been with the old shared password?
  2. Your pipeline's token lives 5 minutes; a run takes 20. What breaks if the client fetches one token at startup and never refreshes — and what's the minimal fix?
  3. A vendor offers you a single API key with full admin scope "for simplicity." Name two concrete risks, and what you'd ask for instead.
  4. The HMAC check rejects a legitimate webhook. Name two innocent causes before you suspect an attack.
  5. Maria will use the operator console; the pipeline will keep pulling at 6 AM. Which gets SSO and which gets a service account — and what's the one-sentence reason?
What good answers look like

1. With client credentials: revoke the secret at the vendor (one call, no deploy), issue a new one into the server's env, done — the blast radius is bounded by the token's 5-minute life and narrow scope. With the old shared password in the deployment notes: everyone who ever saw the file has it, so you must change it everywhere, redistribute it to everyone who needs it, and hope nobody kept a copy — revocation is a broadcast, not an operation. 2. Every API call after minute 5 returns 401 and the run fails (or worse, silently processes nothing). The minimal fix is the resilient-client pattern: cache the token with an expiry margin, and on 401 refresh once and retry once — not infinite retries, not ignoring the error. 3. Risks: (a) a leaked key grants full admin — the blast radius is the entire account, not one integration; (b) revoking it breaks everything at once, so nobody ever revokes it. Ask for scoped keys instead (read-only for the pipeline) and separate keys per environment, so staging and production fail independently. 4. Clock skew between sender and receiver beyond the skew window (the timestamp check rejects a legitimate-but-late request), or the body was modified in transit by something innocent — a proxy re-encoding JSON, whitespace normalization, a load balancer altering headers that were part of the signed payload. Check the clocks and byte-compare the received body first. 5. Maria gets SSO — she's a human, and her identity should come from the company directory with MFA and IT-managed revocation. The pipeline gets the service account — no human is awake at 6 AM to complete a login dance, and a machine needs a credential designed for machines: scoped, short-lived, rotatable without a deploy.