You'll need: the env-var discipline from Secrets & Security Basics, a FastAPI app with dependencies (see Building APIs with FastAPI), and the structured-logging habit from the previous lesson. Everything below runs against a server-side session store and a real secret-store interface — no real credentials anywhere in this post.

Monitoring answers "is it broken?" This lesson answers two harder questions: "who is allowed to do what?" and "whose data is whose?" — the locks, keys, and boundaries that decide whether CityOps survives its second customer.

Thursday, 2:00 PM. Lisa from security has one job this week: the pre-pilot review before City B signs on. She opens with the easy question. "Show me where the production API key lives."

Dev hesitates. "In the .env file." A pause. "Which is... on the shared drive. And probably in the repo history from March." Lisa writes something down without looking up. "Second question. When City B signs up — can a City A user see City B's reports?"

Dev opens the codebase. The reports query filters by report ID. It does not filter by tenant, because until this meeting there was only one tenant and the filter felt like decoration. "I'd have to check every query," he says quietly.

Maria, who has been silent, speaks for the first time: "I need that answer in writing before we sign City B."

The customer problem

Stage 1 taught the team that a secret in code is a secret already shared. Production with a second customer makes that lesson load-bearing: now it's not just your key on the shared drive, it's the key to two cities' data. And the missing tenant filter isn't a style issue — it's the difference between "CityOps serves two cities" and "CityOps lets each city read the other's service reports." The contract City B is about to sign requires tenant isolation; shipping without it wouldn't just be embarrassing, it would violate the client's compliance requirement.

Three gaps, matching Lisa's two questions plus the one she asks next:

  • Secrets with no home. API keys and database passwords live in a .env file on a shared drive, with copies in shell history and repo history. Nobody can say who has read them, and nobody can rotate one without a scavenger hunt.
  • Permissions by convention. "Admins can do admin things" is enforced by which buttons the frontend shows. The API itself checks nothing — anyone with a token can call any endpoint.
  • Data with no boundaries. Every query assumes a single tenant. The day City B's rows land in the same tables, every unscoped query becomes a cross-tenant read.

Clarify the ask

Dev turns Lisa's interrogation into three work orders, because "do security" is a wish, not a plan:

  • "Give every secret a home." — secrets live in a secret store, are loaded once at startup, and never appear in code, logs, or chat. Rotation is a procedure, not an archaeology expedition.
  • "Write down who can do what — in code." — a permission model the API enforces on every request: roles, resources, actions, deny by default.
  • "Prove whose data is whose." — tenant isolation in the data layer, with a test per query that fails loudly if a filter goes missing.

Out of scope, deliberately: single sign-on, and the full identity-provider universe. CityOps authenticates API callers with bearer tokens against a server-side session store — enough to teach the authorization and isolation patterns honestly. The patterns below survive the move to an identity provider; the token-lookup function is the only seam that changes.

Which brings us to the principle. Last lesson's principle was about knowing you're broken. This one's about the other half of production trust:

One principle

"Permissions are code too." If the permission model lives in a wiki page, in frontend button visibility, or in Dev's head, it isn't enforced — it's hoped for. Roles, resources, and tenant boundaries belong in versioned, reviewed, tested code, exactly like the business logic they protect.

The minimum concept: secrets, roles, and tenants

Three ideas, each one sentence, then the ways teams blur them:

A secret is a credential; config is everything else. The database host is config — it can sit in a file, in the repo, on a slide. The database password is a secret — it lives in the secret store, it's loaded at startup, and it never appears in code, logs, or error messages. The test is simple: would reading this value let someone impersonate us? If yes, it's a secret. (Stage 1's lesson, now with production consequences.)

RBAC answers "may this caller act on this resource?" A role is a named bundle of permissions ("operator"). A permission is a (resource, action) pair — ("reports", "delete"), not a vibe. The matrix lives in code; the check runs on every request; unknown roles and unknown actions are denied, not shrugged at. Note the deliberate split from the guardrails lesson: authorization (may this caller act on this resource?) is answered here, by the matrix. Approval (did a named human sign off on this specific consequential action?) is a separate mechanism for the actions that need it — don't merge them into one fuzzy "access control."

Tenant isolation answers "whose data is whose?" A tenant is a customer boundary — here, a city. Three standard models:

ModelHow it isolatesCost
Shared schema, tenant_id columnEvery row carries tenant_id; every query filters on itCheapest; isolation depends on query discipline + tests
Schema per tenantOne database schema per city; search path set per requestMedium; migrations must run N times
Database per tenantPhysically separate databasesHeaviest; strongest blast-radius boundary

CityOps picks the first model — shared schema with tenant_id — because with two cities the operational cost of separate databases buys nothing the query discipline doesn't already provide. That's a judgment call, documented next to the decision, with a revisit trigger: when a tenant's compliance terms demand physical separation, we move that tenant. The model isn't the point; the point is that the boundary is explicit, enforced, and tested — whichever model you choose.

The myth that fails the audit

"The frontend hides the button, so users can't do that." The frontend is a suggestion. The API is the boundary. Every permission the UI hides must also be denied by the API — because the attacker, the curious power user, and the buggy script all skip the UI entirely.

Build: give every secret a home

The rule is mechanical: secrets are loaded once, at startup, from the secret store — and a missing secret prevents boot. "Fail closed" isn't philosophy here; it's the difference between "the deploy failed loudly in staging" and "production has been running for a month with a blank API key and nobody noticed."

# cityops/secrets.py
import abc
import os

class MissingSecretError(RuntimeError):
    """Raised at startup when a required secret is absent. Fail closed."""

def require_secret(name: str) -> str:
    value = os.environ.get(name)
    if not value:
        raise MissingSecretError(
            f"{name} is not set — refusing to start with a missing secret"
        )
    return value

class SecretStore(abc.ABC):
    @abc.abstractmethod
    def get(self, name: str) -> str:
        """Return the secret value, or raise MissingSecretError."""

class EnvSecretStore(SecretStore):
    """Local development ONLY. Production never reads secrets from the shell."""

    def get(self, name: str) -> str:
        return require_secret(name)

class AwsSecretsManagerStore(SecretStore):
    """Production: secrets live in the secret store and rotate on a schedule."""

    def __init__(self) -> None:
        import boto3
        self._client = boto3.client("secretsmanager")

    def get(self, name: str) -> str:
        value = self._client.get_secret_value(SecretId=name).get("SecretString")
        if not value:
            raise MissingSecretError(f"secret {name} has no string value")
        return value

The interface is the point: application code calls store.get(...) and never knows which backend answered. Local dev uses environment variables (convenient, and the blast radius is a laptop); production uses the secret store (audited reads, rotation, no shell history). One seam, chosen by APP_ENV:

# cityops/settings.py
import dataclasses
import os
from cityops.secrets import AwsSecretsManagerStore, EnvSecretStore

@dataclasses.dataclass(frozen=True)
class Settings:
    db_password: str
    api_key: str
    app_env: str

    def __repr__(self) -> str:
        # A stray print(settings) in a log must never leak values.
        return "Settings(***redacted***)"

def load_settings() -> Settings:
    prod = os.getenv("APP_ENV") == "prod"
    store = AwsSecretsManagerStore() if prod else EnvSecretStore()
    return Settings(
        db_password=store.get("cityops/db/password"),
        api_key=store.get("cityops/api/key"),
        app_env=os.getenv("APP_ENV", "dev"),
    )

SETTINGS = load_settings()  # once, at startup — missing secret = no boot

Three properties worth naming, because Lisa will ask about each: no secret in the repo (the .env file keeps local config; secret values for prod were never in it), no secret in the logs (the logging allowlist from the previous lesson drops anything that isn't an approved field — and the redacted __repr__ covers the careless print), and rotation is a procedure: rotate the value in the store, rolling-restart the fleet so every process loads the new value at startup, verify, then revoke the old one. No scavenger hunt, because there's exactly one place the value lives.

Build: who — the permission matrix in code

Authentication ("who are you?") is the token lookup. Authorization ("may you do this?") is the matrix. Keep them separate — a valid login is not a permission.

# cityops/auth.py
import dataclasses
from fastapi import Depends, Header, HTTPException

@dataclasses.dataclass(frozen=True)
class RequestContext:
    user_id: str
    role: str
    tenant_id: str  # from the SESSION — never from the request body

@dataclasses.dataclass(frozen=True)
class Session:
    user_id: str
    role: str
    tenant_id: str
    revoked: bool = False

_sessions: dict[str, Session] = {}  # demo seam; production: sessions table

def lookup_session(token: str) -> Session | None:
    return _sessions.get(token)

def get_request_context(authorization: str = Header(...)) -> RequestContext:
    scheme, _, token = authorization.partition(" ")
    if scheme.lower() != "bearer" or not token.strip():
        raise HTTPException(status_code=401, detail="Missing bearer token")
    session = lookup_session(token.strip())
    if session is None or session.revoked:
        raise HTTPException(status_code=401, detail="Invalid session")
    return RequestContext(
        user_id=session.user_id, role=session.role, tenant_id=session.tenant_id
    )

The critical line is the one that isn't there: nothing in the request — not the body, not a query parameter, not a header the client controls — gets to declare the tenant or the role. The session is looked up server-side; the token is just a random key to the session. Untrusted content is data, not authority — and "I am the admin of City B" arriving in a request body is untrusted content.

Now the matrix — permissions as code, reviewed and tested like any other code:

# cityops/rbac.py — permissions are code: versioned, reviewed, tested.
PERMISSIONS: dict[str, set[tuple[str, str]]] = {
    "viewer": {
        ("reports", "read"),
    },
    "operator": {
        ("reports", "read"),
        ("reports", "create"),
        ("incidents", "ack"),
    },
    "admin": {
        ("reports", "read"),
        ("reports", "create"),
        ("reports", "delete"),
        ("tenants", "manage"),
        ("users", "manage"),
    },
}

def require_permission(resource: str, action: str):
    def checker(
        ctx: RequestContext = Depends(get_request_context),
    ) -> RequestContext:
        allowed = PERMISSIONS.get(ctx.role, set())  # unknown role: no permissions
        if (resource, action) not in allowed:
            audit_deny(ctx, resource, action)  # deny + audit, fail closed
            raise HTTPException(status_code=403, detail="Not permitted")
        return ctx
    return checker

Read that .get(ctx.role, set()) twice: an unknown role gets the empty permission set. A typo in a role name, a role removed last quarter, a session carrying a role the matrix never defined — all denied. And every denial is audited with the request's run_id, joining the structured logs from the previous lesson. Unknown actions fail the same way: if it's not in the matrix, it's not permitted. This is the authorization half of the guardrails lesson's contract — may this caller act on this resource? — answered in code, on every request.

Build: whose — tenant isolation in the data layer

Authorization decided who. Tenant scoping decides whose. Both, on every query — because a query that checks the role but not the tenant answers the wrong question.

# cityops/reports.py
import psycopg
from fastapi import APIRouter, Depends, HTTPException
from cityops.auth import RequestContext, get_request_context
from cityops.rbac import require_permission

router = APIRouter()

def fetch_report(report_id: int, ctx: RequestContext) -> dict:
    # tenant_id comes from the AUTHENTICATED SESSION, not the request.
    # Forget this filter on one query and you have a cross-tenant read.
    with psycopg.connect(DB_DSN) as conn:
        row = conn.execute(
            "SELECT id, title, body FROM reports "
            "WHERE id = %s AND tenant_id = %s",
            (report_id, ctx.tenant_id),
        ).fetchone()
    if row is None:
        # 404 either way: don't reveal whether the row exists in another tenant.
        raise HTTPException(status_code=404, detail="Report not found")
    return {"id": row[0], "title": row[1], "body": row[2]}

@router.get("/reports/{report_id}")
def get_report(
    report_id: int,
    ctx: RequestContext = Depends(require_permission("reports", "read")),
):
    return fetch_report(report_id, ctx)

Two details carry the lesson. First, ctx.tenant_id — the tenant is a property of the session, established at login, never supplied per request. A tenant_id query parameter would be the client choosing whose data to see, which is exactly the vulnerability. Second, the 404: a City A user probing report IDs must not be able to distinguish "doesn't exist" from "belongs to City B" — different status codes would leak the existence of other tenants' rows.

And because "permissions are code too" extends to the boundary itself, the isolation gets a test per query function — the test that would have caught Dev's missing filter before Lisa did:

# tests/test_tenant_isolation.py
import pytest
from fastapi import HTTPException
from cityops.auth import RequestContext
from cityops.reports import fetch_report

def make_ctx(tenant_id: str) -> RequestContext:
    return RequestContext(user_id="u-" + tenant_id, role="viewer",
                          tenant_id=tenant_id)

def test_tenant_cannot_read_other_tenants_report():
    ctx_a, ctx_b = make_ctx("city-a"), make_ctx("city-b")
    report = create_report(tenant_id="city-a", title="Pothole on 5th")
    assert fetch_report(report["id"], ctx_a)["title"] == "Pothole on 5th"
    with pytest.raises(HTTPException):
        fetch_report(report["id"], ctx_b)
flowchart TD A["HTTP request"] --> B["Authenticate: bearer token → server-side session"] B -->|unknown or revoked| D401["401 + audit event"] B --> C["Authorize: role checked against permission matrix"] C -->|denied| D403["403 + audit deny"] C -->|allowed| D["Tenant scope: tenant_id from session"] D --> E["Query carries tenant_id filter"] E --> F["Response"] style D401 fill:#3a1620 style D403 fill:#3a1620

Break: four failures, demonstrated

Dev wires it all up, then Lisa asks him to attack it — because the review that only checks the happy path is how the .env ended up on the shared drive.

Failure 1: the forgotten filter. A new /reports/search endpoint filters by title — and by nothing else. City B onboards; a City A user searches "pothole" and gets City B's reports in the results. The isolation test suite covers fetch_report but nobody wrote one for the new query function. Tenant scoping is a property you test per query, not once per codebase. The fix is the test, written before the endpoint ships.

Failure 2: role-checked, tenant-blind. The /admin/export endpoint requires the admin role — and then dumps every tenant's reports into one CSV. Authorization answered who (an admin) but nobody answered whose. An admin of City A now holds City B's data, with full permission-matrix compliance. Every query needs both checks: may this caller act, and on whose rows?

Failure 3: rotation breaks the night. Saturday, the database password rotates in the secret store — exactly per procedure. Sunday 2 AM, the worker fleet starts failing authentication: months ago, someone "temporarily" baked the old password into a worker's environment to debug a queue issue, and temporary became permanent. The API fleet loaded the new secret at startup; the workers are still carrying the old one. Secrets must have exactly one home, and the deploy checklist must ban them from images and env files — "temporary" credentials are permanent until they break.

Failure 4: the harmless read. A support engineer with a valid session reads a City B report "just to check the formatting." Nothing is modified, nothing is deleted — and it's still an incident. Don't assume reads are harmless: export_tenant_csv and get_report can breach confidentiality without changing a single byte. Govern every action by sensitivity and impact, reads included — which is why denials and sensitive reads both land in the audit log.

Productionize: rotation, least privilege, and the audit trail

The code is the mechanism; production is the discipline around it. Four practices turn the lesson into something Lisa signs off on:

Rotate on a schedule, not on a scare. Short-lived database credentials where the platform supports them; scheduled rotation everywhere else. The runbook is three lines: rotate in the store, rolling-restart so every process loads the new value at startup, verify, revoke the old. Because secrets load once at startup from one home, rotation touches exactly one system — compare with the shared-drive .env, where rotation meant finding every copy.

Least privilege, per environment. Dev keys can't touch prod; the staging secret store holds staging values. A leaked dev API key is a bad afternoon; a leaked prod key is a breach notification. Environment separation is what makes that sentence true.

Audit the decisions, minimize the content. Every allow and every deny records who (user ID, role, tenant), what (resource, action), when, and the run_id — and nothing else. No request bodies, no report contents, no PII beyond what's needed to answer "who did what." Minimization before redaction, again: fields the audit log never receives can't leak from it.

Reviews on a calendar. Lisa runs a quarterly access review: who holds admin, which sessions are stale, which service accounts still exist. Onboarding grants the minimum role for the job; offboarding revokes the session the same day. Permissions rot — today's correct matrix is next year's over-privileged one, unless someone re-reads it on purpose.

Authorization is not approval

The matrix answers "may this caller act on this resource?" — that's authorization, and it runs in code on every request. Some actions additionally need "did a named human approve this exact action?" — that's approval, from the guardrails lesson, with its own records behind a boundary the agent can't write. Don't merge them: a matrix can't capture "Maria signed off on deleting City A's archive," and an approval queue can't answer ten thousand reports:read checks per minute.

Communicate: the one-page access memo

Maria asked for the answer in writing before signing City B. Dev writes it once, Lisa files it, and it becomes the document every future customer review starts from:

The access-control memo

  • Who can do what: the permission matrix — roles, resources, actions — with a pointer to the code file and the commit that last changed it. ("Permissions are code too" means the doc points at the code; it doesn't duplicate it.)
  • Whose data is whose: the tenant isolation model (shared schema, tenant_id on every row), plus the test suite that proves each query enforces it — and the revisit trigger for physical separation.
  • Where secrets live: the secret store name, who can read which secrets, the rotation cadence, and the ban on secrets in repos, images, and chat.
  • What we log: auth decisions (allow/deny) with user, role, tenant, and run_id — minimized, retained per policy, reviewable by Lisa.

The memo's real audience isn't Lisa — it's the next Lisa, at the next customer, asking the same questions. A team that can hand over this page in the first meeting sells enterprise trust; a team that schedules a week to "look into it" sells risk. And notice what the memo never contains: a single secret value, a single customer's data, or a promise the code doesn't keep.

Where this lands in CityOps

City B's signature was blocked on Lisa's review. With the memo filed, it unblocks:

  • Secrets leave the shared drive this week. The .env keeps local config; every production credential moves to the secret store, the repo history gets the key rotation it should have had in March, and the deploy checklist gains its ban.
  • The API enforces the matrix. Every endpoint from Milestone 3: The CityOps API gets its require_permission dependency — including the ones the frontend "already hides."
  • Retrieval respects the boundary. The Ask CityOps pilot's vector search carries a tenant_id metadata filter on every query: City A's residents never retrieve City B's documents, which means the model never sees them either. Tenant isolation in the data layer is what makes the RAG layer's answers tenant-safe.
  • Denials join the audit trail. The run_id from the logging lesson ties a 403 to the exact request, the exact session, and the exact matrix version that denied it — the reconstructability the guardrails lesson demanded, now covering access control.

Lisa's two questions finally have written answers: the production key lives in the secret store, readable by the production role only, rotated on schedule — and a City A user cannot see City B's data, and here's the test that proves it. Maria signs. City B onboards. The boundary holds because it's code, not convention.

Field check

  1. A request arrives with {"tenant_id": "city-b"} in the JSON body, from a session belonging to City A. What tenant's data does fetch_report return — and why is that the only safe answer?
  2. Your new /reports/search endpoint passes the permission check but has no tenant filter. Who can exploit it, what do they learn, and which test would have caught it?
  3. The database password rotates in the secret store on Saturday. List every place a copy of the old password could be hiding, and the one practice that makes the list short.
  4. A support engineer reads another tenant's report "just to check formatting" — no writes, no exports. Is this an incident? What does the audit log record, and what should it not record?
  5. Lisa asks for "a guarantee that no unauthorized access will ever happen." Write the two-sentence honest answer, in the style of the guardrails lesson.
Answers

1. City A's data — because fetch_report takes tenant_id from the authenticated session, never from the request. The body field is untrusted content: it has no authority. If the query honored the body's tenant_id, any caller could read any tenant by editing a JSON field.
2. Any authenticated user of any tenant — the permission check confirms they're someone, but nothing restricts whose rows match. They learn the existence and contents of other tenants' reports. The per-query tenant-isolation test (City B context must 404 on City A's row) would have caught it — which is why the test is written per query function, not once per codebase.
3. Shell history, CI logs, the .env on the shared drive, repo history, chat messages, the old container image's environment, a teammate's notes. The practice that shortens the list to one item: secrets live in exactly one home (the secret store), are loaded once at startup, and are banned from repos, images, env files, and chat by the deploy checklist.
4. Yes — reads are governed by sensitivity and impact, not by whether bytes changed. A cross-tenant read breaches confidentiality. The audit log records who (user, role, tenant), what (resource, action), when, the run_id, and the deny/allow decision. It must not record the report's contents or any PII beyond what's needed to answer "who did what" — minimization before redaction.
5. "No — no system can guarantee that, and anyone selling the guarantee is selling the audit, not the outcome. What we guarantee is the mechanism: permissions enforced in code on every request, tenant boundaries tested per query, secrets in one store with scheduled rotation, every decision audited — and we continuously test all of it, including attempts to break it."