Prerequisites: the CI/CD pipeline (which deploys every merge), the Compose stack, the health checks and monitoring that watch it, and the Milestone 5 incident — the 7:40 fix that shipped to 100% of traffic in one jump.

Friday, two days after launch day. The dashboard is current, the mayor's office is happy, and Dev has a new model artifact burning a hole in his pocket: cityops-classifier v3.2.1, which cuts misrouted reports by 12% on the evaluation set.

Dev: "I'd like to ship the new classifier today." Maria: "The council briefing is Wednesday. Everything works right now. Please don't touch anything." A pause. Dev: "That's exactly the problem. 'Don't touch anything' is not a deployment strategy — it's a deployment freeze, and freezes end the day something urgent needs to ship."

He's right, and they both know it. The 7:40 incident fix went to all traffic at once and it worked — but that was luck plus good CI, not safety. The question on the table: how do you change a running production system without betting the whole system on every change?

CI/CD answers "how does code get to production?" This lesson answers the harder question: "how much of production should it touch at once?" Blue-green deployments, canary releases, and feature flags are three ways to make every deploy a small, reversible bet instead of an all-or-nothing gamble.

The customer problem

Maria's fear is rational. Every deploy she's watched has been binary: the old version runs, then someone flips a switch, then the new version runs everywhere. When that switch flipped at 7:40 AM on launch day, the fix worked — but if it hadn't, the 9 AM demo would have died on a bug introduced by the rescue itself. There was no middle ground between "don't ship" and "ship to everyone."

Three gaps, one per strategy this lesson builds:

  • Deploys have a downtime-shaped hole. Restarting the API to pick up the new image means a window where nobody is serving — or a prayer that the restart is fast enough that nobody notices.
  • New code meets 100% of traffic on its first day. The classifier upgrade was evaluated on a labeled set, but production traffic is weirder than any eval set. The first real-world surprise arrives at full blast radius.
  • Deploying code and releasing behavior are the same event. Merging the quarantine-dashboard UI means it goes live the moment CI finishes. There's no "shipped but dark," no way to turn a feature on for Maria's team first and the public later.

Clarify the ask

Dev turns Maria's "don't touch anything" into three work orders — because "be careful" is a wish, not a plan:

  • "Ship the API without an application restart outage." — a second, warm API version takes traffic while the first stays up. The switch is a router change, not a restart.
  • "Let the new classifier prove itself on real traffic — a little at a time." — 5% of requests first, metrics compared against the current version, promote only on evidence.
  • "Decouple shipping code from releasing behavior." — the quarantine dashboard merges this week but stays dark until Maria says go; if it misbehaves, it's a toggle, not a redeploy.

Out of scope, deliberately: multi-region failover and service-mesh machinery. One reverse proxy and one Compose host are enough to teach all three patterns honestly — the patterns scale up; the lesson doesn't need the scale.

Which brings us to the principle. Last milestone's principle was about integration. This one's about the moment of change itself:

One principle

"Change production gradually, reversibly, and with evidence—or hold." Every deploy is a bet on code untested against reality. You don't eliminate the bet — you size it: a slice of traffic, a toggle, a warm standby. "Hold" is always a valid answer when the evidence isn't there.

The minimum concept: three ways to shrink the blast radius

Three patterns, each one sentence, then when to reach for which:

Blue-green: two API versions behind one proxy, one serves. Blue runs production. You deploy the new version to green, warm it up, point the router at green. Blue stays up, untouched — rollback is pointing the router back, provided the old version remains compatible with any shared state or schema changes. Blue-green can avoid an application restart outage by warming the new stack before routing traffic to it — but "no downtime" is a design goal, not a guarantee: in-flight requests, draining, and shared migrations can still bite.

Canary: a slice of traffic tries the new version first. Approximately 5% of requests go to the new build while the rest stay on the old. You compare error rate, latency, and the business-quality metric between the two slices, then widen the slice — 5% → 25% → 50% → 100% — promoting only while the numbers hold. Rollback is shrinking the slice to zero.

Feature flag: a runtime toggle decouples deploy from release. The code ships to production dark; the behavior turns on when a flag flips. With a production flag service, behavior can change without redeploying — and targeted rollouts (Maria's team first, 10% of users) need one. The minimal file-based implementation below is global on/off and demonstrates the branching pattern; it requires a restart to reload, which is a stated limitation, not a hidden one.

StrategyBest whenRollback looks like
Blue-greenAPI/container upgrades where downtime is unacceptable and the new version is believed goodRouter points back to blue — seconds
CanaryRisky changes (new model, new parser) where real-traffic evidence beats eval setsTraffic slice shrinks to zero — minutes
Feature flagUI/behavior changes, gradual rollouts, kill-switches for anything uncertainToggle flips off — seconds, no deploy

The myth that ships the outage

"CI is green, so the deploy is safe." CI proves the code passes its tests. It says nothing about production traffic, production data shapes, or the 2:14 AM field rename. Green CI is the entry ticket to a release strategy — not the strategy itself.

flowchart TD A["Change proposed"] --> B{"CI passes?"} B -->|no| C["Don't ship"] B -->|yes| D["Choose exposure"] D --> E["BLUE-GREEN
switch version"] D --> F["CANARY
small slice"] D --> G["FEATURE FLAG
dark behavior"] E --> H["Readiness + compatible state"] F --> I["Tagged metrics + enough samples"] G --> J["Controlled enable + owner + expiry"] H --> K["Observe evidence"] I --> K J --> K K --> L{"Acceptable?"} L -->|no| M["Rollback / off — investigate"] L -->|yes| N["Promote — widen exposure"] style M fill:#3a1620

CI decides whether a change is eligible to enter production; release strategy decides how much production is allowed to trust it at once.

Build: blue-green with Compose and a reverse proxy

The idea is mechanical: two API versions behind one proxy, sharing the production database — the proxy decides which version serves. A fuller blue-green can duplicate more components, but shared state is common, and it's exactly why migration compatibility matters (more on that danger in the Break section):

# docker-compose.release.yml — blue-green on one host
services:
  proxy:
    image: caddy:2
    ports: ["80:80", "443:443"]
    volumes:
      - ./Caddyfile:/etc/caddy/Caddyfile:ro
    networks: [cityops]
  api-blue:
    image: registry.internal/cityops/api:${BLUE_TAG:?BLUE_TAG required}
    # Database credentials omitted here — reuse the secrets/config pattern
    # from the production Compose lesson; this lesson is about routing.
    healthcheck:
      # python:3.12-slim ships no curl — probe with the runtime that's there.
      test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/ready', timeout=5)"]
      interval: 15s
      retries: 3
    networks: [cityops]

  api-green:
    image: registry.internal/cityops/api:${GREEN_TAG:?GREEN_TAG required}
    # Database credentials omitted here — reuse the secrets/config pattern
    # from the production Compose lesson; this lesson is about routing.
    healthcheck:
      # python:3.12-slim ships no curl — probe with the runtime that's there.
      test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/ready', timeout=5)"]
      interval: 15s
      retries: 3
    networks: [cityops]

  db:
    image: postgres:16
    volumes: [pgdata:/var/lib/postgresql/data]
    networks: [cityops]

volumes:
  pgdata:
networks:
  cityops:

The Caddyfile is the whole release mechanism — one line decides which stack serves:

# Caddyfile — blue is live; green is warm standby
cityops.internal {
    reverse_proxy api-blue:8000
}

The release procedure, as a script Tom can run at 2 AM without improvising:

#!/bin/bash
# release.sh — cut traffic between the live backend and the warm candidate
# Usage: ./release.sh <new-image-tag>   (immutable tag, already through CI)
#
# State model — the Caddyfile is the source of truth for routing:
#   live      = backend currently serving traffic (api-blue or api-green)
#   candidate = the other backend: warmed, verified, then promoted
# Rollback target = the previous live backend, which stays warm.
set -euo pipefail
COMPOSE="docker compose -f docker-compose.release.yml"
NEW_TAG="$1"

# 1. Who is live? Read it from the proxy config — no guessing, no globals.
LIVE="$(grep -o 'api-blue\|api-green' Caddyfile | head -n 1 | sed 's/api-//')"
if [ "$LIVE" = "blue" ]; then CANDIDATE="green"; else CANDIDATE="blue"; fi

# Resolve the currently deployed image tags from the running containers,
# then EXPORT them — the Compose file requires these variables on EVERY
# invocation, not just the first.
current_tag() {
  local cid
  cid="$($COMPOSE ps -q "api-$1" 2>/dev/null || true)"
  if [ -n "$cid" ]; then docker inspect --format '{{.Config.Image}}' "$cid"; fi
}
BLUE_TAG="${BLUE_TAG:-$(current_tag blue)}"
GREEN_TAG="${GREEN_TAG:-$(current_tag green)}"
# The candidate gets the new image; the live backend keeps its tag.
if [ "$CANDIDATE" = "blue" ]; then BLUE_TAG="$NEW_TAG"; else GREEN_TAG="$NEW_TAG"; fi
export BLUE_TAG GREEN_TAG

# 2. Warm the candidate and WAIT for readiness — with a deadline.
#    Readiness needs a deadline: a single probe can check too early,
#    and an unbounded loop can wait forever.
READY_DEADLINE=180
echo "warming candidate api-$CANDIDATE ($NEW_TAG)..."
$COMPOSE up -d "api-$CANDIDATE"
start="$(date +%s)"
until $COMPOSE exec -T "api-$CANDIDATE" \
    python -c "import urllib.request; urllib.request.urlopen('http://localhost:8000/ready', timeout=5)"; do
  if [ "$(( $(date +%s) - start ))" -gt "$READY_DEADLINE" ]; then
    echo "candidate not ready within ${READY_DEADLINE}s — aborting, live untouched" >&2
    exit 1
  fi
  sleep 5
done

# 3. Flip the router — this is the deploy. Everything before was staging.
#    This routes NEW requests; in-flight connections need a draining policy
#    (see below) — the teaching example focuses on routing.
sed -i "s/api-${LIVE}:8000/api-${CANDIDATE}:8000/" Caddyfile
# Reload the proxy with the updated configuration.
$COMPOSE exec -T proxy caddy reload --config /etc/caddy/Caddyfile

# 4. Watch the monitors for 10 minutes. If anything looks wrong, roll back:
#      sed -i "s/api-${CANDIDATE}:8000/api-${LIVE}:8000/" Caddyfile
#      ...then reload the proxy again. The old backend stayed warm for
#      forensics — rollback is a router change, not a redeploy.
echo "traffic now on api-$CANDIDATE ($NEW_TAG); api-$LIVE kept warm for rollback"

Three details carry the lesson. First, the script reads its state — who's live — from the proxy config instead of assuming a direction, and it resolves the currently deployed image tags from the running containers and exports them, because the Compose file's required interpolation applies to every invocation. Second, step 2 waits on /ready with a deadline — the readiness distinction from the health-checks post — so the candidate never takes traffic it can't serve, and the script can't hang forever. Third, the previous live backend is never torn down during the release window; rollback is a router change, not a redeploy — provided shared state stayed compatible.

One thing the script doesn't do: drain in-flight connections. The cutover routes new requests; a production cutover also needs a draining policy for in-flight and long-lived connections. That's a one-sentence addition to the release plan, not a new mechanism — but don't let "zero downtime" marketing make you forget it.

flowchart TD A["Deploy new image to GREEN"] --> B["Wait for GREEN /ready"] B --> C["Flip proxy: traffic → GREEN"] C --> D["Watch metrics 10 min"] D -->|healthy| E["GREEN is live; BLUE stays warm"] D -->|unhealthy| F["Flip proxy back → BLUE"] F --> G["Debug GREEN offline"] style F fill:#3a1620

Build: the canary — a slice of traffic, judged by numbers

Blue-green answers "how do I switch without downtime." The canary answers "how do I know the new version deserves the traffic." The mechanism is a weighted split at the proxy — approximately 5% of requests to the new build under this simple proxy configuration (weights steer the distribution; they don't produce an exact randomized experiment):

# nginx.conf — canary slice at the reverse proxy
upstream cityops {
    server api-stable:8000 weight=95;
    server api-canary:8000 weight=5;
}
server {
    listen 80;
    location / {
        proxy_pass http://cityops;
    }
}

One thing the proxy config doesn't show: how do you know which requests were canary? The verdict script below reads baseline.jsonl and canary.jsonl — but those files don't exist until your telemetry can separate the two. Every request metric used for canary comparison must carry the served release as a label — release=stable vs release=v3.2.1 — set from an environment variable on each API container and attached to every structured log line and metric. The two files are just filtered views of one telemetry stream:

{"ts": "2026-10-05T09:02:11-07:00", "release": "v3.2.1", "latency_ms": 41,
 "error": false, "misrouted": false, "report_type": "pothole"}

Without the label, the comparison is fiction. This is the observability lesson paying rent: you can only compare the slices you can identify.

Then the part most teams skip: the verdict is a script, not a vibe. Dev writes a small checker that compares the canary slice against the baseline on error rate, p95 latency, and misroute rate — with a minimum-sample guard, because small numbers lie (more on that in the Break section):

# scripts/canary_check.py — the promotion decision, in code
import json, sys

MIN_SAMPLES = 200  # teaching threshold for THIS release plan, not a universal
                   # rule — the evidence you need depends on event rate,
                   # effect size, risk, and traffic mix.
ERR_CEILING = 0.02      # absolute safety ceiling: the canary error rate may
                        # never exceed 2%, whatever the baseline says.
ERR_TOLERANCE = 1.5     # ...and may not exceed 1.5x a NONZERO baseline.
P95_TOLERANCE = 1.3
MISROUTE_CEILING = 0.15 # canary misroute rate may never exceed 15%:
                        # a model can be fast, HTTP-200, and wrong.

def summarize(samples):
    lat = sorted(s["latency_ms"] for s in samples)
    errors = sum(1 for s in samples if s["error"])
    misroutes = sum(1 for s in samples if s.get("misrouted"))
    # simple nearest-rank teaching calculation; production monitoring should
    # use the same percentile definition as your observability system.
    p95 = lat[max(0, int(0.95 * len(lat)) - 1)]
    n = len(samples)
    return {"n": n,
            "error_rate": errors / n,
            "p95_ms": p95,
            "misroute_rate": misroutes / n}

def verdict(baseline, canary):
    # Check sample count BEFORE computing anything — summarize() divides.
    if len(baseline) < MIN_SAMPLES or len(canary) < MIN_SAMPLES:
        return ("HOLD",
                f"too few samples (baseline n={len(baseline)}, "
                f"canary n={len(canary)}) — wait, don't promote")
    b, c = summarize(baseline), summarize(canary)
    reasons = []
    # Absolute ceiling first: a 0% baseline would make ANY relative
    # threshold meaningless — 1.5x of zero is still zero.
    if c["error_rate"] > ERR_CEILING:
        reasons.append(f"canary errors {c['error_rate']:.2%} exceed "
                       f"absolute ceiling {ERR_CEILING:.0%}")
    elif b["error_rate"] > 0 and c["error_rate"] > b["error_rate"] * ERR_TOLERANCE:
        reasons.append(f"canary errors {c['error_rate']:.2%} vs "
                       f"baseline {b['error_rate']:.2%}")
    if c["misroute_rate"] > MISROUTE_CEILING:
        reasons.append(f"canary misroutes {c['misroute_rate']:.2%} exceed "
                       f"ceiling {MISROUTE_CEILING:.0%} — the model is wrong "
                       "even if the API is healthy")
    if c["p95_ms"] > b["p95_ms"] * P95_TOLERANCE:
        reasons.append(f"canary p95 {c['p95_ms']}ms vs baseline {b['p95_ms']}ms")
    if reasons:
        return ("ROLLBACK", "; ".join(reasons))
    return ("PROMOTE", f"canary within tolerance (n={c['n']})")

if __name__ == "__main__":
    baseline = [json.loads(l) for l in open(sys.argv[1])]
    canary = [json.loads(l) for l in open(sys.argv[2])]
    decision, reason = verdict(baseline, canary)
    print(f"{decision}: {reason}")

Note what this script does not claim: it does not prove the canary is better, and the tolerances (2% ceiling, 1.5x, 1.3x, 15% misroute) are team-chosen judgment calls, not universal constants. Operational metrics protect system health; model-quality metrics protect the customer outcome. An AI classifier can be HTTP 200, fast, and wrong all day — the release criterion must include the business metric you actually care about, here the misroute rate. (Misroute labels arrive with a delay in this example, so the checker runs on the latest fully-labeled window. That's realistic, not a flaw in the method — it just means the verdict lags the traffic.) What the script does is replace "looks fine on the dashboard" with a repeatable rule — the same move as the evals lesson's "the rubric defines correctness": the thresholds are written down before the release, not invented after the numbers arrive.

One more design constraint before the schedule: stable and canary should see comparable traffic populations. If the canary only serves office IPs at noon while stable serves the whole city at rush hour, you're comparing populations, not versions. Slice the comparison by release version plus the dimensions that matter — report type, client/device, time window — or make the slice representative by construction.

The rollout schedule for the classifier upgrade is then a series of these verdicts, widening the slice (an illustrative CityOps rollout schedule — a low-traffic service needs longer windows; a high-volume one reaches sufficient evidence faster):

flowchart TD A["5% canary — 30 min"] --> B{"canary_check: PROMOTE?"} B -->|yes| C["25% — 1 hour"] B -->|no| Z["Shrink slice to 0 — investigate"] C --> D{"canary_check: PROMOTE?"} D -->|yes| E["50% — 2 hours"] D -->|no| Z E --> F{"canary_check: PROMOTE?"} F -->|yes| G["100% — canary becomes baseline"] F -->|no| Z style Z fill:#3a1620

Build: feature flags — deploy the code, release the behavior

The canary controls which version serves. The flag controls which behavior runs — independent of deploys entirely. Dev merges the quarantine-dashboard UI on Monday; it stays dark until Maria gives the go-ahead Wednesday. The minimal honest implementation — a JSON file, a lookup, a default:

# cityops/flags.py — minimal feature flags, no service required
import json
import os

_FLAGS = None

def _load():
    global _FLAGS
    path = os.getenv("FLAGS_FILE", "flags.json")
    try:
        with open(path) as f:
            _FLAGS = json.load(f)
    except FileNotFoundError:
        _FLAGS = {}
    return _FLAGS

def enabled(name: str, default: bool = False) -> bool:
    """True if the flag is on. Flags are config, not secrets."""
    flags = _FLAGS if _FLAGS is not None else _load()
    return bool(flags.get(name, default))

def reset() -> None:
    """Tests only — lets each test start from a known flag state."""
    global _FLAGS
    _FLAGS = None
# flags.json — the release plan, as data
{
  "quarantine_dashboard": false,
  "classifier_v3_scoring": false,
  "new_export_csv": true
}
# usage — the flag wraps the new behavior, the old path stays default
from cityops import flags

def get_dashboard(ctx):
    rows = fetch_reports(ctx)
    if flags.enabled("quarantine_dashboard"):
        rows = attach_quarantine_counts(rows, ctx)
    return render(rows)

Three honest caveats, because a flag system you don't understand is a loaded footgun. First, this minimal version loads flags once at startup — flipping flags.json requires a restart; a production flag client polls or subscribes, which is a stated upgrade, not a hidden one. Second, flags are not a permission system: "who can see this" is the RBAC matrix from the secrets post, not a flag. Third, every flag gets an owner and a removal date, tracked in the flag registry — a companion doc or work item, since JSON has no comments — because flags you never remove become the next section's failure mode.

flowchart TD A["Merge code behind flag OFF"] --> B["Deploy via CI/CD — behavior unchanged"] B --> C["Flip flag ON"] C --> D["Watch metrics + feedback"] D -->|good| E["Remove flag + old path — flag is temporary"] D -->|bad| F["Flip flag OFF — no redeploy"] style F fill:#3a1620

Break: four ways releases go wrong anyway

Dev wires up all three patterns, then deliberately attacks them — because the release that only works on the happy path is how the 7:40 fix shipped to everyone at once.

Failure 1: the canary metrics lie. The classifier canary runs at 5% for 30 minutes. Zero errors, p95 looks great — canary_check.py says PROMOTE. Dev promotes to 25%. Errors spike. What happened: the 5% slice served 41 requests, all of them simple "pothole" reports from the morning batch. The eval set had hundreds of tricky cases; the canary slice had dozens of easy ones. A 0% error rate on n=41 is not evidence — it's the absence of evidence wearing a lab coat. The MIN_SAMPLES guard existed and the team overrode it because "the numbers looked good." Small-n evidence is how you SHIP a regression with a green dashboard — the Stage 4 evals rule applies to releases too: prefer a wider slice or a longer window over promoting on thin data.

Failure 2: blue-green meets the shared database. Green runs the new API build, which includes a migration: reports.priority becomes NOT NULL with a new default. Traffic flips to green — all good. Then the monitors twitch, Tom flips back to blue, and blue starts crashing: the old code inserts rows without priority, and the column no longer accepts NULL. The database was the one thing blue and green shared, and the migration only ran in one direction. Blue-green requires backward-compatible migrations: expand first, then contract. The safe sequence has three phases — expand, migrate/backfill, contract — which often span multiple deployments: (1) add the column as nullable, both versions run; (2) backfill and switch writes, both versions run; (3) add the NOT NULL constraint, old version retired. Database changes don't need to be magically reversible, but they must preserve compatibility for the release/rollback window — or have an explicit recovery plan.

Failure 3: flag debt. Six months later, CityOps has 47 flags. Twelve have been on for 100% of traffic for months — they're not flags anymore, they're untested configuration. One, new_export_csv, wraps two divergent code paths that nobody remembers the difference between. During an incident, someone flips quarantine_dashboard off "just in case" and the on-call dashboard goes blank — the flag had become load-bearing infrastructure with no owner. Flags are temporary by definition; a flag without an expiry date is tech debt with a toggle. The fix is the lifecycle from the caveats: every flag ships with an owner and a removal date, and the weekly ops review includes "which flags died this week."

Failure 4: the canary that only served the office. The team dogfoods the new export feature on their internal staging environment, which mirrors production's config. Metrics look perfect for a week. They roll it to everyone — and it breaks on the city's ancient IE11 kiosks, which nobody on the team uses. The canary traffic wasn't representative traffic; it was the most forgiving traffic available. A canary measures the slice it serves, not the population you hope for. Dogfooding catches your own bugs; it doesn't predict the field. The release plan has to say whose traffic the canary sees — and "people like us" is the wrong answer when the users are a city's residents.

Productionize: releases as a written practice

The patterns are mechanisms; production is the discipline around them. Four practices turn release strategies into something Maria can rely on before a council briefing:

Write the promotion criteria before the release. The canary thresholds, the minimum sample counts, the bake times at each slice — all decided up front, in the release plan, not improvised while watching a dashboard at midnight. The canary_check.py thresholds are the starting point; the release plan names the slice schedule (5% → 25% → 50% → 100%) and who has the authority to hold or roll back at each step.

Record the strategy in the deploy record. The Milestone 5 deploy record gets two new fields: which release strategy was used and the current slice. "We deployed the fix" becomes "we deployed the fix as a 5% canary at 07:40, promoted to 100% at 09:15 after three green verdicts." The audit trail can then answer not just what is running, but how cautiously it got there.

Automate the rollback trigger. The canary verdict script runs on a schedule during the release window; a ROLLBACK verdict pages Tom and shrinks the slice automatically. Humans approve promotions; machines execute rollbacks. But automate rollback only for well-tested, bounded triggers — when the rollback target is known-good and shared state remains compatible; otherwise fail closed and page the operator. The 2 AM version of Dev should not be making judgment calls — the judgment was made at 2 PM when the thresholds were written.

Give every flag a death date. Flag created → flag rolled out → flag removed, with the removal tracked like any other work item. The weekly ops review asks "which flags died this week" the same way it asks "what paged us." A flag that survives its purpose becomes configuration, and configuration nobody remembers is how the next incident starts.

Later, not now

  • Automated progressive delivery (flagger/Argo Rollouts style controllers) — the scripted canary teaches the judgment; the controller automates it once the judgment is sound.
  • A dedicated feature-flag service with targeting rules and audit — the JSON file is honest for one team; multi-team flag governance is a later scaling decision.

Communicate: the release brief

Maria doesn't need to know what a canary is. She needs to know, before Wednesday, whether the classifier upgrade can break her briefing. Dev's release brief is three sentences — the release note as a deliverable:

The release brief

  • What changed: the report classifier upgraded to v3.2.1 — 12% fewer misrouted reports on our evaluation set, no API or dashboard changes.
  • Who gets it, and when: 5% of traffic today, widening to 100% by Tuesday only if the error, latency, and misroute checks stay green at every step. The briefing data is unaffected unless we promote.
  • How we know it's fine, and how we undo it: an automated check compares the new version against the current one every 30 minutes; if it fails, traffic returns to the current version — rollback is designed to need only a routing change, not a redeploy (target: under one minute). Worst case for Wednesday: we stay on the current version.

The brief's real audience isn't Maria — it's the next Maria, at the next customer, asking "is it safe to deploy this week?" A team that can hand over this note before every release sells operational maturity; a team that says "trust us, CI is green" sells risk. And notice the last line: the plan explicitly names the fallback. "We stay on the current version" is a complete, respectable outcome — rolling forward is optional, rolling back is always available.

Where this lands in CityOps

The classifier upgrade ships the following week, exactly per the brief:

  • Canary, not coin-flip. v3.2.1 serves 5% of scoring requests Monday morning. The verdict script compares misroute rate and p95 latency against v3.2.0. Two HOLDs on sample size (Monday's traffic is thin before 9 AM — the team waits instead of overriding, having learned Failure 1 the cheap way). Promotion to 25%, 50%, 100% across Monday–Tuesday, all green.
  • The 7:40 fix, replayed. The team re-runs the incident morning as a tabletop: with flags, the tolerant parser would have shipped dark on Friday, and the quarantine path enabled by a flag flip — no redeploy, no war room. The fix that took a war room becomes a toggle plus a canary. The pipeline stays the same; the blast radius doesn't.
  • Blue-green for the boring upgrades. The next FastAPI runtime/base-image patch goes blue-green: new image warmed, router flipped, old stack kept for a day. Nobody pages Tom. That's the point — release strategy should make most deploys unremarkable.
  • Flags for the visible changes. The new public dashboard layout merges behind new_dashboard_layout, dark for two sprints, then flipped on globally. When a council staffer reports a confusing label, it's a flag flip and a fix — not an emergency deploy.

Maria's summary at the Wednesday briefing, unprompted: "They upgraded the system twice this week and I didn't notice either time." That sentence is the whole lesson. The best deploys are the ones the customer never has to think about — and the way you earn "didn't notice" is slices, toggles, and warm standbys, not courage.

Field check

  1. Your canary at 5% shows zero errors after 30 minutes — but only 41 requests hit it. The dashboard looks perfect. Promote, or hold? What does the verdict script say, and why is overriding it dangerous?
  2. You need to add a NOT NULL column to the reports table during a blue-green deploy. Walk through the expand → migrate → contract sequence, and say what breaks if you do it in one step.
  3. A feature flag has been on for 100% of traffic for four months. Nobody remembers what the old code path did. What do you do — and what process change prevents the next one?
  4. Your canary slice serves only your office's IPs and the metrics look perfect for a week. Is this evidence the release is safe for the city's residents? What traffic should the canary see?
  5. Maria asks: "Prove the new classifier won't break Wednesday's briefing." Write the three-sentence release brief — and name the respectable fallback outcome.
Answers

1. Hold. The verdict script says HOLD — 41 samples is below the 200 minimum, so the numbers are noise, not signal. Overriding it is dangerous because a 0% error rate on thin, easy traffic tells you nothing about the hard cases; promoting on it is how you SHIP a regression with a green dashboard. Widen the slice or wait for more traffic, then re-run the verdict.
2. Expand: add the column as nullable and deploy — both versions run. Migrate: backfill existing rows and switch new writes to populate it — both versions still run. Contract: add the NOT NULL constraint only after the old version is fully retired. In one step, the NOT NULL constraint lands while the old version still runs: the old code doesn't populate the new field, so it can no longer perform its old write successfully — and rollback, the whole point of blue-green, crashes on the shared database.
3. Remove the flag and delete the old code path — after confirming the rollout is intentionally permanent and rollback no longer depends on the old path. It's not a flag anymore, it's untested configuration. Verify with tests that the flagged-on path is the only path, then delete. The process change: every flag ships with an owner and a removal date, and the weekly ops review asks "which flags died this week."
4. No — the canary measured your office, the most forgiving traffic available, not the city's residents with their old kiosks and odd report shapes. The canary should see a representative slice of real production traffic: same mix of devices, report types, and times of day as the full population. Dogfooding catches your bugs; it doesn't predict the field.
5. "The classifier upgraded to v3.2.1 — 12% fewer misroutes on our eval set, no API or dashboard changes. It rolls out to 5% of traffic today, reaching 100% by Tuesday only if the error, latency, and misroute checks stay green at every step. If any check fails, traffic returns to the current version — rollback needs only a routing change, not a redeploy." The respectable fallback: not rolling forward at all.