pgdata volume is where everything below starts), secrets discipline from Secrets, RBAC & Tenant Isolation, and the monitoring habit from Health Checks, Logs & Monitoring. The release-strategies lesson (same stage, by title) covers getting new code out safely; this one covers surviving the day the machine itself dies.Tom (ops, 3:07 AM, phone to his ear): "The prod host is gone. Not slow — gone. The provider's status page says 'hardware failure, host unreachable.' Postgres was on that host's local disk."
Dev (already awake, already dreading the question): "The pgdata volume… was on that host's disk."
Maria (customer, on the same call, very calm): "Two questions. One: do we have a backup. Two: when was the last time anyone proved it restores."
Silence. Then Dev: "We have… a cron job. I think."
"We have a cron job. I think." — four words that turn a hardware failure into a career event. This lesson is about replacing them with a better sentence: "Last night's backup restored cleanly at 6 AM; here's the log."
The customer problem
CityOps runs on a single production host. The Compose stack from Milestone 5 keeps Postgres data in the pgdata volume — on that host's disk. If the host dies, the database dies with it: every citizen report, every audit record, every row the security-review evidence pack promised to keep for seven years. The monitoring from the health-checks lesson will page Tom within minutes. Paging is not recovery.
Maria's two questions map to two different failures teams confuse constantly:
- "Do we have a backup?" is a capture question — is something, somewhere, copying the data on a schedule?
- "When did anyone prove it restores?" is a recovery question — has a human or a machine actually rebuilt a working database from that copy, recently?
Most teams can answer the first with "we have a cron job." Almost nobody can answer the second with a date. And the second is the only one that matters at 3 AM, because a backup that has never been restored is not a backup — it's a hypothesis about a backup.
Clarify the ask
Before writing a single line of backup code, the team does what it always does: clarifies. Dev's instinct is "back up everything, constantly." Tom's instinct is "back up whatever's cheapest." Maria settles it with the two numbers every disaster-recovery conversation needs:
| Term | Plain meaning | Who decides |
|---|---|---|
| RPO — recovery point objective | How much data the business can lose, measured in time. "RPO 24h" means the recoverable point must never be more than 24 hours old — which takes more than a nightly schedule: it takes the schedule, the monitor, and operational margin. A failed run or a slow upload can push the recoverable point past 24h, so "nightly" alone doesn't guarantee the number. | Maria, with engineering's cost/feasibility input |
| RTO — recovery time objective | How long the business can wait to get the required service back. "RTO 4h" means the agreed service must be available again within four hours of the failure. | Maria, with engineering's cost/feasibility input |
Maria's answers, after one uncomfortable meeting about what downtime actually costs: RPO 24 hours and RTO 4 hours. The 24h RPO doesn't mean "one nightly backup is enough, job done" — it means successful, recoverable backups must be produced frequently enough, with monitoring and operational margin, that the recoverable point stays within 24 hours. The 4h RTO means the dashboard must be back before the workday's second half. Those two numbers now size everything: backup frequency, retention, staffing, and the drill that proves it.
Notice the shape of this decision: the business owns the tolerance; engineering explains the cost and feasibility of meeting it. Maria doesn't pick RTO/RPO in a vacuum — Dev and Tom tell her what 24h vs. 1h costs in storage, complexity, and staffing — and then Maria makes the business decision with that information. Same pattern as the retention policy in the security-review lesson: the data owner decides, informed by engineering.
One principle
"A backup you haven't restored is a hope, not a backup." The backup job, the encrypted file, the retention policy — all of that is preparation. The only evidence that recovery works is a restore that actually happened, with a timestamp on it. Everything in this lesson exists to produce that timestamp, regularly.
The minimum concept: what "backed up" actually requires
Three ideas, each small, each load-bearing:
1. A backup has six properties, not one. Teams say "we back up nightly" and mean one of six different things. A real backup is scheduled (it runs without a human remembering), off-host (a copy on the same disk as the original protects against nothing), protected (encrypted, access-controlled — it holds citizen PII), retained (old copies age out on a policy, not on disk-full accidents), observable (someone or something notices when it stops happening), and restore-tested (proven, on a schedule). Miss any one and you have five-sixths of a backup, which at 3 AM rounds to zero.
2. The 3-2-1 rule, stated plainly. 3-2-1 is a resilience heuristic: keep multiple independent copies, avoid one shared failure domain, and keep at least one copy off-site/off-host. Don't get hung up on literally "two different kinds of storage" — the point is independent failure domains. For CityOps: the live database (copy 1), the nightly encrypted dump in object storage (copy 2, an independent failure domain), and a weekly copy in a second region (copy 3, off-site). The rule is a heuristic, not a law — but "we have one copy on the dead host's disk" is how you learn why the heuristic exists.
3. Full vs. incremental is an architecture decision, not a toggle. A full backup copies everything; an incremental copies only what changed since the last one. Incrementals are cheaper to store and faster to take, but restoring means replaying the full plus every incremental in order — more moving parts at the worst possible moment. One precision that matters for Postgres: pg_dump is not an incremental mechanism. CityOps starts with nightly full logical dumps: the database is small enough that simplicity beats cleverness, and the restore path is one file, one command. That's a judgment call with a documented revisit trigger: when full logical dumps no longer meet the backup window, RPO, restore-time, or storage requirements, re-evaluate the backup architecture — for example physical backups plus WAL/PITR — not simply "make pg_dump incremental."
pgdata volume"] --> B["pg_dump
nightly, 2 AM"] B --> C["Encrypt
key from secret store"] C --> D["Upload
object storage, off-host"] D --> E["Write manifest
hash, versions, format"] E --> F["Restore test
into scratch DB"] F --> G["Monitor pages on
explicit failure OR
missing expected success"] style G fill:#123324
The pipeline's shape matters more than any single step: every stage produces evidence (the manifest), and the restore test is what turns the whole chain from hope into proof. A backup job without that final box is just a very organized way of writing files nobody will ever read.
Don't memorize
- Vendor-specific backup product names — the six properties matter; the brand of object storage doesn't.
- Exact RTO/RPO numbers from this post — they're Maria's decisions for CityOps, informed by engineering, not universal constants. Your customer's numbers will differ.
Build: the backup job
The job is a shell script with four stages — dump, encrypt, upload, manifest — run by a scheduler, watched by a monitor. Each stage is small enough to read in one sitting, because the person debugging it at 3 AM will be reading it in one sitting.
#!/usr/bin/env bash
# /opt/cityops/bin/nightly-backup.sh — run by systemd timer, 02:00 daily.
# Exits nonzero on ANY failure so the monitor can see it. (set -euo pipefail
# is the whole "alert on failure" strategy for stages 1-3; stage 4 has more.)
set -euo pipefail
STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
BUCKET="s3://cityops-backups-prod"
COMPOSE="docker compose -f /opt/cityops/docker-compose.prod.yml"
APP_VERSION="$(cat /opt/cityops/VERSION)" # release identity, per the release lesson
# --- Stage 1: dump. pg_dump runs against the LIVE database; Postgres
# --- handles concurrent reads, so no downtime and no container stop.
$COMPOSE exec -T db pg_dump -U cityops --format=custom cityops \
> "$WORK/cityops-$STAMP.dump"
# --- Stage 2: encrypt BEFORE upload. The key file comes from the secret
# --- store at deploy time (Stage 5 secrets post) — never in the repo.
gpg --batch --yes --pinentry-mode loopback \
--passphrase-file /run/secrets/backup_key \
--symmetric --cipher-algo AES256 \
-o "$WORK/cityops-$STAMP.dump.gpg" "$WORK/cityops-$STAMP.dump"
# --- Stage 3: upload off-host. The bucket keeps versioning plus object lock
# --- with a defined retention window, so a compromised host cannot
# --- overwrite or delete recent history during the protected window.
aws s3 cp "$WORK/cityops-$STAMP.dump.gpg" "$BUCKET/daily/"
# --- Stage 4: manifest — operational evidence, not data proof. It records
# --- the artifact's identity (hashes, tooling versions) so the restore
# --- test can verify WHAT it downloaded. It deliberately does NOT record
# --- a live row count: production keeps writing between the dump and any
# --- later query, so a count taken after pg_dump is not from the same
# --- snapshot as the dump and must never be used to "prove" the restored
# --- DB matches.
SIZE="$(stat -c%s "$WORK/cityops-$STAMP.dump.gpg")"
SHA="$(sha256sum "$WORK/cityops-$STAMP.dump.gpg" | cut -d' ' -f1)"
cat > "$WORK/cityops-$STAMP.manifest.json" <<EOF
{
"stamp": "$STAMP",
"artifact": "daily/cityops-$STAMP.dump.gpg",
"backup_format": "pg_dump custom",
"source_database": "cityops@prod-db",
"app_version": "$APP_VERSION",
"pg_dump_version": "$($COMPOSE exec -T db pg_dump --version)",
"source_pg_version": "$($COMPOSE exec -T db psql -U cityops -tAc 'SHOW server_version')",
"backup_bytes": $SIZE,
"sha256": "$SHA"
}
EOF
aws s3 cp "$WORK/cityops-$STAMP.manifest.json" "$BUCKET/daily/"
echo "backup-ok stamp=$STAMP bytes=$SIZE"
A healthy run ends with one line in the job log — backup-ok stamp=… bytes=… — and two objects in the bucket. The monitor's job is almost insultingly simple, and it watches two things: alert if no backup-ok line and no fresh manifest appear within 20 hours (inside the 24h RPO, with margin — the RPO is a guarantee the schedule plus the monitor defends), and alert if no restore-ok appears within 30 hours. The SLO isn't "a backup file exists." It's "a recent backup exists and recent restore verification succeeded." Page on explicit failure or missing expected success — silent absence is exactly how backups die.
Two details carry the lesson. First, encryption happens before upload, with a key from the secret store — the uploaded backup object contains ciphertext rather than plaintext PII, no matter who can read the bucket. But encryption creates a dependency the job alone can't solve: a backup is only recoverable if the decryption key survives the same disaster independently. If the production host dies and the only usable copy of the key was injected onto that host, the backup is unrecoverable. Key storage and recovery are part of the DR plan — a sealed copy of the key lives outside the failed host's failure domain, and the runbook says where to obtain it (never the key itself). Second, the manifest records the pg_dump tool version and the source server version. Those fields look like trivia until the restore fails on them — see the Break section.
Manifest integrity: what the hash actually proves
A hash detects accidental corruption or a mismatch — but only when the manifest itself comes from a trusted, protected record. If an attacker could rewrite both the backup and the manifest, they could replace both the artifact and the hash. Here, the bucket's object lock and tight access controls provide part of that trust boundary: the same idea as the audit hash chain from the guardrails lesson, minus the signing. Integrity is a property of the whole system, not of one field.
Retention: a policy, not an accident
Backups accumulate. Without a pruning policy, the bill grows until someone panics and deletes things by hand — which is how you discover, months later, that the only surviving copy is from a Tuesday in March. The policy is grandfather-father-son: keep 7 daily, 4 weekly, 6 monthly copies. The customer owns the numbers (Maria again, informed by Tom's storage bill); the script enforces them.
One architectural note before the code: the bucket's object lock protects recent backups for a defined minimum period, so retention past that window is enforced by storage-side lifecycle/retention rules — because retention should continue even if the application host disappears. The Python below is the teaching planner that computes the same plan: dry-run by default, so the plan is reviewable before anything is destroyed. Deletion deserves a preview.
# prune_backups.py — GFS-style retention planner (teaching version).
# From the timeline of backup timestamps, keep:
# daily: every backup from the most recent 7 days
# weekly: the newest backup of each ISO week, for the 4 weeks before that
# monthly: the newest backup of each calendar month, for ~6 months before that
# Keep = the union. Deletion is the one irreversible action in this post;
# the default is dry-run: print the plan, destroy nothing.
import datetime, subprocess
DAILY_DAYS, WEEKLY_DAYS, MONTHLY_DAYS = 7, 35, 215
def select_gfs(dts, ref):
"""dts: sorted datetimes. Returns the set of indices to keep."""
keep = set()
for i, dt in enumerate(dts): # daily window
if 0 <= (ref - dt).days < DAILY_DAYS:
keep.add(i)
seen = set() # weekly representatives
for i in range(len(dts) - 1, -1, -1):
if DAILY_DAYS <= (ref - dts[i]).days < WEEKLY_DAYS:
wk = dts[i].isocalendar()[:2]
if wk not in seen:
seen.add(wk); keep.add(i)
seen = set() # monthly representatives
for i in range(len(dts) - 1, -1, -1):
if WEEKLY_DAYS <= (ref - dts[i]).days < MONTHLY_DAYS:
m = (dts[i].year, dts[i].month)
if m not in seen:
seen.add(m); keep.add(i)
return keep
if __name__ == "__main__":
out = subprocess.run(["aws", "s3", "ls", "s3://cityops-backups-prod/daily/",
"--recursive"], capture_output=True, text=True,
check=True).stdout
keys = sorted(l.split()[-1] for l in out.splitlines()
if l.endswith(".manifest.json"))
dts = [datetime.datetime.strptime(k.split("-")[-1].split(".")[0],
"%Y%m%dT%H%M%SZ") for k in keys]
ref = datetime.datetime.now(datetime.timezone.utc).replace(tzinfo=None)
keep_idx = select_gfs(sorted(dts), ref)
ordered = sorted(range(len(keys)), key=lambda i: dts[i])
keep = {ordered[i] for i in keep_idx}
delete = [k for j, k in enumerate(keys) if j not in keep]
print(f"manifests: {len(keys)}, keep: {len(keep)}, delete: {len(delete)}")
for d in delete:
print(" DELETE", d)
# --apply path omitted: it deletes both the .gpg and its .manifest.json
# for each DELETE line, then exits nonzero if any deletion fails.
# Production prefers the bucket's lifecycle rules for the real enforcement.
On a sample timeline of 20 manifests — 7 dailies, 4 weekly representatives, 6 monthly representatives, plus 3 mid-week/mid-month duplicates — this prints manifests: 20, keep: 17, delete: 3, dropping exactly the duplicates the policy says are redundant. The script computes the plan from the actual manifests; the selection rule (not a one-shot label per backup) is what makes it genuinely grandfather-father-son.
The restore test: the only proof that counts
Every morning at 6, after the backup lands, a second job restores the latest backup into a scratch Postgres container and verifies it. This is the timestamp Maria asked for:
#!/usr/bin/env bash
# /opt/cityops/bin/verify-restore.sh — runs 06:00 daily, after the backup.
# Verifies the LATEST backup: hash -> decrypt -> restore -> schema + invariants.
# Honesty label: a passing run proves THAT backup restored TODAY.
# It does not prove next week's backup will restore.
set -euo pipefail
BUCKET="s3://cityops-backups-prod"
WORK="$(mktemp -d)"; trap 'rm -rf "$WORK"' EXIT
# Newest manifest. Filenames use UTC YYYYMMDDTHHMMSSZ, so lexical order
# equals chronological order — that is what makes `sort | tail -1` select
# the newest backup rather than a random one.
LATEST="$(aws s3 ls "$BUCKET/daily/" --recursive | grep '\.manifest\.json$' \
| sort -k4 | tail -1 | awk '{print $4}')"
aws s3 cp "$BUCKET/$LATEST" "$WORK/manifest.json"
# One canonical artifact key, read from the manifest — no string surgery
# on filenames in two different forms.
ARTIFACT="$(python3 -c "import json;print(json.load(open('$WORK/manifest.json'))['artifact'])")"
aws s3 cp "$BUCKET/$ARTIFACT" "$WORK/backup.gpg"
# Verify the manifest's SHA-256 BEFORE decrypting. A recorded hash that the
# recovery pipeline never checks is documentation, not evidence.
WANT_SHA="$(python3 -c "import json;print(json.load(open('$WORK/manifest.json'))['sha256'])")"
GOT_SHA="$(sha256sum "$WORK/backup.gpg" | cut -d' ' -f1)"
[ "$GOT_SHA" = "$WANT_SHA" ] || { echo "restore-HASH-MISMATCH" >&2; exit 1; }
gpg --batch --yes --pinentry-mode loopback \
--passphrase-file /run/secrets/backup_key \
-o "$WORK/restore.dump" -d "$WORK/backup.gpg"# Scratch DB at the major version the manifest records. postgres:16
# constrains the major version but is still a moving tag — the manifest's
# recorded pg_dump version is what makes tooling drift visible, not silent.
PGV="$(python3 -c "import json;m=json.load(open('$WORK/manifest.json'));print(m['source_pg_version'].split('.')[0])")"
docker run -d --name restore-test -e POSTGRES_PASSWORD=
trap 'docker rm -f restore-test >/dev/null; rm -rf "$WORK"' EXIT
for i in $(seq 1 30); do # bounded readiness wait
docker exec restore-test pg_isready -U postgres >/dev/null 2>&1 && break
[ "$i" = 30 ] && { echo "restore-scratch-not-ready" >&2; exit 1; }
sleep 2
done
docker exec restore-test psql -U postgres -tAc \
'CREATE DATABASE cityops_restore' >/dev/null
docker exec -i restore-test pg_restore -U postgres -d cityops_restore \
< "$WORK/restore.dump"
# Smoke invariants — computed against the RESTORED backup itself, never
# against the live database, whose rows may have moved since the dump.
TABLES="$(docker exec restore-test psql -U postgres -d cityops_restore -tAc \
"SELECT count(*) FROM information_schema.tables WHERE table_schema='public'")"
REPORTS="$(docker exec restore-test psql -U postgres -d cityops_restore -tAc \
'SELECT count(*) FROM reports')"
echo "restore-ok tables=$TABLES reports=$REPORTS stamp=$(basename "$LATEST" .manifest.json)"
Read the honesty label on this script carefully. Daily verification means: the artifact hash is valid, decryption succeeds, pg_restore succeeds, the expected schema objects exist, and a small set of data invariants holds. The row counts are a smoke invariant, not proof of database equivalence — they don't prove every table, relationship, attachment, or audit record survived; they prove the restored database isn't empty or structurally wrong. The quarterly drill does the stronger validation: full environment, the application boots, a smoke workflow runs, and the RTO is measured. One passing test is one data point — the power is in the daily cadence, which turns "did it restore?" from a question into a time series.
host-down alert"] B --> C["Provision replacement host
runbook step 1"] C --> D["Restore latest verified backup
into fresh Postgres"] D --> E{"Schema + data checks
pass?"} E -->|no| F["Escalate: try previous
daily, then weekly"] E -->|yes| G["Bring up API + worker
Compose stack"] G --> H["Readiness check
+ smoke queries"] H --> I["API + worker + dashboard
usable — RTO clock stops"] style F fill:#3a1620 style I fill:#123324
One definition, used everywhere in this lesson: the RTO clock stops when the agreed recovery service level is available — for CityOps, the API and worker functioning and the dashboard usable — not when a single process answers HTTP 200.
Break: three ways the backup lies to you
Dev builds the pipeline, the manifests accumulate, the restore test goes green for a month. Then the team does what this curriculum always does — tries to break it, on purpose, before 3 AM does it for them.
Failure 1: the backup that never ran. Six weeks after setup, Tom notices the bucket's newest manifest is six weeks old. The systemd timer was defined on the old staging host; when production moved to the new host, nobody re-created the timer. The fix is the monitor from the Build section: alert on the absence of backup-ok, not just on failures. A backup job that can die silently is a backup job that will. (This is the monitoring lesson applied to the backup itself: monitor the outcome, not the process.)
Failure 2: the restore that fails on tooling. During a drill, the restore test dies with pg_restore: error: unsupported version (1.15) in file header. The dump was taken by Postgres 16's pg_dump in custom format; the scratch container ran a hand-written postgres:15 tag, so its older pg_restore couldn't read the newer archive format. The manifest recorded the pg_dump version and the source server version — the evidence was there, the script just wasn't reading it yet. The fix: the verify script now restores with tooling matched to the manifest (as written above). The lesson is precise: record the backup tool version and the source server version; restore compatibility depends on the dump format, the tool version, and the target PostgreSQL's compatibility — don't blindly use latest, and don't assume the server versions must match exactly. (Logical dumps are routinely used across versions for upgrades; what breaks is an older tool reading a newer format.) Version the system, the backup, and the restore tooling — the Stage 4 rule, extended to disaster recovery.
Failure 3: the backup that leaks. Lisa, doing her quarterly review from the security-review lesson, asks who can read the backup bucket. Answer: the reporting service account, which needs "read-only, no PII columns" on the live database — but the bucket holds full database dumps, PII columns included, and an early version of the job uploaded them unencrypted for a week before the gpg stage was added. Two holes, one lesson: backups are PII in concentrated form — every citizen phone number, neatly packaged. The fixes: encrypt-before-upload (already in the job), bucket policy restricted to the backup role only, and object lock with a defined retention window so that properly configured immutable retention can prevent ordinary overwrite or deletion during the protected window. The security boundary includes the backups; the PII inventory from the security-review post gains a row for the bucket itself.
Must know
- Monitor the absence of success, not just the presence of failure. A backup job that dies silently is the common case, not the edge case.
- Record the backup tool version and the source server version, and restore with compatible tooling. The dump format belongs to a tool chain, not to "latest."
- Backups are concentrated PII. Encrypt before upload, restrict the bucket, and put the bucket in the PII inventory.
- A backup is only recoverable if the backup, its decryption key, the restore tooling, and the runbook all survive the same disaster. The key's recovery is part of the DR plan, not an afterthought.
DROP TABLE typo, or a ransomware-style encryption of everything the compromised role could write are all outside the provider's promise. Durability is the provider's promise about their service. Recovery is your responsibility for your data.Productionize: the drill, the runbook, the cadence
The pipeline runs daily. Productionizing means the organization can recover, not just the script — which needs three things the code can't provide: a practiced drill, a written runbook, and a cadence that survives staff turnover.
The quarterly drill. Every quarter, the team runs a planned recovery drill — scheduled, staffed, with some scenarios and details withheld until the exercise begins — and executes the DR runbook for real: provision a fresh host, restore a backup, bring up the Compose stack, run the readiness checks and smoke queries, and stop the clock. The drill measures the actual RTO against Maria's 4-hour target and writes the result into the DR summary — measured, not estimated. And it doesn't always restore the latest backup: periodically the drill restores an older retained backup, because the daily test only ever proves the newest one — it never proves a six-month-old monthly copy is still decryptable and restorable. The first drill found that nobody knew the object-storage credentials for the fresh host (they lived in Tom's shell history); the runbook gained a step. Drills don't test the backup. They test the team.
The runbook. docs/dr-runbook.md in the repo, printed on Tom's desk (Milestone 5's rule: incidents don't wait for Wi-Fi). It names every command in the flowchart above, the bucket name, where to obtain the decryption key and the object-storage credentials (the locations, never the secrets themselves), the expected outputs, and the escalation path when the latest backup fails verification (try the previous daily, then the newest weekly — the retention policy exists partly so there's always a fallback). The runbook is reviewed after every drill and after every real incident, with a date and an owner on each revision — Lisa's rule from the security-review post, applied to recovery.
The cadence. Daily: backup, restore test, manifest. Weekly: Tom eyeballs the backup dashboard (freshness, sizes, any anomalies — a backup that suddenly halves in size is a story worth reading). Quarterly: the full drill. CityOps reviews RTO/RPO annually and after major business or architecture changes — the numbers are Maria's to change with engineering's input, and the mechanism must be re-sized if she does.
| What | Cadence | Proves | Owner |
|---|---|---|---|
| Nightly backup + manifest | Daily, 02:00 | Capture works | Timer + monitor |
| Restore into scratch DB | Daily, 06:00 | Latest backup restores | Verify job |
| Freshness + restore-ok alerting | Continuous | Silence is noticed | Monitoring |
| Full DR drill, timed | Quarterly (planned) | RTO target is real | Tom + Dev |
| RTO/RPO review | Annually + after major changes | Targets still fit the business | Maria + engineering |
One more production detail, because the security review will ask: the audit trail from the security-review post has a 7-year retention requirement. The audit log is data too — it ships append-only to independent object storage, with the same encryption and retention classes as the database backups. But note the scope carefully: at recovery time, nobody "restores" seven years of audit logs before the API serves traffic. Recovery re-establishes the logging pipeline and verifies the historical logs remain accessible; the API's availability doesn't wait for the archive.
Later, not now
- Hot standby / multi-region failover — the quarterly drill against RTO 4h is the current commitment; automatic failover is a later scaling decision with its own cost and complexity.
- Point-in-time recovery with WAL archiving — worth adding when RPO 24h stops being acceptable; until Maria changes the number with engineering's input, nightly full dumps are the honest fit.
Communicate: the one-page DR summary
Maria doesn't read runbooks. She needs one page she can forward to procurement — the same audience as the security-review evidence pack — that answers her original two questions with dates:
The DR summary (updated after every drill)
- RTO 4h / RPO 24h — committed targets, owned by Maria with engineering's cost/feasibility input, reviewed annually and after major business or architecture changes.
- Last proven restore: date of the most recent green restore test and the most recent quarterly drill, with the measured recovery time next to the 4h target.
- What gets recovered, in what order: restore state first, then dependencies and services in the order required to reach the agreed business recovery point — for CityOps: Postgres, then the audit-log pipeline, then the API, then the worker; dashboard last. The order is in the runbook because order matters under pressure.
- What's covered: Postgres dumps (this lesson's pipeline), the audit-log stream (append-only, independent object storage — pipeline re-established at recovery, history verified accessible), and object-storage attachments (versioned bucket + cross-region copy — existing infrastructure, outside this lesson's code sample). What's not covered is listed too — honesty about scope is what makes the page credible.
This page is the Communicate step of the lesson DNA made concrete: the learner's job isn't just to build the pipeline, it's to explain the outcome to a human who signs contracts. "Our RTO is 4 hours, last proven at 2h40m in the October drill" is a sentence Maria can say in a meeting. "We have a cron job, I think" is not.
Note the discipline in that example sentence: the 4h is a committed target (Maria's decision, informed by engineering), the 2h40m is a measured result (the drill's clock). Targets and measurements are different kinds of numbers — the summary keeps them in separate columns, and so should you.
Where this lands in CityOps
- The
pgdatavolume gets its pipeline. The nightly job, the encrypted bucket, the 6 AM restore test, and the freshness + restore-ok alerting from this lesson run against the production Compose stack from Milestone 5. The single host is still a single host — but it's no longer a single copy. - The evidence pack gains a row. The security-review post's PII inventory and access matrix now cover the backup bucket: what's in it (full PII), who can read it (the backup role only), and how long copies live (the retention policy). Lisa's next questionnaire answer on backups is a link to the DR summary.
- The release pipeline and the recovery pipeline are siblings. The release-strategies lesson (same stage, by title) makes deploys safe; this lesson makes disasters survivable. One guards against bad code, the other against a dead machine — CityOps needs both, and the runbook names which one to reach for.
- Backups join the audit trail. Every backup and every restore test writes to the same structured logging the monitoring lesson set up, with the manifest's hash as the correlation ID. When Maria asks "prove October's backups ran," the answer is a query, not a shrug.
from latest verified dump"] B --> C["Re-establish audit-log pipeline
verify history accessible"] C --> D["Bring up API
readiness gates traffic"] D --> E["Bring up worker
re-pull missed 311 window"] E --> F["API + worker + dashboard usable
RTO clock stops"] style B fill:#3b1d1d,stroke:#f87171 style F fill:#123324
The order is the lesson: restore state first, then dependencies and services in the order required to reach the agreed business recovery point — for CityOps, Postgres, then the audit-log pipeline, then the API, then the worker, then the dashboard humans look at. A runbook that restores the dashboard before the database is a runbook written by someone who has never done it at 3 AM.
Zoom out, and the whole lesson is one architecture with two halves:
RPO + RTO"] --> B["CAPTURE"] & C["RECOVER"] B --> B1["consistent backup"] --> B2["encrypt"] --> B3["off-host copy"] --> B4["retention"] --> B5["freshness monitor"] C --> C1["protected backup"] --> C2["verify hash/artifact"] --> C3["decrypt"] --> C4["restore"] --> C5["data/schema checks"] B5 --> D["DR DRILL"] C5 --> D D --> E["application + people +
credentials + runbook"] --> F["measured RTO"] style F fill:#123324
And the distinction the entire lesson exists to teach, in three levels: having a backup is not the same as being able to restore it, and being able to restore it is not the same as being able to recover the business within the promised RTO. Capture gives you the first. The restore test gives you the second. Only the drill — the application, the people, the credentials, the runbook, the measured clock — gives you the third.
Field check
- Tom's host dies at 3 AM. The last manifest is from 26 hours ago and the freshness alert never fired. Name two separate failures this implies, and which one is worse.
- Your quarterly drill restores the database in 2h40m against an RTO target of 4h. Your manager asks you to "commit to 2 hours" in the next customer meeting. What do you say?
- The restore test has been green for 90 days. Lisa asks: "Does that prove next week's backup will restore?" What's the honest answer, and what would strengthen it?
- A teammate suggests skipping encryption because "the bucket is private and only our role can read it." Give the two-sentence rebuttal, using the PII argument from this lesson.
- Maria wants RPO tightened from 24h to 1h. List everything that has to change — and who has to agree first.
Answers
1. Failure one: the backup job hasn't produced a manifest in 26 hours (the job is broken or gone). Failure two: the freshness monitor didn't alert on the absence (the monitor is broken or was never wired to this job). The second is worse — a broken job with a working monitor gets fixed in the morning; a broken job with a broken monitor gets discovered at 3 AM during a disaster.
2. You say no — politely, and with the reason: 2h40m is one measurement under drill conditions, not a commitment. The committed target stays 4h until several drills cluster well under 2h and Maria agrees the business wants the tighter number. Targets are promises; measurements are evidence. Don't convert one into the other in a meeting.
3. No — it proves the backups that were actually tested during those 90 days restored, which is strong evidence about the pipeline's health but not a guarantee about the next one. What strengthens it: the freshness alert (catches silent death), the recorded restore tooling versions (removes drift), the hash verification (catches corruption), and the quarterly full drill (tests the parts the daily test doesn't — provisioning, networking, the humans). Confidence comes from layers, not from a streak.
4. "The bucket holds concentrated PII — every citizen phone number, neatly packaged — so 'private bucket' is one control, not a strategy. Encryption before upload means a bucket-policy mistake or a compromised credential exposes ciphertext, not citizen data."
5. Maria has to agree first — RPO is her business decision, informed by engineering's cost and feasibility analysis, and 1h changes cost and complexity. Then: re-evaluate the backup architecture rather than defaulting to hourly full dumps — WAL/PITR is likely more appropriate for a true one-hour RPO than repeated full logical dumps, which get expensive fast. The backup window, storage bill, and retention policy get re-sized; the restore test runs more often; the drill re-measures against the unchanged-or-tightened RTO; and the DR summary gets new numbers with new dates.