The client sends a cloud architecture diagram — boxes, arrows, acronyms. Your team needs a VM, a database, and a place to put files, but nobody can tell from the diagram where any of that lives. This lesson is about reading the map before you walk the territory.
The customer problem
Monday, 9:12 AM. Maria — the city's program manager — forwards an email from city IT with a PDF attached and one line of her own:
"They sent their cloud architecture diagram. Can you confirm where CityOps will run, where the data lives, and who can access it?"
Tom opens the PDF. It is dense: a VPC with two subnets, an ALB, ECS tasks, an RDS instance, an S3 bucket, a KMS key, IAM roles with names like CityOpsTaskRole. Dev squints at it and asks the question that kills the morning:
"So… do we get a VM or what?"
Nobody on the team can answer Maria's email. Not because anyone is unqualified — because nobody was ever taught to read one of these diagrams. The project stalls for a week while Tom reverse-engineers it box by box.
That week was avoidable. This is the lesson that avoids it.
Clarify the ask
When a customer hands you an architecture diagram, the ask is almost never "redesign our cloud." It is four concrete questions, and your job is to answer them from the diagram:
- Where does the app run? — which compute box is ours, and how do we reach it?
- Where does the data live? — which database, which storage, in which region?
- Who can touch what? — which identities have access to resident data?
- How does traffic flow? — what path does a request take from the internet to the database?
This frames the whole lesson as literacy, not mastery. A Forward Deployed Engineer reads diagrams, requests resources correctly, and spots risks — the city already has cloud architects for the deep work. Be honest about which side of that line you're on, and you'll earn trust fast. Bluff across it, and you'll lose it faster.
Must know (this lesson)
- Read a cloud architecture diagram: compute, storage, network, identity
- Request the right resource from Tom in the right terms (VM vs managed DB vs object storage)
- Spot the three risks Lisa will ask about: public exposure, over-broad access, unencrypted data
- Read a Terraform file and a Terraform plan well enough to review one
Later (not this lesson)
- Designing landing zones, multi-region failover, or cost optimization
- Writing production Terraform modules from scratch
- Kubernetes cluster administration
Minimum concept: the cloud in one page
Cloud marketing has hundreds of services. Nearly all of them are combinations of four ideas: compute (where code runs), storage (where data rests), networking (how they reach each other), and identity (who is allowed to do what). Learn these four and a diagram stops being alphabet soup.
Compute: VMs, containers, serverless
| Option | What it is | You manage | FDE cares because |
|---|---|---|---|
| Virtual machine (VM) | A whole virtual server — your OS, your rules | OS patching, runtime, app | Most common in enterprise diagrams; simple mental model |
| Containers (ECS / Kubernetes) | Your app packaged with its dependencies, scheduled onto shared machines | The image and its config; the platform handles the machines | Where CityOps will most likely run — you read manifests, not servers |
| Serverless functions | Code that runs on events, no server visible at all | Just the function code and its triggers | Common for glue: webhooks, scheduled jobs, image resizing |
The pattern across all three: you trade control for convenience as you move down the list. When Dev asks "do we get a VM or what," the grown-up answer is "it depends what the diagram already provides — let's look."
Storage: three kinds, three jobs
| Kind | Behaves like | Use it for | Example services |
|---|---|---|---|
| Block storage | A hard drive attached to one VM | OS disks, scratch space | EBS, Persistent Disk, Managed Disks |
| Object storage | An infinite filing cabinet addressed by key | Complaint photos, exports, backups, static assets | S3, Google Cloud Storage, Blob Storage |
| Managed database | A database with backups, patching, and failover handled | Anything relational CityOps stores | RDS, Cloud SQL, Azure Database |
The FDE instinct to build here: prefer the managed option unless you have a reason not to. Running your own database on a VM means you own backups, patching, and the 3 AM pages. A managed database moves that to the provider. (Self-hosted costs don't vanish, they shift — hardware, operations, and your sleep.)
Networking in one page
Every enterprise diagram has the same skeleton. Learn it once:
- Region / Availability Zone (AZ) — the geographic area your resources live in. Data residency ("resident data stays in-state") is a region choice. AZs are isolated buildings inside a region; spreading across two is how you survive one burning down.
- VPC (Virtual Private Cloud) — your private network inside the cloud. Everything the city runs sits inside one.
- Subnet — a slice of the VPC. Public subnets can be reached from the internet; private subnets cannot. The single most important line on any diagram: is the database in a private subnet?
- Security group — a firewall attached to a resource. It lists which traffic is allowed in.
0.0.0.0/0means "the whole internet" — on a database, that's the diagram screaming at you. - Load balancer — the front door. Internet traffic hits it; it forwards to your app instances and checks they're healthy.
Identity: IAM in thirty seconds
IAM (Identity and Access Management) answers "who can do what to which resource." Three parts: an identity (a user, or a role the app assumes), an action (read, write, delete), and a resource (the bucket, the database). The principle Lisa will quiz you on is least privilege: every identity gets only the permissions it needs, nothing more. When the diagram shows the app running with a role that can only read one bucket and write to one database, that's least privilege done right. When it shows one admin key shared by everything, that's the opposite.
The Rosetta Stone
AWS, Google Cloud, and Azure all sell the same four ideas with different names. You don't need to memorize SKUs — you need to recognize the shape:
| Concept | AWS | Google Cloud | Azure |
|---|---|---|---|
| Virtual machine | EC2 | Compute Engine | Virtual Machines |
| Object storage | S3 | Cloud Storage | Blob Storage |
| Managed Postgres | RDS | Cloud SQL | Database for PostgreSQL |
| Container orchestration | ECS / EKS | Cloud Run / GKE | Container Apps / AKS |
| Serverless functions | Lambda | Cloud Functions | Functions |
| Key management | KMS | Cloud KMS | Key Vault |
The literacy tour: Kubernetes, Helm, Terraform
Three more names appear on every enterprise diagram. You won't administer them — but you will read them.
Kubernetes runs containers at scale. Its vocabulary is small: a Pod is one running copy of your container; a Deployment says "keep 3 copies of this Pod running"; a Service gives those copies one stable address. If you can read those three resources in a YAML manifest, you can follow any Kubernetes conversation.
Helm is templated Kubernetes YAML — a chart is a package of manifests with the environment-specific values (image tag, replica count, database URL) filled in at install time. When Tom says "we deploy with Helm," he means the YAML is generated, not hand-written.
Terraform describes infrastructure as code in HCL files: you declare "there should be a database with these settings," and Terraform creates it. The two commands that matter for reading: terraform plan shows what would change (review this like a code diff), and terraform apply makes it real. State is Terraform's memory of what it built — whoever holds the state file controls the infrastructure, which is why Lisa cares where it's stored.
Don't memorize
- Every service name on each cloud — recognize the four shapes, look up the rest
- Terraform resource argument lists — read the intent, check the docs for the details
- Kubernetes API versions — they change; the Pod/Deployment/Service model doesn't
Build: reading the city's diagram
Let's rewind to Monday morning and read the PDF properly. Here is the same architecture, redrawn the way you should sketch it in your notebook — traffic flow first, everything else second:
public subnet] ALB --> App[CityOps API
containers, private subnet] App --> DB[(Managed Postgres
private subnet)] App --> Obj[(Object storage
complaint photos)] KMS[KMS key] -. encrypts .-> DB KMS -. encrypts .-> Obj Role[IAM role:
CityOpsTaskRole] -. assumed by .-> App
Now answer Maria's four questions straight from the picture:
- Where does the app run? In containers on the private subnet — no VM to SSH into, no server to patch. We deploy a container image; the platform runs it.
- Where does the data live? Postgres in a managed database (backups included), complaint photos in object storage, both in the city's region.
- Who can touch what? The app assumes
CityOpsTaskRole— one IAM role, scoped to that database and that bucket. - How does traffic flow? Internet → load balancer → app → database/storage. The database has no path from the internet at all.
That's the whole email to Maria, and it took ten minutes once you know what to look for.
Reading the Terraform behind the diagram
Diagrams lie by omission; the infrastructure code doesn't. When Tom shares the Terraform, read it like this — a simplified but structurally honest excerpt:
terraform {
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.0"
}
}
}
provider "aws" {
region = var.region # data residency lives in this one line
}
variable "region" {
description = "Cloud region - must match the city's data-residency requirement"
type = string
default = "us-west-2"
}
resource "aws_vpc" "cityops" {
cidr_block = "10.0.0.0/16"
tags = { Name = "cityops" }
}
resource "aws_subnet" "private_a" {
vpc_id = aws_vpc.cityops.id
cidr_block = "10.0.1.0/24"
availability_zone = "${var.region}a"
tags = { Name = "cityops-private-a" }
}
resource "aws_db_instance" "cityops" {
identifier = "cityops-postgres"
engine = "postgres"
instance_class = "db.t4g.micro"
allocated_storage = 20
storage_encrypted = true
backup_retention_period = 7
skip_final_snapshot = true
tags = { Name = "cityops-postgres" }
}
Reading notes for the FDE eye: the region is a variable (ask what it's set to — that's your data-residency answer); the subnet is named private (good); the database has storage_encrypted and a 7-day backup_retention_period (both answers Lisa will want). You don't need to write this from scratch — you need to read it and ask the right follow-ups.
Reading a Terraform plan
Before anything changes, Tom runs terraform plan. Review it the way you'd review a pull request — what is being added, changed, or destroyed (trimmed excerpt):
Terraform will perform the following actions:
# aws_db_instance.cityops will be created
+ resource "aws_db_instance" "cityops" {
+ allocated_storage = 20
+ backup_retention_period = 7
+ engine = "postgres"
+ identifier = "cityops-postgres"
+ storage_encrypted = true
+ instance_class = "db.t4g.micro"
...
}
Plan: 4 to add, 0 to change, 0 to destroy.
The line to read first is always the summary: "0 to destroy" means nothing existing is being torn down. A plan that says "1 to destroy" on a database deserves a phone call before anyone types apply.
Reading a Kubernetes manifest
When the app runs on Kubernetes, Tom hands you a manifest instead of a server. Here's the shape — a Deployment for the CityOps API (trimmed):
apiVersion: apps/v1
kind: Deployment
metadata:
name: cityops-api
spec:
replicas: 2
selector:
matchLabels:
app: cityops-api
template:
metadata:
labels:
app: cityops-api
spec:
containers:
- name: api
image: registry.city.gov/cityops/api:1.4.2
ports:
- containerPort: 8000
env:
- name: DATABASE_URL
valueFrom:
secretKeyRef:
name: cityops-db
key: url
readinessProbe:
httpGet:
path: /health
port: 8000
What to notice: replicas: 2 (two copies — a deploy won't take the site down); the image uses a versioned tag 1.4.2, not latest (a versioned tag narrows what runs; a digest would pin it exactly — see the Docker lesson); the database URL comes from a secret, not hard-coded (good — matches the Secrets & Security Basics rule); and there's a readinessProbe on /health (the load balancer only sends traffic to copies that pass it — the same liveness-vs-readiness distinction from the deployment lessons).
Break: three ways to misread the diagram
Literacy means knowing what wrong looks like. Each of these has bitten a real deployment:
1. The app in the public subnet. The diagram shows the API containers next to the load balancer in the public subnet "for simplicity." That means every container has a direct path to and from the internet — the load balancer is no longer the only front door. The fix: app and database both live in private subnets; only the load balancer is public. If the subnet label doesn't say private, ask.
2. The database security group allows 0.0.0.0/0. Someone opened the database to the world "temporarily for debugging" and the Terraform still carries it. On the diagram this looks like an arrow from the internet to the database — the single most alarming arrow you can see. The fix: the DB security group allows traffic only from the app's security group, on the database port, nothing else.
3. Credentials baked into the picture. The Terraform (or a startup script) contains a database password in plain text, or the diagram notes "uses the admin key for all services." That password now lives in version control, in CI logs, and in everyone's memory of the repo. The fix: secrets come from the secret manager or Kubernetes secrets at runtime — never from code, never from the diagram.
Notice the pattern: every one of these is visible before anything is deployed, if you read the diagram and the plan with Lisa's questions in mind.
Productionize: the FDE's pre-deploy checklist
You don't need to be the cloud architect. You need to be the person who asks these questions before CityOps goes live on the city's cloud — and gets answers in writing:
- Region confirmed — the region matches the city's data-residency requirement (ask Tom for the value of
var.region, not "where is it roughly"). - Private subnets — app and database in private subnets; only the load balancer is public.
- Database — managed, encrypted at rest (KMS), automated backups with a stated retention period, and a restore that's been tested at least once.
- Object storage — versioning on (so an overwrite doesn't destroy evidence), lifecycle rules for old complaint photos if the retention policy requires it.
- Identity — the app runs as a scoped IAM role; no long-lived access keys in code or config.
- Health checks — the load balancer has a health-check path (
/health) and the app answers it; unhealthy copies get replaced, not traffic. - Runbook contact — you know who to page at the city when the cloud side breaks, because it won't be you with the console access.
And the email to Tom that gets you all of this without sounding lost:
Subject: CityOps infra details for go-live checklist
Hi Tom — before we schedule go-live, could you confirm:
1. The region CityOps will run in (for our data-residency doc)
2. Whether the app + database subnets are private, with only the
load balancer public
3. The database backup retention period + KMS key for encryption
4. The IAM role the app will assume, and what it's scoped to
5. Who we contact on your team if the cloud side has an incident
A terraform plan output or the current diagram is plenty — no
need for anything formal.
Thanks!
That email is the whole lesson in action: literate enough to ask precisely, humble enough not to pretend you're the architect.
Communicate: translating infra into answers
Same facts, three audiences. Maria asks business questions — answer in business terms:
- Maria: "Resident data stays in our region, encrypted, with daily backups. Only the CityOps app itself can reach the database — there's no public path to it. Here's the one-page diagram."
- Lisa: "The app assumes CityOpsTaskRole, scoped to one database and one bucket. Storage is encrypted with the city's KMS key. Here's the IAM policy and the Terraform plan showing zero resources destroyed."
- Tom: "We're targeting the private subnets behind the ALB, image pinned to 1.4.2, DATABASE_URL from the cityops-db secret, readiness on /health. Anything in that list that conflicts with your standards?"
Notice what you never say to anyone: "the cloud handles it." The cloud handles the machines. The answers above are yours.
CityOps: the reference architecture
Here's where everything in this stage lands for CityOps — the same four questions, answered for our system:
public subnet] ALB --> API[CityOps API - FastAPI
2 containers, private subnet] API --> PG[(Managed Postgres
private subnet, encrypted, 7-day backups)] API --> S3[(Object storage
complaint photos, versioned)] API --> LLM[Model endpoint
egress via NAT] Role[IAM role] -. read/write .-> S3 Role -. connect .-> PG
Trace one request: a resident uploads a complaint photo. It hits the load balancer, lands on a healthy API container (readiness probe passed), the record goes to Postgres, the photo goes to object storage — and the model call for triage leaves through controlled egress. Every box on this diagram maps to one line on your pre-deploy checklist. When the city's diagram differs from this shape, the differences are what you ask Tom about.
Field check
- Tom sends you a diagram where the managed database sits in the public subnet "so the team can connect from home." What's wrong, and what do you propose instead?
- The Terraform plan says "2 to add, 1 to change, 1 to destroy" — the destroy is the production database. What do you do before anyone runs
apply? - Lisa asks: "Which identities can read the complaint-photo bucket?" Walk through how you'd answer from the IAM policy and the bucket policy.
- Maria asks: "If the cloud region has an outage, do we lose resident data?" What in the diagram and Terraform gives you the honest answer?
- You see
image: registry.city.gov/cityops/api:latestin a manifest. Why does the pinned tag matter, and what breaks when it'slatest?
1 — Database in the public subnet
Wrong because the database gets a path to and from the internet — anyone who guesses credentials (or finds them leaked) can reach it directly. Propose: move it to the private subnet and give the team access through the city's VPN or a bastion host, with the DB security group allowing only the app's security group plus that admin path.
2 — Plan destroys the production database
Stop. Do not apply. Ask Tom what change caused the replacement (often a renamed identifier or changed engine parameter forces replacement), confirm there's a tested backup and a maintenance window, and get explicit written sign-off. A destroy on stateful data is never routine.
3 — Who can read the bucket
Read both policies: the IAM policies attached to roles/users (who is granted access) and the bucket policy (what the bucket allows, including any public or cross-account statements). The effective access is the union. Report the list of identities, and flag anything broader than the app's role and the data team's read role.
4 — Region outage and data loss
Look for: multi-AZ database deployment (survives one building failing, not the whole region), backup retention and whether backups are copied to another region, and object-storage replication. Honest answer is usually "we survive an AZ failure; a full region loss needs cross-region backups, which we should confirm we have."
5 — The latest tag
latest moves — two deploys of "the same" manifest can run different code, and a rollback can't name what it rolled back to. Versioned tags (1.4.2) narrow what runs and make incidents traceable to a specific release; a digest would pin the exact bytes. What breaks with latest: rollbacks, audits, and your ability to say what was running when the bug happened.