Residents don't search like databases. They type "my street smells" and expect the garbage complaints. Keyword search fails them — so you'll turn words into numbers and make "find similar" an arithmetic problem.
pip install numpy. No API keys, no model downloads — every number in this lesson is computed live below.The customer problem
Lisa from the 311 team has a request that sounds simple: "Can residents search complaints in plain English? Someone types 'my street smells' and it should find the garbage complaints."
You try the obvious thing first — keyword search against the complaint text:
-- the obvious thing
SELECT request_id, complaint_text
FROM complaints
WHERE complaint_text LIKE '%smell%';
-- 0 rows.
-- The actual complaints say "rotting garbage bags on the curb",
-- "overflowing dumpster behind the deli", "foul odor near the park".
-- None of them contain the word "smell".
Keyword search needs the same words. Humans don't use the same words. "My street smells," "rotting garbage," "foul odor," "trash piled up" — four ways to say one thing, zero shared keywords. This is the vocabulary mismatch problem, and no amount of LIKE clauses fixes it.
Clarify the ask
Before reaching for anything called "AI," pin down what Lisa actually needs:
- Input: a plain-English sentence from a resident.
- Output: a ranked list of the most relevant 311 complaints, with the obviously-right ones near the top.
- Not asked for: a chatbot, an answer in sentences, or perfect understanding. Just better retrieval than keywords.
And the honest success bar: the top few results should be relevant most of the time, and when the search fails it should fail visibly — not with confident-looking garbage. You'll hold yourself to that in the Communicate section.
The minimum concept: words become numbers
An embedding is a list of numbers that represents a piece of text. The trick: embedding models map text into vectors so that texts the model considers similar tend to land closer together in that number-space. "My street smells" and "rotting garbage bags on the curb" get number-lists that point in nearly the same direction. "Broken streetlight on 5th" points somewhere else entirely.
Once meaning is geometry, "find similar" becomes arithmetic. The standard ruler is cosine similarity: the cosine of the angle between two vectors. Mathematically it ranges from -1 (opposite directions) to 1 (same direction). Higher generally means more similar under this representation — but don't treat any particular score as a universal relevance threshold; score distributions depend on the embedding model, the domain, and the corpus.
import numpy as np
def cosine(a, b):
return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
# tiny 3-number "embeddings", just to feel the ruler
smell = np.array([0.9, 0.1, 0.2]) # "my street smells"
garbage = np.array([0.8, 0.2, 0.1]) # "rotting garbage bags"
lamp = np.array([0.1, 0.9, 0.2]) # "broken streetlight"
print("smell vs garbage:", round(cosine(smell, garbage), 3))
print("smell vs lamp: ", round(cosine(smell, lamp), 3))
smell vs garbage: 0.987
smell vs lamp: 0.256
Note: our stand-in normalizes its vectors, so the norm division here is redundant — cosine() computes norms anyway, so it stays correct for arbitrary vectors, not just normalized ones.
That's the core retrieval idea. Production systems add better embeddings, indexing, filtering, and evaluation around the same basic similarity operation. (Real embedding models typically produce much higher-dimensional vectors than our 3-number demo — and you'll compare against thousands of complaints instead of two.)
The principle
- "Similarity is a measurement, not a meaning." Cosine similarity tells you two texts landed near each other — it doesn't tell you they're true, relevant, or what the resident actually wants. It's a sibling of the Pydantic lesson's rule: schema validity ≠ business truth. Here: vector closeness ≠ relevance.
Build: a tiny vector search, for real
Teaching stand-in — read this first
- In production, embeddings come from a trained embedding model (later we'll show a real model such as a
sentence-transformersmodel) — no keyword lists anywhere. Below, we use a tiny stand-in embedder instead: 10 hand-made concept axes filled by keyword hits, so every number stays inspectable. - The similarity-ranking concept is the same: embed each text once, embed the query, rank by similarity, sort. At scale, an index avoids brute-force comparison against every vector — but the idea you're learning is the same. Only the numbers' source differs.
- The stand-in is deliberately dumber than a real model: it only knows the synonyms listed, it can't read negation ("wasn't picked up"), and it matches crude substrings. You'll watch all three weaknesses bite below — that's intentional. When they do, remember: a real model fixes the embeddings; nothing fixes sloppy search math, which is what you're really learning here.
Twelve real-ish CityOps complaints, one embedder, one query, ranked by cosine similarity:
import numpy as np
CONCEPTS = ["waste", "odor", "noise", "water", "lighting",
"road", "animals", "service", "vandalism", "litter"]
KEYWORDS = {
"waste": ["garbage", "trash", "dumpster", "refuse"],
"odor": ["smell", "smells", "odor", "stench", "rotting", "foul"],
"noise": ["noise", "noisy", "loud", "construction"],
"water": ["water", "hydrant", "spraying", "flood"],
"lighting": ["streetlight", "lamp", "dark"],
"road": ["pothole", "street", "sidewalk", "curb", "gutter"],
"animals": ["raccoon", "raccoons", "rat", "rats", "dog"],
"service": ["missed", "pickup", "sweeper", "never comes", "weeks"],
"vandalism": ["graffiti", "vandalism", "broken"],
"litter": ["litter", "gutters"],
}
# ============================================================
# TEACHING MODEL — NOT AN EMBEDDING MODEL
# Hand-built vectorizer used only to expose the retrieval math.
# ============================================================
def embed(text):
"""Stand-in embedder: keyword hits per concept axis (crude substring
matching — deliberate), normalized. In production, swap this for a
real model -- the rest is unchanged."""
t = text.lower()
v = np.array([sum(t.count(k) for k in KEYWORDS[c])
for c in CONCEPTS], dtype=float)
n = np.linalg.norm(v)
return v / n if n else v
def cosine(a, b):
# stand-in vectors are pre-normalized; norms computed anyway so this
# stays correct for arbitrary vectors, not just normalized ones.
na, nb = np.linalg.norm(a), np.linalg.norm(b)
if na == 0 or nb == 0:
return 0.0 # no concepts fired: honestly "no idea", not an error
return float(np.dot(a, b) / (na * nb))
complaints = [
"Rotting garbage bags piled on the curb outside 44 Mercer St",
"Overflowing dumpster behind the deli on Grand Ave, trash everywhere",
"Foul odor coming from the alley near the park entrance",
"Streetlight out on 5th Ave and Main, intersection is dark",
"Pothole the size of a bathtub on Riverside Dr",
"Garbage truck missed our block three weeks in a row",
"Loud construction noise starting at 5am near the school",
"Graffiti covering the subway entrance walls",
"Missed trash pickup again, bins still full on Thursday",
"Broken fire hydrant spraying water onto the sidewalk",
"Raccoons tearing open trash bags every night on Elm St",
"Street sweeper never comes down our block, gutters full of litter",
]
vectors = np.array([embed(c) for c in complaints])
print("corpus vectors:", vectors.shape) # 12 complaints, 10 numbers each
corpus vectors: (12, 10)
Now the search itself — embed the query, score everything, sort:
def search(query, vectors, texts, k=5):
q = embed(query)
scores = [cosine(q, v) for v in vectors]
ranked = sorted(zip(texts, scores), key=lambda t: t[1], reverse=True)
return ranked[:k]
for text, score in search("my street smells", vectors, complaints):
print(f"{score:.3f} {text}")
0.894 Foul odor coming from the alley near the park entrance
0.775 Rotting garbage bags piled on the curb outside 44 Mercer St
0.447 Pothole the size of a bathtub on Riverside Dr
0.258 Street sweeper never comes down our block, gutters full of litter
0.200 Streetlight out on 5th Ave and Main, intersection is dark
The original texts share no literal search term — "smells" appears in none of the complaints — but our teaching embedder manually maps several of those words onto the same concept axis, so this demo shows you the retrieval math, not proof that embeddings discover meaning on their own. A real embedding model learns useful relationships from training rather than from a hand-written synonym table. The real value proposition: embeddings can retrieve semantically related wording even when literal terms differ.
And read the full top 5 honestly, because #3 is already teaching you something: the pothole complaint matched on the word "street," not on meaning. The stand-in is crude — and that's the point. A good real embedding model should handle this relationship more intelligently, but you'll verify that with retrieval evaluation rather than assume it. Keep that skepticism handy; the Break section is built on it.
The architecture is almost boring, which is the point:
12 x 10 numbers] C[Resident query
'my street smells'] -->|embed| D[Query vector
10 numbers] B --> E[Cosine similarity
against every row] D --> E E -->|sort| F[Ranked results]
Break: where similarity lies to you
Two failure modes you must see before you trust this in production.
1. Plausible but wrong
Try a query about a garbage problem:
for text, score in search("garbage problem on Elm Street", vectors, complaints):
print(f"{score:.3f} {text}")
0.816 Rotting garbage bags piled on the curb outside 44 Mercer St
0.707 Overflowing dumpster behind the deli on Grand Ave, trash everywhere
0.707 Pothole the size of a bathtub on Riverside Dr
0.408 Street sweeper never comes down our block, gutters full of litter
0.316 Streetlight out on 5th Ave and Main, intersection is dark
Read #2 and #3 carefully: the dumpster complaint (genuinely about garbage) and the pothole complaint (nothing to do with garbage) tie at 0.707. Why? The dumpster matched on the waste axis; the pothole matched on the road axis, because the query says "Elm Street." The ruler can't tell a meaningful match from a coincidental one — both are just "one shared axis."
A real embedding model is subtler than our stand-in, but it has its own version of this failure: any shared context — a street name, a time of day, a mention of "night" — can pull an irrelevant complaint up the ranking. Similarity is a measurement, not a meaning. The score says "these texts point the same way." It cannot say "this solves the resident's problem." Production retrieval needs additional controls and evaluation around similarity ranking — which is exactly what Productionize adds next.
2. Coarse chunks blur the signal
Now chunk badly. Instead of one complaint per vector, embed five complaints mashed into a single vector:
blob = " ".join(complaints[:5]) # garbage + dumpster + odor + streetlight + pothole
print("blob score:", round(cosine(embed("my street smells"), embed(blob)), 3))
blob score: 0.723
The 5-complaint blob scores 0.723 on "my street smells" — a strong-looking score — even though three of its five members (streetlight, pothole, dumpster) have nothing to do with smells. In our stand-in, one matching part drags the whole blob up, and the score can't tell you which part matched. The stand-in illustrates the general risk: large chunks can mix several topics into one representation, making retrieval less precise. Retrieve that blob and you'd hand the resident five complaints, four-fifths noise, stamped "0.723 relevant."
The lesson: chunk at the granularity you want to retrieve. One complaint → one vector. If you want to retrieve paragraphs, embed paragraphs. The vector can only be as precise as the unit you embedded — a coarse chunk is a blurry photograph. Good ranking can't fully recover information you erased by embedding an overly coarse unit; retrieve at a useful granularity, then rerank if needed.
Don't memorize
- The exact dimensionality (10 in our stand-in; much higher in real models) — different models use different sizes; what matters is that query and indexed corpus use compatible embedding representations — normally the same embedding model, version, and configuration.
- Distance formulas beyond cosine — Euclidean and dot-product exist, but cosine similarity is a common choice for text embeddings. Understand it cold — then use the similarity metric recommended for your embedding model and index; look up the rest when you need them.
Productionize: filters plus ranking
A real search isn't just ranking — Lisa's team also filters by borough and status, and those are exact constraints. This is where you learn the division of labor: structured filters belong in SQL; fuzzy matching belongs in vectors. Never ask an embedding whether a complaint is in Queens — that's a column, not a vibe. (Your SQL lesson already taught you the exact-match machinery; use it.)
The production shape: apply exact metadata constraints as part of retrieval, then rank eligible candidates semantically. A column is not a vibe — use exact filters for exact facts and vectors for fuzzy similarity.
class ComplaintSearch:
def __init__(self, embed_fn, records):
# records: list of dicts with text, borough, status
self.embed = embed_fn
self.records = records
self.vectors = [embed_fn(r["text"]) for r in records] # embed ONCE
def search(self, query, borough=None, status="open", k=5):
q = self.embed(query)
scored = []
for rec, vec in zip(self.records, self.vectors):
if borough and rec["borough"] != borough:
continue
if status and rec["status"] != status:
continue
scored.append((rec, cosine(q, vec)))
scored.sort(key=lambda t: t[1], reverse=True)
return scored[:k]
records = [
{"text": t, "borough": b, "status": s}
for t, b, s in zip(
complaints,
["MANHATTAN", "QUEENS", "BROOKLYN", "MANHATTAN", "BRONX", "QUEENS",
"BROOKLYN", "MANHATTAN", "QUEENS", "BRONX", "BROOKLYN", "MANHATTAN"],
["open"] * 9 + ["closed"] * 3,
)
]
cs = ComplaintSearch(embed, records)
for rec, score in cs.search("missed garbage pickup", borough="QUEENS"):
print(f"{score:.3f} [{rec['borough']}/{rec['status']}] {rec['text']}")
1.000 [QUEENS/open] Garbage truck missed our block three weeks in a row
1.000 [QUEENS/open] Missed trash pickup again, bins still full on Thursday
0.447 [QUEENS/open] Overflowing dumpster behind the deli on Grand Ave, trash everywhere
Two things to notice. First, exact metadata constraints reduce the candidate space before semantic ranking — the filter narrows the field, similarity only scores candidates. (The expensive operation is often embedding generation; similarity/index search over stored vectors can be very fast.) Second, the class embeds the corpus once — at startup, in this toy. Re-embedding 10,000 complaints on every query would be the kind of bug that only shows up in production — at 3 AM, as a latency graph shaped like a hockey stick.
In this toy class we precompute at startup. In production, the write path runs once per record: a complaint is created or updated → the pinned embedding model embeds it → you persist the text, metadata, vector, and model version → the vector index picks it up. Queries then embed with the same compatible embedder and search the index — never rebuild the corpus on application startup.
pinned version] B --> C[persist text + metadata
vector + model version] C --> D[vector index] end subgraph Read path E[Resident query] --> F[same compatible embedder] F --> G[metadata constraints
+ vector retrieval] G --> H[ranked complaints] end
Pin the embedding model like you pin the prompt
Every vector in the index carries an invisible assumption: it was produced by one specific embedding model, version, and configuration. Embed a million complaints with model v1, then query with model v2, and the vectors may not inhabit a compatible representation — the geometry changes under you, and similarity scores become meaningless. So the production invariant: pin and record the embedding model and version used for the index. Changing it generally means re-embedding the corpus into a new index or version. Same discipline as the prompt-versioning lesson: what you measured is only valid for the artifact you measured.
When NOT to use vector search
- Exact identifiers — request IDs like
SR-10421. That's aWHERE request_id = ..., not a vibe check. - Dates, ranges, statuses — "open complaints from last week" is SQL. Dates, ranges, and statuses are structured constraints — use structured filters.
- If exact/keyword search already satisfies the retrieval need — don't add embeddings just because they're fashionable. Adding vectors brings a model dependency, latency, and a new failure mode for zero gain.
- Rule of thumb: vectors for "find me things like this"; SQL for "find me things exactly like this."
Communicate: the search-quality report
Lisa doesn't want a demo — she wants to know whether to ship it. So you show her a report, not a screenshot. Five test queries, top result each, honest verdicts:
| Resident query | Top result | Score | Verdict |
|---|---|---|---|
| "my street smells" | Foul odor coming from the alley… | 0.894 | Hit — no keyword overlap, still found |
| "trash wasn't picked up" | Overflowing dumpster behind the deli… | 1.000 | Partial — trash-related, but it's about an overflowing dumpster, not a missed pickup. The stand-in can't read "wasn't picked up"; a real embedding model may handle this better — test it on your actual query set |
| "the noise is terrible at night" | Loud construction noise starting at 5am… | 1.000 | Hit |
| "water everywhere on the sidewalk" | Broken fire hydrant spraying water… | 0.853 | Hit |
| "my landlord won't fix the elevator" | (none — every result scores 0.000) | 0.000 | Miss — and in this stand-in, the score says so honestly: none of our handcrafted concept axes fired, so the guard returns 0.0 instead of a confident wrong answer. That 0.000 is meaningful only because of our toy vectorizer — see below |
Read that table the way Lisa will. Three queries nail it. One is related-but-off in a way the score can't flag — a perfect 1.000 on the wrong meaning, which is exactly why the principle says similarity is a measurement, not a meaning. And the last one fails visibly at 0.000 — in our stand-in, the correct behavior when nothing relevant exists. Your report says exactly that: where it works, where the score lies to you, and what we'd do next.
Five demos are not an eval
This table is a demo — five queries, hand-picked to show range. Before shipping, build a golden query set: a representative list of real resident queries, each with judgments about which complaints count as relevant. Then measure retrieval at a cutoff k: recall@k (of the relevant complaints, how many appear in the top k?) and precision@k (of the top k results, how many are relevant?). Five examples are a demo; an eval set is evidence.
This is also where "no good match" gets its real answer. Our stand-in's 0.000 is honest because of our toy vectorizer — no handcrafted concept fired. A real embedding model often produces nonzero, even fairly high, scores for irrelevant items. So don't hard-code if score < 0.5: no_match on intuition: choose the threshold empirically, from labeled evaluation data. A similarity score becomes useful only after you measure what "good" looks like on your own data.
A search that says "I don't know" beats one that confidently shows raccoons for a dog problem — and "I don't know" is a decision you calibrate, not a constant you inherit.
CityOps: the retrieval layer for what's next
This search is a component, not a product. In the CityOps platform, it becomes the retrieval layer: when a resident asks a question in plain English, the system first finds the relevant complaints with exactly this machinery, then does something useful with them — summarizes them, routes them, drafts the response. The search doesn't answer questions; it finds the material answers are built from.
That "something useful" is the Ask CityOps milestone at the end of this stage — a tool-using assistant over agency procedures. It will need this retrieval layer the way the intake pipeline needed Pydantic: the unglamorous foundation the impressive demo stands on. Build it well now; you'll be glad later.
Useful later
- Real embedding models —
sentence-transformerswithall-MiniLM-L6-v2is the drop-in upgrade for our stand-in: sameembed → cosine → rankcode, but the numbers come from a trained model instead of 10 keyword lists. Whether it fixes the "wasn't picked up" blind spot is something you'll measure, not assume. - Dedicated vector databases (Pinecone, Weaviate, pgvector) — for when the corpus outgrows an in-memory list. The database solves storage, indexing, and search at scale; retrieval quality still comes from the embedding model, the chunking, the metadata, the configuration, and the evaluation.
- Hybrid search — keyword (BM25) scores blended with vector scores, so exact terms like street names still match precisely.
- Re-ranking — a second, slower model that re-scores the top 20 for precision after the fast model recalls the top 200.
Field check
- Lisa asks: "Why can't we just use
LIKE '%smell%'?" What do you tell her?Reveal
Keyword search needs the same words on both sides. Residents write "my street smells"; complaints say "rotting garbage" and "foul odor" — zero shared keywords, zero rows. Embeddings can retrieve semantically related wording even when literal terms differ, so all three surface for one query.
- Your search returns a complaint with score 0.92 and one with 0.31. What does each number actually claim?
Reveal
Only that the two texts point in a similar (0.92) or dissimilar (0.31) direction in embedding space. It claims nothing about truth, relevance to the resident's real problem, or correctness. Similarity is a measurement, not a meaning — 0.92 can still be the wrong answer.
- A teammate wants to filter "complaints in Queens" using embeddings. What's wrong with that?
Reveal
Borough is an exact, structured attribute — a column, not a vibe. Use the metadata filter (SQL-style exact match) before ranking. Asking an embedding "is this Queens?" wastes the model's strength and invites fuzzy errors on something that has a right answer.
- You embedded whole case files (five complaints each) instead of individual complaints. A query about smells returns a 5-complaint blob scoring 0.723 — but only one of the five is about smells. What went wrong, and what's the fix?
Reveal
The chunking is too coarse: in our stand-in, one matching part drags the whole blob's score up, and the score can't tell you which part matched — you'd hand the resident five complaints, four-fifths noise, stamped "0.723 relevant." Fix: chunk at the granularity you want to retrieve — one complaint per vector. Good ranking can't fully recover precision you erased at chunking time.
- The corpus has no elevator complaints and someone searches "my landlord won't fix the elevator." Every result scores 0.000. Ship it or not — and why?
Reveal
Ship the behavior, not the result: for this stand-in, 0.000 means none of our handcrafted concept axes fired, so returning "no good matches" is safer than forcing a result. It doesn't even prove nothing relevant exists — it proves none of our known concepts fired. In a real embedding system, calibrate the no-match policy from evaluation data rather than assuming zero or any universal cutoff. A search that admits ignorance beats one that confidently serves raccoons for a dog problem.