Maria wants every 311 complaint summarized by the LLM and loaded into the dashboard database. The summaries read beautifully — and roughly one in twenty is unusable: prose instead of JSON, a cut-off response, a confident "cat" where the contract says "category". This is the lesson where you stop trusting the model's manners and start checking its work — and the modern answer isn't better parsing, it's constraining what the model can emit in the first place.
pip install pydantic openai. The API-call code in this lesson uses OpenAI's Responses API with provider-native structured output and is syntax-checked, not run (no key in this environment); the validation, repair, and quarantine logic — the actual point of the lesson — runs for real against representative outputs.The customer problem
Wednesday, 2 PM. Maria from the data team has a new request: "The dashboard shows complaint counts, but the mayor's office wants to read what's happening — one-line summaries of each complaint, plus a category and a priority flag. We have forty thousand complaints a month. Can the LLM do it?"
You wire up the API call, feed it complaint text, and ask for JSON. It works. The summaries are good — sharp, accurate, well categorized. You load a thousand of them into the staging table.
Then the dashboard renders a card that says [None] Missed trash pickup for the third week r... priority=normal. One row came back as chatty prose instead of JSON. Another was cut off mid-sentence. A third used "cat" and "pri" as keys and dropped the rest. Maria's question is the same one she asked about the CSV intake: "Which rows did we lose — and why?"
Clarify the ask
The ask isn't "make the LLM behave." You can't — it's a probability machine, and probability machines occasionally roll the wrong number. The ask is the one you've answered twice already in this curriculum: draw a boundary. Every summary that enters the database either satisfies a contract you've defined or is rejected with a clear reason. The model doesn't get to decide what "valid" means. You do.
The principle
- "An LLM's promise is not a schema." The model will happily say it returned JSON with the right keys. Saying and satisfying are different things — and downstream code can only rely on one of them.
The minimum concept
Three separate guarantees, and beginners constantly collapse them into one:
- Syntax — the reply parses as JSON. Is it well-formed?
- Schema — the JSON has the required fields, the right types, the allowed values. Does it match your contract?
- Semantic correctness — the values are actually supported by the source complaint. Is it true?
Pydantic can prove #2 relative to your contract. It cannot prove #3. Provider-native structured output — the 2026 default you'll learn below — helps strongly with #1 and #2, but also cannot prove #3. Watch how deep that goes:
>>> LLMSummary(**json.loads('{"category": "pothole", "summary": "Large pothole on 5th Avenue blocking the bike lane.", "priority": "urgent"}'))
LLMSummary(category='pothole', summary='Large pothole on 5th Avenue blocking the bike lane.', priority='urgent')
Perfectly valid. Also perfectly wrong — the original complaint was about a noisy bar in Brooklyn. The schema is satisfied; the truth isn't. The distinction to carry forward: structured output can enforce output shape. Your application still validates business rules and must never assume that schema-valid content is factually correct. This is the setup for the LLM evaluation lesson later in this stage.
This is the Pydantic lesson's "validate at the boundary; trust inside" — with the boundary moved to the model's output instead of the CSV intake. Same philosophy, new frontier. And the connection runs both ways:
- CSV boundary: untrusted external data → validate → trusted internal object.
- LLM boundary: untrusted model output → validate → trusted internal object.
The 2026 hierarchy
For a modern lesson, teach the solutions in this order — strongest first:
- Provider-native schema-constrained structured output — the schema goes into the request; the provider constrains generation to it. Best available when the provider supports it.
- Application validation with Pydantic — your contract, checked on your side, regardless of what the provider promised.
- Bounded repair/retry if needed — corrective context, limited budget.
- Quarantine — rejected rows, with reasons attached.
JSON mode plus manual parsing is then the fallback/portable pattern — what you reach for when native schema enforcement isn't available or isn't appropriate for the call. The concept never changes: validate model output before side effects. What changes is how much of the enforcement the provider does for you.
schema in the request] --> B[Constrained generation
shape enforced by provider] B --> C[Pydantic boundary
your contract, your side] C -->|invalid| D[Bounded repair
then quarantine] C -->|valid| E[Semantic checks
grounded in the source] E -->|grounded| F[Database
side effects only here] E -->|not grounded| G[Quarantine + human review]
Defense in depth, not either/or: the constraint, the validator, and the semantic check each catch failures the others can't.
Build: the output contract
Before the contract, the most important design decision in this whole lesson. Your source record already knows the request ID and the borough:
source = {"request_id": "SR-10425", "borough": "Queens",
"text": "Missed trash pickup for the third week running."}
Don't ask the model to generate facts your system already knows
request_idandboroughlive in application state — the source 311 record already has them. Keep them there. The model derives only what it must:category,summary,priority.- Every field you ask the model to reproduce is a hallucination opportunity you volunteered for. The model can omit the ID, invent a borough, or swap two rows' identifiers — and nothing about "being a good summarizer" prevents it.
- The authoritative request ID always comes from input/job metadata, never from model output — which is also why quarantine records carry the input's ID, not the model's.
So the model's contract is deliberately small:
from typing import Literal
from pydantic import BaseModel, ConfigDict, Field, ValidationError
import json
class LLMSummary(BaseModel):
"""What the model is allowed to decide. Nothing else."""
model_config = ConfigDict(extra="forbid")
category: Literal["pothole", "noise", "sanitation", "water", "other"]
summary: str = Field(min_length=10)
priority: Literal["low", "normal", "urgent"] = "normal"
Three decisions worth naming. extra="forbid" — the schema-drift lesson from the Pydantic post: if the model invents a new field like "confidence", that's a contract event, not a bonus. Literal types for category and priority — the dashboard renders these as filters, so "Pothole", "pothole", and "road damage" can't all mean the same thing downstream. And Field(min_length=10) — which proves the summary isn't extremely short, and only that:
>>> LLMSummary(category="noise", summary="aaaaaaaaaa", priority="low")
LLMSummary(category='noise', summary='aaaaaaaaaa', priority='low')
Ten a's pass. Shape, not quality — another instance of "schema validation ≠ semantic quality."
schema_version is application metadata, not something the LLM should decide — so it doesn't live in the model's contract at all. The application attaches it after validation, when it joins the model's derived fields back to the source-owned fields:
def join(source, llm_out):
return {
"schema_version": "v1", # application metadata — the model never decides it
"request_id": source["request_id"], # from input metadata — never from the model
"borough": source["borough"], # source-owned, joined here
**llm_out.model_dump(), # model-derived: category, summary, priority
}
The call: schema goes into the request
Written against the real SDK (syntax-checked; wire in your key from the previous lesson to run it):
import os
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from the environment
model = os.environ["LLM_MODEL"] # model choice is configuration, not code
prompt = ("Summarize this 311 complaint. Output only:\n"
"- category: one of pothole, noise, sanitation, water, other\n"
"- summary: one plain sentence\n"
"- priority: low, normal, or urgent\n\n" + source["text"])
resp = client.responses.parse(
model=model,
input=[{"role": "user", "content": prompt}],
text_format=LLMSummary, # the schema goes INTO the request — generation is constrained to it
timeout=20, # per-attempt timeout: bounds YOUR wait, proves nothing about the remote side
)
llm_out = resp.output_parsed # an LLMSummary — or the SDK raises
Note the two deliberate absences. No temperature=0 — temperature doesn't make output reliable; schema constraint + validation is the reliability mechanism, not a sampling knob. And no conflated retry setting — this lesson needs two retry budgets kept separate:
- Transport retry — 429s, temporary 5xx, network failure. The same request may work again. The SDK handles these; the previous lesson's budgets apply.
- Semantic repair — invalid shape, missing field, invalid enum. The model needs corrective context or a new generation. That's this lesson's loop, with its own budget.
Different failures, different budgets, counted separately. Conflating them is how you burn a repair budget on network blips — or retry a doomed prompt until the invoice arrives.
What native structured output buys you: syntax and schema conformance become the starting point, not the achievement. What it doesn't buy: business truth. The pothole-that-was-noise reply above would sail through the provider's constraint — so the boundary still validates, and semantic checks still decide what enters the database.
Break: the ways it fails
Representative failure outputs, modeled on real LLM failure patterns — four shapes, fed through the same boundary. (The first two live mostly in the fallback path; the last two can bite even with structured output.)
# representative failure outputs, labeled as what they are
GOOD = '{"category": "pothole", "summary": "Large pothole reported at the intersection, growing over two weeks.", "priority": "urgent"}'
TRUNCATED = '{"category": "noise", "summary": "Loud construction noise star'
PROSE = 'Sure! Here is the summary you asked for:\n\nThe complaint is about a pothole in Queens.'
EXTRA_FIELD = '{"category": "water", "summary": "Water main leak flooding the sidewalk near the station.", "days_open": "N/A", "priority": "urgent"}'
WRONG_NAME = '{"cat": "sanitation", "summary": "Missed trash pickup for the third week running.", "priority": "normal"}'
def validate(raw):
return LLMSummary(**json.loads(raw)) # may raise either error
>>> validate(TRUNCATED)
json.JSONDecodeError: Unterminated string starting at: line 1 column 34 (char 33)
>>> validate(PROSE)
json.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
The first two never even reach the contract — json.loads rejects them. Keep the distinction you'll reuse in Productionize: a parse failure (not JSON at all) versus a schema failure (JSON, but wrong). Different failures, different repairs.
>>> validate(EXTRA_FIELD)
1 validation error for LLMSummary
days_open
Extra inputs are not permitted [type=extra_forbidden, input_value='N/A', input_type=str]
For further information visit https://errors.pydantic.dev/2.13/v/extra_forbidden
>>> validate(WRONG_NAME)
2 validation errors for LLMSummary
category
Field required [type=missing, input_value={'cat': 'sanitation', 'su...', 'priority': 'normal'}, input_type=dict]
For further information visit https://errors.pydantic.dev/2.13/v/missing
cat
Extra inputs are not permitted [type=extra_forbidden, input_value='sanitation', input_type=str]
For further information visit https://errors.pydantic.dev/2.13/v/extra_forbidden
The last one is the dangerous one — the "looks right" trap. Every value in it is plausible. A human skimming it sees a fine summary. But the keys are wrong, which means downstream code asking for ["category"] gets a KeyError — or worse, code using .get("category") gets None and renders it. Watch:
def render_dashboard_card(d):
return f"[{d.get('category')}] {d.get('summary')[:40]}... priority={d.get('priority')}"
>>> render_dashboard_card(json.loads(WRONG_NAME))
'[None] Missed trash pickup for the third week r... priority=normal'
[None] on the mayor's dashboard. Not a crash — a quiet corruption, the kind nobody notices until someone important does. The contract exists precisely so this shape of data can never reach the render function: with validation, the record either has a real category or it doesn't exist past the boundary.
Myth: "the model said it followed the format, so it's fine." The model also said "cat" was the right key. Confidence is not conformance — the error above came from a reply that looked completely reasonable. If you can't parse it against a contract, you can't trust it. And even when you can — remember the pothole-that-was-noise: conformance isn't truth.
Productionize: repair, quarantine, and the side-effect rule
Rejecting bad output is only half the job. A truncated reply is usually one retry away from a good one — throwing it out immediately wastes money and loses data. So the production shape is a loop: validate → on failure, repair with full context → bounded → give up and quarantine.
Two kinds of repair
Before the loop, the distinction that decides what "repair" even means:
- Mechanical repair — markdown fences, a wrong field alias (
cat→category), formatting. Deterministic code fixes deterministic formatting before you pay for another model call. - Semantic regeneration —
category,priority, the summary's content. Suppose the model returns"priority": "CRITICAL"and the validator rejects it. Asking "fix the validation error" might convert CRITICAL → urgent — but that's now a meaning decision. Did the source complaint justify urgent? Repairingcat→categoryis structural; repairing CRITICAL → urgent may alter meaning.
So: for semantic failures, regenerate from the original complaint — never mutate the previous output just to satisfy the validator. A repair that makes the validator happy while drifting from the source is worse than a quarantine.
The deterministic fix first
Models love wrapping JSON in markdown fences, and json.loads chokes on the backticks. Strip them before you judge — no API call needed:
def strip_fences(raw: str) -> str:
t = raw.strip()
if t.startswith("```"):
t = t.split("\n", 1)[1] if "\n" in t else t[3:]
if t.rstrip().endswith("```"):
t = t.rstrip()[: -3]
return t.strip()
>>> LLMSummary(**json.loads(strip_fences('```json\n{"category": "pothole", "summary": "Pothole swallowing a traffic cone on the corner.", "priority": "normal"}\n```')))
LLMSummary(category='pothole', summary='Pothole swallowing a traffic cone on the corner.', priority='normal')
This lives in the fallback path — JSON mode or plain-text calls where you're parsing manually. Use deterministic code to fix deterministic formatting before paying for another model call. One caveat: if you've asked a provider for strict structured output and it still returns fences, that's not a formatting detail anymore — it's a provider-contract issue worth investigating.
def short_reason(e) -> str:
# one human-readable line per failure, for the quarantine table
if isinstance(e, ValidationError):
first = e.errors(include_url=False)[0]
loc = ".".join(map(str, first["loc"]))
return f"{loc}: {first['msg']}"
return f"not valid JSON: {e}"
The bounded loop
On a schema failure, a validation error can provide useful, specific repair context — it's concrete and tells the model what to fix. But the repair input must be explicit: the original complaint, the expected contract, the previous output, and the failure. Whether a retry works depends entirely on the model seeing all four — "fix this error" alone lets the repair drift away from the source. And send only the minimum repair information needed: don't blindly feed raw internal errors to a model if they could carry sensitive source data, internal schema details, or user-controlled content.
def summarize_with_retry(source, ask_llm, max_repairs=2):
"""source carries request_id, borough, text — the model only derives
category/summary/priority. Every repair regenerates from the ORIGINAL
complaint: previous output + validation failure are context, not the source."""
prompt = ("Summarize this 311 complaint as a JSON object with keys "
"category, summary, priority. "
"category must be one of: pothole, noise, sanitation, water, other. "
"priority must be one of: low, normal, urgent.\n\n" + source["text"])
prev_raw, last_error, attempts = None, None, 0
for attempt in range(max_repairs + 1):
attempts = attempt + 1
if attempt == 0:
raw = strip_fences(ask_llm(prompt))
else:
raw = strip_fences(ask_llm(
f"Original complaint:\n{source['text']}\n\n"
f"Your previous reply:\n{prev_raw}\n\n"
f"Why it failed validation: {last_error}\n"
"Regenerate the JSON for category, summary, priority "
"from the original complaint."))
prev_raw = raw
try:
data = json.loads(raw)
except json.JSONDecodeError as e:
error = f"parse failure: {e}"
else:
try:
return ("ok", join(source, LLMSummary(**data)), attempts)
except ValidationError as e:
error = f"schema failure: {short_reason(e)}"
# same failure twice: another identical attempt has sharply diminishing value
if error == last_error:
return ("quarantine", f"repeated failure, quarantining early: {error}", attempts)
last_error = error
return ("quarantine", last_error, attempts)
The loop logic above runs for real below — with ask_llm replaying representative replies so the mechanics are honest. In production, ask_llm is the client call from the previous lesson; everything else is identical:
>>> # first reply uses wrong keys, second reply fixes them
>>> # (replay() feeds representative replies so the loop runs without an API key;
>>> # in production, ask_llm IS the client call from the previous lesson)
>>> def replay(responses):
... it = iter(responses)
... return lambda prompt: next(it)
>>> source = {"request_id": "SR-10425", "borough": "Queens",
... "text": "Missed trash pickup for the third week running."}
>>> summarize_with_retry(source, replay([
... '{"cat": "sanitation", "summary": "Missed trash pickup for the third week running.", "priority": "normal"}',
... '{"category": "sanitation", "summary": "Missed trash pickup for the third week running.", "priority": "normal"}']))
('ok', {'schema_version': 'v1', 'request_id': 'SR-10425', 'borough': 'Queens',
'category': 'sanitation', 'summary': 'Missed trash pickup for the third week running.',
'priority': 'normal'}, 2)
>>> # three replies, none of them JSON — the identical failure repeats, so quarantine early
>>> summarize_with_retry(source, replay(['nope', 'still not json', 'never json']), max_repairs=2)
('quarantine', 'repeated failure, quarantining early: parse failure: Expecting value: line 1 column 1 (char 0)', 2)
Two things to notice. The repair budget is bounded — and if the same validation failure repeats, another identical attempt has sharply diminishing value, so the loop quarantines early rather than blindly spending the whole budget. And the quarantine is a first-class output, the same philosophy as the Debugging lesson: the record isn't silently dropped, it's set aside with the reason attached, so Maria can see exactly what failed and why.
Quarantine carries identifiers, not just raw text
For Maria to answer "which complaint failed?", the quarantine row needs the source's identifiers — attempt count, failure type, model and config version, run ID — because the model output may omit or hallucinate the very fields that would identify the row:
def parse_or_quarantine(sources, replies, model_id, run_id):
valid, quarantined = [], []
for src, raw in zip(sources, replies):
try:
data = LLMSummary(**json.loads(strip_fences(raw)))
valid.append(join(src, data))
except (json.JSONDecodeError, ValidationError) as e:
quarantined.append({
"request_id": src["request_id"], # from INPUT metadata — never from the model
"attempts": 1,
"failure_type": "parse" if isinstance(e, json.JSONDecodeError) else "schema",
"reason": short_reason(e),
"model": model_id,
"schema_version": "v1",
"run_id": run_id,
})
if len(valid) + len(quarantined) != len(sources):
raise RuntimeError("batch reconciliation failed")
return valid, quarantined
>>> sources = [
... {"request_id": "SR-10421", "borough": "Queens"},
... {"request_id": "SR-10422", "borough": "Brooklyn"},
... {"request_id": "SR-10424", "borough": "Bronx"}]
>>> valid, quarantined = parse_or_quarantine(
... sources, [GOOD, TRUNCATED, EXTRA_FIELD],
... model_id="llm-summarizer-2026-10", run_id="2026-10-03-nightly")
3 received, 1 accepted, 2 quarantined
- SR-10422 | parse | not valid JSON: Unterminated string starting at: line 1 column 34 (char 33)
- SR-10424 | schema | days_open: Extra inputs are not permitted
The reconciliation check is the same one you've used since the intake pipeline — every record accounted for, or the batch fails loudly instead of silently losing rows.
Semantic checks: the third guarantee
Schema validation can't catch the pothole-that-was-noise. That needs a check grounded in the source — cheap guardrails now, real groundedness evals in the later lesson. A minimal honest example:
CATEGORY_HINTS = {
"pothole": {"pothole", "pavement", "road", "street", "asphalt"},
"noise": {"noise", "loud", "music", "party", "construction"},
"sanitation": {"trash", "garbage", "pickup", "sanitation", "waste"},
"water": {"water", "leak", "flood", "hydrant", "main"},
"other": set(),
}
def looks_grounded(complaint_text, llm_out):
# Heuristic, not proof: the chosen category should have SOME support
# in the source text. "other" always passes — the honest answer.
hints = CATEGORY_HINTS[llm_out.category]
return not hints or any(h in complaint_text.lower() for h in hints)
>>> text = "Loud bass from the bar on the corner every single night until 3am."
>>> trap = LLMSummary(category="pothole", summary="Large pothole on 5th Avenue blocking the bike lane.", priority="urgent")
>>> looks_grounded(text, trap)
False
>>> looks_grounded(text, LLMSummary(category="noise", summary="Loud late-night bar noise reported on the corner.", priority="normal"))
True
A heuristic, not a verdict — but it demonstrates the architecture: the semantic check sits after schema validation and before the database, and its failures go to quarantine for human review, not to silent acceptance.
One final production invariant
Validation belongs before side effects
- Never perform the side effect until validation succeeds. Generate → validate → only then write the DB row, send the message, trigger the action.
- This matters twice as much later in this stage, when the model isn't merely producing summaries but choosing actions and tool calls. The rule is the same: no unvalidated model output ever triggers a downstream effect.
Schema versioning: prompts evolve
Next month Maria asks for a "sentiment" field. You update the prompt and the model — and now the database holds rows validated under two different contracts. That's why schema_version is attached at join time: every row carries the version of the contract it satisfied. The rule: version the contract when a change affects compatibility or downstream interpretation — don't silently introduce breaking schema changes. Adding an optional, backward-compatible field may not need the same treatment as a breaking change; silently widening a contract that the dashboard depends on always does. The old version keeps validating old rows, and the migration is a decision, not an accident.
The two paths, and what each guarantees
- Native structured output (preferred): the schema goes into the request; the provider constrains generation. You still validate on your side — constrained generation doesn't guarantee business rules — and semantic checks still decide what enters the database.
- Fallback path (JSON mode / plain text + manual parsing): portable across providers, but every guarantee is yours to build: fence-stripping,
json.loads, the Pydantic boundary, the repair loop. Reach for it when native enforcement isn't available or isn't appropriate for the call. - Constrained decoding / grammar-based generation: the model can't emit invalid JSON. Powerful — but it constrains syntax, not business rules. You still validate.
Communicate: the contract you'd hand the data team
Maria doesn't need to know about repair loops. She needs the contract — what the pipeline guarantees about every row in the summaries table:
"Every summary row in the dashboard satisfies contract v1: one of five fixed categories, a summary of at least 10 characters, and a priority of low, normal, or urgent — joined to the request_id and borough from your source records, which the model never decides. Rows the model couldn't produce correctly within the repair budget are quarantined with the failure reason — they're in the quarantine table, not silently dropped. If the quarantine rate spikes, the summarization boundary is unhealthy and needs investigation — page me, and I'll check the deployment, the config, the model identifier, and the quarantined samples before changing anything. When the contract changes in a way that affects your queries, the version bumps and I'll tell you before it lands."
Notice the shape: what is guaranteed, what happens to failures, what the alert means, and how change is announced. That's the same shape as the bug report from the Debugging lesson — guarantees, not vibes. And notice what it doesn't do: it doesn't diagnose the cause of a spike before the evidence exists.
CityOps: summaries flow into the intake DB
The payoff: the nightly job now has two validated stages. The intake pipeline validates the raw 311 rows with the Stage 2 contract; the summarizer validates the model's derived fields with this lesson's contract, then joins them to the source-owned identifiers the model never sees. Same quarantine philosophy, same reconciliation check — one boundary for CSV rows, one for model replies, and the database only ever receives rows that survived both. The dashboard renders summaries from fully validated rows, and the quarantine table tells Maria exactly which complaints need a human-written summary instead.
Must know
- The 2026 hierarchy: native schema-constrained structured output → Pydantic validation → bounded repair → quarantine; JSON mode + manual parsing is the fallback path
- Three guarantees, not two: syntax → schema → semantic correctness — Pydantic proves schema relative to your contract; nothing proves truth except grounding checks
- Don't ask the model to generate facts your system already knows — source-owned fields stay in application state; the model derives only what it must
- Transport retry and semantic repair are different budgets — 429s/5xx/network vs invalid shape/missing field/wrong enum
- Repair input is explicit: original complaint + expected contract + previous output + validation failure; structural repair fixes formatting, semantic failures regenerate from the source
- The repair loop is bounded, and a repeated identical failure quarantines early — sharply diminishing value
- Quarantine carries identifiers from input metadata — the model may omit or hallucinate the fields that would identify the row
- Validation belongs before side effects — no unvalidated model output ever triggers a downstream effect
- An LLM's promise is not a schema. If you can't parse it against a contract, you can't trust it — and even conformance isn't truth
Don't memorize this
- The exact structured-output parameter names per provider — they differ; look them up
- Every
Literalvalue — those are your contract's vocabulary, not Pydantic's - The repair-loop code line by line — remember the shape: explicit repair input → bounded attempts → quarantine
Field check
- The model returns perfect JSON with all the right keys — but
"priority": "CRITICAL". Does the boundary accept it? Why or why not? - Your repair loop retries 10 times "to be safe." What's wrong with that?
- A reply comes back wrapped in
```jsonfences.json.loadsfails. Is this a model failure or a formatting detail — and where do you fix it? - Maria asks for a new
"sentiment"field next month. You add it to the prompt but forget the model. What breaks, and how doesschema_versionhelp? - The quarantine table shows 40% of tonight's summaries quarantined with "not valid JSON." The model worked fine yesterday. What's your first move?
What good answers look like
1. Rejected — priority is a Literal["low", "normal", "urgent"], and "CRITICAL" isn't in it. Right keys and right JSON aren't enough; the values have to satisfy the contract. That's the difference between syntax and schema. 2. Every repair attempt adds latency, load, and potentially cost; repeated attempts against a systematic failure have sharply diminishing value. Ten retries on a broken prompt is ten times the load for zero new information — bound it (2–3), then quarantine. 3. In the fallback path, a formatting detail — the content is fine, the wrapping isn't. Fix it deterministically with strip_fences before validation, not by re-prompting the model. Don't spend API calls fixing what string code fixes for free. But if you've asked for strict structured output and still get fences, treat that as a provider-contract issue worth investigating. 4. With extra="forbid", every new reply fails validation — the field the prompt now produces isn't in the contract. schema_version lets you introduce "v2" for the new contract while v1 rows keep validating; the database can tell which rows were checked against which contract. Version when the change affects compatibility or downstream interpretation. 5. Treat the spike as evidence that something changed — don't diagnose before the evidence. Check the deployment/config diff, the model identifier actually served, the prompt version, upstream inputs, API behavior/status, and representative quarantined samples before assigning cause. Quarantine first, investigate the boundary — the same evidence-first discipline as the Debugging lesson.