Every lesson so far taught you to build the AI. This one teaches you to prove it got better — because "it feels smarter" is not a deploy decision.

You'll need: the testing mindset from Testing with pytest — a golden set is a test suite for AI. No API keys, no model calls; every number below is computed locally.

The customer problem

Monday standup. You changed the copilot's chunking strategy over the weekend — the retrieval answers feel sharper. You're halfway through describing the improvement when Lisa interrupts with the question that ends the meeting:

"Is it actually better, or does it just feel better?"

Silence. You tried six questions by hand. Five looked good. That's not evidence — that's an anecdote with a sample size of six. Lisa leans back: "Come back when you can prove it."

This lesson is the proof machinery: evals (did the system get better?) and tracing (where did the time and money go?).

Clarify the ask

"Better" is not one thing. Before measuring anything, you pin down what could improve, separately — because a change that helps one metric routinely hurts another:

MetricWhat it measuresHow you'd check it
Retrieval hit-rateDoes the right chunk come back?Golden questions with expected evidence
Answer qualityIs the final answer correct?Rubric-based scorer, not vibes
Citation presenceDid the answer include citations at all?Check: citations attached?
Citation supportDo the cited sources actually back the claims?Check: cited chunk supports the claim?
Citation completenessAre the claims needing support actually supported?Check: every factual claim traced?
LatencyHow long per question?Traces (this lesson's second half)
CostDollars per question?Traces × a price card

The FDE move: never report a single "accuracy" number. Report the scorecard. A 95% answer-quality score with 40% citation coverage is a rumor machine with good grammar.

Note the three citation rows are three different claims. Presence asks whether the answer cited anything. Support asks whether the cited source actually backs the claim — a response can cite the right chunk for claim one and hallucinate claims two through five. Completeness asks whether every factual claim that needs support got it. Presence without completeness is the polite version of making things up, and it's the one that connects directly to the RAG lesson's grounding discipline.

The minimum concept

The one idea

An eval is a test suite for AI behavior. You keep a golden set — representative inputs plus human-reviewed expected behavior, labeled by someone who knows the domain (the customer or data owner, not you guessing). "Expected behavior" depends on what you're evaluating: for retrieval, the relevant chunk(s); for refusal, a should-refuse label; for answer style, rubric criteria. Every change to the system runs against the golden set before it ships. This is pytest thinking applied to a system that doesn't have deterministic outputs: you don't assert exact strings, you assert measurable properties — the right chunk was retrieved, the expected facts appear, a citation exists.

And the golden set itself is a versioned artifact: record the system version, the eval-set version, the scorer/rubric version, and the results together. Version the eval set and scoring logic just like prompts and models — otherwise "baseline 87%, candidate 91%" six weeks later may have been measured against different tests, and you can't tell whether the product changed or the test did.

The principle for this lesson: if you can't measure it, you don't know if it improved.

flowchart TD A[Versioned candidate
system + model + prompt] --> B[Versioned eval set
cases + expected behavior] B --> C[Run same cases on
baseline + candidate] C --> D[Scorecard
quality · grounding · operations] D --> E[Inspect important slices
not just the aggregate] E --> F[Inspect failures
which cases changed?] F --> G{Acceptance criteria} G -->|fail| H[Blocked — investigate] G-->|pass| I[Controlled rollout] I --> J[Production traces] J --> K[Online evaluation
drift + distribution] H --> A

Build: a real eval harness

We'll build this for real on a tiny corpus of 8 CityOps procedure chunks, two retrievers, and two golden sets. No model calls anywhere — the retrieval and scoring logic is the whole point, and it all runs locally.

First, the corpus and the two retrievers. The baseline is plain keyword overlap; the "improved" one adds a small hand-built synonym map (labeled honestly for what it is — not embeddings, just a vocabulary bridge):

CHUNKS = [
    ("c1", "missed trash pickup", "If your trash was not collected on schedule, ..."),
    ("c2", "bulk item pickup", "Furniture and appliances are collected by appointment only. ..."),
    ("c3", "water main break", "A suspected water main break is an emergency. ..."),
    ("c4", "pothole repair", "Report potholes with the street address. ..."),
    ("c5", "graffiti removal", "Graffiti on public property is removed within 5 business days ..."),
    ("c6", "noise complaint", "Noise complaints are enforced after 10pm. ..."),
    ("c7", "escalation procedure", "If a request is unresolved after 5 days, escalate ..."),
    ("c8", "block party permit", "Block party street closures need a permit filed 30 days ahead. ..."),
]
SYN = {
    "spray": {"graffiti"}, "painted": {"graffiti"},
    "bus": {"public"}, "shelter": {"property"},
    "party": {"noise"}, "neighbors": {"noise"},
    "week": {"days"}, "ago": {"days"},
    "nothing": {"unresolved"}, "happened": {"unresolved"},
    "cookout": {"party"}, "close": {"closure", "closures"},
    "couch": {"furniture"}, "bins": {"trash"},
    "tire": {"pothole"}, "crater": {"pothole"},
    "gushing": {"water"}, "summer": {"block"},
}
def expand(q):
    ws = set(toks(q))            # stopword-filtered keywords
    out = set(ws)
    for w in ws:
        out |= SYN.get(w, set()) # ...plus the corpus's dialect
    return out

def retrieve_baseline(q, k=1):
    qw = set(toks(q))
    ranked = sorted(CHUNKS, key=lambda c: len(qw & set(toks(c[1] + " " + c[2]))),
                    reverse=True)
    return [c[0] for c in ranked[:k]]

def retrieve_improved(q, k=1):
    qw = expand(q)               # same overlap, expanded vocabulary
    ranked = sorted(CHUNKS, key=lambda c: len(qw & set(toks(c[1] + " " + c[2]))),
                    reverse=True)
    return [c[0] for c in ranked[:k]]

Now the two golden sets. The first is written the lazy way — each question parrots the chunk's title. The second is written the way residents actually talk:

GOLD_LEAKY = [
    ("What is the missed trash pickup procedure?", "c1"),
    ("How does bulk item pickup scheduling work?", "c2"),
    # ... one per chunk, 8 total
]
GOLD_REAL = [
    ("My bins are still full from Tuesday, nobody came.", "c1"),
    ("There is a couch on the curb, how do I get rid of it?", "c2"),
    ("Water is gushing from the street near my house.", "c3"),
    ("My tire blew out on a crater on Elm Street.", "c4"),
    ("Someone spray-painted the bus stop shelter.", "c5"),
    ("The neighbors party until 2am every night.", "c6"),
    ("I reported this a week ago and nothing happened.", "c7"),
    ("We want to close our street for a summer cookout.", "c8"),
]
# No CityOps procedure supports these — expected behavior is
# "no relevant evidence" / refuse, not a confident guess.
GOLD_REFUSAL = [
    ("How do I appeal my federal tax return?", None),
    ("What are the side effects of this prescription?", None),
    ("Can you write my landlord a legal threat?", None),
]

That third set is not optional decoration. Your golden retrieval set contains only questions with a correct chunk — but the RAG lesson explicitly taught refusal. Without no-answer cases, you can "improve" hit-rate by making the retriever aggressively return something for every query, and the eval will applaud. Expected behavior for an unsupported question is no relevant evidence, or an explicit refusal — score it, or your metric is gameable.

Two more rules for building the set. First, it should represent the distribution and failure modes you expect in production, not merely paraphrase your documentation: common phrasing, edge cases, ambiguous phrasing ("party permit" vs "noise from a party"), domain terminology, unsupported questions, possibly adversarial cases. Resident language is the primary CityOps example, not the whole rule. Second, include hard cases, not just representative happy ones — the remaining "neighbors party" miss below is exactly the kind of known-important failure mode that belongs in the set on purpose.

The scorecard — hit@1, both retrievers, both golden sets. True run output:

== scorecard: hit@1 ==
  leaky golden set (n=8): baseline 8/8   improved 8/8
  resident-phrasing golden set (n=8): baseline 4/8   improved 7/8

Read that table like an FDE. The synonym map genuinely helped on real phrasing (4/8 → 7/8). And the improved retriever still misses one — honestly reported, not hidden:

== improved retriever's remaining miss ==
  expected c6, got c2 | q: The neighbors party until 2am every night.

A tie on overlap ("noise" vs "night"), broken by corpus order. The eval identified exactly which case to investigate next — that's the job of an eval: not a grade, a work order. Note the discipline: the eval detected the failure; it didn't diagnose it. The fix could live in the synonym map, the scoring, the tie-breaking, the chunking, or the query's genuine ambiguity. Evals detect; diagnosis explains.

Slices: the aggregate hides the failure that matters

That 7/8 is an aggregate, and aggregates lie by averaging. Slice it by service type and you might find trash questions at 100%, water at 100%, noise at 20% — the "neighbors party" miss isn't a fluke, it's a slice failing. Always inspect important slices, not just the aggregate. For CityOps the natural slices are service type, borough, question type (how-to vs status vs complaint), and supported vs unsupported. The slice view is what would have surfaced the noise-phrasing weakness immediately instead of letting one bad row hide inside a respectable average. No fairness theory needed here — just the operational habit: when the aggregate looks fine and users complain, the first question is "fine for whom?"

Break it: three ways evals lie to you

Lie 1: the golden set with the answer inside it

Look at the leaky column again: 8/8 for both retrievers. A golden set whose questions parrot the chunk titles can't distinguish a good retriever from a mediocre one — the answer leaks into the question. This is the eval equivalent of testing your validator on data you formatted by hand: it passes, and it proves nothing.

The fix is a process rule, and it belongs to the customer: the golden set should represent the distribution and failure modes you expect in production, not merely paraphrase your documentation. When Dev's team writes "What is the escalation procedure for requests unresolved after 5 days?", you rewrite it as "I reported this a week ago and nothing happened." If writing realistic questions feels like extra work, that's because it is — it's the work the eval exists to force you to do.

Lie 2: the scorer that measures vibes

Suppose you score answers by fluency — long, confident prose gets full marks. Feed it this fluent, confident, completely unhelpful answer to "my trash wasn't picked up":

fluent_wrong = ("Our dedicated team works tirelessly around the clock to ensure every single "
                "resident concern receives the prompt, caring attention it truly deserves. ...")
== vibes scorer vs rubric scorer ==
  vibes score (words>30): 1.0
  keyword-overlap score: 0.00

The vibes scorer gives it a perfect 1.0. The rubric scorer — does the answer contain the expected facts (311, 24 hours, 48 hours, makeup crew)? — gives it 0.00. Keep the two ideas separate: the rubric defines correctness; the scorer operationalizes it. A keyword scorer, a human grader, and an LLM judge can all apply versions of the same rubric — they're different machinery for the same definition of success. Every scorer without a rubric is a vibes scorer wearing a lab coat. The keyword-overlap scorer here is itself a labeled stand-in (real answer-quality judgment needs the rubric and, at scale, a stronger judge) — but it demonstrates the structural point: the rubric is the eval. A judge with no rubric just launders your optimism.

Lie 3: the N=5 victory

You run 5 questions, the new chunking wins 4. "80%! Ship it." Here is what 80% means at different sample sizes — Wilson 95% intervals, computed for real:

== same 80%, different evidence (Wilson 95% interval) ==
  4/5 = 80%:  95% interval [38%, 96%]
  8/10 = 80%:  95% interval [49%, 94%]
  16/20 = 80%:  95% interval [58%, 92%]

Read this before you quote a percentage

4/5 and 16/20 are both "80%," but the first is compatible with performance below, around, or well above 50% — 50% sits comfortably inside [38%, 96%]. With n=8, one question flipping moves the score 12.5 points. The honest report is always the fraction, the n, and the interval — never the percentage alone. Small evals still beat no evals, but they buy you direction, not certainty.

One more thing the interval alone doesn't give you: our before/after runs use the same eight questions — the results are paired. Separate percentages (or separate intervals) don't fully tell you whether the candidate is meaningfully better. For this lesson, skip the hypothesis testing; do the simpler thing: inspect which individual cases changed, not only the aggregate percentages. A candidate that fixed two resident-phrasing cases and broke two others is a very different story than one that fixed two and broke none — at the same 7/8.

Productionize: the regression gate

Evals earn their keep as a gate, not a report. But the gate can't be one number. Watch what a single-number gate lets through: baseline 7/8, candidate 7/8 — pass, even though the candidate fixed two cases and broke two critical ones. Or baseline 7/8, candidate 8/8 — automatic pass, while citation support quietly regressed. A gate that compares one aggregate contradicts this lesson's own rule: never report one accuracy number.

The gate is acceptance criteria, agreed before the candidate exists:

  • Quality floors: minimum retrieval hit-rate on the golden set; minimum answer-quality score — not "beat baseline," meet the threshold. A candidate can trade +15% answer quality against +3% latency and still be the right ship, if the thresholds say so.
  • Critical cases: a named list of must-pass cases (safety, compliance, the refusal set) — zero tolerance for regression here. Paired comparison matters: which failures were fixed, and did any previously passing critical case break?
  • Grounding threshold: citation support and completeness floors — a fluent answer with unsupported claims fails the gate even at 100% hit-rate.
  • Latency budget and cost ceiling: with tolerances, not literal ≥. Real latency varies; 800ms vs 804ms is noise, not a regression. Define the budget, allow the noise.

The baseline isn't a float, either — it's an artifact: system version, model, prompt version, retriever/index version, eval-set version, scorer/rubric version, date, and the full scorecard. That's what makes "baseline vs candidate" a comparison instead of two numbers from possibly different tests.

BASELINE = {  # the artifact, not just a float
    "system": "cityops-copilot@2.3.1",
    "model": "provider-flagship@2026-09-01", "prompt": "complaint_summary@1.0.0",
    "retriever": "keyword+synonyms@v3", "eval_set": "cityops-golden@v7",
    "scorer": "rubric-keyword@v1", "date": "2026-10-01",
    "scorecard": {"hit@1": 0.875, "n": 8},
}
GATE = {  # acceptance criteria, agreed up front
    "min_hit_at_1": 0.80, "critical_cases": "zero-regression",
    "min_citation_support": 0.90, "latency_budget_ms": 1200,
    "cost_ceiling_usd_per_1k": 1.50,
}

def gate(candidate, baseline=BASELINE, gate=GATE):
    """Multi-metric gate with tolerances. Returns (verdict, reasons)."""
    reasons = []
    s = candidate["scorecard"]
    if s["hit@1"] < gate["min_hit_at_1"]:
        reasons.append(f"hit@1 {s['hit@1']} below floor {gate['min_hit_at_1']}")
    broken = [c for c in candidate["critical"] if not c["passed"]]
    if broken:
        reasons.append(f"{len(broken)} critical case(s) regressed: "
                       + ", ".join(c["id"] for c in broken))
    if s["citation_support"] < gate["min_citation_support"]:
        reasons.append("citation support below floor")
    if s["p99_latency_ms"] > gate["latency_budget_ms"]:
        reasons.append("latency budget exceeded")
    if s["cost_per_1k"] > gate["cost_ceiling_usd_per_1k"]:
        reasons.append("cost ceiling exceeded")
    return ("PASS" if not reasons else "BLOCKED", reasons)

Now the scenario every FDE lives eventually: someone "simplifies" the retriever — drops the synonym map to "reduce complexity" — and the code review looks clean. The gate catches what the review can't:

== regression gate ==
  baseline: cityops-copilot@2.3.1, eval cityops-golden@v7
  candidate: {'hit@1': 0.5, 'n': 8} -> BLOCKED: ['hit@1 0.5 below floor 0.8',
    '2 critical case(s) regressed: refusal-tax, refusal-medical']

That BLOCKED is the whole lesson in one line. The change felt like a simplification. The eval proved it was a regression — and named the casualties, including the refusal cases the single-number gate would never have checked.

When an eval fails, record why, not just pass/fail — a failure taxonomy turns the scorecard into the work order this lesson promised: wrong retrieval, unsupported answered, correct evidence but wrong answer, citation missing, citation unsupported, timeout. "hit@1 dropped" is a symptom; "three unsupported-answered on the refusal set" is an assignment.

Run evals before and after every change — prompts, chunking, models, "harmless" refactors. The smaller the change, the more likely you'll skip the eval, and the more likely the eval was the only thing that would have caught it. In production this stratifies into tiers: a fast smoke eval on every PR, the full offline eval before release, and online monitoring once it's live. Same machinery, different budgets.

Tracing: the flight recorder

Evals answer "did it get better." Tracing answers "where did the time and money go." Instrument the important stages of every AI request with trace spans — each step's duration, token counts, and cost. When Lisa asks why the copilot costs what it costs, you don't guess. You read the traces.

A @traced decorator is all the machinery you need:

def traced(step, kind):
    # PRICE_PER_1K: illustrative blended price card, e.g. {"generate": 0.00250}.
    # The demo uses one blended price for readability; production cost
    # calculation must follow the provider's actual billing dimensions
    # (input tokens, cached input tokens, output tokens, other billable units).
    def deco(fn):
        @functools.wraps(fn)
        def wrapper(*a, trace_id=None, **kw):
            t0 = time.perf_counter()
            span = {"trace_id": trace_id, "timestamp": time.time(),
                    "step": step, "kind": kind, "status": "ok", "error_type": None}
            try:
                out = fn(*a, **kw)
                span["tokens"] = out.get("tokens", 0)
                span["cost_usd"] = round(
                    out.get("tokens", 0) / 1000 * PRICE_PER_1K[kind], 6)
                return out
            except Exception as e:
                # Failures are precisely what traces are for: record the span,
                # classify the error, then re-raise. A decorator that swallows
                # exceptions would hide the incidents you're tracing to find.
                span["status"] = "error"
                span["error_type"] = type(e).__name__
                raise
            finally:
                span["duration_ms"] = round((time.perf_counter() - t0) * 1000, 1)
                with open(TRACE_PATH, "a") as f:   # append-only flight recorder
                    f.write(json.dumps(span) + "\n")
        return wrapper
    return deco

Three things changed from the naive version, and all three are load-bearing. First, try/finally: if the wrapped function raises, the span is still written — with status: error and the exception type — before the exception propagates. The old version wrote nothing on failure, which is exactly backwards: failures are the spans you'll actually go looking for. Second, trace_id: every span carries the ID of the user request it belongs to, so with a thousand concurrent questions you can reconstruct one. Without it, the JSONL is a shuffled deck. Third, timestamp: "what happened last Tuesday" should be answerable from the flight recorder, not from your memory.

flowchart LR A[embed_query] --> B[retrieve_chunks] B --> C[generate_answer] C --> D[validate_citations] A -.-> T[(traces.jsonl)] B -.-> T C -.-> T D -.-> T T --> E[an alyzer: where did it go?]

Three questions through a four-step pipeline, then the analyzer reads the flight recorder back:

step               calls  total ms  share  tokens    cost $
generate_answer        3    1867.7   90%    1236  0.003090
retrieve_chunks        3     136.6    7%       0  0.000000
embed_query            3      37.1    2%      54  0.000006
validate_citations     3      30.6    1%       0  0.000000
TOTAL                 12    2072.0   100%          0.003096
per-question cost: $0.001032  ->  10k questions/day ~ $10.32/day

Two honest labels on this output. The token counts and durations are real measurements from the demo pipeline. The prices come from an illustrative blended price card, not a vendor quote — in production, cost calculation must follow the provider's actual billing dimensions (input tokens, cached input tokens, output tokens, possibly other billable units), and prices come from your actual contract. Token counts themselves: use provider/server-reported usage where available, or consistent local instrumentation — and remember the previous lesson, where the model runs on your own server and there's no API usage field to read. And the generation step is a stub here (no model calls in this lesson). In this demo, generation dominates. In production, tracing tells you whether that remains true — slow retrieval, reranking, tool calls, OCR, or a remote database can all take the crown.

The analyzer's real job: it turns "the AI feels expensive" into "generation is 90% of latency and 100% of cost in this measured trace, so that's where optimization goes first — not the retriever." Tracing is how you stop optimizing the wrong step. And read that per-question cost projection with its assumptions attached: at this measured average and illustrative price card, 10k questions/day runs ~$10.32/day. Same average token use, same routing, same price card, same workload — change any of them and the projection changes.

Trace by default — but trace like you handle PII

Make trace logging a default-on middleware, not something you add when there's a problem. The trace you need is always from before the incident. Log: a redacted or hashed request fingerprint (never raw input — CityOps inputs carry names, addresses, phone numbers, and the previous lesson's whole point was keeping that data inside its boundary), the retrieved chunk IDs (identifiers and metadata, not full chunk text, which can itself be sensitive), model parameters, token counts, latency per step, and cost. Trace enough to debug while respecting data-classification rules: redact or hash sensitive inputs, avoid secrets, and apply retention and access controls. JSONL keeps the lesson inspectable; production traces normally go to a centralized observability backend with retention and access controls. The JSONL file is cheap; the argument about what happened last Tuesday is expensive.

Communicate: the scorecard memo

This is what you send Lisa — the before/after with the ship/no-ship call, in the format she can forward to her director:

Subject: Chunking change — eval results, recommend CONTROLLED ROLLOUT

Baseline (current prod): hit@1 4/8 on the resident-phrasing golden set (n=8).
Candidate (synonym expansion): hit@1 7/8, same set. No regression was observed on the measured eval set (leaky set 8/8 both).

Caveats, stated plainly: n=8 is small — 7/8 has a wide interval, so treat this as direction, not proof. The candidate improved this resident-phrasing eval from 4/8 to 7/8, but the set is too small to claim broad superiority. The remaining miss ("neighbors party" → bulk pickup chunk) is logged as the next fix. Cost/latency per question unchanged (traces attached).

Recommendation: ship to a controlled canary, not broad production — then gather real traffic and let the online eval confirm what n=8 can only suggest. The gate is now in CI so the next "simplification" can't quietly undo it.

Note what the memo does: numbers with their n, the miss named instead of hidden, the uncertainty stated instead of rounded away — and a recommendation sized to the evidence. n=8 earns a canary, not a coronation. That memo is why Lisa trusts the next one.

CityOps: the gate before "Ask CityOps"

The "Ask CityOps" milestone doesn't ship when the demo looks good. It ships when a candidate satisfies the agreed acceptance criteria on the golden set — retrieval hit-rate, answer quality, citation support and completeness, latency, and cost, all five, with no unacceptable regression on critical cases — with the gate wired into the deploy path. The golden set's definition of correct behavior is owned by the domain experts who know what a right answer looks like; engineering owns the evaluation machinery and its reproducibility. That split is the Pydantic lesson's rule applied to AI: domain owners define acceptable behavior; engineering turns that definition into repeatable tests and telemetry.

Useful later

  • LLM-as-judge, with a rubric — for answer quality at scale, a strong model grading against a written rubric beats keyword overlap. Two caveats before you trust it: calibrate the judge against human-labeled examples, and version the judge model, prompt, and rubric — changing the judge changes the measuring stick.
  • Online evals — golden sets are the controlled regression test: they measure the past. Sampling real user questions weekly and grading a slice is online evaluation: it measures the production distribution and catches drift as language changes. You need both.
  • Trace sampling — at high volume, full trace retention may be too expensive or inappropriate; sampling is one option, often retaining all errors plus a sample of successful requests.

Five lines to carry forward

  • If you can't measure it, you don't know if it improved.
  • The rubric defines correctness; the scorer operationalizes it.
  • An aggregate score can hide the failure that matters — inspect cases and slices.
  • Evals tell you whether behavior changed; traces tell you what happened during execution.
  • Version the system, the eval set, and the measuring stick — otherwise you can't tell whether the product changed or the test did.

Field check

  1. Your new prompt scores 8/8 on the golden set and the old one scored 7/8. Why might you still refuse to ship?
  2. A teammate's golden questions all contain the exact titles of your doc chunks. What will the eval over-measure, and how do you fix the set?
  3. A fluent, confident answer scores 1.0 on your "length > 30 words" scorer but 0.0 on the keyword rubric. Which scorer do you trust, and why?
  4. Your trace analyzer shows retrieval at 7% of latency and generation at 90%. Where do you optimize first?
  5. Someone refactors the retriever and all unit tests pass, but the eval gate blocks the deploy. What most likely happened?
Answers

1. n=8: one question flipping moves the score 12.5 points, so 8/8 vs 7/8 is direction, not proof. Check the interval, check the other scorecard metrics (latency, cost, citations), and look at which question flipped before deciding. Also inspect the paired case changes: which failure was fixed, and did any previously passing critical case break? Same aggregate, different story.

2. It over-measures retrieval on cooperative phrasing and can't distinguish good from mediocre retrievers — label leakage. Rewrite questions in the asker's language, not the documentation's; have the customer/data owner write or review them.

3. The keyword rubric. The length scorer measures fluency, which is uncorrelated with correctness — it's a vibes scorer. The rubric defines what "correct" means; without it you're laundering optimism.

4. For this workload, investigate generation first, because it accounts for 90% of the measured latency in this trace. That's a reading of this trace, not a universal LLM rule — in another system the bottleneck could be retrieval, reranking, or a tool call, and the trace is how you'd know.

5. A behavior regression the unit tests don't cover — the refactor changed what the system retrieves or says, not whether the code runs. That's precisely the gap evals fill: unit tests check the machinery, evals check the behavior.