The demo took ten minutes. The 2 AM page took one hung request. This is the lesson where your LLM features stop being demos — timeouts, retries that know when to quit, streaming, fallbacks, and a bill you can explain.

You'll need: APIs from the Python Side (requests, timeouts, retries), Secrets & Security Basics (your API key lives in the environment, never in code), and Pagination, Rate Limits & Retries (backoff, already in your toolbox). pip install openai for the real SDK shape — every demo below runs against a stubbed stand-in so the engineering is what's actually being exercised.

The customer problem

Thursday afternoon. Lisa watches the new operator console you shipped in Stage 3 and points at a complaint thread three screens long: "Can we get a Summarize button? My team reads two hundred of these a day."

Dev's intern builds it in an afternoon. One API call, straight to the model, works beautifully in the demo. Friday morning Lisa asks for it on the nightly batch job too — summarize all 10,000 complaints, every night.

Saturday, 2:14 AM, your phone buzzes. The batch job has been stuck for four hours on complaint number 312. Monday, the invoice arrives: the intern's script retried the same malformed request 4,000 times over the weekend — failing identically every time, burning worker hours and request capacity on a request that was doomed from attempt one.

The model was never the problem. The wrapper was.

Clarify the ask

Notice what nobody is asking: "can the model summarize text?" Of course it can — that's the demo. The production questions are different:

  • Reliability: what happens when the API is slow, rate-limits you, or goes down for ten minutes?
  • Cost: what does 10,000 summaries a night actually cost — and who notices before the invoice?
  • Latency: does the user stare at a spinner for nine seconds, or see words appear immediately?

Your job isn't to call a model. It's to build the wrapper that makes calling a model safe to run unattended, at volume, on someone else's budget.

The principle

  • "Every external call is a failure mode with a price tag." An LLM call can hang, fail, or succeed expensively. The wrapper's job is to bound all four: how long you'll wait (per-attempt timeout + overall deadline), which failures deserve a retry (classified, backoff with jitter, bounded globally), how many run at once (concurrency budget), and what each call costs (tracked).

The minimum concept: how LLMs actually work (the operator's cut)

You don't need transformer math to operate these APIs. You need five facts — everything in this lesson hangs off them:

1. Next-token prediction. At a useful level of abstraction: at the core, a language model generates output token by token based on the preceding context. A summary isn't retrieved from anywhere — it's generated one token at a time.

2. Tokens are a key unit. A token is roughly a word-piece (about 4 characters of English, very roughly). For text generation, tokens are a key unit for context limits, usage, and cost — output tokens usually cost more than input tokens, because generating is more expensive than reading.

3. The context window is finite. Every model has finite context/output limits — prompt plus response, combined. Oversized requests must be rejected, truncated, chunked, or otherwise handled according to your API configuration; blindly retrying the same oversized request won't fix it. That matters for retries later: a too-long prompt will never succeed no matter how many times you retry it.

4. Sampling controls affect variation. Controls like temperature affect how much outputs vary on models that support them — but temperature 0 is not a determinism guarantee, and newer reasoning models don't all expose temperature the same way. Check the chosen model's supported parameters. For production summaries and classification, reliability comes from clear instructions, structured outputs, validation, evaluation, and a versioned model/config — not from a single parameter.

5. Streaming delivers output incrementally. Many LLM APIs can stream generated output incrementally instead of making you wait for the complete response. That's streaming — same task, but the first words arrive in a fraction of a second.

flowchart TD A[Your prompt
as tokens] --> B[Predict next token] B --> C[Append it to the context] C -->|repeat| B B --> D[Stream tokens
as they're made]

That's the whole primer. If you understand "tokens in, tokens out one at a time, billed per token, finite window," you can operate these APIs. Everything else is prompt engineering, and that's a later lesson.

Build: the wrapper, piece by piece

First, the shape of a real production call. The wrapper architecture below is provider/API independent; this example uses OpenAI's current Responses API (import-checked against the installed openai package — OpenAI() reads OPENAI_API_KEY from the environment, and responses.create() accepts timeout and stream):

import os
from openai import OpenAI

def real_call_llm(messages, timeout=20):
    """Production-shaped call. Model choice is configuration,
    not application logic — it comes from the environment."""
    model = os.environ["LLM_MODEL"]
    client = OpenAI()  # reads OPENAI_API_KEY from the environment
    resp = client.responses.create(
        model=model,
        input=messages,
        timeout=timeout,
    )
    return {
        "text": resp.output_text,
        "input_tokens": resp.usage.input_tokens,
        "output_tokens": resp.usage.output_tokens,
    }

One honesty note before we go further: I have no API key in this environment, so every demo below runs against a stubbed call_fn returning a recorded sample. In production that argument is real_call_llm hitting the live API. The stub exists so the retry logic, timeouts, and cost tracking — the actual subject of this lesson — are what's being exercised. The outputs below are real runs of that logic.

Piece 1: know which failures deserve a retry

Not every error should be retried. A 503 (server hiccup) might succeed in ten seconds; a 400 (your request was malformed) will fail forever. Retries consume time and request capacity — and may consume billable work depending on when and how the request fails — so the retry budget for a 400 is zero. This is the same discipline as the Debugging lesson's evidence-vs-story rule, applied to failure classification:

class LLMError(Exception):
    def __init__(self, message, status=None):
        super().__init__(message)
        self.status = status

def is_retryable(err):
    """Which failures deserve another attempt?"""
    if isinstance(err, TimeoutError):
        return True
    if isinstance(err, LLMError):
        return err.status in (429, 500, 502, 503, 504)
    return False

One more refinement before the table: HTTP status alone isn't the whole story. Production retry classification looks at the provider's error type/code and the operation's semantics — and a 429 should first honor provider retry guidance (a Retry-After header) before falling back to your own backoff. Retry policy is provider-specific configuration, not a universal hard-coded HTTP table — treat the table below as simplified starting examples:

StatusMeaningRetry?
429Rate limited — you're calling too fastYes, with backoff
500 / 502 / 503 / 504Their server stumbledYes
Timeout / connection resetThe network, not the requestYes
400Your request was malformed (too long, bad params)No — fix the request
401 / 403Bad key or no permissionNo — fix the credentials
404Model name doesn't existNo — fix the config
429 (rate limited): retryable = True
503 (server hiccup): retryable = True
400 (bad request): retryable = False
401 (bad key): retryable = False
404 (bad model): retryable = False
TimeoutError: retryable = True

Piece 2: retries with backoff and jitter

When a retry is deserved, don't hammer the server — wait longer each time (exponential backoff), plus a small random jitter so a fleet of your workers don't all retry in the same millisecond (the thundering-herd problem from the Retries lesson):

import time, random

def call_with_retry(fn, max_attempts=4, base_delay=0.2):
    last = None
    for attempt in range(1, max_attempts + 1):
        try:
            return fn(), attempt
        except Exception as e:
            last = e
            if not is_retryable(e) or attempt == max_attempts:
                raise
            # First honor provider retry guidance; otherwise back off with jitter.
            delay = getattr(e, "retry_after", None) or \
                (base_delay * (2 ** (attempt - 1)) + random.uniform(0, 0.1))
            print(f"  attempt {attempt} failed ({e}); retrying in {delay:.2f}s")
            time.sleep(delay)
    raise last

Watch it handle a flaky server — two 503s, then success on the third attempt:

attempt 1 failed (service unavailable); retrying in 0.26s
attempt 2 failed (service unavailable); retrying in 0.40s
result: SUMMARY: Pothole on 5th Ave; crew dispatched. | attempts: 3 | api calls: 3

Piece 3: always set a timeout

The timeout parameter caps how long you'll wait. Without it, a hung request holds your worker forever — that's the 2 AM page. With it, the hang becomes a TimeoutError, which is retryable:

returned after 3.1s with no answer until then — a hung worker

That's the no-timeout version, bounded at 3 seconds just for the demo — in production it would have sat there until someone killed it. Now with timeout=1:

attempt 1 failed (request exceeded timeout of 1s); retrying in 0.13s
request exceeded timeout of 1s after 2.1s — and timeouts ARE retryable

A timeout converts an unbounded hang into a bounded, retryable failure — but only from your caller's perspective. A client timeout bounds how long your worker waits; it does not prove the remote system did nothing. The provider may have received and processed the request anyway. That's exactly why the idempotent downstream write matters later. Production callers should have an explicit timeout/deadline policy.

Piece 5: the whole operation gets a deadline too

The per-attempt timeout is only half the story. Retries multiply latency — and once a fallback enters the picture, one logical request can keep Lisa's batch job waiting far longer than she agreed to. Run the arithmetic before you choose the numbers:

def worst_case_seconds(timeout, max_attempts, base_delay):
    backoff = sum(base_delay * (2 ** i) for i in range(max_attempts - 1))
    return timeout * max_attempts + backoff

print(f"one request can take up to ~{worst_case_seconds(20, 3, 0.5):.0f}s before fallback")
one request can take up to ~62s before fallback

A 20-second timeout with 3 attempts plus backoff — and then fallback attempts on top — turns one logical request into over a minute of waiting. The wrapper needs two clocks: a per-attempt timeout and an overall deadline for the logical operation. Set the deadline from what the caller was promised, then size attempts and backoff to fit inside it — not the other way around.

Piece 4: track the cost of every call

The API response reports usage metadata: how many tokens went in and out. That gives you the inputs to estimate or attribute call cost under your provider's pricing rules — log it every time, and the invoice can never surprise you. (The simple formula below is our teaching example; real pricing can include cached input, batch tiers, tool calls, long-context multipliers, and other provider-specific dimensions.)

PRICES = {  # USD per 1M tokens — ILLUSTRATIVE, check your provider's pricing page
    "m-primary": {"input": 2.50, "output": 10.00},
    "m-backup":  {"input": 0.50, "output": 1.50},
}

class CostTracker:
    def __init__(self):
        self.calls = 0
        self.input_tokens = 0
        self.output_tokens = 0
        self.cost_usd = 0.0
    def record(self, model, input_tokens, output_tokens):
        p = PRICES[model]
        c = (input_tokens * p["input"] + output_tokens * p["output"]) / 1_000_000
        self.calls += 1
        self.input_tokens += input_tokens
        self.output_tokens += output_tokens
        self.cost_usd += c
        return c

Lisa's nightly batch — 10,000 summaries, 120 input tokens and 25 output tokens each:

per summary: $0.00055
10,000/day: $5.50/day -> ~$165/month

Half a tenth of a cent per summary; $165 a month at volume. That's either obviously fine or obviously worth optimizing — but only if someone measured it. The intern's version measured nothing.

Break: three ways the naive version dies

1. The hang. No timeout on the batch job. One slow response at 2 AM, and the worker sits forever — no error, no log, no alert. You saw it above: 3.1 seconds of silence in the demo, four hours in production.

2. Retrying the unretryable. The intern's script retried everything, including a 400 caused by a prompt that exceeded the context window. Watch what a disciplined wrapper does instead:

gave up immediately: invalid request (status=400) | api calls: 1

One call, then it stops — because a malformed request will never succeed on attempt 4,000 either. Every retry of a 400 wastes time and request capacity with zero chance of success — and may also consume billable work depending on the provider and where the failure occurred.

3. The surprise bill. No cost tracking, no per-call logging. The $165/month above is cheap — but swap in a larger model at 10× the price, or let a bug double your output tokens, and nobody knows until finance forwards the invoice. Untracked spend is unbudgeted spend.

flowchart TD A[Call fails] --> B{Retryable?
429 / 5xx / timeout} B -->|Yes| C[Wait: backoff + jitter] C --> D{Attempts left?} D -->|Yes| E[Retry] E --> A D -->|No| F[Try compatible fallback] B -->|No: 400 / 401 / 404| G[Raise immediately
retrying can't help]

Productionize: the LLMClient

Now assemble the pieces into the wrapper every AI feature in this curriculum will use. Timeouts, classified retries, a fallback model for capacity failures, cost tracking on every call:

class LLMClient:
    def __init__(self, primary="m-primary", backup="m-backup",
                 timeout=20, max_attempts=3, call_fn=None, tracker=None):
        self.primary = primary
        self.backup = backup
        self.timeout = timeout
        self.max_attempts = max_attempts
        self.call_fn = call_fn          # stub in demos; real_call_llm in production
        self.tracker = tracker or CostTracker()
        self.fallbacks = 0

    def summarize(self, text):
        prompt = [{"role": "user",
                   "content": "Summarize in one sentence: " + text}]
        try:
            resp, _ = call_with_retry(
                lambda: self.call_fn(prompt, model=self.primary, timeout=self.timeout),
                max_attempts=self.max_attempts)
            model_used = self.primary
        except Exception as e:
            if not is_retryable(e):
                raise  # a 401 won't fix itself on the backup model either
            print(f"  primary exhausted ({e}); falling back to {self.backup}")
            self.fallbacks += 1
            resp, _ = call_with_retry(
                lambda: self.call_fn(prompt, model=self.backup, timeout=self.timeout),
                max_attempts=self.max_attempts)
            model_used = self.backup
        cost = self.tracker.record(model_used, resp["input_tokens"], resp["output_tokens"])
        return {"summary": resp["text"], "model": model_used, "cost_usd": round(cost, 5)}

Happy path — one call, tracked:

{'summary': 'SUMMARY: Pothole on 5th Ave; crew dispatched.', 'model': 'm-primary', 'cost_usd': 0.00055}

Primary down with 503s — retries exhaust, then the cheaper backup takes over:

attempt 1 failed (service unavailable); retrying in 0.22s
primary exhausted (service unavailable); falling back to m-backup
{'summary': 'SUMMARY: Pothole on 5th Ave; crew dispatched.', 'model': 'm-backup', 'cost_usd': 0.0001}
fallbacks used: 1 | 1 calls, 120 in / 25 out tokens, total $0.0001

But a 401 does not fall back — and that's deliberate:

raised invalid API key (status=401); fallbacks used: 0

A fallback model helps with capacity failures (rate limits, outages, timeouts). It cannot help with your failures — a bad key, a malformed request, a wrong model name will fail identically on the backup. Falling back on a 401 just spends a second call learning what the first one already told you. Debug with causality: change only something related to the failure. A model fallback doesn't repair credentials.

And a fallback is only safe if you've verified it satisfies the same application contract: the inputs it accepts, context limits, output shape, quality level, latency, cost, and safety behavior. A cheaper backup that writes twice as many tokens, truncates your context, or drifts on format isn't a fallback — it's a different product. One more trap: if your primary and backup sit on the same provider and account, a provider-wide outage or rate limit may take both down together. A fallback is not redundancy by itself.

Finally, zoom out from one request to the whole fleet. Retries can turn a provider outage into your outage: 10,000 jobs × 3 primary attempts × 3 fallback attempts, every worker retrying independently during an incident, is a retry storm aimed at a system that's already down. Bound retries globally as well as per request — concurrency limits, queueing, and circuit breakers are the full answer (circuit breakers are in "Useful later" for a reason).

Streaming: words now, not nine seconds from now

For the Summarize button in the operator console, nobody wants a spinner. Streaming delivers each token as it's generated — the user reads the first words while the rest is still being written:

def stub_stream(messages, model="m-primary", timeout=None):
    for chunk in ["SUMMARY: ", "Pothole on ", "5th Ave; ", "crew dispatched."]:
        time.sleep(0.05)
        yield chunk

parts, first_at = [], None
t0 = time.time()
for chunk in stub_stream([{"role": "user", "content": "x"}]):
    if first_at is None:
        first_at = time.time() - t0   # time to first token
    parts.append(chunk)
print("".join(parts))
first token after 0.05s; full text after 0.20s
assembled: SUMMARY: Pothole on 5th Ave; crew dispatched.

Streaming changes how output is delivered, not the task itself — the first token lands in 0.05s instead of after the full 0.20s. In production that gap is seconds, and it's the difference between "the button works" and "the button is broken" in a user's mind. (In the real SDK: stream=True on the call — a parameter we verified exists.) One production failure mode to design for from day one: a stream can fail after you've already shown partial output, so the UI needs to distinguish partial from complete results.

Idempotency: retries must not double-apply

One more production hazard, and it's subtle. Suppose the summary write to the downstream system succeeds, but the network drops the response. Your retry fires again — and now the complaint has two summaries. The fix is an idempotency key: one key per logical operation, reused across retries, so the downstream system recognizes the duplicate:

class FakeSummaryStore:
    """Stand-in for the downstream system receiving summaries."""
    def __init__(self):
        self.seen = {}
    def write(self, idempotency_key, summary):
        if idempotency_key in self.seen:
            return {"status": "duplicate-ignored", "summary": self.seen[idempotency_key]}
        self.seen[idempotency_key] = summary
        return {"status": "stored", "summary": summary}
attempt 1: response lost after write — retrying with the SAME key
final: duplicate-ignored | attempts: 2 | stored rows: 1

The first attempt's write landed but its response was lost; the retry carried the same key, the store recognized it, and exactly one row exists. Generate the idempotency key once per operation, not once per attempt — a fresh key per retry defeats the entire mechanism. (You met this idea in Webhooks: receivers must tolerate redelivery. This is the sender's half of the same contract.)

One clarification the architecture needs: this idempotency protects the downstream write, not the LLM call itself. If the LLM call is repeated, you may still pay for and recompute the summary even though the eventual database write deduplicates. That's the correct split to keep in your head: LLM call → summary produced → idempotent downstream write. Each layer handles its own duplicates.

Piece 6: bound concurrency for the batch

Lisa wants 10,000 summaries a night. The lesson never asked the obvious production question: how many at once? Launching 10,000 simultaneous requests would slam the provider's rate limits, trigger 429s across the fleet, and turn your own retry logic into the retry storm from the fallback section. Bound concurrency to respect the provider's limits and your own cost/latency budget:

from concurrent.futures import ThreadPoolExecutor

MAX_WORKERS = 8  # bound concurrency: respect provider limits and your budget
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
    results = list(pool.map(client.summarize, complaints))

No async tutorial required — the concept is one sentence: 10,000 jobs ≠ 10,000 simultaneous requests. Size the worker pool from the provider's rate limits and the per-request cost you measured, and let the queue absorb the rest.

flowchart TD A[Work item] --> B[Concurrency / rate budget] B --> C[Overall deadline] C --> D[Primary call
per-attempt timeout] D -->|transient failure| E[Bounded retry
backoff + jitter] E --> F{Compatible
fallback?} F -->|yes| G[Fallback call] F -->|no| H[Fail fast] D --> I[Usage + cost
+ latency metrics] G --> I I --> J[Idempotent
downstream write]

Communicate: the report Lisa actually needs

You don't send Lisa the wrapper. You send her what the wrapper measured — the shape of every production AI report from here on:

"Summarization is live on the nightly batch. The numbers from this week: $0.55 per 1,000 summaries on the primary model, so the full 10,000-a-night run costs about $5.50/day (~$165/month). Fallback model kicked in 3 times during Tuesday's provider hiccup — those summaries cost $0.10 per 1,000 instead. Every call has a 20-second timeout with 3 retries and an overall per-request deadline; nothing hung all week. Two guardrails I'd like your sign-off on: (1) a monthly spend policy — alert at 80%, and a decision at 100%: do we stop, degrade to a cheaper model, queue work, or continue and page someone? (2) a fallback-rate alert — if the backup handles more than 5% of traffic, something's wrong upstream."

Notice the shape: unit cost, total cost, reliability evidence, and decisions framed for the stakeholder. The per-call numbers are real — they came out of the CostTracker, not a pricing page. That's what "no surprise bills" looks like operationally.

CityOps: where this lands

Two homes, both already built in earlier stages:

  • The operator console (Stage 3): the Summarize button calls LLMClient.summarize() with streaming, so agents see words immediately. The API key comes from the environment — the Secrets lesson's three-file pattern, not pasted into the console code.
  • The intake pipeline (Milestone 2): the nightly batch summarizes new complaints with the fallback model configured, idempotency keys on every write, and the cost tracker feeding a daily spend log.

One thing this lesson deliberately doesn't solve: trusting the summary's content. The wrapper bounds waiting and retries, tracks cost, and makes delivery failures observable — it still says nothing about whether the summary is correct or shaped the way your pipeline needs. That's the next lesson: structured output, validated with Pydantic, so "SUMMARY: ..." becomes data your code can rely on. And after that comes the third question — can we prove the answer is actually good?

Must know

  • Every external call is a failure mode with a price tag — bound the wait (per-attempt timeout + overall deadline), the retries (classified), the concurrency, and the spend (tracked).
  • Retry transient failures according to provider guidance — honor Retry-After where available, otherwise bounded exponential backoff with jitter. Retry policy is provider-specific configuration, not a universal hard-coded HTTP table.
  • Always set an explicit timeout/deadline policy. A client timeout bounds how long your worker waits; it does not prove the remote system did nothing.
  • For text generation, tokens are a key unit for context limits, usage, and cost — input + output, output usually costs more. Every model has finite context/output limits; handle oversized requests explicitly, because blindly retrying them won't fix anything.
  • Reliability comes from clear instructions, structured outputs, validation, evaluation, and a versioned model/config — not from temperature 0. Sampling controls affect variation on models that support them; they don't guarantee determinism.
  • Streaming mainly improves time to first output; don't assume it reduces generation work or cost. A stream can fail mid-delivery — the UI must distinguish partial from complete results.
  • Use a fallback only for failures it can plausibly resolve, and only if it satisfies the same application contract (inputs, context, output shape, quality, latency, cost, safety).
  • Idempotency keys are generated once per operation and reused across retries — they protect the downstream write, not the LLM call itself.

Useful later

  • Response schemas / JSON mode — making the model's output machine-shaped (next lesson).
  • Caching identical prompts — the cheapest call is the one you don't make.
  • Provider status pages and client-side circuit breakers — when the fallback is down too.

Don't memorize this

  • Exact per-token prices — they change quarterly; the CostTracker pattern is the durable part.
  • The full HTTP status catalog — remember the three buckets: retry with backoff / fix and don't retry / fail over.
  • SDK method names verbatim — remember the shape (client → responses.create → output_text/usage) and check the docs.

Field check

  1. You have timeouts and retries. Why is classifying errors into retryable vs not-retryable still necessary?
    Reveal

    Because retries consume time and request capacity, and may consume billable work depending on when and how the request fails. Retrying a 503 can succeed; retrying a 400 (say, a prompt over the context window) fails identically every time. The deeper lesson is causality: nothing about another identical attempt fixes the cause. Classification is what stops the wrapper from burning thousands of attempts on a request that was doomed from attempt one.

  2. The primary model returns 503s for ten minutes. When is falling back to the backup model the right move — and when would it be wrong?
    Reveal

    Right for capacity failures: 503s, 429s, timeouts — the backup is a different pool of capacity. Wrong for your own failures: a 401 (bad key), 400 (bad request), or 404 (wrong model name) will fail identically on the backup, so falling back just buys a second failure. Check is_retryable before failing over.

  3. What must be true about the backup model before you enable that fallback?
    Reveal

    It must satisfy the same application contract: the inputs it accepts, context limits, output shape, quality level, latency, cost, and safety behavior — plus your operational requirements. A cheaper backup that truncates context or drifts on format isn't a fallback, it's a different product. Verify the contract before the outage, not during it.

  4. Streaming shows the user words sooner. What does it not change about the call?
    Reveal

    The total tokens and the cost stay roughly the same — streaming reorders delivery, it doesn't reduce work. It mainly improves time to first output (perceived latency), which is a UX win, not an efficiency win. Don't assume it changes total generation time either — implementation and network buffering can shift it.

  5. A retried summary write created two rows downstream for one complaint. What was missing?
    Reveal

    An idempotency key — one key per logical operation, reused across retries, so the downstream store recognizes the second attempt as a duplicate. And the key must be generated once per operation: minting a fresh key per attempt defeats the mechanism.

  6. The monthly bill doubled after switching the batch to a larger model. Which two numbers explain it, and where do you read them?
    Reveal

    Input/output token counts and the per-token price — read both from the tracker's per-call records (resp.usage.input_tokens / output_tokens × the price table). Output tokens usually cost more than input, so a model that writes longer summaries can double the bill even at the same request count.