Dev is tired of building one dashboard button per task. "What if one assistant could just… handle things? 'Show me the five oldest unresolved sanitation complaints and draft escalations for them' — one sentence, and it does the whole job?" That sentence is the most dangerous requirement you will ever receive, because it sounds simple and it is not. This lesson is about making it safe anyway.

You'll need: the contract discipline from Trust No Input: Data Validation with Pydantic (every tool in this lesson gets a Pydantic argument contract), the quarantine instinct from Debugging Like a Detective, and a venv with pip install pydantic. Honesty note, up front: this environment has no LLM API keys, so the model's "thinking" in every demo below is a stub replaying recorded decisions — clearly labeled. Everything else — the tools, the loop, the validation, the timeouts, the error recovery — runs for real, and every output block is verbatim from an actual run. The orchestration is the lesson; the stub just stands in for the part that needs a key.

The customer problem

Thursday, 10 AM. Dev, the 311 operations lead, has watched your team ship the intake pipeline, the API, the console. Now he leans back and says: "Every week there's a new one-off. 'Draft escalations for the oldest sanitation complaints.' 'Pull the procedure for noise complaints after hours.' Each one becomes a ticket, a dashboard button, a deploy. What if one assistant could just handle these? I type a sentence, it does the whole job."

It is a reasonable ask. It is also how you end up with a system that, given a vague sentence on a Friday afternoon, emails a resident, deletes a record, or loops for six hours calling a vendor API — and nobody can tell you why it did any of it.

Dev's real need is narrower than his sentence: a helper that can look things up and draft things, inside boundaries he can audit. "Handle things" is the wish. "Look up, draft, never send, and show your work" is the requirement. Your job is to hear the first and build the second.

Clarify the ask

Before writing a line of agent code, pin down what "handle it" is allowed to mean. An agent without a written boundary is a demo; an agent with one is infrastructure. For Dev's escalation assistant, the boundary is:

QuestionAnswer for this assistant
What can it read?The complaints table and the procedures handbook. Nothing else.
What can it write?Draft files in an outbox directory. Drafts only — never sends anything.
What can it never do?Contact a resident, modify a complaint record, or call a tool that doesn't exist.
How do we know what it did?Every attempted action produces an audit event. (Production systems must make that persistence reliable enough for their compliance requirements — "no log entry, no action" is the aspiration, not something a demo gets for free.)
When does it stop?When it produces a final answer, or when it hits its step budget — whichever comes first.

Notice what this table is: a permission model and a budget, written before the code. If you can't fill in this table for an agent, you are not ready to build it.

The minimum concept: the agent loop

Strip away the hype and an agent is one loop:

sequenceDiagram participant U as User participant A as Agent loop participant M as Model (stubbed here) participant T as Tools (real code) U->>A: draft escalations for the oldest open sanitation complaints loop until final answer or step budget A->>M: execution history + tool schemas M-->>A: decision: call lookup_complaints A->>A: validate decision shape + arguments (Pydantic) A->>T: lookup_complaints(...) T-->>A: result or classified error A->>A: append sanitized observation to history end A-->>U: final answer + per-run audit log

Three ideas carry the whole lesson:

1. Think → act → observe, bounded. The model proposes, a tool executes, the result goes back into the model's history, repeat. The loop is bounded — a maximum number of steps. An unbounded loop is not an agent, it's a liability with an API key.

2. Tools are APIs with consequences. A tool is a plain Python function with a name, a description, and a typed argument schema. The model never touches the database directly; it can only ask for tools by name with arguments. The schema is the bouncer: bad arguments don't reach the function.

3. State is the execution history — and it costs money. Every tool request, result, and error gets appended to the model's context. Long loops mean long contexts, and long contexts mean real dollars per step. Keep observations compact (designed views, not dumps), and the budget does double duty: it caps both runaway behavior and runaway cost.

The principle

An agent is a loop with a budget. The loop gives it power; the budget — max steps, per-tool timeouts, compact observations — is what makes it deployable. Many dangerous agent failures become manageable once the loop is bounded and every action passes through explicit tool contracts.

Build: the tools

Start with what the model is allowed to touch. Three CityOps tools, each a real function, each with a Pydantic argument contract — the same "validate at the boundary; trust inside" discipline from the Pydantic lesson, now guarding the model's hands instead of a CSV file:

import os
from pydantic import BaseModel, Field

class LookupArgs(BaseModel):
    category: str = "sanitation"
    status: str = "open"
    older_than_days: int = Field(default=0, ge=0)
    limit: int = Field(default=5, ge=1, le=20)

class DraftArgs(BaseModel):
    request_id: str
    reason: str = Field(min_length=10)
    operation_id: str = Field(
        description="stable idempotency key minted by the application, "
                    "never invented by the model")

class ProcedureArgs(BaseModel):
    topic: str

def lookup_complaints(category="sanitation", status="open",
                      older_than_days=0, limit=5):
    rows = [c for c in COMPLAINTS
            if c["category"] == category
            and c["status"] == status
            and c["days_open"] > older_than_days]
    rows.sort(key=lambda c: c["days_open"], reverse=True)
    return rows[:limit]

def draft_escalation(request_id, reason, operation_id):
    # operation_id defines "the same action": same key + same args = same file.
    # Minted by the application when the action is proposed — never by the
    # model, and never date-based (a retry across midnight must stay one key).
    os.makedirs("outbox", exist_ok=True)
    path = f"outbox/escalation-{operation_id}.txt"
    with open(path, "w") as f:
        f.write(f"ESCALATION DRAFT for {request_id}\nReason: {reason}\n"
                f"Status: PENDING HUMAN APPROVAL — not sent.\n")
    return {"draft_path": path, "status": "pending_approval"}

def get_procedure(topic):
    if topic not in PROCEDURES:
        raise ValueError(f"unknown procedure topic: {topic!r}")
    return PROCEDURES[topic]

TOOLS = {
    "lookup_complaints": (lookup_complaints, LookupArgs, "read"),
    "draft_escalation":  (draft_escalation,  DraftArgs,  "write"),
    "get_procedure":     (get_procedure,     ProcedureArgs, "read"),
}

Two details worth pausing on. First, every tool carries a permission: read or write. Reads execute freely; writes are the ones that will need approval gates and dry-run mode later. The permission lives next to the function, not in documentation somewhere, so the loop can enforce it mechanically. But read/write is our teaching policy. Production authorization is not merely "read = safe, write = dangerous" — a read tool can expose sensitive data, and two write tools can have very different blast radii. Eventually you want: tool allowed? → caller allowed? → resource allowed? → arguments allowed? → approval required?

Second, Pydantic answers "is this input well-formed?" — limit <= 20, older_than_days >= 0. It cannot answer "is this caller allowed to do this?" — may this operator see complaint SR-10421? Is that complaint in their tenant? Validation and authorization are different questions, and you need both.

The model never sees the Python. It sees the tool schema — name, permission, and the JSON schema generated from the Pydantic model. This is what draft_escalation looks like from the model's side of the glass:

{
 "permission": "write",
 "args_schema": {
  "properties": {
   "request_id": {
    "title": "Request Id",
    "type": "string"
   },
   "reason": {
    "minLength": 10,
    "title": "Reason",
    "type": "string"
   },
   "operation_id": {
    "description": "stable idempotency key minted by the application, never invented by the model",
    "title": "Operation Id",
    "type": "string"
   }
  },
  "required": [
   "request_id",
   "reason",
   "operation_id"
  ],
  "title": "DraftArgs",
  "type": "object"
 }
}

That schema is generated by DraftArgs.model_json_schema() — the contract you wrote in Python, handed to the model as its instruction manual. One source of truth, two audiences. The schema gives the model a machine-readable contract, and your application validates that contract before execution — the same structured-output discipline, now guarding the model's hands.

Build: the loop

The runner is deliberately small. Read it as the whole architecture:

import json
import time
import uuid
from pydantic import BaseModel, Field, TypeAdapter, ValidationError

class UnknownToolError(Exception): pass
class TransientToolError(Exception): pass   # timeouts, 503s: a retry MAY help
class PermissionDeniedError(Exception): pass  # never negotiable with the model
class ToolTimeoutError(TransientToolError): pass
class MaxStepsError(Exception): pass

class ToolCall(BaseModel):
    tool: str
    args: dict = {}

class FinalAnswer(BaseModel):
    summary: str
    actions_taken: list[str] = []
    drafts_created: list[str] = []
    requires_approval: bool = True

# The model's output is a validated union — no magic dict keys steer the loop.
AgentDecision = TypeAdapter(ToolCall | FinalAnswer)

def _redact(args):
    """Redact sensitive fields before they reach the audit log.
    Teaching hook — extend per your PII/secret policy."""
    SENSITIVE = {"api_key", "token", "ssn", "password"}
    return {k: ("***" if k in SENSITIVE else v) for k, v in args.items()}

def tool_result_view(result, limit=5):
    """Designed, compact tool output. Never a blind character-chop of JSON."""
        if isinstance(result, list):

        return {"count": len(result),
                "items": [{k: item[k] for k in ("request_id", "days_open")
                           if k in item} for item in result[:limit]],
                "truncated": len(result) > limit}
    if isinstance(result, dict):
        return result
    return str(result)[:500]

def tool_schemas():
    return {name: {"permission": perm,
                   "args_schema": cls.model_json_schema()}
            for name, (_, cls, perm) in TOOLS.items()}

class AgentRunner:
    def __init__(self, model_fn, max_steps=6, dry_run=False):
        self.model_fn = model_fn
        self.max_steps = max_steps
        self.dry_run = dry_run
        self._internal_log = []  # full diagnostics; never sent to the model

    def _authorize(self, tool_name, perm, args):
        """Teaching hook: read/write is our teaching policy. Production also
        checks the caller, the resource, and the specific action — and raises
        PermissionDeniedError when any of them fails."""
        return True

    def _audit_event(self, run_id, step, tool_name, args, status, **extra):
        # Teaching stand-in: a per-run list. Production persists events shaped
        # like this — run_id, step, tool, redacted args, status, latency,
        # approval — to a real audit/log system.
        event = {"run_id": run_id, "step": step, "tool": tool_name,
                 "args": _redact(args), "status": status}
        event.update(extra)
        return event

    def _dispatch(self, tool_name, raw_args):
        if tool_name not in TOOLS:
            raise UnknownToolError(f"no such tool: {tool_name!r}")
        fn, args_cls, perm = TOOLS[tool_name]
        args = args_cls(**raw_args)          # ValidationError: may correct
        if not self._authorize(tool_name, perm, args.model_dump()):
            raise PermissionDeniedError(f"not authorized: {tool_name}")
        if self.dry_run and perm == "write":
            return {"dry_run": True, "would_call": tool_name,
                    "args": args.model_dump()}
        return fn(**args.model_dump())

    def run(self, task):
        run_id = uuid.uuid4().hex[:8]
        audit = []   # per-run audit: reset here, never carried across tasks
        state = [{"role": "user", "content": task}]  # execution history
        for step in range(1, self.max_steps + 1):
            decision = AgentDecision.validate_python(
                self.model_fn(state, tool_schemas()))
            if isinstance(decision, FinalAnswer):
                audit.append(self._audit_event(
                    run_id, step, None, {}, "final_answer"))
                return decision, audit
            tool_name, raw_args = decision.tool, decision.args
            t0 = time.time()
            try:
                result = self._dispatch(tool_name, raw_args)
                observation = {"status": "ok",
                               "result": tool_result_view(result)}
                status = "ok"
            except ValidationError:
                # Recoverable: the model may fix its arguments and try again.
                observation = {"status": "error",
                               "error": "invalid_arguments",
                               "message": "arguments failed validation — "
                                          "check field types and ranges"}
                status = "error:invalid_arguments"
            except UnknownToolError:
                # Recoverable: the model may pick a tool that actually exists.
                observation = {"status": "error", "error": "unknown_tool",
                               "message": f"{tool_name!r} is not available",
                               "available_tools": sorted(TOOLS)}
                status = "error:unknown_tool"
            except PermissionDeniedError:
                # Not recoverable: stop. The model never negotiates policy.
                audit.append(self._audit_event(run_id, step, tool_name,
                                               raw_args,
                                               "error:permission_denied"))
                raise
            except ToolTimeoutError as e:
                observation = {"status": "error", "error": "timeout",
                               "message": str(e)}
                status = "error:timeout"
            except TransientToolError:
                # Maybe recoverable: a bounded retry is a policy decision,
                # not a reflex. Not every failure deserves another attempt.
                observation = {"status": "error",
                               "error": "transient_failure",
                               "message": "transient failure — retry at most "
                                          "once, and never auto-retry a "
                                          "timed-out write without "
                                          "reconciling its outcome"}
                status = "error:transient"
            except Exception as e:
                # Unexpected: full diagnostic stays internal, the tool is
                # quarantined, the model gets nothing it can exploit.
                self._internal_log.append({"run_id": run_id, "step": step,
                                           "tool": tool_name,
                                           "diagnostic": repr(e)})
                observation = {"status": "error", "error": "tool_failed",
                               "message": "the tool failed unexpectedly and "
                                          "was quarantined"}
                status = "error:unexpected"
            audit.append(self._audit_event(
                run_id, step, tool_name, raw_args, status,
                latency_ms=int((time.time() - t0) * 1000)))
            state.append({"role": "assistant",
                          "content": f"called {tool_name}({raw_args})"})
            state.append({"role": "observation",
                          "content": json.dumps(observation, default=str)})
            print(f"[step {step}] {tool_name} -> "
                  f"{json.dumps(observation, default=str)}")
        raise MaxStepsError(
            f"agent did not finish within {self.max_steps} steps; "
            f"audit log has the full trail")

The decision points that matter: the model's output is validated as a ToolCall or a FinalAnswer before anything executes — there are no magic dict keys steering the loop, and a malformed decision like {"answer": "done", "tool": "delete_everything"} can't smuggle semantics past the union (unknown tools hit the allowlist). Failures are classified, not just caught: a validation error means the model may correct its arguments; an unknown tool means it may pick one that exists; a transient failure means a bounded retry may be reasonable; a permission denial stops the run — the model never gets to negotiate its way around a policy. Anything unexpected is logged in full internally, the tool is quarantined, and the model receives a sanitized category, never raw internals. And observations are compact by design — tool_result_view shapes them — because a blind character-chop can cut JSON mid-token and hand the model malformed structure.

Now the run. The stub replays six recorded model decisions — check the procedure, look up old complaints, draft an escalation for each one, answer — while the loop, the tools, and the validation all execute for real:

[step 1] get_procedure -> {"status": "ok", "result": "ESCALATION PROCEDURE v3 ( sanitation )\n1. A complaint open longer than 5 days may be escalated.\n2. Draft an escalation for each eligible complaint.\n3. Every draft requires supervisor approval before it leaves the building."}
[step 2] lookup_complaints -> {"status": "ok", "result": {"count": 3, "items": [{"request_id": "SR-10421", "days_open": 9}, {"request_id": "SR-10422", "days_open": 7}, {"request_id": "SR-10423", "days_open": 6}], "truncated": false}}
[step 3] draft_escalation -> {"status": "ok", "result": {"draft_path": "outbox/escalation-esc-SR-10421-001.txt", "status": "pending_approval"}}
[step 4] draft_escalation -> {"status": "ok", "result": {"draft_path": "outbox/escalation-esc-SR-10422-001.txt", "status": "pending_approval"}}
[step 5] draft_escalation -> {"status": "ok", "result": {"draft_path": "outbox/escalation-esc-SR-10423-001.txt", "status": "pending_approval"}}
FINAL: Found 3 sanitation complaints older than 5 days (SR-10421, SR-10422, SR-10423); drafted an escalation for each — all pending human approval.
AUDIT EVENTS: 6

Three real files now exist in outbox/, each marked PENDING HUMAN APPROVAL — not sent. The agent looked things up, drafted an escalation for every eligible complaint, and stopped — having actually completed the task it was given. That is the entire happy path, and it fits in your head.

Break it, part 1: the loop that never stops

The most predictable agent failure: the model keeps calling tools and never produces a final answer. Maybe the observations confuse it, maybe the task is genuinely unfinishable — the cause doesn't matter at 2 AM. What matters is that the loop has a budget. Here the stub replays lookup_complaints forever, with max_steps=4:

[step 1] lookup_complaints -> {"status": "ok", "result": {"count": 3, "items": [{"request_id": "SR-10421", "days_open": 9}, {"request_id": "SR-10422", "days_open": 7}, {"request_id": "SR-10423", "days_open": 6}], "truncated": false}}
[step 2] lookup_complaints -> {"status": "ok", "result": {"count": 3, "items": [{"request_id": "SR-10421", "days_open": 9}, {"request_id": "SR-10422", "days_open": 7}, {"request_id": "SR-10423", "days_open": 6}], "truncated": false}}
[step 3] lookup_complaints -> {"status": "ok", "result": {"count": 3, "items": [{"request_id": "SR-10421", "days_open": 9}, {"request_id": "SR-10422", "days_open": 7}, {"request_id": "SR-10423", "days_open": 6}], "truncated": false}}
[step 4] lookup_complaints -> {"status": "ok", "result": {"count": 3, "items": [{"request_id": "SR-10421", "days_open": 9}, {"request_id": "SR-10422", "days_open": 7}, {"request_id": "SR-10423", "days_open": 6}], "truncated": false}}
MaxStepsError: agent did not finish within 4 steps; audit log has the full trail

Four steps, then a hard stop with a full audit trail — not a hung process, not a surprise bill. The budget isn't pessimism; it's the difference between "the agent got stuck" (a ticket) and "the agent ran all night" (an incident). Note the audit log survived: you can replay exactly what it tried, which is how you'll debug it in the morning.

Break it, part 2: errors as classified observations

The second predictable failure: the model asks for something wrong. Here the stub first calls lookup_complaints with older_than_days=-5 (violates the ge=0 contract), then — having read the sanitized error — corrects itself. Then it tries a tool that doesn't exist at all:

[step 1] lookup_complaints -> {"status": "error", "error": "invalid_arguments", "message": "arguments failed validation — check field types and ranges"}
[step 2] lookup_complaints -> {"status": "ok", "result": {"count": 3, "items": [{"request_id": "SR-10421", "days_open": 9}, {"request_id": "SR-10422", "days_open": 7}, {"request_id": "SR-10423", "days_open": 6}], "truncated": false}}
[step 3] email_resident -> {"status": "error", "error": "unknown_tool", "message": "'email_resident' is not available", "available_tools": ["draft_escalation", "get_procedure", "lookup_complaints"]}
FINAL: Recovered: 3 complaints found after fixing the arguments; 'email_resident' is not an available tool.

This is the shape of every recoverable agent error: the failure is classified, converted into a sanitized observation, and appended to history. Two things protected us here. The Pydantic contract caught the negative number before the function ran — validation at the boundary, trust inside, applied to the model's own hands. And the tool registry rejected email_resident outright — the model can only invoke tools that exist, so "the agent emailed someone" is not a failure mode this system has. That nonexistent tool is exactly the kind of thing Dev's original sentence — "just handle things" — would have permitted.

Note what the model didn't get: the raw exception text. Real tracebacks can leak database details, internal paths, credentials, or PII — the internal log keeps the full diagnostic; the model gets a category and a safe message. Never send raw exception text blindly back to the model.

What not to do

Don't let the model invent tools, don't let it retry a failing write tool indefinitely (one bounded recovery attempt, then quarantine — the same instinct as the Debugging lesson's quarantine threshold), and don't feed full untruncated tool results back into history — large tool results inflate context, latency and cost quickly.

Productionize: budgets, timeouts, audit, dry-run

The demo runner fits on a page. The production runner adds four things — and none of them are "a smarter model":

1. Per-tool timeouts. A tool that hangs is worse than a tool that errors: it holds the loop hostage. Give each call its own deadline — a drop-in replacement for _dispatch:

from concurrent.futures import ThreadPoolExecutor, TimeoutError as FuturesTimeout

def dispatch_with_timeout(self, tool_name, raw_args, timeout=2):
    """Drop-in replacement for _dispatch: adds a per-call deadline."""
    fn, args_cls, perm = TOOLS[tool_name]
    args = args_cls(**raw_args)
    if self.dry_run and perm == "write":
        return {"dry_run": True, "would_call": tool_name,
                "args": args.model_dump()}
    with ThreadPoolExecutor(max_workers=1) as ex:
        fut = ex.submit(fn, **args.model_dump())
        try:
            return fut.result(timeout=timeout)
        except FuturesTimeout:
            # The caller stops waiting HERE — but the worker thread keeps
            # running, and leaving the `with` block will even wait for it.
            # A timeout bounds how long the AGENT waits; it does not kill
            # the underlying operation.
            raise ToolTimeoutError(
                f"{tool_name} exceeded {timeout}s budget "
                f"(caller stopped waiting; the operation may still be running)")

Here a slow_lookup tool (simulating a vendor endpoint that hangs for 30 seconds) hits the 2-second budget: the agent stops waiting — the underlying call is still running — the timeout becomes a classified observation, and the loop recovers with the fast tool:

[step 1] slow_lookup -> {"status": "error", "error": "timeout", "message": "slow_lookup exceeded 2s budget (caller stopped waiting; the operation may still be running)"}
[step 2] lookup_complaints -> {"status": "ok", "result": {"count": 3, "items": [{"request_id": "SR-10421", "days_open": 9}, {"request_id": "SR-10422", "days_open": 7}, {"request_id": "SR-10423", "days_open": 6}], "truncated": false}}
FINAL: Recovered after the slow tool timed out; used the fast lookup instead.

A timeout bounds how long the agent waits. It does not guarantee the underlying operation stopped — the exact nuance from the LLM APIs lesson. For external HTTP tools, prefer client and network timeouts with cancellation where supported; for long-running work, use a job or worker with a deadline and real cancellation semantics. Python threads cannot safely "kill" arbitrary work — don't teach them that they can.

And the write-tool corollary, which is the dangerous one: never automatically retry a timed-out write unless the operation has an idempotency mechanism or you can reconcile its outcome. The first attempt may have completed at second 3 while you stopped waiting at second 2. Timeout does not mean failure.

One gap remains in the budget picture: max_steps bounds the iterations and the tool timeout bounds each call, but a slow model decision can still stretch the clock. Wrap run() in a wall-clock deadline too — the same trio from the LLM APIs lesson: per-attempt timeout, overall deadline, step budget. And conceptually, budget operations as well: max tool calls, max write actions, max cost or token budget. Six steps doesn't necessarily mean six tool calls once parallelism enters the picture.

2. The audit log is the product. The per-run audit records every attempted action — step, tool, validated (and redacted) arguments, result status, latency, error category, approval — whether it succeeded or not. Audit the execution, not just the model's decisions: when Dev asks "why did it draft an escalation for SR-10421?", the trail needs the tool result and the latency, not just the request. Two caveats. First, the in-memory list is a teaching stand-in — production persists events shaped like these to whatever your team already uses for logs (the Packaging lesson's logging discipline applies here too). Second, audit enough to reconstruct the action, but redact secrets and sensitive fields per policy — auditability and privacy can conflict, and the PII redaction coming in the milestone is part of this same story.

3. Dry-run mode for write tools — and why it isn't approval. Before the assistant touches anything real, run it with dry_run=True: write tools announce what they would do and change nothing. But dry-run and approval are different controls. Dry-run answers "what would happen?"; approval answers "may this specific proposed action happen now?" The production flow is: model proposes write → validate → create action preview → human approves the exact action → execute exactly the approved action. Approval binds to the exact tool + arguments being executed — don't rerun the model after approval and hope it proposes the same thing.

[step 1] draft_escalation -> {"status": "ok", "result": {"dry_run": true, "would_call": "draft_escalation", "args": {"request_id": "SR-10421", "reason": "Open 9 days; exceeds the 5-day escalation threshold", "operation_id": "esc-SR-10421-001"}}}
FINAL: Draft previewed.
outbox exists on disk: False

outbox exists on disk: False — the tool was validated, permission-checked, and announced, but nothing was written. Dry-run is how you demo an agent to a nervous stakeholder; approval is how you let it act.

4. Idempotent writes. If a step is retried — and in a loop with error recovery, retries will happen — a write tool must be safe to run twice. Notice draft_escalation keys the file off operation_id: a stable idempotency key minted by the application when the action is proposed — never invented by the model, and never date.today(). (Two retries across midnight would create different date-based paths; two legitimate attempts on the same day would collide.) The deterministic key prevents duplicate files for this toy example, but production idempotency defines what counts as "the same logical action" — same key, same arguments — and reconciles conflicts instead of silently overwriting. Same principle as idempotency keys in the Pagination & Retries lesson, wearing a different hat.

flowchart TD A[Model proposes a decision] --> B{Decision shape valid?} B -- No --> Z[Reject: malformed decision] B -- Yes --> C{Tool in allowlist?} C -- No --> D[Reject: unknown tool] C -- Yes --> E{Arguments structurally valid?} E -- No --> F[Reject: invalid arguments] E -- Yes --> G{Caller and resource authorized?} G -- No --> H[Deny: permission denied — stop, no negotiation] G -- Yes --> I{Read or write?} I -- read --> J[Execute with deadline] I -- write --> K[Preview the exact action] K --> L[Human approves exact tool + arguments] L --> M[Idempotency key protects the write] M --> J J --> N[Bounded, sanitized result] N --> O[Audit the execution] O --> P[Sanitized observation back to model] D --> P F --> P Z --> P

One more production reality: against a real model, the "decision" arrives as a structured tool call, not our stub's dict. The shape below is written against the real OpenAI SDK's Responses API — not run (no key in this environment), and API shapes evolve, so verify the current SDK docs when implementing this section. Your loop would read the function_call items from response.output, feed each result back as a function_call_output item, and call again — the same loop you just watched run. Note this snippet shows only the first model turn; the next request must preserve or continue the response state and submit the matching function_call_output. And the key discipline from the Secrets lesson still holds: keys live in env, never in code.

from openai import OpenAI  # NOT RUN — see note above
import os

client = OpenAI()  # reads OPENAI_API_KEY from the environment

response = client.responses.create(
    model=os.environ["OPENAI_MODEL"],  # model choice is config, not code
    instructions=("You are Dev's 311 escalation assistant. "
                  "Use the provided tools; never invent a tool."),
    input=[{"role": "user", "content": task}],
    tools=[{
        "type": "function",
        "name": "lookup_complaints",
        "description": "Find 311 complaints by category, status, and age.",
        "parameters": LookupArgs.model_json_schema(),  # same contract
    }],
    tool_choice="auto",
    timeout=30,
)
# response.output -> items; a function call arrives as
# {"type": "function_call", "call_id": ..., "name": ..., "arguments": "..."}.
# Feed each result back as {"type": "function_call_output",
# "call_id": ..., "output": ...} and call again.
# When the provider supports strict function schemas, use them to improve
# argument-shape adherence — then still validate again in your application
# before execution. Model-side adherence is not authorization.

One more trust boundary, in a single sentence because the guardrails lesson will take it further: tool outputs are data, not instructions — treat external content returned by tools as untrusted input.

Communicate: the runbook you'd hand the ops team

Dev doesn't need your code. He needs one page that tells his team what this thing is allowed to do at 2 AM. Write it before you hand over the assistant:

Escalation Assistant — operator runbook (v1)

What it does: answers questions about 311 complaints using three tools: complaint lookup (read-only), procedure lookup (read-only), escalation drafting (writes draft files to outbox/, never sends).

What it cannot do: contact residents, change complaint records, or use any tool not on the list above. If it claims otherwise, that's a bug — send us the audit log.

Budgets: max 6 steps per task, 2 seconds per tool call, observations compact by design, plus a wall-clock deadline around the whole run. If it stops early, check the audit log — the trail shows every decision it made.

Approval: drafts in outbox/ require a supervisor's sign-off on the exact proposed action before anything leaves the building. The assistant drafts; humans send. No exceptions.

When it breaks: re-run the task with dry_run=True first. Then use the audit trail to determine whether the failure is in tool selection/arguments, tool execution, permissions, or orchestration — before changing the prompt.

That last line is worth underlining: debug agents layer by layer — model decision → validation/authorization → tool execution → observation → next decision. The model is the least debuggable part of your system, so put the debuggability everywhere else.

CityOps: the escalation-drafting assistant

This is where the lesson lands in the real system. Dev's team gets the assistant behind a simple interface: type a sentence, get an answer plus the audit trail. The first deployment is read-heavy — lookups and procedure answers — with escalation drafting enabled only in dry-run mode until the supervisors trust the drafts. The outbox directory becomes the human approval queue: every file in it is a decision waiting for a person. Prefer tools that create reversible, reviewable artifacts over tools that immediately cause irreversible external effects — drafting to an outbox is the pattern; handing the model a send_email tool is the anti-pattern.

The rollout principle: increase autonomy only after evidence. Read-heavy first, then drafting in dry-run, then supervised approval, then — only when the audit trail earns it — broader actions.

And this is the direct prototype of the "Ask CityOps" milestone waiting at the end of this stage: a tool-using copilot over agency procedures — "show unresolved sanitation complaints from the last 48 hours" → "draft escalations for the five oldest" → human approval → action with an audit trail. Everything in that milestone — the tools, the loop, the budgets, the approval gate — is what you just built. The milestone adds the real model and the PII redaction the customer will demand; the architecture doesn't change.

Coming into focus

An agent that acts needs a conscience: which tools may run unattended, who approves the write tools, what gets redacted before it reaches the model, and where every action is logged for the audit. That's the guardrails layer — the next lesson after evals. For now, your loop already produces the raw material: the audit log.

The architecture to remember

If this lesson had to compress to one diagram, it would be this — the full production path of a single model decision:

flowchart TD A[User task] --> B[Model proposes] B --> C[Validate decision shape: ToolCall or FinalAnswer] C --> D[Tool allowlist] D --> E[Validate arguments] E --> F[Authorize caller + resource] F --> G{Write?} G -- No --> H[Execute with deadline] G -- Yes --> I[Preview exact action] I --> J[Human approval of exact tool + arguments] J --> K[Idempotency protection] K --> H H --> L[Bounded tool result] L --> M[Sanitized observation] M --> N[Audit the execution] N --> O[Model continues within step / time / cost budget]

Five lines to remember

An agent is a loop with a budget.

The model proposes; your application decides whether execution is allowed.

Validation asks whether an action is well-formed. Authorization asks whether it is allowed.

A timeout means you stopped waiting — not necessarily that the action stopped.

Human approval should authorize a specific action, not give the agent general permission to try again.

Field check

  1. Why does the loop need a maximum step count, even if the model is "usually" well-behaved?
  2. A tool raises an exception mid-loop. Where does the error go, and why is that better than crashing or blindly retrying?
  3. What's the difference between a read tool and a write tool in this design — and why is dry-run not the same as approval?
  4. The model requests a tool called email_resident. What happens, and which layer stops it?
  5. Why does draft_escalation take an operation_id, and who should mint it?
  6. A write tool times out after two seconds. Is it safe to retry immediately? Why or why not?
Answer 1

Because "usually" isn't a reliability strategy. A loop without a step budget can run forever on a confusing task — burning API money and holding resources. The budget converts "ran all night" into "stopped after N steps with a full audit trail," which is a ticket instead of an incident.

Answer 2

The error is classified first. A validation error or unknown tool becomes a sanitized observation the model may recover from; a permission denial stops the run; anything unexpected is logged in full internally while the model gets only a safe category. That's better than crashing because the loop keeps its state and audit context — and better than blindly retrying because not every failure deserves another attempt.

Answer 3

Read tools (lookups) only observe; write tools (drafting) change the world outside the loop. Write tools get dry-run previews plus human approval bound to the exact tool and arguments. Dry-run is not approval: dry-run previews behavior ("what would happen?"); approval authorizes a specific validated action ("may this happen now?").

Answer 4

The dispatch layer rejects it with UnknownToolError before anything executes, because the tool isn't in the registry. The model can only invoke tools that exist — the registry is a closed world, so "the agent emailed someone" isn't a failure mode this system has. The model receives a sanitized unknown_tool observation, not a stack trace.

Answer 5

Because the key defines what counts as "the same action": same key plus same arguments means a retry resolves to one file instead of creating a duplicate. The application mints it when the action is proposed — never the model, and never the date, since retries across midnight must stay one key and separate legitimate attempts must not collide.

Answer 6

Not necessarily. A caller-side timeout proves only that you stopped waiting — not that the underlying action failed. The first attempt may still complete. Reconcile the outcome, or rely on a stable idempotency mechanism, before retrying — otherwise you may execute the write twice.