Every AI lesson so far assumed you could send data to someone else's computer. This is the lesson where the client takes that option off the table — and you ship the copilot anyway.
The customer problem
Thursday. The "Ask CityOps" copilot prototype works — point it at a complaint, get a clean summary, draft an escalation. Lisa is excited. Then Legal reviews the architecture diagram and kills it in one sentence:
"Complaint text contains resident names, addresses, and phone numbers. It cannot leave our network. Find another way."
In the current prototype, complaint text containing PII is sent to an external provider on every call. The current design violates the client's legal/compliance requirement — not wrong, prohibited. (Note the precision: Legal told us the data cannot leave the network. That's enough to kill the design; we haven't established that a statute was broken, and an FDE doesn't claim one was.)
Lisa asks the question this lesson answers: "Can the copilot still work if the model has to live inside our walls?"
Clarify the ask
"Can't leave our network" is a constraint, not a technology choice. Before reaching for tools, pin down what it actually rules out:
- Data residency: complaint text must not cross the company boundary. This kills external APIs for the PII-bearing calls — full stop.
- But "local" is a spectrum, not a laptop: the allowed boundary could be customer-controlled on-prem infrastructure, a private cloud/VPC, approved regional hosting, an isolated managed deployment, or fully air-gapped inference. Don't let "data residency" collapse in your head to "a laptop under the desk."
- The strict assumption for this lesson: nothing PII-bearing goes out, period. Where a client's policy permits multiple data classes, routing can be per workload — but under CityOps' strict assumption here, PII-bearing workloads stay entirely inside the approved boundary.
- The real question: can we run a capable-enough model inside the approved boundary, on hardware the client can actually operate — and prove it before we recommend it?
Notice this isn't your preference — it's the client's constraint. FDEs don't argue with Legal about data residency; they design inside it. The FDE principle for this lesson:
One principle
"Data gravity decides where the model is allowed to run." You don't pick local models because they're fashionable. You pick them when the data can't move — and then you engineer around everything the on-prem model can't do as well. In this lesson our implementation choice is Ollama serving a quantized model, but the principle covers any approved boundary: on-prem, VPC, regional hosting, air-gapped.
The minimum concept: quantization, honestly
A local model is the same kind of neural network as the API ones. Two independent dimensions decide whether it fits on your hardware — don't blur them together:
1. Parameter count — how big the model is: 8B vs 70B parameters. This is about capability and size.
2. Quantization precision — how many bits store each weight: FP16/BF16, 8-bit, 4-bit. This is purely a storage/accuracy tradeoff.
An 8B model may be unquantized. A 70B model may be quantized. A smaller model's capability difference isn't caused solely by quantization — keep the two axes separate in your head.
Quantization stores each weight in fewer bits. Lower precision substantially reduces weight memory; depending on the model and quantization method, it can also reduce output quality or accuracy somewhat — so evaluate the quantized model on your actual workload. What quantization does not do: 4-bit quantization does not change the model's configured context window, and "wandering" isn't a technical property of quantization at all.
The arithmetic is useful first-order estimation — but label it correctly. This computes raw weight storage only, not total runtime memory:
# raw WEIGHT storage estimates (run for real) — not total runtime memory
for name, params_b, bits in [("8B model", 8, 16), ("8B model", 8, 4), ("70B model", 70, 4)]:
gb = params_b * 1e9 * (bits / 8) / 1e9
print(f"{name} at {bits}-bit: {gb:.1f} GB of weights")
8B model at 16-bit: 16.0 GB of weights
8B model at 4-bit: 4.0 GB of weights
70B model at 4-bit: 35.0 GB of weights
Don't memorize — internalize
Actual inference memory is weights plus: KV cache, runtime/backend overhead, working buffers, context-dependent memory — and quantization formats themselves carry metadata overhead. An 8B 4-bit model has roughly 4 GB of raw quantized weights, but actual runtime memory is higher and depends on the serving backend, context length, and hardware. If you have exactly 4 GB free and load a "4 GB model," you will have a bad day. Quantization reduces precision and memory; it does not magically make runtime memory equal to the model file size.
For this lesson, we'll use an 8B-class quantized model because it illustrates a common laptop/small-server deployment tradeoff. Many quantized models can run CPU-only — but interactive latency and concurrent throughput may be unacceptable, so benchmark on the actual deployment hardware. "Runs on a decent laptop" is a claim you test, not a slogan you print.
Build: Ollama and the provider abstraction
The practical path is Ollama: one install, one command to fetch a model (ollama pull llama3.1:8b), and it serves an OpenAI-compatible HTTP endpoint on localhost:11434. That compatibility is genuinely useful — but read it precisely: Ollama exposes OpenAI-compatible endpoints for common API patterns, which lets your code talk to it like an API. Don't assume every provider feature behaves identically: tool calling, structured output, streaming, multimodality, error semantics, and provider-specific options can differ across compatibility layers. That's exactly why the pattern below matters more than the compatibility itself.
Your application never calls a provider. It calls a ModelBackend, and configuration picks the implementation. The same code runs against the cloud API in dev and the on-prem model at the client:
from abc import ABC, abstractmethod
class ModelBackend(ABC):
@abstractmethod
def generate(self, system: str, user: str) -> str:
"""Return the model's raw text reply."""
# Set from the workload's latency SLO and measured model performance —
# not a magic constant. A local CPU model may legitimately be slow,
# but the application still needs an overall deadline.
REQUEST_TIMEOUT_S = 60
class LocalBackend(ModelBackend):
"""Talks to Ollama's OpenAI-compatible chat endpoint."""
def __init__(self, model, base_url="http://localhost:11434/v1", post_fn=None):
import requests
self.model = model
self.url = base_url.rstrip("/") + "/chat/completions"
self._post = post_fn or requests.post # injectable for tests
def generate(self, system, user):
resp = self._post(
self.url,
json={"model": self.model,
"messages": [{"role": "system", "content": system},
{"role": "user", "content": user}]},
timeout=REQUEST_TIMEOUT_S,
)
resp.raise_for_status()
return resp.json()["choices"][0]["message"]["content"]
Two deliberate choices in that code. First, no temperature=0: lower sampling variability is not a reliability mechanism — our reliability comes from schema-constrained output when the serving stack supports it, plus application validation, quarantine, and evaluation. Second, the post_fn injection: in production it's requests.post; in tests — and in this lesson — it's a stub returning Ollama's real response shape. (No multi-GB downloads in this environment, so the lesson exercises everything around the model: routing, parsing, validation.)
One principle
"Compatibility reduces adapter work; your own interface prevents compatibility quirks from becoming application architecture." Notice the code calls ModelBackend.generate(...), never raw OpenAI-shaped HTTP. If Ollama's streaming or error semantics differ from the provider's, the difference is contained in one adapter — not smeared across the copilot.
Backend selection is configuration, not code — an env var, per the Secrets lesson's discipline (env vars are one good mechanism; config files, deployment systems, and secret managers are others):
import os
def choose_backend() -> ModelBackend:
mode = os.environ.get("CITYOPS_MODEL_BACKEND", "local").lower()
if mode == "local":
# The friendly tag is for humans. The deployment identity —
# exact artifact, quantization, serving config — is recorded
# separately (see "Productionize").
return LocalBackend(model=os.environ.get("CITYOPS_LOCAL_MODEL",
"llama3.1:8b"))
key = os.environ.get("CITYOPS_API_KEY")
if not key:
raise SystemExit("CITYOPS_API_KEY is not set (and "
"CITYOPS_MODEL_BACKEND is not 'local').")
return APIBackend(model="provider-flagship", api_key=key)
# run for real: the same call site, two worlds
os.environ["CITYOPS_MODEL_BACKEND"] = "local"
b = choose_backend()
print(type(b).__name__, "->", b.model) # LocalBackend -> llama3.1:8b
os.environ["CITYOPS_MODEL_BACKEND"] = "api"
os.environ["CITYOPS_API_KEY"] = "demo"
print(type(choose_backend()).__name__) # APIBackend
LocalBackend -> llama3.1:8b
APIBackend
One call site, two deployments. The copilot code never changes — only the environment does. That's what "data gravity decides" looks like in practice.
calls ModelBackend.generate] --> B{Env config
CITYOPS_MODEL_BACKEND} B -->|api| C[External API
data crosses the boundary] B -->|local| D[Ollama inside the approved boundary
localhost:11434] D --> E[Quantized model
weights ~4 GB + runtime overhead]
Don't memorize — internalize
That diagram says "inside the approved boundary" — not "data never leaves." For this single-machine configuration, requests stay local. But production on-prem often means application server → internal model server, and you still have to secure the serving stack: network controls, authentication, TLS where appropriate, access control, and logging/telemetry policy. A local deployment can still leak data through logs, telemetry, crash reporting, backups, monitoring, or misconfigured networking. Local inference can keep model inputs within the approved boundary, but you still have to secure the serving stack and verify its network/telemetry behavior. "Local" is not automatically private, secure, or compliant.
The decision framework: local vs API, honestly
Neither option wins everywhere. Here's the comparison you'd actually put in front of Lisa — with the honest "it depends" left in:
| Dimension | External API | Local model |
|---|---|---|
| Data privacy | Data leaves your network — a non-starter for PII under strict policy | Can be operated entirely within the approved environment when configured accordingly — but you still secure the serving stack |
| Model quality | Generally stronger, especially on reasoning and long context | May be weaker on your workload — measure the exact model and configuration rather than assume parity |
| Cost shape | Per-token billing; scales with usage, surprises at scale | No provider per-token charge; costs shift toward hardware, capacity, electricity, and operations |
| Latency | Network round-trip + shared capacity; usually fine | No network hop — but CPU-only inference can be too slow for interactive use; benchmark |
| Ops burden | Lower infrastructure burden: the provider operates inference hardware, while you still own integration, reliability, governance, and cost controls (credentials, rate limits, timeouts, retries, provider outages, version changes, observability) | You own it: server health, model updates, memory, monitoring |
| Capacity / scale | Provider capacity, subject to quotas and rate limits | Throughput and concurrency limited by the hardware you provision |
| Change control | Provider and model behavior can change under you; pin model/version where possible | You control when artifact and runtime updates happen |
| Offline / air-gapped | Unavailable | External hosted APIs are unavailable; inference must run inside the air-gapped environment |
The FDE read: match the backend to the constraint. Under CityOps' strict assumption, PII-bearing summarization runs local; where a client's policy permits multiple data classes, routing can be per workload — but that's a policy decision, not an engineering shortcut. The abstraction makes either choice a config decision, not a rewrite.
Break: smaller models wander — the contract catches it
Here's what lower instruction-following reliability looks like in practice. The stub below replays Ollama's real response shape with a recorded reply from a small local model asked to classify a complaint as JSON. Watch what comes back:
import json
from typing import Literal
# BaseModel, Field, ValidationError: from pydantic, per the validation lesson
class ComplaintSummary(BaseModel):
category: Literal["sanitation", "noise", "roads", "water", "other"]
urgency: Literal["low", "medium", "high"]
summary: str = Field(min_length=10, max_length=280)
def summarize_with_backend(backend, complaint):
raw = backend.generate(
system=("Classify the 311 complaint. Reply with ONLY a JSON object "
"with keys category, urgency, summary."),
user=complaint,
)
try:
data = json.loads(raw)
except json.JSONDecodeError:
return ("quarantined", f"not JSON: {raw[:60]!r}")
try:
return ("ok", ComplaintSummary(**data))
except ValidationError as e:
return ("quarantined",
"; ".join(f"{err['loc'][0]}: {err['msg']}"
for err in e.errors()))
A note on the asking: "Reply with ONLY a JSON object" plus manual parsing is a valid fallback. The real hierarchy is: provider/server-supported constrained structured output first, then application validation, then bounded repair or fallback if appropriate, then quarantine — never "please output JSON → hope → json.loads." One line to carry forward: use schema-constrained generation when the chosen local serving stack/model supports it; validate again in the application either way. (And note the schema itself: Literal enums, not regex patterns — the same contract style the Pydantic lesson taught, producing a clearer generated schema.)
# recorded local-model replies, run through the contract for real
complaint = "Trash bins overflowing on 5th Ave, second report this week."
# a good day for the small model:
('ok', ComplaintSummary(category='sanitation', urgency='high', summary='Overflowing trash bins on 5th Ave reported twice this week.'))
# a normal day for the small model:
('quarantined', 'not JSON: "Well, the complaint... it\'s about, uh, garbage? or maybe the"')
The full recorded reply behind that quarantine: "Well, the complaint... it's about, uh, garbage? or maybe the noise from the trucks? Hard to say really, could be either honestly." The small model ignored the instruction and mused out loud. Smaller/local models may have lower instruction-following reliability on this workload — behavior depends on the specific model, quantization, prompt, decoding, structured-output support, and task, so measure the exact model and configuration rather than assume parity with a big API model. This is why the structured-output discipline from the Pydantic lesson isn't optional for local deployments — it's the compensating control. The less reliable the model is on your task, the more visible the value of strong contracts becomes — but the contract belongs around every model. Quarantine, don't crash: the complaint goes to a review queue with the reason attached, exactly like the Debugging lesson's quarantine philosophy.
Don't memorize — internalize
Two more failure modes you'll meet with local models. Memory exhaustion: model/runtime memory must fit the available GPU/system-memory configuration — local inference can split and offload between GPU and system RAM depending on the backend. If it doesn't fit, loading may fail outright, or offloading may make latency unacceptable. (Offloading is literally the case where it gets slower rather than failing — size the model to the hardware before demo day.) And silent drift on updates: pulling a newer model can change behavior without warning. A friendly tag like llama3.1:8b is for humans; what you record is the model deployment identity — immutable artifact identifier or digest where the serving stack supports it, plus quantization, context configuration, serving backend/version, and prompt version. When quality changes next month, that identity is how you find out what actually changed — the same instinct as pinning the embedding model version with its index from the embeddings lesson.
Productionize: own the server like you own the code
Running the model means operating it. Three pieces, all runnable:
1. Health checks — liveness vs readiness. Ollama answers GET / with Ollama is running when the process is up. That's liveness: the server responds. It does not prove the model you need is loaded, that inference works, that latency is acceptable, or that memory isn't exhausted. That's readiness: the required model can actually serve requests. Teach both, check both — a copilot that silently talks to a live-but-unready server is worse than one that's loudly down:
import logging
log = logging.getLogger("cityops.local")
def local_liveness(base_url="http://localhost:11434", get_fn=None):
"""Liveness: does the server process respond?"""
import requests
get = get_fn or requests.get
try:
r = get(base_url + "/", timeout=5)
return r.status_code == 200 and "Ollama is running" in r.text
except Exception as e:
# A boolean probe is the right shape for a health boundary —
# but log the real diagnostic, or operators get "False" with no idea why.
log.warning("local model liveness probe failed: %r", e)
return False
def local_readiness(model, base_url="http://localhost:11434", get_fn=None):
"""Readiness: is the required model actually available to serve?"""
import requests, json
get = get_fn or requests.get
try:
r = get(base_url + "/api/tags", timeout=5)
r.raise_for_status()
loaded = [m["name"] for m in r.json().get("models", [])]
return any(model in name for name in loaded)
except Exception as e:
log.warning("local model readiness probe failed: %r", e)
return False
# stubs stand in for requests.get so the lesson runs with no server:
class _FakeResp:
def __init__(self, text=None, payload=None, status=200):
self.text, self._payload, self.status_code = text or "", payload, status
def json(self):
return self._payload
def raise_for_status(self):
pass
def _refused(url, timeout):
raise ConnectionError("refused")
up = lambda url, timeout: _FakeResp(text="Ollama is running")
ready = lambda url, timeout: _FakeResp(payload={"models": [{"name": "llama3.1:8b"}]})
print("liveness, server up: ", local_liveness(get_fn=up))
print("liveness, server down:", local_liveness(get_fn=_refused))
print("readiness, model loaded:", local_readiness("llama3.1:8b", get_fn=ready))
liveness, server up: True
liveness, server down: False
readiness, model loaded: True
2. Graceful degradation. If the local server is unreachable, the copilot must not crash mid-shift. Don't catch one exception type and call it done — think in failure categories (unreachable, timeout, server error, capacity/OOM, invalid model response), then decide per category: retry, queue, quarantine, or alert:
import requests
def failure_category(exc):
"""Map a transport failure to an operational category."""
if isinstance(exc, requests.exceptions.Timeout):
return "timeout"
if isinstance(exc, requests.exceptions.ConnectionError):
return "unreachable"
if isinstance(exc, requests.exceptions.HTTPError):
return "server_error"
return "unknown"
try:
summary = summarize_with_backend(local_backend, complaint)
except requests.exceptions.RequestException as e:
category = failure_category(e)
# not a model failure — an ops failure. Queue, don't crash.
print(f"queued for retry later (local model {category}: {e})")
queued for retry later (local model unreachable: localhost:11434 refused)
One sentence the lesson can't skip: queueing also needs a retry budget, backpressure, and operator alerting — a dead model server for six hours can otherwise build an enormous silent queue. This connects straight back to the API reliability lessons.
3. Env discipline. Backend selection, model tag, and base URL all come from the environment — never hardcoded. The Secrets lesson's rule applies to config generally, not just keys: what varies between deployments lives outside the code.
liveness + readiness?} B -->|yes| C[Generate summary] B -->|no| D[Queue for retry
operator alerted] C --> E{Passes Pydantic
contract?} E -->|yes| F[Write to dashboard] E -->|no| G[Quarantine + reason
human review]
The evaluation gate: privacy decides where, evaluation decides what
Before Lisa gets a recommendation, the lesson's missing FDE step: benchmark the candidate model on the actual target hardware, on representative CityOps tasks. Privacy determines the deployment boundary; evaluation determines whether the model inside that boundary is good enough. The gate is small but non-negotiable — measure quality, latency, throughput, memory, failure rate, and structured-output adherence:
quality + operations eval] D --> E{Meets quality and latency
thresholds?} E -->|no| F[Try a smaller/different
model or config] E -->|yes| G[Deploy inside the boundary] G --> H[Schema-constrained output] H --> I[Application validation] I -->|pass| J[Product] I -->|fail| K[Quarantine / review]
Run this gate and two things happen: you find out whether the 8B quantized model is actually good enough for complaint summarization (not assumed — measured), and you get the numbers the options memo below needs. A recommendation without this gate is a guess with a logo.
Communicate: the options memo
Lisa doesn't need a lecture on quantization. She needs a decision she can defend to Legal. The memo is three options, tradeoffs, one recommendation:
To: Lisa · Re: Copilot under the data-residency constraint
Option A — fully on-prem (recommended): Ollama + an 8B-class quantized model on the CityOps server. Complaint text is processed inside the approved boundary. Tradeoff: a smaller model with lower instruction-following reliability — we compensate with schema-constrained output, strict application contracts, and a quarantine queue. We don't yet know the review rate; we'll measure it on a representative CityOps evaluation set before setting the production review/SLA expectation.
Option B — hybrid: on-prem for PII-bearing calls, external API for redacted analytics where policy permits. More moving parts, two bills, but best quality where it's allowed.
Option C — wait: pause the copilot until Legal approves external processing. Costs nothing, delivers nothing.
Recommendation: A, subject to the benchmark gate: local deployment is acceptable only if it meets both the privacy requirement and the agreed quality/latency thresholds on the target server. It appears feasible for customer-controlled hardware — the benchmark confirms it.
Note the honesty rule in action: no invented review percentages, no "runs on hardware we have" — per the curriculum's QA standard, you don't assert a number your system hasn't produced, and a recommendation without measured evidence is a guess.
CityOps: the copilot that stays inside the walls
With the abstraction in place, "Ask CityOps" gets a deployment profile: CITYOPS_MODEL_BACKEND=local on the on-prem server, liveness and readiness checks in the startup path, quarantine queues wired to the operator console, and the full model deployment identity recorded alongside every release. The same codebase, the same contracts — and the architecture gives Legal and Security a design they can actually evaluate, because inference remains inside the approved boundary.
The same data-residency logic applies to the embeddings and vector search from the earlier lesson — local embedding models are cheap to run, and the embedding model version stays pinned with the index. And this is the dress rehearsal for the Milestone's harder constraint: it's not just "data can't leave" — it's "PII can't reach the model at all," even the local one. Redaction before inference. But that's the Milestone's fight.
Useful later
- vLLM / llama.cpp server mode — when Ollama's simplicity isn't enough and you need throughput tuning, batched inference, or custom serving.
- Fine-tuning small models — for narrow, well-defined tasks, fine-tuning may improve a smaller model's task performance; evaluate whether the gain justifies the data and operational cost. A project, not an afternoon.
- Embeddings locally too — the same data-residency logic applies to the vector search from the embeddings lesson; local embedding models are cheap to run, and the embedding model version stays pinned with the index.
Field check
- Legal says complaint text can't leave the network. Why is "just use the API anyway and hope" not an option, technically and professionally?
- An 8B model at 16-bit needs ~16 GB of weights; at 4-bit ~4 GB. What did you trade away for that 4× saving?
- Why does the lesson put the model behind a
ModelBackendinterface instead of calling Ollama directly in the copilot code? - The small model returns chatty prose instead of JSON. What catches it, and where does the complaint go?
- Your on-prem server has 8 GB of RAM free. A teammate suggests the 70B model "for better quality." What do you tell them?
Answers
1. Technically, in the current prototype, complaint text containing PII is sent to an external provider — using the API anyway violates the constraint by design, not by accident. Professionally, FDEs don't bypass Legal; constraints are design inputs, and violating data policy can end the engagement and create liability.
2. Numerical precision per weight — 4-bit stores each weight as a coarser approximation, which substantially reduces weight memory. Depending on the model and quantization method, some task quality may degrade, so you evaluate the quantized model on your workload. What you did not trade away: the configured context window isn't automatically quartered, and "wandering" isn't a property of quantization.
3. So backend selection is configuration, not a rewrite: the same copilot code runs against the cloud API in dev and the on-prem model at the client. Compatibility reduces adapter work, but your own interface prevents compatibility quirks — differing structured-output support, error semantics, streaming — from becoming application architecture. It also makes the code testable via stub injection.
4. The application-side contract catches it: json.loads fails on the prose, so the reply is quarantined with the reason attached and routed to human review. The less reliable the model is on your task, the more visible the value of strong contracts becomes — but the contract belongs around every model, including the big API ones.
5. No. Even the rough 4-bit weight estimate for 70B is ~35 GB — before KV cache, runtime overhead, buffers, and context memory — so it won't fit in 8 GB; you'll get load failure or unusably slow offloading, not better quality. Choose a model that fits the actual memory budget and benchmark it on the target hardware.