Prerequisites: Stages 1–5, the CityOps spine through Milestone 5: CityOps Goes Live, and the companion finale CityOps Under Fire. That chapter attacked the product from the inside — the pager, the incidents, the trial. This one works the other side of the table: a real client, Acme, signing on. Same curriculum, mirrored: the outcome is the destination, but this time the customer is in the room writing the done-line with you.

The companion finale attacked the product — six incidents, one trial, judgment under pressure. This finale builds the relationship: Acme, a facilities-services company, is onboarding onto the platform. Discovery interviews, the real ask behind the ask, a 30-day plan, one dataset with contracts, success metrics defined together, and a stakeholder review that proves the outcome. The one principle for the whole chapter: onboard the workflow, not the data.

Dev (Monday, 9:00 AM): "New engagement. Acme — facilities services, six sites, maintenance tickets in three spreadsheets and a shared inbox. They've signed a pilot. Your job: onboard them. Not 'load their data.' Onboard them. Four weeks from today, their ops director decides whether this was worth it."

Maria: "And the ops director is the one who decides what 'worth it' means. Not us. Her signature is on the done-line before we touch a single ticket."

How to read this chapter

This chapter is a simulation. Acme is fictional; every number in it — 12,400 tickets, 424 quarantined records, 34 technicians, 30 days — is a scenario constant, not a measured fact. The client is teaching fiction; the machinery is real. Every artifact below runs against the actual platform components from the milestones: the validation contracts from the data-validation lesson, the API authorization from Milestone 3, the briefing grammar from CityOps Under Fire. The code is written to be read and adapted, not copy-pasted into a client engagement.

The shape of the engagement

Four weeks, four artifacts. Each week ends with something the client can hold — and each artifact has an owner on the client side, because an onboarding owned only by the vendor is a demo with a longer timeline:

WeekThe workThe artifactOwned by
1Discovery interviewsThe real ask, written downAcme's ops director
2Scope the first engagementThe 30-day plan, signedAcme's ops director
3Onboard the first datasetContracts, reconciliation, the live workflowAcme's data owner
4Measure success, then review itMeasured metrics + the stakeholder reviewAcme's ops director

Notice what's missing: there is no "migrate all the data" week and no "train everyone" week. The engagement is scoped to one dataset, one workflow, one review. Everything else is explicitly refused in the plan — with the refusal written down, the way the borough-pack tradeoff memo wrote down its refusals. Scope discipline isn't saying no to the client; it's saying no to the second, third, and fourth engagements hiding inside the first.

The minimum concept: onboard the workflow, not the data

Every failed onboarding Dev has seen failed the same way: the team treated the engagement as a data-migration project — extract, load, transform, celebrate row counts. The client's actual problem was never "our data isn't in your system." It was a workflow that hurt: a supervisor building tomorrow's backlog by hand at 11 PM, a standup that starts with guesses, a spreadsheet nobody trusts. The data is the fuel; the workflow is the engine. Onboard the engine.

The protocol, in order — it mirrors the triage protocol's shape (scope, then evidence, then the smallest safe change), but its instruments are people:

flowchart TD ASK[Stated ask
“load our data”] --> DIS[1. Discovery
who decides,
what a good week looks like] DIS --> REAL[The real ask
the workflow that hurts] REAL --> DONE[2. Done-line
written with the client,
owned by the client] DONE --> PLAN[3. 30-day plan
one dataset, one workflow,
one review] PLAN --> DATA[4. First dataset
validate, reconcile, authorize] DATA --> BRK{Reconciliation clean?} BRK -- No --> Q[Quarantine — client decides
the policy; map, version,
re-ingest] BRK -- Yes --> MET Q --> MET[5. Metrics
defined together, versioned,
baselined] MET --> REV[6. Stakeholder review
what happened, what's true now,
what changed permanently] REV --> OUT[Customer outcome]

Three things about this order. First, discovery before the done-line: you cannot write what "done" looks like until you've watched the workflow that hurts. Second, the client owns the done-line: "worth it" is their judgment, and the artifact that records it carries their signature — because the person who decides whether the pilot continues is the person who defined success. Third, the review is a deliverable, not a ceremony: it runs on the same grammar as an incident briefing — what happened, what's true now, what changed permanently — because an onboarding review and an incident briefing are the same skill: bad news delivered early, with a plan attached, to someone who needs to act.

The one question that starts every engagement

Before the first interview, ask: "Who decides whether this was worth it, and what does their best week look like?" The answer names the owner of the done-line and the workflow you're actually onboarding. Everything after it is execution.

Week 1: the discovery interviews

Acme's ops director (kickoff call): "Here's what we need: get our maintenance tickets into your platform. Three spreadsheets, one shared inbox, six sites. We want to see everything in one place — dashboards, trends, the works."

Maria (after the call, to you): "Write down exactly what she said. Now forget the nouns and keep the verbs. What did she actually ask for?"

This is the Stage 1 skill — clarify the ask — applied to a client who is paying for the answer. "See everything in one place" is a stated ask; it describes a dashboard, not a problem. The discovery interviews exist to find the workflow that hurts. Two 30-minute interviews and one shadowed shift:

  • The night-shift supervisor (the person closest to the work): "Every night around 11, I build tomorrow's backlog. I pull the inbox, the three spreadsheets, and my notebook. Takes about 40 minutes. If I miss an urgent ticket, the morning standup finds out at 6 AM."
  • The ops director (the person who decides): "My best week is a week where the 6 AM standup starts with the backlog already prioritized and nobody's surprised. I don't need trends. I need Tuesday to not start with a surprise."
  • The shadowed shift: the supervisor's notebook has urgency codes the spreadsheets don't — "will flood by morning" isn't a column anywhere, but it decides the order of the backlog.

Interviews tell you what people say the workflow is. Shadowing shows you what the workflow actually depends on. And the notebook leaves the engagement's first open data question on the table: the urgency codes live in a paper notebook — in no spreadsheet, no inbox export. The new workflow cannot reproduce the supervisor's prioritization unless urgency enters the data deliberately. So discovery ends with a decision, not a finding: the ops director chooses where urgency enters the new workflow before Week 3. Her answer: the supervisor tags urgency, in her own codes, when the site's daily CSV is produced — turning the notebook's tacit knowledge into an explicit, validated field at the source.

The real ask, written down and read back to the client for correction: the night-shift supervisor needs tomorrow's prioritized backlog before the 6 AM standup, built from Acme's tickets, with nothing urgent missed. Not a data lake. Not dashboards. Not trends. One workflow, one person, one deadline. The ops director corrects one word — "prioritized by our urgency codes, not yours" — and that correction is the most valuable sentence of the engagement, because it becomes a contract requirement in Week 3.

The trap: the requirements document. The fancy fix is a 20-page requirements doc — every field, every report, every "nice to have," signed in triplicate. It feels rigorous and it answers the wrong question: it catalogs the stated ask. The disciplined fix is three conversations and one corrected sentence. Discovery isn't documentation; it's finding the person whose week you're about to change and learning what "better" means to them.

The done-line: written with the client

Before any plan, any code, any data — the done-line, written with the ops director, carrying her signature. It has the same shape as the acceptance lines from the trial: specific, dated, and naming who decides.

The done-line (Week 1, signed by Acme's ops director): "By November 30, the night-shift supervisor uses an Acme-prioritized backlog at the 6 AM standup, every in-scope ticket is reconciled according to Acme's agreed policy, and we review the agreed outcome metrics against their recorded baselines. I decide whether the pilot continues."

Study what the done-line does. It names the person (the supervisor), the moment (the 6 AM standup), the evidence (reconciliation under the agreed policy; metrics reviewed against recorded baselines), and the owner of the judgment (the ops director). The "agreed outcome metrics" it references were sketched in Week 1's discovery and formalized — definitions versioned, baselines planned — in the Week 2 plan, before the pilot ran. It does not mention dashboards, trends, or row counts. If the engagement ended with 12,400 tickets loaded and the supervisor still building the backlog by hand at 11 PM, the done-line says it failed — and the done-line is right. Write the verdict before the work, with the person who delivers it.

Week 2: the 30-day plan

The plan is one page. It fits on one page because anything that doesn't fit is a second engagement. Every row names the deliverable, the owner, and — the load-bearing column — what is explicitly refused:

WeekDeliverableExplicitly refused
1Discovery notes + the signed done-linePlatform access before the plan is signed
2The 30-day plan (this page), signedMulti-site rollout; the mobile app they mentioned once
3First dataset live (pilot site): tickets → prioritized backlog, contracts enforcedHistorical analytics; backfill beyond 90 days; the "trends" dashboard
4Measured metrics against the versioned definitions; the stakeholder reviewNew features; a second workflow

The refused column is what protects November 30. Each refusal was a real sentence someone said — "while you're at it, can we get the mobile app," "we should really do all six sites at once" — and each one is answered in the plan with the reason: the pilot proves one workflow, then earns the right to the next. The ops director signs the plan including the refusals, which is what makes them stick when the requests resurface in Week 3. They will resurface. That's what the signature is for.

The signed plan also carries the success-metric definitions (v1) and the baseline collection plan — because acceptance is defined before the pilot runs, the way eval acceptance is defined before the system runs. The RAG-evals lesson's rule transfers directly: define the metric after seeing the outcome and you can always choose the flattering one. Week 4 measures against these definitions; it does not invent them.

The trap: the big-bang onboarding. The fancy fix loads everything, trains everyone, and goes live at all six sites on the same Monday. It maximizes the blast radius of every misunderstanding discovered in Week 3 — and Week 3 always discovers misunderstandings. The disciplined fix proves one workflow with one dataset, then earns the next scope. Small blast radius, fast learning, a signature that holds.

Week 3: the first dataset is a conversation, not a load

Week 3 opens with the pilot site's first data drop: 90 days of maintenance tickets, 12,400 rows, one CSV — one site, one workflow, the small blast radius the plan promised. The temptation is to load it. The discipline is to interrogate it — because the first dataset is where the client's mental model of their own data meets reality, and they are never the same thing.

Before any loading, the data-boundary decision — the PII thinking from the security lessons, applied to someone else's data. The tickets contain free-text notes, and free text contains names: technicians, tenants, "called Maria about the leak." The contract states the classification up front: raw free text stays inside Acme's approved sensitive-data boundary — the storage and processing environment the ops director's security contact approved for sensitive text, with access and retention controls, defined once in Week 2 and used consistently from here on. Only the minimized fields the workflow needs (ticket id, site, status, urgency, opened time) cross to the shared platform; anything uncertain is refused into the review queue — which is an approved sensitive-data store with access controls, or the data doesn't go there. Minimize first; detect what remains; block according to policy. The boundary is agreed in Week 2, before the data exists.

Worked example: contracts, then reconciliation

The contract is the client's vocabulary, written as code — the validation lesson's Pydantic contract, with the curriculum's rule applied twice: the client's vocabulary is validated, never guessed at — and the Week 2 urgency decision is where that rule gets teeth:

from pydantic import BaseModel, field_validator, ValidationInfo

KNOWN_STATUSES = {"OPEN", "IN_PROGRESS", "ON_HOLD", "CLOSED"}  # Acme's vocabulary,
    # confirmed with the ops director in discovery — "prioritized by OUR codes"
KNOWN_URGENCIES = {"routine", "urgent", "will_flood_by_morning"}  # the supervisor's
    # notebook codes ("will flood by morning", normalized) — the codes the old
    # workflow actually sorted by. Week 2 decision: the supervisor tags urgency
    # when the site's daily CSV is produced, so every row carries Acme's
    # classification at the source.

class MaintenanceTicket(BaseModel):
    ticket_id: str
    site: str
    status: str
    urgency: str  # no default: a blank urgency is an unknown, and unknowns
                  # quarantine — see the Week 2 decision below
    opened_at: str

    @field_validator("status", "urgency")
    @classmethod
    def known_vocabulary(cls, v: str, info: ValidationInfo):
        known = KNOWN_STATUSES if info.field_name == "status" else KNOWN_URGENCIES
        if v not in known:
            raise ValueError(f"unknown {info.field_name} {v!r}: quarantine, don't coerce")
        return v

The blank-urgency question nearly became a default. The ops director's first instinct was urgency: str = "routine" — her call, her business meaning, and on the surface the curriculum's rule ("defaults encode business meaning, so the client sets them") seemed to bless it. Dev asked the question the pilot promise demanded: "If urgency is blank, are you comfortable treating it as routine — or could a blank ever mean an urgent ticket nobody tagged?" The pilot's core promise is "nothing urgent missed." A blank defaulted to routine could hide exactly the ticket the customer fears most. She changed her answer: a blank urgency is an unknown, and unknowns quarantine — the supervisor classifies it before the row enters the workflow. Client approval makes a default legitimate; it doesn't make it safe. Safety depends on what the default hides if it's wrong. Ask that before encoding any default: what does this default hide if it's wrong?

The load runs. The reconciliation — the same explicit accounting the trial demanded. Note the last column: the 424 are unresolved until the client decides their fate. A quarantine count is not a resolution:

$ python onboard.py --drop acme_pilot_site_90d.csv
received: 12400 | accepted: 11976 | quarantined: 424 | resolved: 0 | unresolved: 424
$ psql -c "SELECT raw->>'status' AS s, COUNT(*) FROM quarantine GROUP BY 1;"
    s     | count
----------+-------
 CANCELLED |   301
 DUPLICATE |   123
-- 301 + 123 = 424. Every quarantined record is accounted for by name —
-- and every one of them is still unresolved, pending the client's policy.

This is the break — and it's a good break, because the contract caught it instead of the standup. The client's vocabulary has two statuses nobody mentioned in discovery. The fix is not to add them to KNOWN_STATUSES and re-run. The fix is a conversation, because the quarantine policy belongs to the data owner, not the lesson: the ops director decides what CANCELLED means for the backlog (excluded from the active backlog, but reported) and what DUPLICATE means (linked to its surviving ticket and deduplicated, with the link recorded). Her decisions go into the versioned mapping table with effective dates — the same pattern that absorbed the codebook change in the trial — and the re-ingest closes the books with the categories preserved: 424 quarantined → 424 dispositioned under the agreed policy → 0 unresolved. The final accounting never pretends all 12,400 became live backlog records: 11,976 accepted as active, 301 excluded-but-reported, 123 deduplicated-and-linked, 0 unresolved. Received, accepted, excluded, deduplicated, unresolved — every ticket in exactly one bucket.

One more contract, and it's the one the security lessons insisted on: the platform proposes the prioritized backlog; the supervisor approves it. These are two different controls. Authorization answers "may this caller act on this resource?" — the platform's service account, scoped to the pilot site, may generate the backlog. Human approval answers "did an authorized human approve this exact snapshot, bound to its content?" — the supervisor approves the specific backlog by its content-addressed snapshot ID (snapshot ID plus payload hash), so the approval binds to the data it was given, not merely to an ID that could later point somewhere else. The approval decision is written through an authenticated authorization path the proposing component cannot forge — the platform can request an approval, but it cannot write the approved state. The backlog becomes the night-shift plan only in the approved state; re-runs mint new snapshots with new content hashes, so an approval can never silently attach to different data. Idempotent executed states, stable unique IDs, and no single model decision with enough authority to turn one failure into a serious incident.

flowchart LR C[Client boundary
raw tickets, free text] --> MIN[Minimize
only the fields
the workflow needs] MIN --> VAL[Validate
the client's vocabulary] VAL -- unknown --> QU[Quarantine
client decides the policy] VAL -- clean --> GEN[Generate backlog
authorized service account] QU --> MAP[Map, version, re-ingest] MAP --> GEN GEN --> APP{Supervisor approves
this exact snapshot?} APP -- Yes --> PLAN[The night-shift plan] APP -- No --> REJ[Rejected — reason recorded
correct the cause,
mint a new snapshot] REJ --> GEN
The trap: "the data is dirty, clean it." The fancy fix coerces the unknowns — maps CANCELLED to CLOSED in the loader, guesses at DUPLICATE, invents a priority for the blank rows so validation passes. It makes the numbers green and the meaning wrong. The disciplined fix treats every unknown as a question for the data owner, because the mapping is data about the client's world, not logic. You don't clean a client's data; you reconcile it, with them, in the open.

Week 4: measure success, then review it

The metric definitions were written and co-signed in Week 2, before the pilot ran — acceptance defined before the outcome, the same rule the evals lesson drilled for RAG systems. Week 4 is where those definitions get their numbers, and where the evals lesson's durable rule gets its client-side form: version the metric definitions, the baseline, and the measuring method. A metric nobody wrote down is a feeling; a metric written down without its definition is an argument waiting to happen; a metric defined after the results are in is a choice of the flattering one. The one-page metric sheet, measured:

MetricDefinition (v1, 2026-11-09)BaselineMeasured how
Standup readinessA standup-ready backlog exists by 5:30 AMNot applicable under the previous workflow — no approval event existed; the supervisor built the backlog manually (~40 min/night, her estimate)Approval timestamp in the audit trail
Urgent-ticket coverageShare of the week's urgent tickets — per Acme's agreed urgency classification — present in the approved backlogUnknown — never measuredReconciliation query against the tagged set, weekly
Time-to-readyMinutes from the 11 PM ticket export to an approved backlog~40 (supervisor's estimate of the manual build)Export timestamp → approval timestamp
Supervisor effortMinutes the supervisor actively spends preparing the backlog each night~40 (supervisor's estimate)Supervisor-reported weekly — self-reported, labeled as such

Read the honesty in that table. "Unknown — never measured" is a complete and respectable baseline; inventing one would be the midnight-commit behavior the trial condemned. "~40 (supervisor's estimate)" is labeled as an estimate, not a fact. And the definitions are versioned — v1, dated — because the ops director will refine what "urgent" means after watching the backlog for two weeks, and when she does, the change goes through a review with her sign-off, not a quiet edit. The number follows the evidence, never the deadline — even when the deadline is a pilot renewal.

Then the adoption ladder, measured honestly at the review — the same ladder the trial used, because the question is the same: is the system being used, and does the use reach the outcome? One correction the trial taught: the denominator is the intended population — the 6 supervisors and ops leads who run the night-shift workflow at the pilot site — not the 34 technicians company-wide. Adoption measured against everyone with an account is a vanity denominator.

RungWhat it means hereWeek 4 (scenario constants)
ReachIntended workflow users with platform access6 of 6
ActivationOpened the backlog view once6
Meaningful useOpened the backlog during the standup window (5:30–6:30 AM, audit-logged)5 of 6
Repeat useDid so two weeks running4
Customer outcome — readinessStandup started with the approved backlog9 of 10 standups
Customer outcome — qualityUrgent tickets missed; surprise escalations0 missed; 1 surprise, documented with cause

Two honesties in that table. First, the audit log proves a view during the window — not that the backlog drove the meeting; the supervisor confirms meeting use at the review, and the table would say "client-reported" if she couldn't. Don't claim the instrumentation proves more than it observes. Second, readiness and quality are separate rungs: "the standup started with the approved backlog" is auditable; "no surprises" is a quality claim that needs its own evidence — the missed-urgent count and the surprise log.

The ladder is reported with the same neutrality as the 429s: 4 of 6 in repeat use is not a failure to hide, it's the next friction to investigate — and the re-measurement date is set before the review ends. And the honest caveat is stated up front: "9 of 10 standups" can't isolate the platform from everything else that changed in November. Measure anyway, and stay honest about what the measurement can't prove — that's the difference between evidence and theater.

The trap: the vanity dashboard. The fancy fix reports logins and page views — numbers that go up when you send the "check out the new dashboard!" email and mean nothing about the standup. The disciplined fix measures the rung that counts: the standup starting with the approved backlog. Production green is not success. Adoption is the next test; customer outcome is the destination — and the customer is the one holding the measuring stick.

The first stakeholder review

Week 4, Friday. The ops director, the supervisor, and you — the whole engagement compressed into the format a customer can actually use. Study the grammar: it's the incident briefing from the trial, because an onboarding review and an incident briefing are the same skill.

You to the ops director: "Three things. One — what happened: we interviewed your team, found the 11 PM backlog build, and scoped the pilot to one site and the 6 AM standup. The first data drop surfaced two statuses nobody had mentioned — CANCELLED and DUPLICATE — and you set the policy for both before we loaded a single row into the workflow. Two — what's true now: 12,400 tickets accounted for under your disposition policy — 11,976 active, 301 excluded-but-reported, 123 deduplicated-and-linked, zero unresolved; the backlog is generated under your urgency codes and approved by your supervisor every night; 9 of the last 10 standups started with it, with zero urgent tickets missed. Three — what changed permanently: the validation contract carries your vocabulary, the quarantine policy is written down with your name on it, and the metric definitions are versioned — so the next site onboards on rails, not on memory."

The ops director: "And the mobile app?"

You: "Still refused — it's in the plan, with your signature. We earn it with the second site."

She renews the pilot. Not because the platform is impressive — the evidence against her done-line supports the renewal decision, in her words. The briefing is a deliverable: she forwards its shape to her leadership, the way Maria forwarded the postmortem. The customer sells the renewal internally; your job was to give her the evidence, in her vocabulary, with nothing softened.

The scorecard, from the client's side

The curriculum's milestone scorecard, applied to the engagement. Each dimension names what "good" looked like this month — and the trap it punishes:

DimensionWhat "good" looked likeThe trap it punishes
Technical correctnessContracts preserve Acme's status and urgency semantics; unknowns are quarantined rather than guessedCoercing unknowns so the load goes green
Data reconciliation12,400 received → 11,976 accepted + 424 quarantined → 424 dispositioned under the client's policy → 0 unresolved"Mostly loaded"; uncounted losses
ReliabilityQuarantine policy written down, review SLA owned, backlog approved every nightThe load that "usually works"
Security & complianceData classified before loading; raw free text stays in Acme's approved sensitive-data boundary; approval behind its own boundary"We'll sort out access later"
Time-to-valueFirst standup on the backlog in Week 3 — not a six-month migrationThe big-bang onboarding
Customer outcome9 of 10 standups start with the approved backlog; 0 urgent tickets missed; the ops director states the metrics herselfA loaded data lake nobody opens
Scope disciplineThe refused column in the 30-day plan, signed — and held in Week 4The second engagement inside the first
CommunicationDiscovery notes read back for correction; briefing grammar at the reviewThe requirements document nobody reads
OperabilityThe data owner can run the reconciliation query; the runbook names herVendor-operated magic
AdoptionLadder measured honestly; next friction named; re-measurement datedLogins as "engagement"

Read the right column as a whole: every trap is the same mistake wearing different clothes — optimizing for impressiveness over outcomes. The companion finale graded the same scorecard under fire; this one grades it across the table. The central lesson, stated one final time: the best onboarding isn't the most complete migration — it's the smallest reliable workflow that creates the required customer outcome and survives the client's reality.

Acme, experienced

Step back and name what the engagement proved. Not that the platform works — the milestones proved that. That it lands: in a client's vocabulary, inside their boundary, on their clock, measured by their judgment. Each week converted a lesson into a relationship property — the corrected sentence, the signed done-line, the refused column, the quarantine policy with a name on it, the versioned metrics.

And the arc of the full curriculum lands here, from the other side. Stage 1 taught you to scope the vague ask — here the asker is a client. Stage 2 taught you to move data without losing it — here the data is someone else's. Stage 3 taught you to integrate without trusting — here the trust is negotiated, not assumed. Stage 4 taught you to put contracts around intelligence — here the contracts carry the client's vocabulary. Stage 5 taught you to operate what you shipped — here the operator is the client's data owner. Stage 6 taught you to present the evidence — here the audience decides the renewal. The companion finale was the exam for judgment under fire. This one is the exam for judgment across the table: what to ask, what to refuse, what to measure, and what to tell the customer on Friday at 4 PM — when the customer is the one who wrote the done-line.

Five lines to leave with. Onboard the workflow, not the data. The first dataset is a conversation, not a load. Write the verdict before the work, with the person who delivers it. Every unknown is a question for the data owner, not a guess for the loader. And the last one, unchanged: production green is not success — adoption is the next test, customer outcome is the destination.

The chapter principle

Onboard the workflow, not the data. The hierarchy: the data loads → the workflow runs → the people use it → the customer outcome improves. A fully migrated data lake with an unchanged 11 PM spreadsheet routine is a failed onboarding wearing a success metric. Build for the last rung.

Field check

  1. Acme's stated ask was "see everything in one place — dashboards, trends, the works." The engagement onboarded one workflow instead. Why was the stated ask the wrong target, and what did the discovery actually find?
  2. The first data drop contained two statuses — CANCELLED and DUPLICATE — that discovery never mentioned. Why was "quarantine, don't coerce" the right default, and who decides the quarantine policy?
  3. The ops director proposes defaulting missing urgency to "routine". Why isn't client approval alone enough to make that safe, and what question should you ask before encoding it?
  4. The platform generates the backlog (authorization), and the supervisor approves the exact snapshot (human approval). Why are these two different controls, and where must the approval record live?
  5. The metric sheet versions its definitions and labels one baseline "unknown — never measured." Why version the definitions, and why is an honest "unknown" better than an estimated baseline?
  6. The stakeholder review uses the same grammar as an incident briefing: what happened, what's true now, what changed permanently. Why does an onboarding review need incident-briefing grammar?
Answers

1. The stated ask described a dashboard, not a problem — "see everything" has no owner, no moment, and no verdict. The discovery found the workflow that hurt: the supervisor building tomorrow's backlog by hand at 11 PM, with the 6 AM standup as the deadline that mattered. Onboarding the workflow gave the engagement a person, a moment, and a done-line; onboarding the data would have produced row counts and an unchanged 11 PM routine.

2. Coercing unknowns makes the load green and the meaning wrong — CANCELLED mapped to CLOSED by a loader's guess would silently drop tickets the client still tracks. Quarantine preserves the question until the person who owns the meaning answers it. The quarantine policy belongs to the data owner (the ops director), not the engineer: she decided CANCELLED is excluded-but-reported and DUPLICATE is deduplicated with the link recorded — her business meaning, her call, recorded in the versioned mapping table.

3. Because the pilot's core promise is that no urgent ticket is missed — and a blank field defaulted to routine can hide exactly the ticket the customer fears most. Client approval makes a default legitimate, not safe: safety depends on what the default hides if it's wrong. Ask: "If urgency is blank, could it ever mean urgent — or unknown?" If yes, quarantine the row or require classification instead of defaulting. A default that can silently violate the business objective is invented risk wearing a signature — and the right answer here was quarantine, not routine.

4. Authorization answers "may this caller act on this resource?" — the platform's service account, scoped to the pilot site, may generate the backlog. Approval answers "did an authorized human approve this exact snapshot, bound to its content?" — the supervisor approves the specific backlog by its content-addressed snapshot ID (ID plus payload hash), so the approval binds to the data it was given. They are different because a system can be authorized to propose and still must not be authorized to decide. The approval decision must be written through an authenticated authorization path the proposing component cannot forge — otherwise the proposer could write its own approval. Re-runs mint new snapshots with new content hashes, so an approval can never silently attach to different data.

5. Definitions are versioned because the client's understanding of "urgent" will evolve after two weeks of watching the backlog — and when it does, the change needs a review and a sign-off, not a quiet edit; otherwise nobody can tell whether the metric or the world moved. An honest "unknown" is better than an estimated baseline because an estimate invented to fill the cell becomes the anchor every future number is judged against. "Never measured" is a complete baseline: it says exactly what we know, which is nothing yet — and the measuring method is written down so the next pull is real.

6. Because both are the same skill: consequential news, delivered early, with a plan attached, to someone who needs to act. The incident briefing tells a stakeholder what broke and what's true now; the onboarding review tells the client what was learned and what's true now — including the bad news (two missed statuses, 4 of 6 in repeat use) with the plan attached. Softening either one destroys the trust both depend on. The grammar forces the discipline: no item ends at "we fixed it" — each ends at the mechanism that makes it cheaper next time.