The companion finale attacked the product — six incidents, one trial, judgment under pressure. This finale builds the relationship: Acme, a facilities-services company, is onboarding onto the platform. Discovery interviews, the real ask behind the ask, a 30-day plan, one dataset with contracts, success metrics defined together, and a stakeholder review that proves the outcome. The one principle for the whole chapter: onboard the workflow, not the data.
Dev (Monday, 9:00 AM): "New engagement. Acme — facilities services, six sites, maintenance tickets in three spreadsheets and a shared inbox. They've signed a pilot. Your job: onboard them. Not 'load their data.' Onboard them. Four weeks from today, their ops director decides whether this was worth it."
Maria: "And the ops director is the one who decides what 'worth it' means. Not us. Her signature is on the done-line before we touch a single ticket."
How to read this chapter
This chapter is a simulation. Acme is fictional; every number in it — 12,400 tickets, 424 quarantined records, 34 technicians, 30 days — is a scenario constant, not a measured fact. The client is teaching fiction; the machinery is real. Every artifact below runs against the actual platform components from the milestones: the validation contracts from the data-validation lesson, the API authorization from Milestone 3, the briefing grammar from CityOps Under Fire. The code is written to be read and adapted, not copy-pasted into a client engagement.
The shape of the engagement
Four weeks, four artifacts. Each week ends with something the client can hold — and each artifact has an owner on the client side, because an onboarding owned only by the vendor is a demo with a longer timeline:
| Week | The work | The artifact | Owned by |
|---|---|---|---|
| 1 | Discovery interviews | The real ask, written down | Acme's ops director |
| 2 | Scope the first engagement | The 30-day plan, signed | Acme's ops director |
| 3 | Onboard the first dataset | Contracts, reconciliation, the live workflow | Acme's data owner |
| 4 | Measure success, then review it | Measured metrics + the stakeholder review | Acme's ops director |
Notice what's missing: there is no "migrate all the data" week and no "train everyone" week. The engagement is scoped to one dataset, one workflow, one review. Everything else is explicitly refused in the plan — with the refusal written down, the way the borough-pack tradeoff memo wrote down its refusals. Scope discipline isn't saying no to the client; it's saying no to the second, third, and fourth engagements hiding inside the first.
The minimum concept: onboard the workflow, not the data
Every failed onboarding Dev has seen failed the same way: the team treated the engagement as a data-migration project — extract, load, transform, celebrate row counts. The client's actual problem was never "our data isn't in your system." It was a workflow that hurt: a supervisor building tomorrow's backlog by hand at 11 PM, a standup that starts with guesses, a spreadsheet nobody trusts. The data is the fuel; the workflow is the engine. Onboard the engine.
The protocol, in order — it mirrors the triage protocol's shape (scope, then evidence, then the smallest safe change), but its instruments are people:
“load our data”] --> DIS[1. Discovery
who decides,
what a good week looks like] DIS --> REAL[The real ask
the workflow that hurts] REAL --> DONE[2. Done-line
written with the client,
owned by the client] DONE --> PLAN[3. 30-day plan
one dataset, one workflow,
one review] PLAN --> DATA[4. First dataset
validate, reconcile, authorize] DATA --> BRK{Reconciliation clean?} BRK -- No --> Q[Quarantine — client decides
the policy; map, version,
re-ingest] BRK -- Yes --> MET Q --> MET[5. Metrics
defined together, versioned,
baselined] MET --> REV[6. Stakeholder review
what happened, what's true now,
what changed permanently] REV --> OUT[Customer outcome]
Three things about this order. First, discovery before the done-line: you cannot write what "done" looks like until you've watched the workflow that hurts. Second, the client owns the done-line: "worth it" is their judgment, and the artifact that records it carries their signature — because the person who decides whether the pilot continues is the person who defined success. Third, the review is a deliverable, not a ceremony: it runs on the same grammar as an incident briefing — what happened, what's true now, what changed permanently — because an onboarding review and an incident briefing are the same skill: bad news delivered early, with a plan attached, to someone who needs to act.
The one question that starts every engagement
Before the first interview, ask: "Who decides whether this was worth it, and what does their best week look like?" The answer names the owner of the done-line and the workflow you're actually onboarding. Everything after it is execution.
Week 1: the discovery interviews
Acme's ops director (kickoff call): "Here's what we need: get our maintenance tickets into your platform. Three spreadsheets, one shared inbox, six sites. We want to see everything in one place — dashboards, trends, the works."
Maria (after the call, to you): "Write down exactly what she said. Now forget the nouns and keep the verbs. What did she actually ask for?"
This is the Stage 1 skill — clarify the ask — applied to a client who is paying for the answer. "See everything in one place" is a stated ask; it describes a dashboard, not a problem. The discovery interviews exist to find the workflow that hurts. Two 30-minute interviews and one shadowed shift:
- The night-shift supervisor (the person closest to the work): "Every night around 11, I build tomorrow's backlog. I pull the inbox, the three spreadsheets, and my notebook. Takes about 40 minutes. If I miss an urgent ticket, the morning standup finds out at 6 AM."
- The ops director (the person who decides): "My best week is a week where the 6 AM standup starts with the backlog already prioritized and nobody's surprised. I don't need trends. I need Tuesday to not start with a surprise."
- The shadowed shift: the supervisor's notebook has urgency codes the spreadsheets don't — "will flood by morning" isn't a column anywhere, but it decides the order of the backlog.
Interviews tell you what people say the workflow is. Shadowing shows you what the workflow actually depends on. And the notebook leaves the engagement's first open data question on the table: the urgency codes live in a paper notebook — in no spreadsheet, no inbox export. The new workflow cannot reproduce the supervisor's prioritization unless urgency enters the data deliberately. So discovery ends with a decision, not a finding: the ops director chooses where urgency enters the new workflow before Week 3. Her answer: the supervisor tags urgency, in her own codes, when the site's daily CSV is produced — turning the notebook's tacit knowledge into an explicit, validated field at the source.
The real ask, written down and read back to the client for correction: the night-shift supervisor needs tomorrow's prioritized backlog before the 6 AM standup, built from Acme's tickets, with nothing urgent missed. Not a data lake. Not dashboards. Not trends. One workflow, one person, one deadline. The ops director corrects one word — "prioritized by our urgency codes, not yours" — and that correction is the most valuable sentence of the engagement, because it becomes a contract requirement in Week 3.
The done-line: written with the client
Before any plan, any code, any data — the done-line, written with the ops director, carrying her signature. It has the same shape as the acceptance lines from the trial: specific, dated, and naming who decides.
The done-line (Week 1, signed by Acme's ops director): "By November 30, the night-shift supervisor uses an Acme-prioritized backlog at the 6 AM standup, every in-scope ticket is reconciled according to Acme's agreed policy, and we review the agreed outcome metrics against their recorded baselines. I decide whether the pilot continues."
Study what the done-line does. It names the person (the supervisor), the moment (the 6 AM standup), the evidence (reconciliation under the agreed policy; metrics reviewed against recorded baselines), and the owner of the judgment (the ops director). The "agreed outcome metrics" it references were sketched in Week 1's discovery and formalized — definitions versioned, baselines planned — in the Week 2 plan, before the pilot ran. It does not mention dashboards, trends, or row counts. If the engagement ended with 12,400 tickets loaded and the supervisor still building the backlog by hand at 11 PM, the done-line says it failed — and the done-line is right. Write the verdict before the work, with the person who delivers it.
Week 2: the 30-day plan
The plan is one page. It fits on one page because anything that doesn't fit is a second engagement. Every row names the deliverable, the owner, and — the load-bearing column — what is explicitly refused:
| Week | Deliverable | Explicitly refused |
|---|---|---|
| 1 | Discovery notes + the signed done-line | Platform access before the plan is signed |
| 2 | The 30-day plan (this page), signed | Multi-site rollout; the mobile app they mentioned once |
| 3 | First dataset live (pilot site): tickets → prioritized backlog, contracts enforced | Historical analytics; backfill beyond 90 days; the "trends" dashboard |
| 4 | Measured metrics against the versioned definitions; the stakeholder review | New features; a second workflow |
The refused column is what protects November 30. Each refusal was a real sentence someone said — "while you're at it, can we get the mobile app," "we should really do all six sites at once" — and each one is answered in the plan with the reason: the pilot proves one workflow, then earns the right to the next. The ops director signs the plan including the refusals, which is what makes them stick when the requests resurface in Week 3. They will resurface. That's what the signature is for.
The signed plan also carries the success-metric definitions (v1) and the baseline collection plan — because acceptance is defined before the pilot runs, the way eval acceptance is defined before the system runs. The RAG-evals lesson's rule transfers directly: define the metric after seeing the outcome and you can always choose the flattering one. Week 4 measures against these definitions; it does not invent them.
Week 3: the first dataset is a conversation, not a load
Week 3 opens with the pilot site's first data drop: 90 days of maintenance tickets, 12,400 rows, one CSV — one site, one workflow, the small blast radius the plan promised. The temptation is to load it. The discipline is to interrogate it — because the first dataset is where the client's mental model of their own data meets reality, and they are never the same thing.
Before any loading, the data-boundary decision — the PII thinking from the security lessons, applied to someone else's data. The tickets contain free-text notes, and free text contains names: technicians, tenants, "called Maria about the leak." The contract states the classification up front: raw free text stays inside Acme's approved sensitive-data boundary — the storage and processing environment the ops director's security contact approved for sensitive text, with access and retention controls, defined once in Week 2 and used consistently from here on. Only the minimized fields the workflow needs (ticket id, site, status, urgency, opened time) cross to the shared platform; anything uncertain is refused into the review queue — which is an approved sensitive-data store with access controls, or the data doesn't go there. Minimize first; detect what remains; block according to policy. The boundary is agreed in Week 2, before the data exists.
Worked example: contracts, then reconciliation
The contract is the client's vocabulary, written as code — the validation lesson's Pydantic contract, with the curriculum's rule applied twice: the client's vocabulary is validated, never guessed at — and the Week 2 urgency decision is where that rule gets teeth:
from pydantic import BaseModel, field_validator, ValidationInfo
KNOWN_STATUSES = {"OPEN", "IN_PROGRESS", "ON_HOLD", "CLOSED"} # Acme's vocabulary,
# confirmed with the ops director in discovery — "prioritized by OUR codes"
KNOWN_URGENCIES = {"routine", "urgent", "will_flood_by_morning"} # the supervisor's
# notebook codes ("will flood by morning", normalized) — the codes the old
# workflow actually sorted by. Week 2 decision: the supervisor tags urgency
# when the site's daily CSV is produced, so every row carries Acme's
# classification at the source.
class MaintenanceTicket(BaseModel):
ticket_id: str
site: str
status: str
urgency: str # no default: a blank urgency is an unknown, and unknowns
# quarantine — see the Week 2 decision below
opened_at: str
@field_validator("status", "urgency")
@classmethod
def known_vocabulary(cls, v: str, info: ValidationInfo):
known = KNOWN_STATUSES if info.field_name == "status" else KNOWN_URGENCIES
if v not in known:
raise ValueError(f"unknown {info.field_name} {v!r}: quarantine, don't coerce")
return v
The blank-urgency question nearly became a default. The ops director's first instinct was urgency: str = "routine" — her call, her business meaning, and on the surface the curriculum's rule ("defaults encode business meaning, so the client sets them") seemed to bless it. Dev asked the question the pilot promise demanded: "If urgency is blank, are you comfortable treating it as routine — or could a blank ever mean an urgent ticket nobody tagged?" The pilot's core promise is "nothing urgent missed." A blank defaulted to routine could hide exactly the ticket the customer fears most. She changed her answer: a blank urgency is an unknown, and unknowns quarantine — the supervisor classifies it before the row enters the workflow. Client approval makes a default legitimate; it doesn't make it safe. Safety depends on what the default hides if it's wrong. Ask that before encoding any default: what does this default hide if it's wrong?
The load runs. The reconciliation — the same explicit accounting the trial demanded. Note the last column: the 424 are unresolved until the client decides their fate. A quarantine count is not a resolution:
$ python onboard.py --drop acme_pilot_site_90d.csv
received: 12400 | accepted: 11976 | quarantined: 424 | resolved: 0 | unresolved: 424
$ psql -c "SELECT raw->>'status' AS s, COUNT(*) FROM quarantine GROUP BY 1;"
s | count
----------+-------
CANCELLED | 301
DUPLICATE | 123
-- 301 + 123 = 424. Every quarantined record is accounted for by name —
-- and every one of them is still unresolved, pending the client's policy.
This is the break — and it's a good break, because the contract caught it instead of the standup. The client's vocabulary has two statuses nobody mentioned in discovery. The fix is not to add them to KNOWN_STATUSES and re-run. The fix is a conversation, because the quarantine policy belongs to the data owner, not the lesson: the ops director decides what CANCELLED means for the backlog (excluded from the active backlog, but reported) and what DUPLICATE means (linked to its surviving ticket and deduplicated, with the link recorded). Her decisions go into the versioned mapping table with effective dates — the same pattern that absorbed the codebook change in the trial — and the re-ingest closes the books with the categories preserved: 424 quarantined → 424 dispositioned under the agreed policy → 0 unresolved. The final accounting never pretends all 12,400 became live backlog records: 11,976 accepted as active, 301 excluded-but-reported, 123 deduplicated-and-linked, 0 unresolved. Received, accepted, excluded, deduplicated, unresolved — every ticket in exactly one bucket.
One more contract, and it's the one the security lessons insisted on: the platform proposes the prioritized backlog; the supervisor approves it. These are two different controls. Authorization answers "may this caller act on this resource?" — the platform's service account, scoped to the pilot site, may generate the backlog. Human approval answers "did an authorized human approve this exact snapshot, bound to its content?" — the supervisor approves the specific backlog by its content-addressed snapshot ID (snapshot ID plus payload hash), so the approval binds to the data it was given, not merely to an ID that could later point somewhere else. The approval decision is written through an authenticated authorization path the proposing component cannot forge — the platform can request an approval, but it cannot write the approved state. The backlog becomes the night-shift plan only in the approved state; re-runs mint new snapshots with new content hashes, so an approval can never silently attach to different data. Idempotent executed states, stable unique IDs, and no single model decision with enough authority to turn one failure into a serious incident.
raw tickets, free text] --> MIN[Minimize
only the fields
the workflow needs] MIN --> VAL[Validate
the client's vocabulary] VAL -- unknown --> QU[Quarantine
client decides the policy] VAL -- clean --> GEN[Generate backlog
authorized service account] QU --> MAP[Map, version, re-ingest] MAP --> GEN GEN --> APP{Supervisor approves
this exact snapshot?} APP -- Yes --> PLAN[The night-shift plan] APP -- No --> REJ[Rejected — reason recorded
correct the cause,
mint a new snapshot] REJ --> GEN
Week 4: measure success, then review it
The metric definitions were written and co-signed in Week 2, before the pilot ran — acceptance defined before the outcome, the same rule the evals lesson drilled for RAG systems. Week 4 is where those definitions get their numbers, and where the evals lesson's durable rule gets its client-side form: version the metric definitions, the baseline, and the measuring method. A metric nobody wrote down is a feeling; a metric written down without its definition is an argument waiting to happen; a metric defined after the results are in is a choice of the flattering one. The one-page metric sheet, measured:
| Metric | Definition (v1, 2026-11-09) | Baseline | Measured how |
|---|---|---|---|
| Standup readiness | A standup-ready backlog exists by 5:30 AM | Not applicable under the previous workflow — no approval event existed; the supervisor built the backlog manually (~40 min/night, her estimate) | Approval timestamp in the audit trail |
| Urgent-ticket coverage | Share of the week's urgent tickets — per Acme's agreed urgency classification — present in the approved backlog | Unknown — never measured | Reconciliation query against the tagged set, weekly |
| Time-to-ready | Minutes from the 11 PM ticket export to an approved backlog | ~40 (supervisor's estimate of the manual build) | Export timestamp → approval timestamp |
| Supervisor effort | Minutes the supervisor actively spends preparing the backlog each night | ~40 (supervisor's estimate) | Supervisor-reported weekly — self-reported, labeled as such |
Read the honesty in that table. "Unknown — never measured" is a complete and respectable baseline; inventing one would be the midnight-commit behavior the trial condemned. "~40 (supervisor's estimate)" is labeled as an estimate, not a fact. And the definitions are versioned — v1, dated — because the ops director will refine what "urgent" means after watching the backlog for two weeks, and when she does, the change goes through a review with her sign-off, not a quiet edit. The number follows the evidence, never the deadline — even when the deadline is a pilot renewal.
Then the adoption ladder, measured honestly at the review — the same ladder the trial used, because the question is the same: is the system being used, and does the use reach the outcome? One correction the trial taught: the denominator is the intended population — the 6 supervisors and ops leads who run the night-shift workflow at the pilot site — not the 34 technicians company-wide. Adoption measured against everyone with an account is a vanity denominator.
| Rung | What it means here | Week 4 (scenario constants) |
|---|---|---|
| Reach | Intended workflow users with platform access | 6 of 6 |
| Activation | Opened the backlog view once | 6 |
| Meaningful use | Opened the backlog during the standup window (5:30–6:30 AM, audit-logged) | 5 of 6 |
| Repeat use | Did so two weeks running | 4 |
| Customer outcome — readiness | Standup started with the approved backlog | 9 of 10 standups |
| Customer outcome — quality | Urgent tickets missed; surprise escalations | 0 missed; 1 surprise, documented with cause |
Two honesties in that table. First, the audit log proves a view during the window — not that the backlog drove the meeting; the supervisor confirms meeting use at the review, and the table would say "client-reported" if she couldn't. Don't claim the instrumentation proves more than it observes. Second, readiness and quality are separate rungs: "the standup started with the approved backlog" is auditable; "no surprises" is a quality claim that needs its own evidence — the missed-urgent count and the surprise log.
The ladder is reported with the same neutrality as the 429s: 4 of 6 in repeat use is not a failure to hide, it's the next friction to investigate — and the re-measurement date is set before the review ends. And the honest caveat is stated up front: "9 of 10 standups" can't isolate the platform from everything else that changed in November. Measure anyway, and stay honest about what the measurement can't prove — that's the difference between evidence and theater.
The first stakeholder review
Week 4, Friday. The ops director, the supervisor, and you — the whole engagement compressed into the format a customer can actually use. Study the grammar: it's the incident briefing from the trial, because an onboarding review and an incident briefing are the same skill.
You to the ops director: "Three things. One — what happened: we interviewed your team, found the 11 PM backlog build, and scoped the pilot to one site and the 6 AM standup. The first data drop surfaced two statuses nobody had mentioned — CANCELLED and DUPLICATE — and you set the policy for both before we loaded a single row into the workflow. Two — what's true now: 12,400 tickets accounted for under your disposition policy — 11,976 active, 301 excluded-but-reported, 123 deduplicated-and-linked, zero unresolved; the backlog is generated under your urgency codes and approved by your supervisor every night; 9 of the last 10 standups started with it, with zero urgent tickets missed. Three — what changed permanently: the validation contract carries your vocabulary, the quarantine policy is written down with your name on it, and the metric definitions are versioned — so the next site onboards on rails, not on memory."
The ops director: "And the mobile app?"
You: "Still refused — it's in the plan, with your signature. We earn it with the second site."
She renews the pilot. Not because the platform is impressive — the evidence against her done-line supports the renewal decision, in her words. The briefing is a deliverable: she forwards its shape to her leadership, the way Maria forwarded the postmortem. The customer sells the renewal internally; your job was to give her the evidence, in her vocabulary, with nothing softened.
The scorecard, from the client's side
The curriculum's milestone scorecard, applied to the engagement. Each dimension names what "good" looked like this month — and the trap it punishes:
| Dimension | What "good" looked like | The trap it punishes |
|---|---|---|
| Technical correctness | Contracts preserve Acme's status and urgency semantics; unknowns are quarantined rather than guessed | Coercing unknowns so the load goes green |
| Data reconciliation | 12,400 received → 11,976 accepted + 424 quarantined → 424 dispositioned under the client's policy → 0 unresolved | "Mostly loaded"; uncounted losses |
| Reliability | Quarantine policy written down, review SLA owned, backlog approved every night | The load that "usually works" |
| Security & compliance | Data classified before loading; raw free text stays in Acme's approved sensitive-data boundary; approval behind its own boundary | "We'll sort out access later" |
| Time-to-value | First standup on the backlog in Week 3 — not a six-month migration | The big-bang onboarding |
| Customer outcome | 9 of 10 standups start with the approved backlog; 0 urgent tickets missed; the ops director states the metrics herself | A loaded data lake nobody opens |
| Scope discipline | The refused column in the 30-day plan, signed — and held in Week 4 | The second engagement inside the first |
| Communication | Discovery notes read back for correction; briefing grammar at the review | The requirements document nobody reads |
| Operability | The data owner can run the reconciliation query; the runbook names her | Vendor-operated magic |
| Adoption | Ladder measured honestly; next friction named; re-measurement dated | Logins as "engagement" |
Read the right column as a whole: every trap is the same mistake wearing different clothes — optimizing for impressiveness over outcomes. The companion finale graded the same scorecard under fire; this one grades it across the table. The central lesson, stated one final time: the best onboarding isn't the most complete migration — it's the smallest reliable workflow that creates the required customer outcome and survives the client's reality.
Acme, experienced
Step back and name what the engagement proved. Not that the platform works — the milestones proved that. That it lands: in a client's vocabulary, inside their boundary, on their clock, measured by their judgment. Each week converted a lesson into a relationship property — the corrected sentence, the signed done-line, the refused column, the quarantine policy with a name on it, the versioned metrics.
And the arc of the full curriculum lands here, from the other side. Stage 1 taught you to scope the vague ask — here the asker is a client. Stage 2 taught you to move data without losing it — here the data is someone else's. Stage 3 taught you to integrate without trusting — here the trust is negotiated, not assumed. Stage 4 taught you to put contracts around intelligence — here the contracts carry the client's vocabulary. Stage 5 taught you to operate what you shipped — here the operator is the client's data owner. Stage 6 taught you to present the evidence — here the audience decides the renewal. The companion finale was the exam for judgment under fire. This one is the exam for judgment across the table: what to ask, what to refuse, what to measure, and what to tell the customer on Friday at 4 PM — when the customer is the one who wrote the done-line.
Five lines to leave with. Onboard the workflow, not the data. The first dataset is a conversation, not a load. Write the verdict before the work, with the person who delivers it. Every unknown is a question for the data owner, not a guess for the loader. And the last one, unchanged: production green is not success — adoption is the next test, customer outcome is the destination.
The chapter principle
Onboard the workflow, not the data. The hierarchy: the data loads → the workflow runs → the people use it → the customer outcome improves. A fully migrated data lake with an unchanged 11 PM spreadsheet routine is a failed onboarding wearing a success metric. Build for the last rung.
Field check
- Acme's stated ask was "see everything in one place — dashboards, trends, the works." The engagement onboarded one workflow instead. Why was the stated ask the wrong target, and what did the discovery actually find?
- The first data drop contained two statuses — CANCELLED and DUPLICATE — that discovery never mentioned. Why was "quarantine, don't coerce" the right default, and who decides the quarantine policy?
- The ops director proposes defaulting missing urgency to
"routine". Why isn't client approval alone enough to make that safe, and what question should you ask before encoding it? - The platform generates the backlog (authorization), and the supervisor approves the exact snapshot (human approval). Why are these two different controls, and where must the approval record live?
- The metric sheet versions its definitions and labels one baseline "unknown — never measured." Why version the definitions, and why is an honest "unknown" better than an estimated baseline?
- The stakeholder review uses the same grammar as an incident briefing: what happened, what's true now, what changed permanently. Why does an onboarding review need incident-briefing grammar?
Answers
1. The stated ask described a dashboard, not a problem — "see everything" has no owner, no moment, and no verdict. The discovery found the workflow that hurt: the supervisor building tomorrow's backlog by hand at 11 PM, with the 6 AM standup as the deadline that mattered. Onboarding the workflow gave the engagement a person, a moment, and a done-line; onboarding the data would have produced row counts and an unchanged 11 PM routine.
2. Coercing unknowns makes the load green and the meaning wrong — CANCELLED mapped to CLOSED by a loader's guess would silently drop tickets the client still tracks. Quarantine preserves the question until the person who owns the meaning answers it. The quarantine policy belongs to the data owner (the ops director), not the engineer: she decided CANCELLED is excluded-but-reported and DUPLICATE is deduplicated with the link recorded — her business meaning, her call, recorded in the versioned mapping table.
3. Because the pilot's core promise is that no urgent ticket is missed — and a blank field defaulted to routine can hide exactly the ticket the customer fears most. Client approval makes a default legitimate, not safe: safety depends on what the default hides if it's wrong. Ask: "If urgency is blank, could it ever mean urgent — or unknown?" If yes, quarantine the row or require classification instead of defaulting. A default that can silently violate the business objective is invented risk wearing a signature — and the right answer here was quarantine, not routine.
4. Authorization answers "may this caller act on this resource?" — the platform's service account, scoped to the pilot site, may generate the backlog. Approval answers "did an authorized human approve this exact snapshot, bound to its content?" — the supervisor approves the specific backlog by its content-addressed snapshot ID (ID plus payload hash), so the approval binds to the data it was given. They are different because a system can be authorized to propose and still must not be authorized to decide. The approval decision must be written through an authenticated authorization path the proposing component cannot forge — otherwise the proposer could write its own approval. Re-runs mint new snapshots with new content hashes, so an approval can never silently attach to different data.
5. Definitions are versioned because the client's understanding of "urgent" will evolve after two weeks of watching the backlog — and when it does, the change needs a review and a sign-off, not a quiet edit; otherwise nobody can tell whether the metric or the world moved. An honest "unknown" is better than an estimated baseline because an estimate invented to fill the cell becomes the anchor every future number is judged against. "Never measured" is a complete baseline: it says exactly what we know, which is nothing yet — and the measuring method is written down so the next pull is real.
6. Because both are the same skill: consequential news, delivered early, with a plan attached, to someone who needs to act. The incident briefing tells a stakeholder what broke and what's true now; the onboarding review tells the client what was learned and what's true now — including the bad news (two missed statuses, 4 of 6 in repeat use) with the plan attached. Softening either one destroys the trust both depend on. The grammar forces the discipline: no item ends at "we fixed it" — each ends at the mechanism that makes it cheaper next time.