# Worked Example — one team, every template filled in

*Everything else in this toolkit is brackets waiting for your details.
This is the same set of documents with the brackets filled in, for one
team, start to finish — so you can see what a real answer looks like
before you write your own.*

**This team is invented.** Northgate Claims Operations is a composite,
not a real company, and the numbers in it are illustrative rather than
measured. Don't cite them. Do copy the shape.

Read it whichever way suits you: straight through as a story, or jump
to the one artifact you're about to write yourself.

| You're about to write | Jump to |
| --- | --- |
| The one-page policy | [2. The policy they published](#2-the-policy-they-published) |
| A use-case shortlist | [3. Scoring the candidates](#3-scoring-the-candidates) |
| Guidance for a borderline call | [4. Two borderline calls](#4-two-borderline-calls) |
| A pilot plan | [5. The pilot, as it actually ran](#5-the-pilot-as-it-actually-ran) |
| A verification step | [6. Checking the output before anyone relied on it](#6-checking-the-output-before-anyone-relied-on-it) |
| An ownership record | [7. Graduating it](#7-graduating-it) |
| An update for your boss | [8. The update they sent](#8-the-update-they-sent) |

---

## 1. The situation

Northgate is a commercial insurance broker, about 400 people. In
January, IT rolled out Microsoft Copilot to the whole company. The
launch email went out on a Tuesday. By March, the admin console showed
that 31% of licenses had been used at least once, and 9% in the last
30 days.

Claims Operations is 14 people, reporting to Dana Okafor. Dana's
situation was the ordinary one: a tool they didn't choose, no budget,
no mandate, a team already busy, and a VP who had started asking in
skip-levels what people were doing with "the AI."

Dana ran the toolkit's sequence over 13 weeks. What follows is what
they actually filled in.

**Two things Dana got right that are easy to skip.** They took a
baseline in week one — before anything launched — and they named an
owner who wasn't themselves. Everything downstream depends on both.

**One thing they got wrong.** See [section 9](#9-what-went-wrong).

---

## 2. The policy they published

*Template: [ai-policy-template.md](./ai-policy-template.md)*

Dana emailed Northgate's security lead in week one asking for the data
rules in writing (the email is in
[getting-approval.md](./getting-approval.md)). The answer came back in
nine days. This is what Dana published on the team's SharePoint page
and pinned in the team channel.

---

### AI Use Policy — Claims Operations, Northgate

**Effective 2 February 2026 · Owner: Marcus Feld · Review: quarterly
(next 1 May)**

We want you using AI at work. This page says how, and where the limits
are, so you don't have to guess.

#### 1. What we want you using it for

Go ahead and use **Microsoft Copilot (your Northgate account)** for:

- Drafting — emails, docs, updates, first passes at anything
- Summarizing long documents, threads, and meeting notes
- Rewriting and editing your own work
- Brainstorming and thinking out loud
- Explaining things you don't understand — jargon, contracts,
  spreadsheets, acronyms
- Turning a claims export into a readable weekly summary
- First-pass replies to routine broker queries
- Reading a carrier policy wording and telling you where the relevant
  clause is

You don't need permission for any of this. You don't need to disclose
that you used AI to help write an internal email.

#### 2. What never goes in

Do not paste into any AI tool:

- **Claimant or insured data** that identifies real people — names,
  policy numbers, addresses, dates of birth, claim reference numbers
- **Employee data** — compensation, performance, health, personal
  details
- **Anything under NDA** or marked confidential
- **Regulated data** — medical reports, financial account details,
  legal case specifics, anything from a litigated file
- **Credentials** — passwords, API keys, tokens
- **Carrier pricing or commission terms** that aren't already public
- **Anything from a file flagged as being in litigation or under
  regulatory review** — no exceptions, ask Marcus

The simple version: **if you'd hesitate to post it publicly, it doesn't
go in.**

Why: with many AI tools, what you type can be retained and, depending
on the product and its settings, used to train future models. Our
Copilot account is on the Microsoft 365 commercial tier, which contracts
that away and keeps data inside our tenant. We have no such agreement
covering anything else, so personal and free accounts are not approved
for Northgate material. Use your Northgate login, not a personal ChatGPT
or Gemini account.

**Sanitized claims data — the one permitted exception.** Northgate
security confirmed in writing (S. Iyer, 28 January 2026) that a claims
export **with the claimant name, address, date of birth, policy number
and claim reference columns deleted** may be pasted into our Copilot
tenant. Delete the columns in Excel and save a copy — don't just hide
them, and don't paste and then edit. This exception covers claims
exports and nothing else: it does not extend to medical reports,
litigated files, or correspondence, no matter how much you remove.

**Need to work with anything else on the never-list?** Ask Marcus Feld.
There's often a safe way — aggregate numbers instead of records, a
made-up example with the same shape, or a different approach entirely.
Ask before, not after.

#### 3. You own what you send

AI drafts. A person approves. Specifically:

- **Verify anything that leaves the team or drives a decision.** These
  tools state wrong things with complete confidence — wrong reserve
  figures, wrong dates, invented clause numbers, plausible-looking
  carrier names. Check the facts, not just the tone.
- **You are responsible for what you send**, regardless of what drafted
  it. "The AI wrote it" is not a defense we can offer a broker or a
  regulator.
- **Disclose where it matters.** Not on a reworded internal email. Yes
  on analysis presented as your own careful work, and anywhere the
  reader would reasonably want to know.

#### 4. Some things stay human

Don't hand these to a tool: performance feedback, disciplinary matters,
condolences, any communication to a claimant about a death or serious
injury claim, or anything where the point is that a person spent the
time.

#### 5. Ask, and tell us what you learn

This policy will change as the tools do. If you're unsure whether
something is okay — **ask Marcus in #claims-ops-ai**. Nobody has ever
been in trouble here for asking. A question asked out loud is a problem
prevented.

Found something that works well? Post it in #claims-ops-ai so the rest
of the team gets it too.

---

**What to notice about that policy.** It's specific to claims work —
the "never goes in" list names claim reference numbers and litigated
files, not generic "confidential data," because generic lists don't
change anyone's behaviour at 4pm on a Thursday. It names a person twice.
And section 1 is longer than section 2, on purpose.

---

## 3. Scoring the candidates

*Template: [operating-templates.md](./operating-templates.md), sheet 1*

Dana ran the "what part of your week do you dread, that's basically the
same every time?" question in a team meeting and got eleven answers on
a whiteboard. Five were plausible. Here's the scoring, 1–5 on each
axis, higher total = better first bet.

| Candidate use case | Freq | Stakes | Data | Review | Reuse | Total |
| --- | --- | --- | --- | --- | --- | --- |
| Monday claims-status report for the VP | 5 | 4 | 4 | 5 | 5 | **23** |
| First-draft replies to routine broker chasers | 5 | 3 | 2 | 4 | 5 | **19** |
| Summarizing carrier policy wordings to find a clause | 4 | 4 | 4 | 4 | 4 | **20** |
| Drafting the quarterly reserve commentary | 2 | 1 | 1 | 2 | 2 | **8** |
| Auto-triaging the shared claims inbox | 5 | 2 | 1 | 2 | 4 | **14** |

**How Dana read it.**

The Monday report won on the total, and it was also the obvious answer
in hindsight — Amara had been building it by hand for two years and
complained about it roughly monthly.

The reserve commentary scored an 8 and was the thing Dana had privately
been most excited about. Quarterly, high-stakes, needs confidential
figures, and nobody can tell quickly whether the reasoning is right.
That combination is the classic trap: it's the task where AI would feel
most impressive and where being wrong would cost the most. It got cut.

Inbox triage scored 14 but has a **1 on data risk** and a **2 on ease
of review** — near-vetoes regardless of the total. Every email in that
inbox has a claimant's name in it, and a mis-triaged claim isn't
noticed until someone chases. Also cut.

The broker-chaser replies scored 19 but carry a 2 on data. Dana kept
them as a candidate for later, once the team had practice sanitizing.

**They started with the Monday report and the policy-wording lookups.**
Two, not five — Dana had one owner with two hours a week, and running
five pilots badly is worse than two properly.

---

## 4. Two borderline calls

*Template: [operating-templates.md](./operating-templates.md), sheet 2*

The decision aid isn't interesting in the abstract. Here are two real
questions the team asked in #claims-ops-ai, and how the five steps
resolved them.

### "Can I paste the loss run to get a summary?"

Asked by Wesley in week 4.

1. **Approved account?** Yes — Northgate Copilot.
2. **Identifies a real person / confidential / regulated?** Yes. A loss
   run has claimant names and claim references. → step 3.
3. **Do the written rules allow sanitized material?** Yes — and this is
   the step Wesley actually had to check rather than reason about. The
   policy names claims exports specifically, with the exact columns to
   delete, because security confirmed that in writing. Wesley needed the
   *pattern* anyway — claim counts by cause and month, not the records —
   so he deleted those columns in Excel, saved a copy, and pasted from
   that.
4. **Leaves the team / drives a decision?** It fed a broker
   conversation, so: verify the numbers against the source before
   using them. He did. One month's total was wrong in the summary — see
   [section 6](#6-checking-the-output-before-anyone-relied-on-it).
5. **Stays human?** No.

**Result: yes, within the written exception.** Marcus added "delete the
name and reference columns and save a copy first" to the team's prompt
library the same day.

**Note what did the work there.** Not Wesley's judgment that removing
names made it safe — that judgment is the thing people get wrong, and
de-identified records can often be re-identified. What made it a yes was
a sentence in the policy, traceable to a named person on a date. Before
28 January, the same question would have been a **STOP, ask Marcus**.

### "Can I use it to draft a decline letter?"

Asked by Rosa in week 6.

1. **Approved account?** Yes.
2. **Identifies a real person?** Yes — a decline letter is by
   definition about one claimant. → step 3.
3. **Do the written rules allow sanitized material?** Not for this. The
   claims-export exception covers exports, not correspondence about an
   individual claimant, and a decline letter is the latter. What *is*
   always available: drafting the reasoning and structure generically,
   with nothing real in it.
4. **Leaves the team / drives a decision?** Emphatically yes.
5. **Stays human?** Borderline, and this is where it stopped. A decline
   is bad news that affects someone's money and often their year.

**Result: a qualified yes that became a house rule.** Draft the
*structure* of a decline letter generically — the standard sections,
the clarity of the reasoning, the tone. Never paste the claimant's
details. A person writes the specifics and a person signs it. Dana
added the rule to section 4 of the policy at the quarterly review.

**What to notice.** Both answers were "yes, but differently than you
asked." The decision aid's job isn't to say no — it's to find the safe
version of the thing someone wanted to do.

---

## 5. The pilot, as it actually ran

*Template: [operating-templates.md](./operating-templates.md), sheet 3*

**Pilot:** Monday claims-status report
**Owner:** Marcus Feld
**Window:** 16 February → 13 March 2026 (four weeks)
**People:** Amara Nwosu (builds the report today), Rosa Delgado
(11 years in claims, vocally unconvinced), Wesley Chan
**Approved tool:** Microsoft Copilot · **Data class allowed:**
internal, non-identifying only — no claim references or claimant names

**The workflow we're testing:** Amara's Monday morning claims-status
report to the VP — a summary of open claims by age, cause, and
escalation status, built from a Wednesday-night system export.

**Baseline, before we start:**
- Time it takes today: about 2 hours every Monday morning, occasionally
  3 when the export is messy
- Current quality / pain: it's copy-paste from a pivot table into a
  Word template. It's never wrong, it's just tedious, and Amara can't
  start real work until it's out.

**Hypothesis:** With Copilot, the Monday report gets substantially
faster without losing the accuracy of the numbers or the "what changed
since last week" commentary the VP actually reads.

**Success threshold:** under an hour, and the VP doesn't notice a
difference in quality.

**Weekly check-ins:**

| Week | Who | What they showed me | Blocked on | Time now |
| --- | --- | --- | --- | --- |
| 1 | Amara | First attempt. Output was generic, "could have been any company" | Nothing — she was just pasting the export cold | ~2 hrs (no gain) |
| 1 | Rosa | Nothing. Hadn't opened it | "I don't see the point yet" | — |
| 1 | Wesley | Loss-run summary (not the pilot task) | Wanted to know if pasting it was allowed | — |
| 2 | Amara | Same task with a saved prompt: last week's report pasted in as the format to match | — | ~70 min |
| 2 | Rosa | Tried once. "It made up a claim cause that isn't one of ours" | Correct, and important — see below | — |
| 2 | Wesley | Sanitized loss-run workflow, working | — | 40 min → 10 min |
| 3 | Amara | Added a "lead with what changed" instruction and a rule to leave blanks rather than guess | — | ~40 min |
| 3 | Rosa | Reran hers with our cause-code list pasted in first. No invented causes | — | — |
| 3 | Wesley | Showed Rosa his sanitizing trick | — | — |
| 4 | Amara | Stable. Same prompt three weeks running | — | ~35 min |
| 4 | Rosa | Using it for policy-wording lookups daily. Still won't use it for anything a claimant sees | — | — |

**Examples worth showing the room:**
- Amara's Monday report: ~2 hrs → ~35 min, and she now writes the
  commentary instead of assembling the tables.
- Rosa found the invented-cause-code problem in week 2 and the fix in
  week 3 — paste the real list in first. That fix is in the prompt
  library and it's the thing everyone else needed to know.
- Wesley's sanitized loss-run summary: ~40 min → ~10 min.

**Decision:** ☑ Continue as-is
**Why:** Two workflows beat the threshold, and the skeptic found the
failure mode and the fix, which is worth more than the time saved.

**What Dana learned from the check-in table.** Week 1 produced nothing
— that's normal and it's why the pilot is four weeks, not two. The
whole gain arrived in week 2 when Amara stopped pasting cold and
started giving the tool last week's version to match. If Dana had
called it after two weeks, the conclusion would have been "this doesn't
help."

---

## 6. Checking the output before anyone relied on it

*Template: [operating-templates.md](./operating-templates.md), sheet 4*

**What we're checking:** the Monday claims-status report
**Reviewer:** Marcus Feld · **Date:** 16 March 2026 · **Tool / prompt
version:** Copilot, prompt v3 (the one with the cause-code list)

Run against three test cases, including one deliberately awkward.

| Test case | Facts | Complete | No invention | Tone/format | Verdict |
| --- | --- | --- | --- | --- | --- |
| Typical week (9 March export) | ✓ | ✓ | ✓ | ✓ | ☑ pass |
| Week with a missing column (export failed mid-run) | ✗ | ✓ | ✓ | ✓ | ☐ **fail** |
| Week with an unusually large single claim | ✓ | ✗ | ✓ | ✓ | ☐ **fail** |

**The two failures, which are the useful part:**

**Missing column.** When the export dropped the "days open" column, the
output didn't say so — it produced a confident ageing summary derived
from something else, and the totals looked entirely plausible. This is
the exact failure the checklist exists to catch: not a wrong-looking
answer, a right-looking one.
*Fix:* the prompt now ends with "If any expected column is missing from
the input, say so at the top and do not estimate that section." Re-ran:
pass.

**Large single claim.** A £1.2m claim skewed the totals and the summary
reported the average without flagging it. Technically accurate,
practically misleading — the VP would have drawn the wrong conclusion.
*Fix:* added "flag any single claim above £250k separately from the
totals." Re-ran: pass.

**Unacceptable failures for this output** — the list Marcus wrote
before testing:
- Never states a reserve figure that isn't in the input
- Never names a claimant
- Never invents a cause code outside our list
None occurred in any test case.

**Re-test trigger:** whenever the prompt changes, and on the first
Monday of each quarter regardless.

**Cleared for use?** ☑ Yes — after the two prompt fixes above.

**What to notice.** The first version passed the typical case and would
have shipped. Both real defects only appeared in the awkward cases,
which is the entire argument for testing more than one.

---

## 7. Graduating it

*Template: [operating-templates.md](./operating-templates.md), sheet 5*

Filled in the week the VP started expecting the report in the new
format — the moment it stopped being an experiment.

**What it is:** The saved Copilot prompt that drafts the Monday
claims-status report from the weekly export
**Owner:** Amara Nwosu
**Graduated on:** 23 March 2026

- **Who can use it:** Amara, Marcus, and Rosa as backup
- **Where it lives:** SharePoint → Claims Ops → AI prompt library →
  `monday-report-v3.md` (not in Amara's chat history)
- **What breaks if it stops working:** the VP's Monday 10am review has
  no numbers. Recoverable, but visibly.
- **Fallback if it's down or wrong:** the old pivot-table method. The
  Word template still exists and Amara can still do it in 2 hours.
- **Sensitive-data constraints:** the export must have the name and
  reference columns stripped before it's pasted. No exceptions.
- **Escalation contact:** Marcus Feld, #claims-ops-ai
- **Re-check cadence:** first Monday of each quarter, using the
  output-evaluation checklist

**Single-point-of-failure check:**
- ☑ It lives somewhere shared, not in one person's private account.
- ☑ A second person has actually run it, start to finish, successfully.
  **Rosa Delgado, 30 March 2026** — while Amara was on leave.

**What to notice.** The second box is the one that gets ticked
optimistically everywhere. Dana made Rosa actually run it during
Amara's leave, which surfaced that the SharePoint folder permissions
only included Amara. Ten minutes to fix, and it would have been a bad
Monday to discover it any other way.

---

## 8. The update they sent

*Template: [operating-templates.md](./operating-templates.md), sheet 6*

**To:** Sam Whitfield, VP Claims · **From:** Dana Okafor · **Period:**
Q1 2026

**Where we are:**
- **Access vs. use:** 14 people have licenses; 11 used Copilot last
  month, up from 3 in January.
- **What's survived:** 6 things were built or set up during the
  quarter; 4 are still in real use at 30 days.
- **Confidence:** team self-rating moved from **2.4 to 3.6** on our
  five-question survey (1–5), same wording both times, 13 of 14
  responding.

**The clearest single result:**

> Amara's Monday claims-status report took about 2 hours and now takes
> about 35 minutes — her estimate. She spends the difference writing
> the commentary you actually read rather than assembling the tables.

**The honest caveat:**

> We have **not** tried to measure aggregate productivity, and I'd be
> skeptical of anyone who claims to. Time saved on a task isn't the
> same as time returned to the business. What we're reporting is
> adoption and concrete examples, not a bottom-line number.
>
> I'd also flag that our verification testing found two ways the report
> could be confidently wrong before we caught them. Both are fixed and
> we now re-check quarterly. I'd rather you hear that from me.

**What we need next:** two hours a week of Marcus's time formally
protected in his objectives through Q3 — it's currently goodwill — and
your attendance at the team showcase on 24 April. That second one costs
nothing and does more than the first.

---

**Why that landed.** Two reasons, and neither is the numbers. Dana
reported a **failure they'd found and fixed**, which is what made the
rest of the report believable. And the ask was small, specific, and
mostly free.

Sam approved both. The Marcus time was formalized; more relevantly,
Sam repeated the "2 hours to 35 minutes" line in an operations meeting
and two other department heads asked Dana how to do it.

---

## 9. What went wrong

Left in deliberately, because a worked example where everything worked
is exactly the failure the showcase section warns about.

**The workshop was too big.** Dana scheduled a full day and a half for
14 people. Six weeks out, two people went on leave, a carrier audit
landed, and it collapsed to a single morning. They ran the compressed
version in [compressed-rollout.md](./compressed-rollout.md) instead —
three hours, everyone built one thing, five of the fourteen built
something they still use. Dana's assessment afterward: the day and a
half would have been better, and the three hours were worth roughly 70%
of it for 20% of the calendar cost. **If you're not sure you can hold
the time, book the short one.**

**The prompt library nearly died.** It was a SharePoint page that
Marcus updated and nobody else did. It got real in week 10 only when
Dana started ending each team meeting by asking "what went in the
library this week?" — a 30-second ritual that turned out to be the
whole mechanism.

**Two people never used it at all.** Dana asked once, got "I'm fine,
thanks," and left it. That's the right call and it should be stated
plainly: 11 of 14 is a good outcome, and chasing the last three would
have cost more goodwill than it was worth.

**The broker-reply use case never happened.** It scored 19 and was
still on the list at the end of the quarter. It turns out three
workflows is what a team of 14 absorbs in a quarter, not five.

---

## 10. How to use this

Don't copy Northgate's answers. Copy the level of specificity.

The difference between a policy that changes behaviour and one that
doesn't is whether section 2 says "confidential data" or "claim
reference numbers and anything from a litigated file." The difference
between a pilot that produces a result and one that produces a shrug is
whether you wrote the baseline down in week zero. The difference
between a leadership update that gets you the next thing and one that
doesn't is whether you volunteered a failure.

Practical order, if you're starting Monday:

1. Read the policy in [section 2](#2-the-policy-they-published)
   alongside the blank
   [ai-policy-template.md](./ai-policy-template.md) and fill yours in.
   Send the approval email from
   [getting-approval.md](./getting-approval.md) first — it's the long
   pole.
2. Run the dread question at your next team meeting and score what
   comes back, using [section 3](#3-scoring-the-candidates) as the
   worked version.
3. Pick two. Not five.
4. Write the baseline down before anyone touches anything.

*Related: the reasoning behind all of this is in
[managers-toolkit.md](./managers-toolkit.md); the blank versions of
every template above are in
[operating-templates.md](./operating-templates.md).*
