The Manager's AI Toolkit
↓ Download

Worked Example — one team, every template filled in

Everything else in this toolkit is brackets waiting for your details. This is the same set of documents with the brackets filled in, for one team, start to finish — so you can see what a real answer looks like before you write your own.

This team is invented. Northgate Claims Operations is a composite, not a real company, and the numbers in it are illustrative rather than measured. Don’t cite them. Do copy the shape.

Read it whichever way suits you: straight through as a story, or jump to the one artifact you’re about to write yourself.

You’re about to writeJump to
The one-page policy2. The policy they published
A use-case shortlist3. Scoring the candidates
Guidance for a borderline call4. Two borderline calls
A pilot plan5. The pilot, as it actually ran
A verification step6. Checking the output before anyone relied on it
An ownership record7. Graduating it
An update for your boss8. The update they sent

1. The situation

Northgate is a commercial insurance broker, about 400 people. In January, IT rolled out Microsoft Copilot to the whole company. The launch email went out on a Tuesday. By March, the admin console showed that 31% of licenses had been used at least once, and 9% in the last 30 days.

Claims Operations is 14 people, reporting to Dana Okafor. Dana’s situation was the ordinary one: a tool they didn’t choose, no budget, no mandate, a team already busy, and a VP who had started asking in skip-levels what people were doing with “the AI.”

Dana ran the toolkit’s sequence over 13 weeks. What follows is what they actually filled in.

Two things Dana got right that are easy to skip. They took a baseline in week one — before anything launched — and they named an owner who wasn’t themselves. Everything downstream depends on both.

One thing they got wrong. See section 9.


2. The policy they published

Template: ai-policy-template.md

Dana emailed Northgate’s security lead in week one asking for the data rules in writing (the email is in getting-approval.md). The answer came back in nine days. This is what Dana published on the team’s SharePoint page and pinned in the team channel.


AI Use Policy — Claims Operations, Northgate

Effective 2 February 2026 · Owner: Marcus Feld · Review: quarterly (next 1 May)

We want you using AI at work. This page says how, and where the limits are, so you don’t have to guess.

1. What we want you using it for

Go ahead and use Microsoft Copilot (your Northgate account) for:

  • Drafting — emails, docs, updates, first passes at anything
  • Summarizing long documents, threads, and meeting notes
  • Rewriting and editing your own work
  • Brainstorming and thinking out loud
  • Explaining things you don’t understand — jargon, contracts, spreadsheets, acronyms
  • Turning a claims export into a readable weekly summary
  • First-pass replies to routine broker queries
  • Reading a carrier policy wording and telling you where the relevant clause is

You don’t need permission for any of this. You don’t need to disclose that you used AI to help write an internal email.

2. What never goes in

Do not paste into any AI tool:

  • Claimant or insured data that identifies real people — names, policy numbers, addresses, dates of birth, claim reference numbers
  • Employee data — compensation, performance, health, personal details
  • Anything under NDA or marked confidential
  • Regulated data — medical reports, financial account details, legal case specifics, anything from a litigated file
  • Credentials — passwords, API keys, tokens
  • Carrier pricing or commission terms that aren’t already public
  • Anything from a file flagged as being in litigation or under regulatory review — no exceptions, ask Marcus

The simple version: if you’d hesitate to post it publicly, it doesn’t go in.

Why: with many AI tools, what you type can be retained and, depending on the product and its settings, used to train future models. Our Copilot account is on the Microsoft 365 commercial tier, which contracts that away and keeps data inside our tenant. We have no such agreement covering anything else, so personal and free accounts are not approved for Northgate material. Use your Northgate login, not a personal ChatGPT or Gemini account.

Sanitized claims data — the one permitted exception. Northgate security confirmed in writing (S. Iyer, 28 January 2026) that a claims export with the claimant name, address, date of birth, policy number and claim reference columns deleted may be pasted into our Copilot tenant. Delete the columns in Excel and save a copy — don’t just hide them, and don’t paste and then edit. This exception covers claims exports and nothing else: it does not extend to medical reports, litigated files, or correspondence, no matter how much you remove.

Need to work with anything else on the never-list? Ask Marcus Feld. There’s often a safe way — aggregate numbers instead of records, a made-up example with the same shape, or a different approach entirely. Ask before, not after.

3. You own what you send

AI drafts. A person approves. Specifically:

  • Verify anything that leaves the team or drives a decision. These tools state wrong things with complete confidence — wrong reserve figures, wrong dates, invented clause numbers, plausible-looking carrier names. Check the facts, not just the tone.
  • You are responsible for what you send, regardless of what drafted it. “The AI wrote it” is not a defense we can offer a broker or a regulator.
  • Disclose where it matters. Not on a reworded internal email. Yes on analysis presented as your own careful work, and anywhere the reader would reasonably want to know.

4. Some things stay human

Don’t hand these to a tool: performance feedback, disciplinary matters, condolences, any communication to a claimant about a death or serious injury claim, or anything where the point is that a person spent the time.

5. Ask, and tell us what you learn

This policy will change as the tools do. If you’re unsure whether something is okay — ask Marcus in #claims-ops-ai. Nobody has ever been in trouble here for asking. A question asked out loud is a problem prevented.

Found something that works well? Post it in #claims-ops-ai so the rest of the team gets it too.


What to notice about that policy. It’s specific to claims work — the “never goes in” list names claim reference numbers and litigated files, not generic “confidential data,” because generic lists don’t change anyone’s behaviour at 4pm on a Thursday. It names a person twice. And section 1 is longer than section 2, on purpose.


3. Scoring the candidates

Template: operating-templates.md, sheet 1

Dana ran the “what part of your week do you dread, that’s basically the same every time?” question in a team meeting and got eleven answers on a whiteboard. Five were plausible. Here’s the scoring, 1–5 on each axis, higher total = better first bet.

Candidate use caseFreqStakesDataReviewReuseTotal
Monday claims-status report for the VP5445523
First-draft replies to routine broker chasers5324519
Summarizing carrier policy wordings to find a clause4444420
Drafting the quarterly reserve commentary211228
Auto-triaging the shared claims inbox5212414

How Dana read it.

The Monday report won on the total, and it was also the obvious answer in hindsight — Amara had been building it by hand for two years and complained about it roughly monthly.

The reserve commentary scored an 8 and was the thing Dana had privately been most excited about. Quarterly, high-stakes, needs confidential figures, and nobody can tell quickly whether the reasoning is right. That combination is the classic trap: it’s the task where AI would feel most impressive and where being wrong would cost the most. It got cut.

Inbox triage scored 14 but has a 1 on data risk and a 2 on ease of review — near-vetoes regardless of the total. Every email in that inbox has a claimant’s name in it, and a mis-triaged claim isn’t noticed until someone chases. Also cut.

The broker-chaser replies scored 19 but carry a 2 on data. Dana kept them as a candidate for later, once the team had practice sanitizing.

They started with the Monday report and the policy-wording lookups. Two, not five — Dana had one owner with two hours a week, and running five pilots badly is worse than two properly.


4. Two borderline calls

Template: operating-templates.md, sheet 2

The decision aid isn’t interesting in the abstract. Here are two real questions the team asked in #claims-ops-ai, and how the five steps resolved them.

”Can I paste the loss run to get a summary?”

Asked by Wesley in week 4.

  1. Approved account? Yes — Northgate Copilot.
  2. Identifies a real person / confidential / regulated? Yes. A loss run has claimant names and claim references. → step 3.
  3. Do the written rules allow sanitized material? Yes — and this is the step Wesley actually had to check rather than reason about. The policy names claims exports specifically, with the exact columns to delete, because security confirmed that in writing. Wesley needed the pattern anyway — claim counts by cause and month, not the records — so he deleted those columns in Excel, saved a copy, and pasted from that.
  4. Leaves the team / drives a decision? It fed a broker conversation, so: verify the numbers against the source before using them. He did. One month’s total was wrong in the summary — see section 6.
  5. Stays human? No.

Result: yes, within the written exception. Marcus added “delete the name and reference columns and save a copy first” to the team’s prompt library the same day.

Note what did the work there. Not Wesley’s judgment that removing names made it safe — that judgment is the thing people get wrong, and de-identified records can often be re-identified. What made it a yes was a sentence in the policy, traceable to a named person on a date. Before 28 January, the same question would have been a STOP, ask Marcus.

”Can I use it to draft a decline letter?”

Asked by Rosa in week 6.

  1. Approved account? Yes.
  2. Identifies a real person? Yes — a decline letter is by definition about one claimant. → step 3.
  3. Do the written rules allow sanitized material? Not for this. The claims-export exception covers exports, not correspondence about an individual claimant, and a decline letter is the latter. What is always available: drafting the reasoning and structure generically, with nothing real in it.
  4. Leaves the team / drives a decision? Emphatically yes.
  5. Stays human? Borderline, and this is where it stopped. A decline is bad news that affects someone’s money and often their year.

Result: a qualified yes that became a house rule. Draft the structure of a decline letter generically — the standard sections, the clarity of the reasoning, the tone. Never paste the claimant’s details. A person writes the specifics and a person signs it. Dana added the rule to section 4 of the policy at the quarterly review.

What to notice. Both answers were “yes, but differently than you asked.” The decision aid’s job isn’t to say no — it’s to find the safe version of the thing someone wanted to do.


5. The pilot, as it actually ran

Template: operating-templates.md, sheet 3

Pilot: Monday claims-status report Owner: Marcus Feld Window: 16 February → 13 March 2026 (four weeks) People: Amara Nwosu (builds the report today), Rosa Delgado (11 years in claims, vocally unconvinced), Wesley Chan Approved tool: Microsoft Copilot · Data class allowed: internal, non-identifying only — no claim references or claimant names

The workflow we’re testing: Amara’s Monday morning claims-status report to the VP — a summary of open claims by age, cause, and escalation status, built from a Wednesday-night system export.

Baseline, before we start:

  • Time it takes today: about 2 hours every Monday morning, occasionally 3 when the export is messy
  • Current quality / pain: it’s copy-paste from a pivot table into a Word template. It’s never wrong, it’s just tedious, and Amara can’t start real work until it’s out.

Hypothesis: With Copilot, the Monday report gets substantially faster without losing the accuracy of the numbers or the “what changed since last week” commentary the VP actually reads.

Success threshold: under an hour, and the VP doesn’t notice a difference in quality.

Weekly check-ins:

WeekWhoWhat they showed meBlocked onTime now
1AmaraFirst attempt. Output was generic, “could have been any company”Nothing — she was just pasting the export cold~2 hrs (no gain)
1RosaNothing. Hadn’t opened it”I don’t see the point yet”
1WesleyLoss-run summary (not the pilot task)Wanted to know if pasting it was allowed
2AmaraSame task with a saved prompt: last week’s report pasted in as the format to match~70 min
2RosaTried once. “It made up a claim cause that isn’t one of ours”Correct, and important — see below
2WesleySanitized loss-run workflow, working40 min → 10 min
3AmaraAdded a “lead with what changed” instruction and a rule to leave blanks rather than guess~40 min
3RosaReran hers with our cause-code list pasted in first. No invented causes
3WesleyShowed Rosa his sanitizing trick
4AmaraStable. Same prompt three weeks running~35 min
4RosaUsing it for policy-wording lookups daily. Still won’t use it for anything a claimant sees

Examples worth showing the room:

  • Amara’s Monday report: ~2 hrs → ~35 min, and she now writes the commentary instead of assembling the tables.
  • Rosa found the invented-cause-code problem in week 2 and the fix in week 3 — paste the real list in first. That fix is in the prompt library and it’s the thing everyone else needed to know.
  • Wesley’s sanitized loss-run summary: ~40 min → ~10 min.

Decision: ☑ Continue as-is Why: Two workflows beat the threshold, and the skeptic found the failure mode and the fix, which is worth more than the time saved.

What Dana learned from the check-in table. Week 1 produced nothing — that’s normal and it’s why the pilot is four weeks, not two. The whole gain arrived in week 2 when Amara stopped pasting cold and started giving the tool last week’s version to match. If Dana had called it after two weeks, the conclusion would have been “this doesn’t help.”


6. Checking the output before anyone relied on it

Template: operating-templates.md, sheet 4

What we’re checking: the Monday claims-status report Reviewer: Marcus Feld · Date: 16 March 2026 · Tool / prompt version: Copilot, prompt v3 (the one with the cause-code list)

Run against three test cases, including one deliberately awkward.

Test caseFactsCompleteNo inventionTone/formatVerdict
Typical week (9 March export)☑ pass
Week with a missing column (export failed mid-run)fail
Week with an unusually large single claimfail

The two failures, which are the useful part:

Missing column. When the export dropped the “days open” column, the output didn’t say so — it produced a confident ageing summary derived from something else, and the totals looked entirely plausible. This is the exact failure the checklist exists to catch: not a wrong-looking answer, a right-looking one. Fix: the prompt now ends with “If any expected column is missing from the input, say so at the top and do not estimate that section.” Re-ran: pass.

Large single claim. A £1.2m claim skewed the totals and the summary reported the average without flagging it. Technically accurate, practically misleading — the VP would have drawn the wrong conclusion. Fix: added “flag any single claim above £250k separately from the totals.” Re-ran: pass.

Unacceptable failures for this output — the list Marcus wrote before testing:

  • Never states a reserve figure that isn’t in the input
  • Never names a claimant
  • Never invents a cause code outside our list None occurred in any test case.

Re-test trigger: whenever the prompt changes, and on the first Monday of each quarter regardless.

Cleared for use? ☑ Yes — after the two prompt fixes above.

What to notice. The first version passed the typical case and would have shipped. Both real defects only appeared in the awkward cases, which is the entire argument for testing more than one.


7. Graduating it

Template: operating-templates.md, sheet 5

Filled in the week the VP started expecting the report in the new format — the moment it stopped being an experiment.

What it is: The saved Copilot prompt that drafts the Monday claims-status report from the weekly export Owner: Amara Nwosu Graduated on: 23 March 2026

  • Who can use it: Amara, Marcus, and Rosa as backup
  • Where it lives: SharePoint → Claims Ops → AI prompt library → monday-report-v3.md (not in Amara’s chat history)
  • What breaks if it stops working: the VP’s Monday 10am review has no numbers. Recoverable, but visibly.
  • Fallback if it’s down or wrong: the old pivot-table method. The Word template still exists and Amara can still do it in 2 hours.
  • Sensitive-data constraints: the export must have the name and reference columns stripped before it’s pasted. No exceptions.
  • Escalation contact: Marcus Feld, #claims-ops-ai
  • Re-check cadence: first Monday of each quarter, using the output-evaluation checklist

Single-point-of-failure check:

  • ☑ It lives somewhere shared, not in one person’s private account.
  • ☑ A second person has actually run it, start to finish, successfully. Rosa Delgado, 30 March 2026 — while Amara was on leave.

What to notice. The second box is the one that gets ticked optimistically everywhere. Dana made Rosa actually run it during Amara’s leave, which surfaced that the SharePoint folder permissions only included Amara. Ten minutes to fix, and it would have been a bad Monday to discover it any other way.


8. The update they sent

Template: operating-templates.md, sheet 6

To: Sam Whitfield, VP Claims · From: Dana Okafor · Period: Q1 2026

Where we are:

  • Access vs. use: 14 people have licenses; 11 used Copilot last month, up from 3 in January.
  • What’s survived: 6 things were built or set up during the quarter; 4 are still in real use at 30 days.
  • Confidence: team self-rating moved from 2.4 to 3.6 on our five-question survey (1–5), same wording both times, 13 of 14 responding.

The clearest single result:

Amara’s Monday claims-status report took about 2 hours and now takes about 35 minutes — her estimate. She spends the difference writing the commentary you actually read rather than assembling the tables.

The honest caveat:

We have not tried to measure aggregate productivity, and I’d be skeptical of anyone who claims to. Time saved on a task isn’t the same as time returned to the business. What we’re reporting is adoption and concrete examples, not a bottom-line number.

I’d also flag that our verification testing found two ways the report could be confidently wrong before we caught them. Both are fixed and we now re-check quarterly. I’d rather you hear that from me.

What we need next: two hours a week of Marcus’s time formally protected in his objectives through Q3 — it’s currently goodwill — and your attendance at the team showcase on 24 April. That second one costs nothing and does more than the first.


Why that landed. Two reasons, and neither is the numbers. Dana reported a failure they’d found and fixed, which is what made the rest of the report believable. And the ask was small, specific, and mostly free.

Sam approved both. The Marcus time was formalized; more relevantly, Sam repeated the “2 hours to 35 minutes” line in an operations meeting and two other department heads asked Dana how to do it.


9. What went wrong

Left in deliberately, because a worked example where everything worked is exactly the failure the showcase section warns about.

The workshop was too big. Dana scheduled a full day and a half for 14 people. Six weeks out, two people went on leave, a carrier audit landed, and it collapsed to a single morning. They ran the compressed version in compressed-rollout.md instead — three hours, everyone built one thing, five of the fourteen built something they still use. Dana’s assessment afterward: the day and a half would have been better, and the three hours were worth roughly 70% of it for 20% of the calendar cost. If you’re not sure you can hold the time, book the short one.

The prompt library nearly died. It was a SharePoint page that Marcus updated and nobody else did. It got real in week 10 only when Dana started ending each team meeting by asking “what went in the library this week?” — a 30-second ritual that turned out to be the whole mechanism.

Two people never used it at all. Dana asked once, got “I’m fine, thanks,” and left it. That’s the right call and it should be stated plainly: 11 of 14 is a good outcome, and chasing the last three would have cost more goodwill than it was worth.

The broker-reply use case never happened. It scored 19 and was still on the list at the end of the quarter. It turns out three workflows is what a team of 14 absorbs in a quarter, not five.


10. How to use this

Don’t copy Northgate’s answers. Copy the level of specificity.

The difference between a policy that changes behaviour and one that doesn’t is whether section 2 says “confidential data” or “claim reference numbers and anything from a litigated file.” The difference between a pilot that produces a result and one that produces a shrug is whether you wrote the baseline down in week zero. The difference between a leadership update that gets you the next thing and one that doesn’t is whether you volunteered a failure.

Practical order, if you’re starting Monday:

  1. Read the policy in section 2 alongside the blank ai-policy-template.md and fill yours in. Send the approval email from getting-approval.md first — it’s the long pole.
  2. Run the dread question at your next team meeting and score what comes back, using section 3 as the worked version.
  3. Pick two. Not five.
  4. Write the baseline down before anyone touches anything.

Related: the reasoning behind all of this is in managers-toolkit.md; the blank versions of every template above are in operating-templates.md.