The Manager's AI Toolkit
↓ Download

AI Post-Mortem — [TEAM / COMPANY NAME]

Something went out with AI’s help and it was wrong. This is the half hour afterwards.

Fill it in during the meeting, on screen, where everyone can see it. Thirty minutes, six questions, one change at the end.

  • Date: [ ]
  • Facilitator: [ ]
  • In the room: [ ]

Before you start — read this out (1 minute)

We’re here to work out which step let this through, not who typed it. The person closest to it is the person who knows most about it, so they’re helping us, not answering for it. We’ll leave with one change and a name against it.

Say it even if it feels unnecessary. It is the thing that determines whether anyone tells you about the next one — and the next one is the one you’d actually like to hear about early.

Two rules for the half hour:

  • No “should have.” If a sentence starts that way, rephrase it as “there was no step that would have.”
  • The artifact is on screen. Not described from memory. The specific wrong sentence is the whole subject of question 1.

1. What was produced, and what exactly was wrong with it? (5 min)

“It was inaccurate” is not a finding. Name the claim, number, or passage.

What was the piece of work?
What was the specific error? (quote it)
What should it have said?
Who found it, and how?

How was AI used here? Tick what applies:

  • Wrote the first draft
  • Summarised or rewrote something we supplied
  • Did research or found sources
  • Did a calculation or analysis
  • Extracted data from documents
  • Ran as an agent across several steps
  • Other: [ ]

2. Where did it get to? (2 min)

This sets what the fix is worth. Nothing else in the meeting depends on it, so don’t spend longer.

  • Caught in draft — nobody outside the team saw it
  • Reached colleagues internally
  • Reached upper management or the board
  • Reached a client, customer, or the public
  • Reached a regulator, auditor, or went on the record
Has it been corrected? By whom?
Does anyone still need telling?

A near miss counts. If the answer is “caught in draft,” run the rest of this anyway. A near miss is the cheapest data you will ever get about how your work fails — same lesson, no damage.


3. Which step should have caught this? (10 min)

The heart of it. Walk the path the work actually took, not the one on the process diagram. List every stage between the prompt and the person who received it.

StepWho did itDid it happen?Would it have caught this?

Then answer the question directly:

The step that should have caught it
What happened at that step instead

4. Was the check missing, skipped, or fooled? (5 min)

Three answers, three completely different fixes. This is the question that earns the meeting, so don’t let the room settle on the first one that sounds right.

  • Missing — nobody was ever supposed to check this. It isn’t anyone’s job and never was. → You need a rule, not a conversation.
  • Skipped — the check existed and didn’t happen. Almost always because it cost more than the deadline allowed, or nobody specifically owned it. → Fix the workflow. Disciplining the person leaves the incentive exactly as it was, and it will happen again.
  • Fooled — the check happened and the error walked through it. → The check is aimed at the wrong failure.

On “fooled,” the most common version: people brace for factual errors — made-up figures, invented citations. Those are roughly 58% of reported AI mistakes, but around half are failures of context and nuance, and a context error reads as completely correct. It is fluent, plausible, internally consistent, and wrong about something the model was never told. Reading it over will not catch that. Something has to test it.

Which of the three, and why
If “fooled” — what was the check actually looking for?

5. What is the one change, who owns it, and by when? (5 min)

Work through these four in order. They’re ordered cheapest-and-most-durable first, and the answer is usually the first one you can honestly say yes to.

A. Was AI the right tool for this task at all?

Ask this first and answer it honestly, because every other fix assumes the answer is yes. Sometimes the real finding is that the task was a bad fit and no amount of better prompting or training would have changed the outcome. That is a legitimate result, not a failure of nerve — and it is the one conclusion a room full of people who have just been told to use AI will avoid reaching on its own.

  • Yes — the task suits it; this was an execution problem
  • Yes, but not unsupervised — it needs a human step we didn’t have
  • No — it needs facts we can’t put in front of the model (internal, current, or confidential)
  • No — the cost of being wrong here is higher than the time it saves
  • No — the work needs accountability a person has to hold
  • Not sure — we’ll try one more time with a change from B, then decide

The four signs the answer is really “no”: there is no source of truth you can hand it; the output is hard to check but expensive to get wrong; the task needs judgement about people; or verifying it properly takes as long as doing it yourself.

If any “No” is ticked, that is the change — stop here. Don’t also write a training plan for a task you’ve just decided shouldn’t use AI.

Decision
Which tasks does this apply to (just this one, or a class of them)?
What we’ll do instead
Who needs telling, so it doesn’t quietly restart

B. Can we change how the tool is set up, so this can’t recur?

The most durable fix, because it works even when everyone is busy and nobody remembers this meeting. Ask what would have had to be in front of the model for it to get this right.

  • Add the missing context to the prompt, project, or instructions
  • Give it the real source document instead of letting it recall
  • Add an explicit “say so if you don’t know” instruction
  • Split one long task into steps with a checkpoint between them
  • Turn the good version into a saved template or reusable prompt
  • Restrict what the tool can reach or do
  • Nothing here would have helped

Notes: [ ]

C. Does anyone need training?

Be specific. “More AI training” is not a change; it’s a feeling.

  • No — the person knew what to do, the process didn’t let them
  • One person needs a specific thing shown to them: [ ]
  • The team shares a gap: [ ]
  • The gap is checking, not prompting — most often it is
  • The gap is knowing when not to use it

Who runs it, and when: [ ]

D. Does a written rule or check need to change?

  • Add a check to an existing step
  • Add this to the AI policy or ground rules
  • Change who signs off, or at what stage
  • Add it to the “never put this in” list
  • No rule change needed

Notes: [ ]

Now pick one

One change. A list of seven improvements is a list of zero — the room feels productive and nothing is different in a month.

The one change
Owner (a person, not a team)
Done by
How we’ll know it worked

6. Who else is exposed to this? (3 min)

The team next door is almost certainly doing the same thing and has not hit it yet.

Who else does work like this?
Who tells them, and by when?

After the meeting — the one line that matters

Write down the failure mode, not the incident. A failure mode is reusable, specific to how your team works, and survives the person who found it leaving. An incident report does none of that, which is why nobody rereads one.

Format: “[What kind of AI-assisted work] has [done what wrong], [how many] times — [what to do about it].”

Example: “AI-drafted client summaries have invented a source twice — check every citation before it leaves.”

Ours:

Failure mode
Added to (policy / ground rules / checklist)
Added by

Put it wherever your AI ground rules already live. A rule in the document everyone was handed on day one gets followed; a rule in the memory of a meeting six people attended does not.


Facilitator notes

Traps this meeting falls into, in order of likelihood:

  1. Everyone agrees it was obvious. Once you know the output was wrong, the error sits there in plain sight and the room concludes someone was careless. Before the fact it was one of four hundred plausible-looking sentences nobody had reason to distrust. If the room reaches “careless” inside five minutes, that’s hindsight, not analysis — go back to question 3.
  2. The fix is “be more careful.” That’s not a change, it’s a mood. If question 5 produces nothing concrete, the honest output is “we accept this risk for now,” which is at least true and reviewable.
  3. Seven action items. Pick one. Note the others somewhere else if you must; they will not happen either way.
  4. The person who used AI does most of the talking. They should be answering questions about the work, not accounting for themselves. If it starts feeling like a hearing, say the ground rule again.
  5. It becomes a debate about AI in general. Whether the team should use AI at all is a real question and this is the wrong meeting for it. Question 5A answers it for this task, which is as far as one bad output can honestly take you.

When not to run this. If the same failure mode has already come up twice and the change from last time never happened, don’t run a third post-mortem. The problem isn’t understanding — it’s that nothing was owned. Go and fix that instead.


From Practical AI Guide. The reasoning behind this worksheet is at https://www.practicalaiguide.com/blog/ai-mistake-post-mortem/ — the rollout sequence it fits into is at https://www.practicalaiguide.com/toolkit/