AI Mistakes at Work: Run a 30-Minute Post-Mortem, Not a Witch Hunt
Two findings from the same survey of 2,272 finance and risk professionals, published this month. 26% say internal audits have caught AI errors that reached external audiences or board members. And 84% say they are at least somewhat confident in the accuracy of AI output without human review.
Those are the same organizations. A quarter have already had something escape, and five in six are comfortable not looking.
Your team will produce a wrong thing with AI at some point, if it hasn’t already. What happens in the half hour after you find out determines whether you get a fix or a folk tale.
The one you hear about is not the first one
In a January 2026 survey of 1,146 US managers, 70% said a direct report had made a mistake using AI in the past year: 43% several times, 12% many times. The damage travelled: it landed on the manager in 58% of cases, on coworkers in 44%, on clients in 40%, and on upper management in 24%. Nearly one in five managers reported losses over $10,000, and 5% put them above $50,000.
So if your honest answer is “we haven’t had any AI mistakes,” the useful question is whether you’d hear about one. Amy Edmondson’s study of hospital nursing units found the units with better management and stronger relationships recorded higher medication error rates, and her reading was that this reflected willingness to report, not sloppier nursing. The teams that looked worst on paper were the ones actually telling you things.
That makes silence the least informative signal you have. What you want to hear about most is the An error that was caught before it reached anyone: the wrong number spotted in review, the invented citation someone happened to check. Near misses are the cheapest possible data about how your work goes wrong, because they cost nothing and tell you exactly the same thing an expensive failure would. Teams that only discuss the disasters are throwing away most of their evidence. , and that only happens if the last person who raised one wasn’t punished for it.
Why “who did this?” is the expensive question
The instinct after a bad output is to find the person who sent it. Resist it for thirty minutes, for two practical reasons.
The first is that you’ll get less information next time. Blame is a very effective way to stop hearing about problems, and you cannot fix a failure mode you never learn about.
The second is that the error usually isn’t where it looks. Someone was handed a tool, no rule about what to verify, and a deadline, and produced exactly what that arrangement produces. Once you know the output was wrong, the mistake looks obvious in a way it never was beforehand, which is The tendency to see an outcome as predictable once you already know it happened. After the fact, the wrong figure sits there in plain sight and the reviewer concludes the person must have been careless. Before the fact it was one of four hundred plausible-looking sentences in a document nobody had reason to distrust. Reviews that don’t correct for this reliably diagnose “carelessness” and reliably change nothing. doing the work, not analysis.
A A review that treats the failure as a property of the process rather than the person: the convention borrowed from aviation and software operations. It is not the same as “no accountability”: the team still owns a change with a name and a date on it. What’s off the table is attributing the cause to someone’s character, because that explanation is always available and never actionable. isn’t softness. It’s what you do when you actually want the next one prevented.
The 30-minute version
You don’t need an incident-management process. Six questions, timeboxed, with the artifact on screen and the person who produced it in the room as a participant rather than a defendant.
There’s a worksheet you can run the meeting from: a Word file, because you fill it in while the meeting is happening and send it round afterwards. It has all six questions with room to write, the tick-lists below, and facilitator notes on the traps. (Read it as a page first if you’d rather see what you’re getting.)
- What was produced, and what exactly was wrong with it? (5 min) Name the specific claim, number or passage. “It was inaccurate” isn’t a finding.
- Where did it get to? (2 min) Draft, internal, client, board. This sets how much the fix is worth, and nothing else in the meeting depends on it.
- Which step should have caught this? (10 min) The heart of it. Walk the path the work actually took, not the one on the process diagram.
- Was the check missing, skipped, or performed and fooled? (5 min) Three different problems: see below.
- What is the one change, who owns it, and by when? (5 min) One. A list of seven improvements is a list of zero.
- Who else is exposed to this same failure? (3 min) The team next door is almost certainly doing the same thing.
Question 4 is the one that earns the meeting, because the three answers point in completely different directions:
- Missing: nobody was ever supposed to check this. That’s a rule to write, and it belongs with your other ones rather than in a meeting memory.
- Skipped: a check existed and didn’t happen. Nearly always because it cost more than the deadline allowed or nobody specifically owned it. Fix the workflow; disciplining the person leaves the incentive exactly as it was.
- Fooled: the check happened and the error walked straight through it. This is the most useful and most overlooked outcome, and usually means people are braced for the wrong failure. Managers expect factual errors: 58% of reported mistakes were factual, but half were failures of context and nuance, and a context error reads as completely correct. How to Check AI’s Work covers what a check has to do to catch that kind.
Write down the failure mode, not the incident
The output of the half hour is one line in a place people will look again: “AI-drafted client summaries have invented a source twice, check every citation before it leaves.”
That’s a failure mode. It’s reusable, it’s specific to how your team works, and it survives the person who found it leaving. An incident report doesn’t do any of that, which is why almost nobody rereads one.
Put it wherever your AI ground rules already live: the toolkit’s setting the ground rules guide is the place if you’re using it, and Write a One-Page AI Policy Your Team Will Actually Follow is the short version if you’re not. A rule that lives in the document everyone was handed on day one gets followed. One that lives in the memory of a meeting six people attended does not.
The takeaway
Somebody on your team will publish something wrong that AI helped write. The half hour afterwards is worth more than any amount of policy written in advance, because it’s the only time you’ll have a real example of how your work actually fails.
Spend it on the process, not the person. Ask which step should have caught it, and whether that check was missing, skipped, or fooled. Those three words point at three different fixes, and picking the right one is most of the value. Ask, too, whether AI was the right tool for that job at all; sometimes the honest finding is that it wasn’t, and no amount of better prompting was going to fix it. Then write down the failure mode, not the incident, and go and warn the team next door.
The post-mortem worksheet has the whole thing laid out to run from.
Sources: The 26% and 84% figures are from Workiva’s 2026 Midyear Executive Benchmark Survey (opens in a new tab), released 11 August 2026, 2,272 finance, risk and sustainability professionals including 847 C-level executives, across North America, Latin America, Europe and Asia Pacific. Note it surveys those functions specifically, not all managers. The manager figures are from Resume.org’s January 2026 survey (opens in a new tab) of 1,146 US managers; it’s a resume-industry panel rather than academic research, the same caveat this site noted when using it before. Amy Edmondson’s finding is from Learning from Mistakes Is Easier Said Than Done (opens in a new tab) (Journal of Applied Behavioral Science, 1996), a study of eight units at one university hospital. The higher reported error rates in better-run units are a correlation, and the reporting explanation is Edmondson’s interpretation of it rather than a measured result. The six questions and the missing/skipped/fooled split are mine, built on the blameless post-mortem convention from aviation and software operations rather than on any study of AI specifically.
Related: How to Check AI’s Work: You’re Looking for the Wrong Mistake is the check itself, what has to happen before the work goes out, and why reading it over doesn’t catch anything. Workslop: The Polished Report That Wastes Everyone’s Time covers the failure that doesn’t announce itself as an error at all, which is the kind least likely to reach a post-mortem.