Tutorial

How to Review Your Team's AI Work (Even When They Hide It)

Something arrives in your inbox from someone who reports to you. A summary, a recommendation, an analysis. It reads well.

How hard you look at it should depend on what happens if it is wrong. In practice, it mostly depends on something else.

Why Punishing AI Disclosure Backfires

Founder Reports surveyed 2,078 employed US adults in April 2026 and found that 77% review a coworker’s work more carefully when they know AI was used, with 36% saying much more carefully. Trust moves the same way: 43% trust a colleague’s output less when they know AI was involved, against 20% who trust it more.

The load-bearing words are when they know.

That makes disclosure the trigger for the entire review, and disclosure is the one input you do not control, cannot verify, and have given your team no particular reason to volunteer. The same survey found that people trust the work less when they are told. Telling you is a cost the person paid voluntarily, and you responded by reading their work more suspiciously.

So the incentive runs the wrong way, and a review process built on Where something came from and who or what made it. In this context it means whether a document was written by a person, drafted by AI and edited, or generated and passed straight on. The word matters here because provenance is what most review instincts reach for first, and it is the one property of a document you usually cannot establish by looking at it. is built on the one thing you are least likely to get.

You are already paying for this

This is not a hypothetical risk to manage. It is a cost you are probably already absorbing, and it lands hardest on you specifically.

In that same survey, 57% of people at manager level and above have had to fix or redo a coworker’s work that leaned too heavily on AI, against 38% of individual contributors. Above senior manager it passes 60%. Seniority is not protecting you from this. It is the exposure.

Resume.org’s January 2026 survey of 1,146 US managers puts the same thing from the other side: 70% had seen at least one AI-related error from a direct report in the previous year, 40% said clients felt the effects, and 5% reported a single incident costing more than $50,000.

Why “did you use AI for this?” is the wrong question

It is the question that comes to mind, and it fails three ways.

You cannot verify the answer. Detection tools do not work reliably, and a document that was drafted by AI and then genuinely reworked by a person is not meaningfully “AI work” anyway.

Asking costs you something. Given that disclosure already gets punished with lower trust, a manager who opens with the provenance question teaches the team that mentioning AI invites scrutiny. You will get less information next quarter, not more.

And the answer would not tell you much. “I used AI” covers everything from fixing the grammar to generating the recommendation you are about to act on. Our own spectrum of AI use runs from reformatting to substantive analysis, and only the far end changes what you should do.

The useful question is not whether AI was involved. It is whether the work has the specific weaknesses that AI-assisted work tends to have. Those are known, they are narrow, and you can check for them without accusing anybody.

How to Spot AI Errors (Without AI Detection Tools)

Resume.org’s managers named what they had caught. 58% had seen factual inaccuracies. More than 50% had seen work that failed to account for critical contextual factors. Those two are not equally hard to catch.

A factual error is findable. You look it up. A Work that is accurate in every checkable particular and wrong about the situation. The number is right for last year rather than this one. The recommendation is sound in general and ignores the constraint your team has been working around for six months. The tone is correct for a customer and wrong for this customer. Nothing in it is false, so proofreading finds nothing. is the one that gets through, because there is nothing on the page to catch. It reads as finished.

So the review is four questions, and none of them mention AI.

  • “What would have to be true for this to be wrong?” Ask it of the one or two claims the whole thing rests on. Most documents have fewer load-bearing claims than they appear to, and the rest is scaffolding.
  • “What does this assume about our situation?” This is the direct hunt for the context failure. You know things the document’s author may not have supplied to anything or anyone: the constraint, the history, the person on the other end who reacted badly last time.
  • “Walk me through how you got to this number.” The single most useful sentence available to you, because it works identically on work a person did alone and work they did with AI. Somebody who understands the number can walk you through it either way. Somebody who cannot is a problem whether or not a model was involved.
  • Recompute exactly one thing. Pick one figure and derive it yourself. One is enough to tell you whether the document was assembled or checked, and one is cheap enough that you will actually do it.

The three checks for your own AI output are the individual version of this. What changes when the work came from a person is that you cannot interrogate the source, so the second and third questions have to be asked out loud, to somebody who will read your tone.

When you find one, talk about the work and nothing else

Finding the error is the easy half. The conversation afterwards is where a manager usually undoes everything above, because the instinct when something incorrect turns up is to reach straight for the tool. Did ChatGPT write this? That question feels like getting to the root cause. What it actually does is make the subject of the meeting the person’s AI use, which is the one subject guaranteed to produce a defensive answer and a quieter employee next quarter.

Keep it on the output. The Situation, Behavior, Impact: a feedback structure from the Center for Creative Leadership. You name where and when it happened, describe only what was observably done, and state the consequence. The discipline is in what it leaves out, which is any claim about the person’s motives, attitude or competence. It works here for a reason that has nothing to do with AI: motives are the part you cannot observe and the part people argue with. is the ordinary tool for this and it needs no adaptation:

  • Situation. The report that went to the client on Tuesday.
  • Behavior. It stated a delivery constraint we stopped working under in March, and it went out without anyone checking that against what we actually do now.
  • Impact. The client is planning around a commitment we cannot meet, and I have to walk that back myself.

Notice what is absent. No mention of a tool, no theory about why it happened, no adjective about the person. Everything there is true and checkable whether the document was drafted by a model, a contractor, or the person themselves at midnight, which is exactly why it cannot be argued with and does not need to be defended against.

If AI comes up, it will come up from the other side of the table, which is the version you want. Somebody explaining what they did and did not check is giving you the information in the previous section for free.

Make it predictable, or it reads as suspicion

A review that appears only when you suspect AI is a review that announces your suspicion. That is how you get a team that stops mentioning it.

The fix is to attach the depth of review to the stakes rather than to your hunch, and to say so before anybody submits anything. Something that goes to a client gets the four questions. An internal summary gets one. Nobody has to guess whether being reviewed means being doubted, because the rule was published in advance and applies to work you are certain no AI touched.

Then, when somebody does tell you they used AI, the thing to do is nothing different. That is the whole discipline. Given that 43% of people trust disclosed work less, the manager who visibly does not is buying information cheaply.

Trust has to move from the output to the method

There is something underneath all of this that the four questions do not touch. Reviewing a document is not the same as trusting the person who sent it, and trust in the person is what actually broke.

A finished, well-argued document used to be evidence of the thinking behind it. That inference is the one every manager runs on, usually without noticing: this person’s work comes back good, so this person has good judgment, so I can review them lightly and spend my attention elsewhere. Polish is now available without the judgment. The signal is still there and it is no longer attached to the thing you were reading it for.

Seen that way, the 43% who trust disclosed work less are not being unreasonable. They have lost a shortcut and nobody has given them another one.

The replacement is not more scrutiny of documents, which does not scale and which your team will correctly read as suspicion. It is making judgment visible somewhere other than the finished artifact. Three questions, asked once about how somebody works rather than every time they send you something:

  • “What did you hand over, and what did you keep?” The people who get the most out of AI hand over less, not more. Somebody who kept the analysis and delegated the drafting is telling you exactly where their attention went.
  • “What did you check before sending it?” Not whether. What.
  • “What did it get wrong that you caught?” The best of the three, because it is almost impossible to answer without having done the work. Somebody who can name what the model got wrong on this task has demonstrably read it against something they knew. Somebody who cannot has either not looked, or not looked in a way that would have found anything.

The good news is that this is recoverable, and the research says so more clearly than you might expect. The BetterUp Labs and Stanford study behind workslop found that 42% of people saw a colleague as less trustworthy and around half saw them as less capable. That penalty attaches to receiving bad AI work, not to AI work. Put next to the 43% figure, the picture is a prior rather than a verdict: disclosure costs you the benefit of the doubt, and demonstrated judgment buys it back.

Which is the point of asking. Once you can see how somebody works, you can go back to reviewing them lightly, a sort of return to normalcy.

If you mandated AI, budget the review

One finding in the Founder Reports data is worth sitting with. In companies that require AI use, 73% of workers have had to fix a coworker’s AI output, and 17% say it happens regularly. Where there is no AI policy at all, that figure is 30%.

Be careful with the direction of that. It is a correlation, and the obvious explanation is volume: mandate AI and there is simply more AI-assisted work in circulation to find fault with. It is not evidence that requiring AI makes people worse at their jobs.

What it does establish is that the rework does not disappear when adoption goes up. It moves. If you are the manager who told the team to use AI, the honest version of that instruction includes who absorbs the checking, and the time that checking costs is the part that never makes it into the business case.

The takeaway

Most review effort is currently allocated by whether someone told you, and telling you is a thing people are quietly punished for.

Review for the failure modes, not the provenance. Ask what the work assumes about your situation, ask somebody to walk you through the number, recompute one thing, and set the depth by what it costs to be wrong rather than by who you suspect. Every one of those questions is a good question to ask about work no AI ever touched, which is exactly why they survive not knowing.

Then do the part that lasts. Ask how somebody works, not what they used, because the thing you really lost was the ability to read a person’s judgment off a finished page. Get that back and you can go back to reviewing lightly, which is the only sustainable answer here. Reviewing everything forever is not a plan.


Sources: the 77%, 43%, 57%/38% and 73%/30% figures are from Founder Reports’ AI at Work survey (opens in a new tab), 2,078 employed US adults fielded via Prolific in April 2026, 80% full-time, across 22 work functions. Prolific is a genuine research panel, but this is a media company’s own survey rather than academic work, so treat the figures as directionally useful rather than precise. The manager figures are from Resume.org’s “AI Slop Crisis” survey (opens in a new tab), 1,146 US managers with at least one direct report, fielded via An online survey panel, where respondents opt in and are paid to answer. Panels are fast and cheap, which is why so much workplace AI research uses them, and they are not random samples of the workforce: the people who join are not the people who do not. This site keeps flagging them for that reason. The findings are worth reading and the decimal places are not. in January 2026, the same caveat this site has applied to it before. The trust figures for people who send poor AI work are from the BetterUp Labs and Stanford Social Media Lab research reported in “AI-Generated ‘Workslop’ Is Destroying Productivity” (opens in a new tab) (Harvard Business Review, September 2025), a survey of 1,150 US full-time employees. The Situation-Behavior-Impact model is the Center for Creative Leadership’s (opens in a new tab), not mine; applying it to an AI error specifically is my extension, and CCL sells feedback training built on it. The claim that mandating AI raises rework is not something either survey tested and I have deliberately not made it; the 73% against 30% is reported as a correlation with the volume explanation stated. The four review questions and the three questions about method are mine, adapted from this site’s checks for your own output rather than drawn from any study of manager review specifically. So is the reading of the two trust findings as a prior rather than a verdict: no study here compared disclosure against demonstrated judgment, and nobody has tested whether the penalty is actually recoverable.

Related: Workslop: The Polished Report That Wastes Everyone’s Time is this problem from the sender’s seat, and worth handing to anyone whose work keeps coming back. Do You Have to Tell People You Used AI? is the disclosure decision from the other side of the desk, which is the decision your team is making about you.