Applied / Business

How to Run an AI Pilot: Aim for Evidence, Not Proof

Two numbers get quoted at managers about to run their first AI pilot.

MIT’s Project NANDA reported that 95% of enterprise generative-AI pilots delivered no measurable impact on profit and loss. Gartner’s survey work put 48% of AI projects reaching production, taking about eight months to get there.

Those can’t both be describing the same thing, and the gap between them is more instructive than either figure.

Why the numbers disagree

Partly it’s definitions. “No measurable P&L impact” and “didn’t reach production” are different failures, and a pilot can do the second while succeeding at something useful.

Partly it’s rigor. The NANDA report rests on 52 interviews, 153 survey responses and a scan of 300 public announcements, gathered over six months: a modest, largely qualitative base for a number as precise-sounding as 95%. It drew public criticism on exactly that point when it landed. The Gartner figure is better grounded but dates from 2024, which in this field is a long time ago.

So treat both as weather reports rather than measurements. What’s worth taking from them isn’t the percentage. It’s the consistent shape underneath: pilots rarely fail during the pilot. They go fine, everyone is mildly impressed, and then nothing happens. The failure is at the transition.

The goal most pilots are quietly aiming at is the wrong one

Here’s the trap. Those studies measure enterprise pilots pointed at production deployment: a system, integrated, running for everyone. That’s a programme, with infrastructure and a budget line, and it’s the shape of thing that stalls in The state where something works well enough that nobody wants to kill it, but not well enough that anyone will commit to rolling it out. It runs on as a permanent trial, consuming attention and producing no decision. .

If you manage a team rather than an enterprise, that goal is wrong for you. You are not trying to deploy a system. You are trying to find out whether this helps, and to come out holding something you can show people.

Your pilot’s job is evidence, not proof.

That distinction changes the design completely. A pilot aiming at proof has to be big enough to be convincing, which makes it slow, expensive and fragile. A pilot aiming at evidence needs two or three people and two weeks, and it succeeds even when the answer is no, because “we tried it on this and it didn’t help” is genuinely useful to know, and it cost almost nothing to find out.

How to design one that can’t fail in the usual way

Five things, and the order matters more than any of them individually.

Two or three people, not the team. They don’t even need to be on your team. Often it’s you and one willing colleague. A pilot big enough to need coordination is big enough to need justifying, and that’s how it becomes a programme.

Ask what people dread. The prompt that finds good candidates is: what part of your week do you dread, that’s basically the same every time? Repetitive, low-stakes and checkable is the shape you want. Then score what comes back and pick two or three. Not five.

Write down the How long the task takes today, measured before anyone touches anything. It feels like bureaucracy and takes ten minutes. Without it you have no way to say what changed, and “it feels faster” is exactly the claim that gets dismissed by the person you need to convince. before anyone starts. This is the step people skip, and it’s the one that decides whether the pilot produces anything usable. You cannot reconstruct it afterwards.

Expect week one to produce nothing. The first week is people working out where the tool is and what it’s bad at. Judge the pilot on it and you’ll kill something that was about to start working.

Pick people for credibility, not enthusiasm. The colleague who is already excited will produce a result nobody believes. The mild skeptic whose opinion carries weight will produce a result that travels.

What the evidence is actually for

The output of a good pilot isn’t a system. It’s three or four before-and-after examples from people your team knows by name, this took ninety minutes and now it takes twenty, plus an honest list of what didn’t work.

That’s what makes the next step land. A team session that opens with real examples from the building is a different event from one that opens with a vendor demo. The pilot exists to produce that opening.

It’s also why the pilot comes before the workshop rather than after, which surprises people. You aren’t asking the pilot group to be expert. You’re asking them to generate the evidence that makes everyone else turn up.

The honest caveat

None of this scales to an enterprise deployment. If your job is rolling AI out to nine thousand people, the studies above are describing your problem rather than mine, and you need the infrastructure conversation they imply.

But most people reading this manage a team, not a company. For them the enterprise pilot model is borrowed trouble. It imports a failure mode that a two-person, two-week evidence-gathering exercise simply doesn’t have.

The takeaway

Stop asking whether the pilot will prove AI works. Ask what evidence you want to be holding in three weeks, then design the smallest thing that produces it.

Two or three people. Two or three tasks they already dread. A baseline written down before anyone starts. Then two weeks, and an honest account of what happened, including the parts that didn’t.

A pilot pointed at proof can fail. A pilot pointed at evidence returns something useful either way, which is why it’s the one worth running.


Sources: the 95% figure is from MIT Project NANDA’s “The GenAI Divide: State of AI in Business 2025” (opens in a new tab) (July 2025), based on 52 structured interviews, 153 survey responses and an analysis of 300+ public AI initiatives, gathered January to June 2025. The report attracted substantial public criticism of its methodology and sample after publication, which is why it appears here as contested rather than settled. The 48%-to-production and eight-month figures come from a Gartner survey press release (opens in a new tab) of May 2024, Gartner blocks automated access, so I’m relying on secondary reporting of that release rather than my own reading of it, and it is over two years old. The pilot method is from this site’s own running a pilot guide, not from either study.

Related: Why Most Workplace AI Rollouts Quietly Fail, and What the Winners Do Differently is the organization-wide version of this problem. AI Workshop in Three Hours: You Won’t Get a Day and a Half is the step the pilot’s evidence feeds into, and the reason the order runs this way round.