The Manager's AI Toolkit

Running a Small Pilot

You have an owner and a policy. Nobody is guessing about what’s allowed any more. The obvious next move is to get everyone in a room and teach them — and it’s the wrong one, because you’d be walking in with nothing to show them.

That’s what the pilot is for. It is not a small version of the rollout, and it is emphatically not a pilot of your team. It’s a pilot of the use case — a few weeks of finding out whether a particular piece of work is amenable, and what the before-and-after actually is.

Two or three people. Not the team.

This is the part that gets misread, so it’s worth being blunt about the numbers.

Two or three people, and one of them is you. Not eight. Not the department. If you can properly support five, fine — but the constraint is your attention, not enthusiasm, and one owner with two protected hours a week can genuinely sit with about three people. A “pilot” involving everybody is just an unannounced rollout.

They don’t have to be on your team. Ideally one or two are, because recognition is what makes the evidence land later — the room believes Maria’s Friday report in a way it will never believe a stranger’s. But if you can’t get volunteers, a peer in another team with the same shape of work does the job, and so does you on your own work. Weaker evidence that exists beats stronger evidence that doesn’t.

Nobody needs to be skilled at this yet. That’s the whole reason the pilot is small: you are sitting next to two people, weekly. You cannot sit next to fourteen, which is exactly why the workshop comes after this and not before it. The skill bar here is lower than the workshop’s, not higher — in the workshop people build alone.

Two or three people. Two or three use cases. Four weeks.

The pilot’s job is evidence, not proof

You already believe AI is useful; that isn’t in question and the pilot won’t settle it either way.

Its job is to produce three or four concrete examples of real work getting faster that you can show the room later. “Here’s how Maria cut her Friday report from 90 minutes to 20” beats any vendor demo, because the room knows Maria and knows the report. A stranger’s case study, however impressive, convinces nobody who has to do the work on Monday.

Take the baseline before anyone touches anything

How long the task takes today, and what’s annoying about it.

Ten minutes of work, and you cannot recover it afterwards. A pilot with no “before” produces a shrug no matter how well it goes — you’ll have a person saying it feels faster and no way to say how much. The pilot tracker has a place for it, along with a success threshold you write down at the start so you can’t move the goalposts later.

Pick people for credibility, not enthusiasm

One respected skeptic who comes around is worth five early adopters.

The enthusiasts convince nobody, because everyone already knows they were going to like it. The person whose opinion the team quietly checks against is the one whose verdict will travel — and if that person concludes it doesn’t help, you’ve learned something genuinely valuable before spending everyone’s afternoon on a workshop.

This is the one argument for pulling your two or three from inside the team where you can. Credibility is transitive: it attaches to the person, not the result.

Check in weekly, individually. Not “how’s it going” — you’ll get “fine” from exactly the people who are stuck. Ask to see what they have working right now.

Expect week one to produce nothing

That’s the normal shape, and it’s why four weeks rather than two.

People start by pasting things in cold and getting generic output back. The gain arrives in week two or three, when they start giving the tool last week’s version to match, or their team’s real terminology to work from. A pilot called after two weeks reliably concludes “this doesn’t help,” and that conclusion is an artefact of the timeline, not a finding.

Pick the right use cases, and score them

Most teams pick badly on the first try, in a predictable direction: too ambitious, too rare, and too far from the person choosing.

The cheapest way to find good ones. Ask each person: what part of your week do you dread, that’s basically the same every time? Write the answers on a wall. Your first use cases are on that wall.

Then score what’s on the wall instead of arguing about it. The use-case scoring sheet rates each candidate 1–5 on frequency, stakes if wrong, data risk, ease of review, and likely reuse. It takes five minutes and it settles the argument, which matters because the thing you’re personally most excited about is usually the worst first bet.

That’s not a throwaway line. The task that feels most impressive — quarterly, high-stakes, needs confidential data, and hard for anyone to check quickly — is the one where being wrong costs most and nobody notices for a month. Two axes are near-vetoes regardless of the total: data risk and ease of review. A frequent, high-reuse task you can’t safely feed or can’t check is a trap, not a starter.

The worked example has this scoring filled in with five real candidates, including why the two that scored highest on enthusiasm got cut.

Avoid at the start: anything touching production, anything reaching customers unread, anything that has to be right 100% of the time, and anything where “did it work?” takes a month to answer.

Engineering teams: tests for a module nobody wants to touch, docs for an undocumented service, a small migration, an ops script, alert triage. Same criteria — frequent, owned, reviewable.

Pick two or three, not five. One owner with two protected hours a week can support two or three pilots properly. Five run badly teach you nothing except that this doesn’t work. Three workflows is about what a team absorbs in a quarter, whatever the plan said.

What to resist

  • Org-wide launch with no pilot. Produces a spike of curiosity and a trough of nothing.
  • Mandating usage. You get compliance theater — people running work through the tool to be seen doing it. Mandate the outcome if you must; never the tool.
  • Waiting for the perfect tool. The gap between the leading tools is much smaller than the gap between using one and using none.
  • Buying more tools when the first isn’t used. Tool count is not adoption.

Where this goes next

Four weeks in, you should be holding three or four real before-and-after examples with names attached, plus a shortlist of the failure modes your team hit — which are exactly the things to warn the room about.

That’s what you need to walk into a workshop — the next guide. Skip this step and you arrive with a vendor demo, which is the single most common reason a workshop falls flat.

The takeaway: the pilot isn’t a trial of your team, it’s a hunt for evidence. Come out of it with one number somebody in the room can verify, and the hard part of the workshop is already done.

All toolkit guides