Making It Stick
Four weeks after a good workshop, one of two things is true. Either three or four of the things people built are quietly still in use — or none of them are, and the whole thing has become a story about a fun week in the spring.
The difference is almost entirely what happened in the weeks immediately after. This guide is that part.
The first week after
Three things, none of them expensive:
Publish within a week. Notes, the prompt library, and the showcase recording, somewhere the whole company can find them. A week later and the energy is gone.
Office hours. Thirty minutes weekly for a month. Attendance will drop to zero by week four, and that’s the correct outcome — it means people are unblocked, not that nobody cares.
Name owners. For each thing someone wants to keep using, write down who maintains it.
Then add one 30-second ritual to a meeting you already have: what went into the prompt library this week? It’s the cheapest mechanism in this whole toolkit, and it’s usually the difference between a library that one person maintains and a library the team maintains.
Ownership, because things rot
Anything built in a workshop and adopted by a team without an owner becomes the team’s problem in six months. The moment something graduates from experiment to dependency — the moment people start relying on it — fill in an ownership record. Two minutes per item, and if you can’t answer the questions on it, it isn’t ready to be relied on.
The one line on it that people tick optimistically: a second person has actually run this, start to finish, successfully. Not “could” — has, with a name and a date. Make somebody do it while the builder is on leave and you’ll find the permissions problem, the undocumented step, or the file that only exists in one person’s account. Finding it that way costs ten minutes; finding it the other way costs a bad Monday.
The maintenance realities people miss:
- Model updates change output. A prompt tuned to one version can behave differently after an update. Anything load-bearing needs an occasional spot-check, not set-and-forget.
- The builder leaves. If one person’s private prompt is now in the team’s critical path, that’s a single point of failure. Move it somewhere shared and make sure someone else has actually run it.
- Silent degradation. Wrong output that looks right is the failure mode to design against. Keep a human check on anything consequential.
- Cost drift. Usage-based tools get expensive quietly.
Sunset deliberately. Once a quarter, ask which of these nobody uses anymore, and delete them. Four things that work beat twelve things where four work.
Skeptics and over-enthusiasts
Both groups fail, in opposite directions, and both need managing.
The skeptic is usually right about something. “It makes things up” — yes, confidently, which is exactly why verification is in the policy. “It’s not good at my job” — often true of the hard 20%, so concede that and aim at the boring 80%. “I tried it and it was useless” — almost always a one-line prompt with no context, so sit with them once. “This is about replacing us” — the fear is legitimate; answer it straight and don’t over-promise, because vague reassurance reads as confirmation.
The move that works is pairing them with the chore they personally hate most. Never argue about AI in the abstract; that argument has no end state. And genuinely leave room for “we tried, it doesn’t help here” — a manager who can’t hear that gets told nothing at all.
The over-enthusiast is less discussed and more damaging when unmanaged. They ship unverified output, automate things that shouldn’t be automated, build fragile machinery only they can debug, and paste things they shouldn’t because speed and caution trade off.
You want their energy — they’re your pilot group and your best demonstrators. Aim it at building things other people can use, and make “someone else can run this” the standard they’re measured against.
Neither loud group is the real target. The quiet majority tries it twice, hits friction, and silently reverts. That’s why the owner exists and why office hours matter more than any launch email.
Measuring honestly
Measure so you can steer, and so you can answer leadership without overclaiming. Both are much harder if you never took a baseline — so take one before anything launches.
Worth tracking: active usage (distinguishes access from adoption), artifacts still in use at 30 days (the number that survives scrutiny), self-reported confidence from the same five questions before and after, and per-task estimates from the person who actually does the task.
The survey is five questions, 1–5, anonymous, and the only rule that matters is that the wording never changes between runs — a survey you reword is a survey you can’t compare.
Verify before you report. Anything you’re about to cite has to have been checked against more than one lucky example, and the output-evaluation checklist is how. Run it against a typical case and at least one awkward one — missing data, an edge case, an unusual input. The awkward case is where the real defects live, and the failure mode you’re designing against is not output that looks wrong. It’s output that looks right.
Be honest about what this isn’t. Self-reported confidence is not productivity. Time saved on a task is not time returned to the business — the saved time goes somewhere, and it’s usually not tracked. Studies claiming dramatic organization-wide gains are mostly measuring task-level speed under controlled conditions.
If you inflate the claim, the first person who checks discredits the whole effort, including the parts that genuinely worked. The version that holds up sounds like this:
Twelve people have access; nine used it last month. Seven things were built in the workshop; five are still in weekly use. The clearest case: Maria’s weekly report, which took about 90 minutes and now takes about 20 — her estimate. We haven’t tried to measure aggregate productivity, and I’d be skeptical of anyone who claims to.
Specific, checkable, and it makes whatever you ask for next credible. There’s a fill-in version in the operating templates, and a complete one — including the ask it was built to support — in the worked example.
The counterintuitive move: volunteer a failure. Report something your verification caught before it shipped. It costs you nothing — you found it and fixed it, which is the system working — and it is the single thing that makes the rest of the numbers believable. A report where everything went well reads like marketing, and gets discounted like marketing.
That matters more than it sounds, because a leader who wants a productivity number will get one from somewhere. Better it comes from you, honestly bounded, than from a vendor deck. “We haven’t measured that, and here’s what we have measured” is a stronger position than a figure you can’t defend — but only if the things you did measure are specific enough to check.
Your own AI use
The fastest way to kill adoption is to sponsor it without doing it. Teams read that instantly: this is a thing that happens to us, not with us.
- Use it weekly on your own work, visibly. Meeting notes, first drafts, long documents. Say when you did.
- Share your failures out loud. The most useful sentence a manager can say is “I tried it for this and it was worse than doing it myself.” It gives everyone permission to report honestly, which is the only way you find out what’s actually working.
- Don’t delegate the learning. If your entire understanding arrives through one enthusiastic person, you can’t evaluate their claims.
- Keep the human parts human. Feedback, hard conversations, praise that’s meant to mean something. The point of that work is that you spent the time.
- Model the verification you’re asking for. Send around AI output with an error in it and the policy is dead the same day.
You don’t need to be the most skilled user on the team. You need to be visibly a user, and honest about where it falls short.
Back to the start
That’s the toolkit: the overview, the ground rules, the pilot, the workshop, and this.
If you’d like to see all of it landed at once — policy, scoring, pilot, verification, ownership record, and leadership update, filled in for one team, along with the four things that went wrong — that’s the worked example.
The takeaway: the workshop is the visible part, but the four weeks after it are what determine whether any of it was real. Publish, hold office hours, name owners, and report numbers you’d be comfortable having someone check.