Don't Pick Your First AI Pilot by Performance Alone
Your best performer can belong in your first AI pilot, but performance alone isn't a reason to choose them. I look for people willing to test a different way of working, then check their feedback against the finished work.
When you bring an AI tool into a service firm, the instinct is to hand it to the person you trust most with a client, the senior, the best performer on the team.
That instinct is half right. You do want the tool near real work, but being good at the job doesn't tell you whether someone wants to help work out a new way of doing it.
At Berry, I never picked people for a pilot by who was best. I picked by who wanted it. If your best person wants to try it, include them, the mistake is picking them just because they're your best.
And by enthusiasm, I mean willingness to try a change and help get it working, not just liking AI. Let me tell you what that looked like in our sales-to-consultant handoff.
At Berry, I started with the people who wanted it
The handoff was supposed to tell the consultant what the client needed and the objective of the work. But the notes weren't doing that well, and there were two ways that played out.
In the worst case, the consultant trusted what was written, went into the first onboarding call with the client, and led the work in the wrong direction. The person taking over was working from a bad account of what the client needed.
Or the consultant went back to the sales conversation to understand it for themselves. They started watching the entire sales meeting, which sometimes ran over an hour, or reading the transcript so they didn't have to depend on the handoff.
So either the bad notes shaped the work, or somebody had to reconstruct the conversation before they could use it. That wasn't a consultant failing to care about the client, we were leaving them with a handoff they couldn't rely on.
The AI enthusiasts noticed that problem themselves. They created a template in ChatGPT, tested it, found that it worked, and spread it to the rest of the team.
The AI would read the sales call transcript, analyze the points they had specified in the template, and generate a standard response based on it. So the handoff had a format to follow, rather than depending on whatever somebody happened to put in the notes.
That standard improved the work. They were changing the handoff, not just trying to write the same notes faster.
Those people had already been trying AI tools and seeing better output in their own work. They gave fast feedback and were willing to reorganize how they worked, which helped us move from trying a tool to using it in the process.
That's the help I want in a first pilot, because an implementation often doesn't work at first. You have to try it, see what broke, and work on it for a while before you get real results.
What I've seen is that people who don't see the improvement or don't understand the technology go back to their old workflows more easily. That doesn't make them stubborn, they still have a job to finish, and if you want them to help test a new way of doing it, you have to make room for that work too.
Strong performance doesn't guarantee a benefit
The research gives you a reason not to assume your best people will benefit most, but it answers a different question from the one I answered at Berry.
A preprint about Alibaba's Taobao after-sales support followed 5,940 agents for four weeks, randomly giving some access to an AI assistant and keeping the others as a comparison group. On average, agents identified issues faster and customer ratings improved.
But when the researchers split agents by their customer ratings before the experiment, the bottom fifth gained 0.812 points on a five-point scale, while the top fifth lost 0.283 points. Those are effects of access compared with agents in the same performance group who weren't given it.
The top group also had a higher rate of customers coming back with the same issue within three days. And they got faster at identifying the issue, so looking only at that part of the job would have missed the worse result for the customer.
The easy read is that the strongest people got lazy and trusted the AI. But they used it less, and a later survey found more skepticism among high performers. The authors found evidence against over-trust and harder cases explaining the decline.
Their analysis points toward a problem with attention, the top group spent more time away from the current chat and answered more slowly while resolving the problem. That's a possible explanation, not a proven cause, because the researchers didn't randomly assign how agents divided their attention.
A separate study by Brynjolfsson, Li, and Raymond, published in the Quarterly Journal of Economics, found a 15% average increase in issues resolved per hour across 5,172 support agents. The most skilled workers had no productivity gain, alongside small but statistically significant decreases in resolution rates and customer satisfaction.
Neither result means strong people are the wrong people to include. And giving agents access at random in the Taobao experiment didn't test whether starting with enthusiasts works better, or show that unwillingness caused the harm. My reason for starting with enthusiasts comes from running the transition at Berry, not from a comparison those researchers made.
Check how the work feels and how it turns out
Even with people who want to try it, the transition can look worse before it looks better. What I've seen is that sometimes you work on two systems at the same time, you're still learning the new workflow, and the work takes longer during that switch.
That has taken two to four weeks in my experience, maybe more if we keep redesigning the process. I haven't seen strong performers get worse in the way the Taobao paper describes, what I'm describing is the transition mess, not a reason to assume its quality decline would disappear with more time.
So a bad second week doesn't have to be the final verdict. But learning a new process and sending worse work to a client are different problems, and you can't ask the client to accept the second because you're still doing the first.
I use two signals. The person's own feedback, and the hard numbers.
Ask whether the new way makes their day easier, and look at how long the whole task takes, whether quality improved, how much work came back to be done again. A faster step doesn't tell you whether the finished work is better, which is the distinction I make in The Extra Time AI Buys You Is Already Gone.
And the person's feeling isn't enough on its own either. In METR's 2025 study, 16 experienced open-source developers believed afterward that AI had cut their completion time by 20%, but the measured result was 19% longer.
METR's 2026 update weakly suggests a speedup with later tools, though changes in which developers and tasks entered the study make that estimate unreliable. The earlier slowdown isn't a verdict on today's tools, but the gap between how the work felt and how long it took is a reason to check both.
When feedback and the finished work improve together, or start moving that way, I take that as a reason to keep going. If people say work is getting harder, keep switching back to the old system, or aren't seeing their day improve, quit it or redesign or try a different approach.
Ask why they went back before treating it as a training problem. Running both systems can be part of learning, but if they still need the old one to get the job done, the switch isn't finished.
What to do in the first month
Start by asking who is already trying AI and wants to help test a change in the work. Keep people with a working method and no appetite for that testing out of round one, not forever, out of the prototype.
Before you start, write down how the work is going now. Pick measures you can compare later, time for the whole task, rework or comebacks, and a sign of quality the client would notice, rather than trying to remember the old process after you've changed it.
Then have a conversation and look at the finished work every week. Keep the results visible for each person, because a better group average can hide someone whose work got worse.
Use those checks to decide whether to widen the group, change the process, or stop. Don't make expansion automatic at the end of the month, and don't leave clients with worse work while you wait for it to end.
Where this stops
The Alibaba agents all had under one year on the job and handled several chats at once under fixed procedures. Being the best in that group isn't the same as being a fifteen-year veteran, and a firm where one person owns a client end to end has a different working day.
A review of human-AI experiments by Vaccaro, Almaatouq, and Malone also found that, on average, the combination beat either working alone when humans performed better than AI alone. That compares humans with AI, not employees with one another, so it doesn't settle what happened at Taobao or support a rule against including experts.
Two to four weeks is my experience of the transition, not a deadline by which your pilot has to work. And picking enthusiasts is how I started at Berry, not a guarantee that their results will carry over to everyone else.
If nobody is trying AI on their own, ask what would let a willing person start. If the missing piece is time, give them hours to learn inside the working week before expecting them to help redesign the job.


