What ROI Does AI Actually Produce in a Service Firm?
I've watched AI help service firms improve retention and margins, but the return came with changes to delivery. A faster task is only part of the calculation.
If you're asking what AI will return in your service firm, I can't give you an honest percentage that applies to your company. I can tell you what I've watched improve, what had to change before that happened, and why buying the tool was only a small part of it.
In firms I've watched make that change, retention improved, margins improved, and revenue grew without headcount growing with it. But those firms were changing how they delivered the service, not just giving people a faster way to do one task.
And by the time the business numbers moved, they had also improved their processes, written down more of what they knew and made that information easier to find. So separating the return from AI from the return on all that other work became difficult.
That doesn't mean you should stop asking where the money went. It means you need to be clear about what you're measuring, because the AI ROI numbers people put in front of you often answer a different question.
Take the surveys. Wharton's 2025 report says three out of four leaders see positive returns on their generative AI investments, while McKinsey's 2025 survey says 39 percent of respondents attribute any impact on earnings before interest and taxes to AI.
Those aren't two researchers looking at the same set of accounts and disagreeing about the answer. They're asking different people different questions, and a positive return on one use of AI is not the same thing as a visible change in the whole company's profit.
The same caution applies to the Microsoft-commissioned IDC study reporting an average return of $3.70 for each dollar invested in generative AI. That's a vendor-commissioned survey result, not a promise that your next dollar will produce the same return.
You may also have seen the 95 percent failure headline attached to MIT's NANDA project. I wouldn't use that as the odds your project will fail, its methodology has been contested, and a widely repeated headline doesn't become a reliable benchmark for your firm because it has a university's name beside it.
There is a more basic measurement problem too. In Thomson Reuters' 2026 professional services survey, only 18 percent of respondents said they knew their organization was tracking the return on AI tools.
That doesn't prove the other 82 percent track nothing, some respondents may simply not know. But it does mean you should be careful about treating a person's impression of the return as if someone had checked the financial result.
Let me give you a concrete example of how far those impressions can drift. In METR's early-2025 trial, experienced open-source developers expected AI to make them faster, and afterward they still thought it had made them about 20 percent faster, but the measured work took 19 percent longer.
That's a historical result in a particular setting, not a verdict on today's coding tools. METR now explicitly says those results no longer reflect current impact, and its February 2026 follow-up found signs of improvement while explaining why changes in who joined the study made the size of the gain hard to estimate.
The reason I keep the older example is much narrower, people can be wrong about how much time they saved even when they've just finished the work. You need something better than that impression before you build a budget around it.
Now, none of this means the task gains aren't real. In the study of 758 BCG consultants, people using AI finished suitable tasks 25.1 percent faster and produced work rated more than 30 percent higher in quality.
But the same study found that, on a task outside what the model could handle, AI users were 19 percentage points more likely to get the answer wrong. So even inside one study, the answer depended on what work you gave the system.
And a published study of 5,172 customer-support agents found a 15 percent increase in issues resolved per hour, with larger benefits for less experienced workers. That's useful evidence for that kind of work, it still doesn't tell you what happens to the profit of a service firm with a different delivery process.
Think about a report your team prepares for a client. Writing the first draft faster helps, but somebody still has to check the numbers, fix the mistakes, approve it and send it, and the price you charge may still depend on how many hours went into it.
If the draft gets done sooner but sits in the same review queue for three days, the client doesn't get it sooner. If checking it takes more work than before, some of the writing gain has already been spent before you've delivered anything.
And please don't answer that by telling people to check less. The report still has to be right, what you're trying to find out is whether the new way of doing the whole job is better.
There is another place the return can disappear, even when you genuinely free up time. I've watched the early gains stay small or absent because nobody had prepared the surrounding process or decided what the team should do with the recovered hours.
The person finishing a task sooner isn't automatically the person who can decide which other work matters most to the company. That's a management decision, and if you leave it to them to guess, you haven't finished the implementation.
I've written about where the extra time goes, but the practical point here is to decide what that time is for before you count it as a return. It can go into more delivery, fewer mistakes or better care for existing clients, whatever your firm actually needs, but someone has to assign that work and check that it happened.
This is part of why the change takes longer than the software setup. In franchise units where everyone wanted AI to speed up delivery, I've watched it take about six months, and in firms where people weren't engaged or a bad first implementation had done damage, I've watched it take over 18 months.
Those are timelines I've observed, not a payback guarantee. Some firms were still working through the change, and getting people to use something after a bad first experience was part of the work they had to pay for.
You need leadership to support it, people inside the team who help it along, and a clear account of how the work is changing. If the team thinks the purpose is to replace them, you can't expect a tool demonstration to settle that fear for you.
They need to know what they're responsible for, what they still have to check and what the company expects them to do with the new capacity. You can ask them to help work that out, but you can't hand them a login and call the uncertainty their resistance.
As firms get further into this, the return spreads across work that wasn't in the original tool pitch. Better information helps someone prepare a proposal, brief a meeting, follow up with a prospect and plan the next piece of client work.
And when the proposal improves, how much came from the AI and how much came from finally having the client's history in one place? You can measure that the process improved without pretending you can separate every cause cleanly.
I wouldn't take that as permission to stop measuring individual tools, either. A narrow tool can have a perfectly useful, measurable return, and an implementation that is hard to measure can still be a bad investment.
My point is that demanding one isolated AI percentage for a whole-company change can push you toward measuring the easiest part instead of the part you actually wanted to improve.
There is an older version of this problem in Paul David's work on factory electrification. Replacing the power source wasn't the whole change, factories also had to reorganize work around what electric motors made possible.
And Erik Brynjolfsson, Daniel Rock and Chad Syverson's research on the productivity J-curve explains how spending on things like new processes and skills can be poorly captured in the early productivity figures. The work is real, but the accounts don't necessarily show its benefit when you're doing it.
That gives us a reason to expect a difficult early stretch, it doesn't prove every AI project is about to pay off. You still need to distinguish a firm that is doing the necessary work from a firm that keeps spending and calling the lack of progress a transition.
And you also need to look at how you charge. If you bill by the hour and the same deliverable now takes fewer hours, you can improve the work and reduce the revenue from that engagement at the same time.
Under a fixed fee, the same saving can help your margin if quality holds and the other costs don't eat it. So two firms can use the same tool, save the same amount of time and end up with different business results because their prices work differently.
Clients have a view on this too. Bloomberg Law's reporting on an ACC and Everlaw survey found that nearly 60 percent of the in-house counsel surveyed saw no noticeable savings from their outside firms' AI use.
My read is that you shouldn't assume a client will keep paying for visible effort that is getting smaller. That doesn't mean every client will demand a discount, or that every service should be priced by outcome, it means the pricing decision belongs in your plan rather than arriving as a surprise after the task gets faster.
So here is how I would approach the money question in your firm. Pick one repeated piece of work, write down how it runs today and count the whole cost of getting an acceptable result, including the person who reviews it and anything that comes back to be done again.
Then include the cost of the change, the tool, the setup, the training and the time people spend getting the new process to work. Those hours don't become free because they're on an existing employee's calendar.
Decide which business number you expect to move and why. The Four Numbers are cost to acquire a client, cost to deliver, retention and price, and you don't need all four to move at once to learn something useful.
If you're trying to reduce delivery cost, start there, and keep checking quality beside it. If the aim is better client service, decide what improvement you'd expect to see before you start calling every saved minute profit.
Then look at margin and revenue per person over time to see whether those changes reached the business. A quarterly review can help you keep that longer view, while the people running the work need a closer check on errors and rework as the new process settles in.
You may find a good task-level gain that hasn't reached the business yet. You may find that the tool isn't earning its cost, or that the gain is real but your pricing model gives it to the client, and those are different decisions.
I can't promise you the return will arrive on the same schedule I've seen elsewhere. What I would ask you to do before the next rollout is sit down with the person who owns the work, agree what should improve, count what changing it will cost, and decide how you'll know whether that improvement happened.


