# What ROI Does AI Actually Produce in a Service Firm?

Published: 2026-07-10T13:00:00.000Z · Updated: 2026-08-01T13:24:16.000Z · Author: Rod Amora · Canonical URL: https://rodamora.com/blog/what-roi-does-ai-actually-produce-in-a-service-firm/

> MIT says 95% of AI pilots return nothing. Wharton says 75% of firms already see positive ROI. Both are true, and neither answers the owner's question. What AI returns in a service firm, in what order, and on what timeline, from watching hundreds of firms make the climb.

Four AI ROI numbers made headlines between mid-2025 and mid-2026. [MIT researchers found](https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf?ref=g.rodamora.com) that 95% of generative AI pilots return nothing. [Wharton found](https://ai.wharton.upenn.edu/wp-content/uploads/2025/10/2025-Wharton-GBK-AI-Adoption-Report_Full-Report.pdf?ref=g.rodamora.com) that 75% of firms already have a positive return. [A Microsoft-commissioned IDC study](https://info.microsoft.com/ww-landing-business-opportunity-of-ai.html?ref=g.rodamora.com) says every dollar invested returns $3.70. [McKinsey found](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai?ref=g.rodamora.com) that 39% of organizations see any profit impact at all, and most of those put it under 5% of earnings. All four were published within months of each other, about the same economy.

So here is the honest answer to the question in the title. AI installed as a tool returns close to nothing. AI adopted as a rebuild of how the firm runs returns better retention, better margins, and revenue that grows without matching headcount, on a schedule I have watched run from 6 months to well over 18. And by the time those numbers move, you can no longer cleanly credit the AI, because the rebuild changed everything around it too.

## Why is every AI ROI number different?

The headline figures measure different things and call them all ROI. One counts pilots, another counts deployments. Most ask about perception rather than measured financials. And they average together firms at different points of the same climb, which is how one survey finds [95% failure](https://rodamora.com/blog/your-ai-is-just-a-better-google-search/?ref=g.rodamora.com) while another finds 75% success.

Headline number

Who published it

What it measured

95% of pilots return nothing

MIT NANDA, July 2025

300 public deployments plus interviews, scored on P&L impact

75% of firms see positive ROI

Wharton, October 2025

\~800 enterprise leaders, self-reported returns

$3.7 back per dollar

IDC, commissioned by Microsoft

4,000+ leaders, self-reported estimates

39% see any profit impact

McKinsey, November 2025

1,993 respondents, EBIT attribution

The deeper problem is that almost nobody is measuring anything. In professional services, [Thomson Reuters found](https://www.thomsonreuters.com/en-us/posts/technology/ai-in-professional-services-report-2026/?ref=g.rodamora.com) that only 18% of firms track whether their AI use returns anything. [S&P Global found](https://www.spglobal.com/market-intelligence/en/news-insights/research/ai-experiences-rapid-adoption-but-with-mixed-outcomes-highlights-from-vote-ai-machine-learning?ref=g.rodamora.com) the share of companies abandoning most of their AI initiatives jumped from 17% to 42% in a single year. When four out of five firms have no baseline and no metric, the survey answer is a feeling.

Feelings about AI productivity turn out to be unreliable in a specific direction. In [a 2025 randomized trial by METR](https://arxiv.org/pdf/2507.09089?ref=g.rodamora.com), experienced developers expected AI to make them 24% faster. Afterward, they estimated it had made them 20% faster. The measured result: 19% slower. A survey built on self-reported gains can produce any number you like.

## The ROI that is real: what controlled studies found

Task-level gains are real and repeatable. The strongest evidence for service work is [a randomized trial with 758 BCG consultants](https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838?ref=g.rodamora.com). Consultants using AI completed 12.2% more tasks, finished them 25.1% faster, and produced work rated more than 30% higher in quality. [A study of 5,172 support agents](https://academic.oup.com/qje/article/140/2/889/7990658?ref=g.rodamora.com) found 15% more issues resolved per hour, with the biggest gains going to the newest agents. [A controlled study of writing tasks](https://www.science.org/doi/10.1126/science.adh2586?ref=g.rodamora.com) found the work done 40% faster at higher quality.

The BCG study carries a warning in the same data. On tasks outside the model's competence, consultants using AI were 19 percentage points more likely to get the answer wrong. The gains are task-shaped. They do not generalize to "everything my firm does."

A task gain is the vendor's number, not yours. For the vendor, the distance from AI use to revenue is zero. For your firm, the gain has to survive review, approval chains, rework, scheduling, and your pricing model before it reaches profit. Every step it passes through takes a cut. A 25% faster first draft that then waits three days for a partner's review changed nothing the client or the P&L can see.

![Drawing of a large blue block shrinking as it passes through a row of gates and a bridge, ending as a small block far away](https://rodamora-uploads-347628392068-sa-east-1-an.s3.sa-east-1.amazonaws.com/illustrations/2026-07-31-ai-roi-service-firms/fig-illus-1.jpg "The distance between a task gain and your profit")

## Where the gain dies between the task and the P&L

[Freed time evaporates](https://rodamora.com/blog/the-extra-time-ai-buys-you-is-already-gone/?ref=g.rodamora.com) unless someone decides in advance where it goes. I've written about this pattern before, and [The Four Numbers](https://rodamora.com/four-numbers?ref=g.rodamora.com) covers how to catch it, so here is the short version. When AI cuts a three-hour task to 40 minutes, the recovered time does not sit and wait for instructions. It gets absorbed by the work around it, and the firm ends up with the same costs, the same revenue, and a quieter afternoon.

[One 2026 benchmark of 255 enterprise leaders](https://optimumpartners.com/insight/59-of-your-ai-productivity-will-never-reach-revenue/?ref=g.rodamora.com) put a number on the leak: the average company converts 41% of its AI time savings into measurable business value. The rest never reaches the P&L. People spend the saved hours hand-checking AI output. Approval chains built for the old pace stay unchanged. Recovered time drifts into meetings and email. The top performers in that benchmark converted 71%, and they got there by redesigning the workflow before deploying the AI, not after.

That matches what I've watched across hundreds of service firms adopting AI through a franchise network's delivery data. The first-stage gain is small to none if the processes around it aren't lined up to support it. You need a strategy that guides people to spend the recovered time improving the rest of the work: better output overall, faster delivery overall, not a faster version of one task. Without that, the firm has [a production gap](https://rodamora.com/production-gap?ref=g.rodamora.com), a real gain that is invisible everywhere it counts.

## How long does AI take to pay off in a service firm?

The return on AI lives in the climb from faster tasks to [a rebuilt delivery model](https://rodamora.com/blog/what-is-an-ai-native-service-business/?ref=g.rodamora.com), and the climb is painful. It takes coordination, buy-in from leadership top down, and evangelists inside the team. It takes a strategy that answers the fear every rollout triggers, the quiet "I will be replaced." Anxious teams do not adopt, they comply. And someone has to manage the new time and output, or the firm is just wasting time faster.

The timeline varies more than any vendor will tell you. In franchise units that were aligned on one goal, speeding up delivery, I've watched the climb take about six months. In firms where employees were not engaged, or where a bad first implementation did more damage than it helped, the same climb took over 18 months, and some are still on it. [Deloitte's survey of 1,854 executives](https://www.deloitte.com/nl/en/issues/generative-ai/ai-roi-the-paradox-of-rising-investment-and-elusive-returns.html?ref=g.rodamora.com) found most organizations report a 2-to-4-year payback on AI, against 7-to-12 months for typical technology, and only 6% see ROI inside a year. [Bain found](https://www.bain.com/insights/your-ai-budget-is-growing-your-returns-arent-heres-why/?ref=g.rodamora.com) that 40% of enterprises that measured their AI savings came in below 10%, after targeting 11% to 20%. Slow paybacks and missed targets are what a rebuild looks like from the inside.

What arrives when the climb works is worth the wait. Across the firms I've watched get there: retention improves, margins improve, and revenue grows without headcount growing with it. The numbers move in [Ladder](https://rodamora.com/delivery-model-ladder?ref=g.rodamora.com) order. Cost to deliver drops first. Retention and margin follow when the freed capacity gets aimed at the client. Revenue per person moves last, and it is the one that tells you the model changed.

## An AI ROI you can isolate is a sign you did too little

The firms that get the best results from AI are the ones least able to tell you its ROI. That sounds backwards until you watch what a correct implementation involves. The firms doing it right are also fixing their processes, documenting how work gets done, centralizing information that lived in people's heads, and making their data reachable. That shows up as better proposals, better sales presentations, quicker follow-ups, better meeting briefs, better client planning. Which line of that improvement belongs to the AI? There is no clean answer, and the owners who run these firms stopped asking.

A firm that can isolate its AI return usually can, for a bad reason. It installed a tool, changed nothing around it, and measured the tool. The attribution is clean because the change was small. That firm's number is also, in most cases, close to zero.

This is an old pattern with a hundred years of evidence behind it. Electric motors were commercially available in the 1880s. US factory productivity did not jump until the 1920s. [Paul David's research on that gap](https://www.jstor.org/stable/2006600?ref=g.rodamora.com) found the reason: factories were built around a central steam shaft, and electricity paid off only after plants were redesigned around the new power source. The return came from the reorganization, not the motor.

![Three-panel drawing: a steam engine on a central branching shaft, an electric motor on the same shaft, then small motors spread out](https://rodamora-uploads-347628392068-sa-east-1-an.s3.sa-east-1.amazonaws.com/illustrations/2026-07-31-ai-roi-service-firms/fig-illus-2.jpg "Swapping the power source pays only after the layout changes")

Erik Brynjolfsson calls the modern version [the productivity J-curve](https://www.aeaweb.org/articles?id=10.1257/mac.20180386&ref=g.rodamora.com): early in a general-purpose technology, firms pour effort into process redesign, training, and data, none of which shows up in output statistics, so measured returns look terrible right before they turn. Robert Solow's 1987 line described the same moment for computers: you can see the computer age everywhere but in the productivity statistics.

We are at that point in the curve with AI. Asking "what is the ROI of AI" in 2026 is asking what the ROI of the steam engine was at the start of the industrial revolution. The question felt reasonable and had no useful answer. AI is not a tool that you install in your company. It is a philosophy on how to build companies.

## Waiting at the tool stage is not a neutral position

You could read the last section as permission to wait. It is the opposite, because the gains you bank at the tool stage get competed away. Clients can see the inputs shrinking. The [Financial Times reports](https://www.ft.com/content/8318a754-7b63-4f1f-8978-7691e7ed3511?ref=g.rodamora.com) clients now asking a simple question: if AI allows projects to be completed faster and with fewer people, why should fees remain the same? Some are demanding automatic discounts where AI is used. About a quarter of McKinsey's global fees now come from outcome-based contracts, a model its own partners describe as taking hold only in the last few years.

Legal shows where this goes. [A survey of roughly 650 in-house counsel](https://news.bloomberglaw.com/in-house-counsel/ai-does-little-to-reduce-law-firm-billable-hours-survey-shows?ref=g.rodamora.com) found nearly 60% see no noticeable savings from their law firms' AI use, and only 13% saw fewer billable hours. Neither side knows how to price the change yet. That standoff will not hold. As [one analysis of consulting pricing](https://www.afr.com/companies/professional-services/why-artificial-intelligence-won-t-eliminate-the-value-of-consulting-20260422-p5zpyy?ref=g.rodamora.com) put it, top-tier firms can defend fees on reputation, but mid-market pricing was built on visible effort, and the visible effort is shrinking. Meanwhile AI lowers the cost of entry for new competitors who never priced on effort at all.

A firm that stays at the tool stage keeps its old cost structure while the market reprices around firms that rebuilt theirs. The efficiency you didn't capture becomes the discount your client asks for.

## The number I refuse to give you

I could invent a headline percentage here. I won't. Nobody has a good number for what AI returns in a service firm, and the honest ones admit it. The return differs by service line, by billing model, by how much of the firm's knowledge is written down, and by how the climb is run. When the pattern data I work with supports a real figure, I will update this post with it. Not before.

What I can give you is what the evidence supports doing while the number doesn't exist:

- **Pick one workflow, not six.** [BCG found](https://www.bcg.com/publications/2025/closing-the-ai-impact-gap?ref=g.rodamora.com) the companies getting real value run 3.5 use cases against an average of 6.1. Focus wins.
- **Budget small, train heavy.** The same BCG research puts the value split at 10% algorithms, 20% technology and data, 70% people and process. Spend like that split is true, because it is.
- **Decide where the freed hours go before the tool ships.** More capacity for delivery, better deliverables, faster turnaround. Pick one and assign it. Unassigned hours leak.
- **Name the number you expect to move.** Cost to deliver, retention, margin, or revenue per person. Then watch it quarterly. [The Four Numbers](https://rodamora.com/four-numbers?ref=g.rodamora.com) covers how.

Measure the business, not the AI. A firm that does only that is ahead of the 82% that track nothing.

Nobody could price the steam engine in 1785, and it did not matter. The factories that rebuilt around it anyway owned the next fifty years. The service firms rebuilding around AI now are making the same bet, and from what I've watched so far, it is paying the same way.

## Common questions about AI ROI

**What is a good AI ROI benchmark?**

There isn't a defensible one for service firms yet. Published figures run from 95% failure to $3.7 back per dollar because they measure different things on different samples. Until the data matures, benchmark against your own baseline: pick one of the four numbers and compare it quarter over quarter.

**How do I measure AI ROI in a service firm?**

Measure the business, not the AI: cost to deliver, retention, margin, and revenue per person. Hours saved is an estimate that rarely reaches the P&L. [The Four Numbers](https://rodamora.com/four-numbers?ref=g.rodamora.com) covers how to track each one.

**Why do most AI pilots show no ROI?**

Because the firm installs a tool and changes nothing around it. The freed time gets absorbed by adjacent work, review chains keep their old pace, and the gain never lands on a number the P&L reports. [The Production Gap](https://rodamora.com/production-gap?ref=g.rodamora.com) names the specific failure modes.
