Calculator · Rod Amora ·
AI ROI Calculator for Service Firms
Most AI ROI calculators multiply hours saved by a billing rate and call it done. This one starts from controlled studies, then walks the ceiling through the five places value leaks in a service firm.
AI ROI leak calculator
Your delivery capacity and work mix
Pick the closest delivery model to pre-fill a typical work mix, then adjust the sliders to match your firm. Task-level gains come from controlled studies. Time not allocated below is synchronous and judgment work: meetings, workshops, on-site, courtrooms, where AI compresses little.
Your operating model: where gains survive or die
Five statements about how your firm runs today. Answer honestly: every "not yet" is a modelled leak, not a moral failing. Measured pass-through data says the average firm keeps 41% of its task gains; the other 59% never reaches the P&L. This calculator allocates that unpassed-through value across five provisional leak categories so you can see where to act.
Defined which outputs need line-by-line review versus spot-checks
BetterUp / Stanford: 40% of desk workers get AI output needing about two hours of rework
Review and approval steps were redesigned since adopting AI, not just drafting steps
Faros: teams merging 98% more pull requests saw review time rise 91%
Written plan for where freed hours go
Optimum Partners: most of the 59% that evaporates never reaches the P&L
Pricing keeps the gain through fixed-fee, retainer, or outcome work, not pure time-billing; retainer scope has not quietly grown
ACC / Everlaw: about 60% of in-house counsel see no savings; 13% see fewer billable hours
Firm knowledge is centralized where AI can reach it
The better-Google-search argument: purchased tools cannot close a context gap they cannot reach
What reaches your P&L
Set your mix and operating model above. The live ranges appear here: ceiling versus what reaches the P&L. Always ranges, never a single ROI percentage.
Capacity value, not profit. Dollar ranges update with your mix, answers, and blended cost.
How fast the climb runs: three enablers
The leak profile above is what an audit finds for real, with your numbers and your workflows in the room. Here is how I work.
The model, in full
Everything the calculator above computes is written out here: every gain range, leak weight, and pass-through anchor, with the source it came from. The tables in this article render from the same module the tool runs on, so what you audit is what runs.
The average firm converts 41% of AI time savings into measurable business value; this calculator estimates your ceiling from controlled studies, then shows where the other 59% leaks.
Why this calculator refuses to give you an ROI percentage
Between mid-2025 and mid-2026, four serious research teams published four contradictory headline numbers about AI returns. MIT's NANDA project found that 95% of generative AI pilots return nothing. Wharton found that 75% of firms already report a positive return. A Microsoft-commissioned IDC study put the figure at $3.70 back per dollar invested. McKinsey found that 39% of organizations see any profit impact at all, and most of those put it under 5% of earnings.
They are not measuring the same thing. One counts pilots that returned nothing. Another surveys firms that already report a positive return. One multiplies dollars invested by a claimed return. Another asks whether organizations see any profit impact at all. The headlines disagree because the questions, samples, and success definitions disagree. The point is not to pick a winner. The point is that a single ROI percentage cannot settle the question for a service firm.
So this calculator will not hand you a single ROI percentage. No honest one exists for a service firm yet. What exists is better raw material: controlled, task-level measurements of what AI compresses, and measured data on how much of that compression survives contact with a real firm. The refusal is the method. You get a ceiling, a leak profile, and a range.
The gains that are real
The task-level evidence is the most usable part of this field for a calculator, because it comes from randomized controlled trials rather than surveys. Consultants, writers, and support agents using AI finished measured work faster, at equal or better quality, under experimental conditions. The table below is the complete set of gain ranges the calculator uses, with the study behind each range.
| Work category | Task-level gain | Evidence |
|---|---|---|
| Drafting & writing | 25–40% | BCG consultants RCT (25.1% faster); writing-tasks RCT (40% faster) ( BCG consultants RCT ; Noy & Zhang RCT ) |
| Research, analysis & review of material | 12–25% | BCG RCT (12.2% more tasks, 25.1% faster) ( BCG consultants RCT ) |
| Client comms & support | 10–15% | 5,172-agent support study (15% more issues per hour) ( Brynjolfsson et al. ) |
| Structured production work | 10–30% | Scoped-task experiments (~55% faster) versus METR (19% slower for expert complex work); mixed band ( Copilot RCT ; METR RCT ) |
| Admin, data entry & processing | 5–15% estimate | Weakest evidence; this range is an estimate, not a controlled finding (Model estimate) |
| Synchronous & judgment work | 0% | Meetings, workshops, on-site, courtroom. No controlled study shows compression here, and this is where the new pile-ups form. |
Two honest caveats sit inside that table. The structured-production band is wide on purpose: one controlled experiment found developers about 55% faster on a scoped task, while METR's 2025 trial found experienced developers 19% slower on complex work they already knew well. If your production work is expert-heavy, slide that category down. And the admin range is an estimate, not a controlled finding; the tool says so wherever it appears.
Treat every range in this table as the vendor's number, not yours. For the vendor, the distance from "the task got faster" to revenue is zero, so the task gain is the whole story. Inside your firm, that gain still has to survive review, approvals, scheduling, and your pricing model before the P&L can see it. The ceiling is real. What matters is how much of it you keep.
The eight delivery-model presets in stage 1 are my informed defaults from watching service firms adopt AI: starting points, not findings. Adjust the mix until it looks like your firm. The slider is the truth; the preset is a guess at it.
Where the value leaks
Between the task and the P&L, the gain passes through your operating model. The calculator prices five leaks. The sources below support the leak mechanisms. The exact weights are provisional model judgment allocations for distributing loss across those leaks, not measured causal shares from a study. They sum to one, and they stay marked for a later sanity check against franchise-network observation.
| Leak | Weight | Evidence |
|---|---|---|
| Validation | 22% | BetterUp / Stanford: 40% of desk workers get AI output needing about two hours of rework ( BetterUp/Stanford ) |
| Approval chain | 24% | Faros: teams merging 98% more pull requests saw review time rise 91% ( Faros AI benchmark ) |
| Reallocation | 22% | Optimum Partners: most of the 59% that evaporates never reaches the P&L ( Optimum Partners ) |
| Pricing | 18% | ACC / Everlaw: about 60% of in-house counsel see no savings; 13% see fewer billable hours ( ACC/Everlaw survey ) |
| Context | 14% | The better-Google-search argument: purchased tools cannot close a context gap they cannot reach ( MIT NANDA report ) |
Validation (22% of the leak model)
AI produces a plausible draft in seconds. Someone then has to find out whether it is right. Often that is a senior reviewer, reading line by line because nobody decided which outputs need that and which need a spot-check. BetterUp Labs and Stanford found that 40% of desk workers receive AI output shoddy enough to need about two hours of rework per incident. The draft was free. The checking is not. The weight is 22% because unscoped review sits on every AI-touched deliverable, so the rework tax hits the whole mix. That is an allocation judgment, not a measured share. The fix is classification, not more care: sort tasks by check speed and mistake cost, and write the answer down. AI only counts when you can prove it in the work. That is the idea behind Proofwork.
Approval chain (24% of the leak model)
Drafting got faster. The approval steps around it did not. Faros measured engineering teams merging 98% more pull requests while review time rose 91%. The constraint moved from writing to reviewing, and the queue grew at the new bottleneck. This leak carries the heaviest weight at 24% because the bottleneck moves to a step every deliverable still has to pass, so the queue tax compounds. That ranking is model judgment, not a causal percentage Faros measured.
Reallocation (22% of the leak model)
Suppose a consultant finishes a draft in 40 minutes instead of three hours. Where do the recovered hours go? Often, nowhere anyone chose. They dissolve into the surrounding work, and the firm ends the quarter with the same costs, the same revenue, and a quieter afternoon. The Optimum Partners benchmark puts unassigned time at the heart of the 59% that never reaches the P&L. The weight is 22%, matching validation, because unassigned hours can erase the kept gain even when review and approvals improve. The benchmark supports the mechanism; the exact split is still an allocation. Freed capacity needs a written destination before it exists, or it evaporates.
Pricing (18% of the leak model)
If you bill hours, a task that takes half the time bills half the fee. The gain left your firm entirely and arrived at your client as a discount nobody negotiated. A survey of roughly 650 in-house counsel by the ACC and Everlaw found about 60% see no savings from their law firms' AI use, and 13% report fewer billable hours. Firms that keep the gain price the outcome: fixed fee, re-scoped retainer, outcome work. The weight is 18%, below the first three, because not every service line bills pure hours and pricing moves on a slower cycle than daily review. Pure time-billing can still hand the gain to the client, so the allocation stays material. It is not a measured share of firm-wide loss.
Context (14% of the leak model)
AI that cannot reach your firm's knowledge is a better Google search. The context it needs (client history, past deliverables, pricing decisions, house style) lives in inboxes, heads, and shared drives the tool cannot see. When the official tool lacks that material, people stop trusting it for real work and route around it, which is how shadow AI starts. MIT's NANDA report is the citation the model keeps for this gap. This leak carries the smallest weight at 14% because a context gap mostly lowers the quality of the ceiling rather than the pass-through rate after a usable draft exists. It still caps everything above it. The percentage is an allocation judgment, not a measured causal share.
How the math works
Stage 2 scores your five answers. Not yet counts 0, Partly counts 0.5, Yes counts 1. Multiply each answer by its leak's weight and add the five products. That score, between 0 and 1, is the share of the gain your operating model currently keeps.
The pass-through midpoint is then 0.27 plus the score times (0.71 minus 0.27). A firm that answered "not yet" to everything sits at the floor. A firm that answered "yes" to everything sits at the top anchor.
Two of the three numbers in that formula are measured, not chosen. The 41% anchor is the average firm in the Optimum Partners benchmark of 255 enterprise leaders. The 71% anchor is the top 7% in the same data, the firms that redesigned the workflow before deploying the AI. Your score interpolates between a firm with no practices in place and the best measured performers.
The floor is the one judgment call in the model, and I want it admitted rather than buried. 0.27 is not a measured anchor. It is my estimate of a below-average firm with none of the five practices in place, set below the measured average because a firm that answered "not yet" five times is worse than average by construction. If you disagree with it, nothing else breaks. Slide your mental floor and every range the tool shows moves with it.
The output is always a range. The midpoint carries a band of 6 points either side, clamped between 15 and 80. The input gains are ranges and the benchmark is a distribution, so printing a single point would be fake precision. The waterfall uses midpoints only to keep the picture readable; the ranges in the verdict are the answer. The dollar line multiplies hours by your blended cost over 46 working weeks and stays labeled capacity value, not profit, because freed hours become margin only when someone points them at billable work, retention, or offers.
What the timeline depends on
The observed range is 6 to 18 months, and sometimes longer. Across the firms I have watched adopt AI through a franchise network's delivery data, aligned units climbed in about six months. Firms with passive adoption took longer, sometimes well over eighteen months, and some are still climbing. Deloitte's survey of 1,854 executives matches the slow end: only 6% of organizations see any ROI inside a year.
Speed comes from three enablers, and they are unglamorous. One named owner for AI adoption. Funded learning time on the calendar. A short written sandbox policy, so people know what is allowed before they experiment. The calculator maps those counts onto bands the model uses for planning: all three in place, 6 to 9 months; two, 9 to 12; one, 12 to 18; none, 18 or more, if at all. Those bands are planning profiles from observation, not a proven month-by-month law. The enabler firms skip is the middle one: learning a new delivery method while still delivering the old way costs real hours, and unfunded hours do not happen. The Production Gap names that the unfunded hour, and it is a common reason firms stall.
What to do with your result
The verdict ranks your top three leaks and suggests one recovery move per leak. Sequence the moves in the order The Delivery Model Ladder pays out, not in the order that feels easiest.
- Validation and reallocation first. These are Stage 1 moves: they cost decisions, not rebuilds, and they can show up in cost to deliver inside a planning quarter when the firm actually follows through.
- Approval chain next. Redesigning the review step is a Stage 2 move, because it changes how the work itself flows, not just how fast one task runs.
- Pricing once the gain is real. Repricing is also Stage 2 territory, and doing it early, on a gain you have not captured yet, just renegotiates your revenue downward.
- Context throughout. Centralizing firm knowledge is slow groundwork for an AI-native Stage 3. Starting earlier leaves more calendar for the rest of the climb.
Whatever the calculator shows, the instrument that matters is not on this page. Pick one of the four numbers (cost to deliver, retention, margin, revenue per person), name which one your top leak is blocking, and watch it quarterly. The Four Numbers covers how to track each one, and it works whether or not a single AI workflow is involved. Measure the business, not the AI. A firm that does only that is measuring the right thing, including when other firms can quote a very precise ROI.
Sources
Every external source the model cites, in the order the model cites it. The four headline surveys from the first section are linked where they appear.
- BCG consultants RCT (N=758), Organization Science 2026
- Noy & Zhang writing-tasks RCT (N=453), Science 2023
- Brynjolfsson et al. support-agent study (N=5,172), QJE 2025
- GitHub Copilot productivity RCT
- METR developer RCT (2025)
- BetterUp Labs / Stanford workslop study (September 2025)
- Faros AI software-engineering benchmark
- Optimum Partners benchmark (N=255), May 2026
- ACC / Everlaw survey of approximately 650 in-house counsel
- MIT NANDA, The GenAI Divide report
- Deloitte's AI ROI survey (N=1,854 executives)
FAQ
What ROI does AI actually produce in a service firm?
It pays out in stages, not as one number. First the task gets faster, which is the gain vendors quote. Then, if the firm plugs its leaks, cost to deliver drops. Retention and margin follow when the freed hours get aimed at clients, and revenue per person moves last. There is no single honest percentage, which is why this page gives ranges and a leak profile instead.
Why do most AI ROI calculators mislead?
They multiply hours saved by a billing rate and stop there. That math assumes every saved hour reaches the P&L. Measured pass-through data says the average firm keeps only 41% of its task gains. A precise multiplication of a wrong assumption is still wrong. This page separates the sourced aggregate from five provisional leak categories the model uses to allocate the rest.
What is the 41% pass-through benchmark?
It comes from a 2026 Optimum Partners benchmark of 255 enterprise leaders. The average firm converts 41% of its AI time savings into measurable business value. The other 59% never reaches the P&L. That figure is an aggregate pass-through result, not a measured split into specific leak mechanisms. The top 7% of firms in that benchmark convert 71%, and they got there by redesigning the workflow before deploying the AI, not after.
How long until AI shows up in a service firm's P&L?
Plan on 6 to 18 months, sometimes more. The observed range runs from about six months in aligned firms to well over eighteen where adoption stays passive. Three enablers set the speed: one named owner, funded learning time on the calendar, and a short written sandbox policy. Deloitte's survey of 1,854 executives found only 6% of organizations see ROI inside a year.
How should a service firm measure AI ROI instead?
Measure the business, not the AI. Track cost to deliver, retention, margin, and revenue per person, before and after, and treat the AI as the explanation for movement rather than the thing being measured. The Four Numbers framework on this site covers how to track each one.