Framework · Rod Amora ·
The Production Gap
The demo working and the P&L changing are two different events. Most AI projects inside service firms never make the trip between them. This is the taxonomy of how they die along the way.
I've watched hundreds of service firms adopt AI through a franchise network's delivery data, and the failures repeat. Almost none of them fail because the model was wrong. They fail because nobody closed the distance between a working demo and a workflow the business actually depends on.
That distance is the production gap, and it has a small number of recurring shapes. I count eleven as of July 2026. Each has an early-warning sign an operator can check before the quarter closes. Each has a documented way out. None of the eleven are about model quality. They are about ownership, measurement, and what happens to the hours AI frees up.
A firm can have excellent AI and still fall into every one of these. The taxonomy below maps each failure mode to where it tends to strike on the Delivery Model Ladder, so an operator can check their own stage against their own risk.
The Production Gap is a failure taxonomy. It names the recurring ways AI pilots stall or decay inside service firms before reaching the P&L. It is built from a growing case bank of adoption patterns. As of July 2026 it holds eleven named modes, each with an early-warning sign and a documented recovery path. It gives operators a diagnostic vocabulary, not a scorecard, so a stalled rollout gets named and fixed instead of quietly abandoned.
Why do AI pilots never make it to production?
Most AI pilots inside a service firm die of neglect. Rejection is rare. A proof of concept works. Everyone is pleased. Nothing makes it load-bearing. No one assigns an owner. No deadline forces a decision.
Across the firms I've observed, the most common stall point sits at Stage 0 or Stage 1 of the Delivery Model Ladder: individuals experimenting, or AI bolted onto an existing workflow without redesigning it. Both stages are comfortable, because nothing is visibly broken. The eleven modes below are what happens next, or instead.
The stall is now visible in population-level data, not just the case bank. In a 2026 NBER survey of nearly 6,000 executives, 89% reported no labor-productivity impact from AI over the past three years: a realized gain of +0.29%, against the +1.4% those same executives still expect from the next three. The number keeps moving to next year. And in an April 2026 Census Bureau working paper, 64% of AI-using firms made no organizational changes at all, and 66% use AI solely to augment existing tasks. That 66% is Stage 1, measured across the economy.
What are the eleven failure modes that stall AI adoption?
The table below is the summary. The sections after it go deeper on each mode.
| Failure mode | What it looks like | Early-warning sign | What survivors did |
|---|---|---|---|
| The Perpetual Pilot | Demos forever, never ships | No owner, no date by week 4 | Named an owner and an outcome before building |
| The Invisible Win | Hours drop, margin doesn't move | "Saved 15 hrs/week," org chart unchanged | Decided in advance where capacity goes |
| The Shadow Rollout | Private use, nothing repeatable | High usage, no two people work alike | Standardized one workflow at a time |
| The Unowned Adoption | Works at launch, then decays | Usage falls 3 weeks, unnoticed | Measured adoption weekly, like revenue |
| The Quality Debt Spiral | Volume up, review capacity flat | Reviewers bottleneck within a month | Redesigned review into the workflow |
| The Wrong-Number Dashboard | Tracks activity, not economics | No AI line in the P&L | Put AI results in the financials meeting |
| The Headcount Reflex | Demand grows, firm hires anyway | Job posts for roles AI was meant to absorb | Hiring conditional on proving AI couldn't absorb it |
| The Owner Ceiling | Owner's adoption pace becomes the firm's | Official adoption far below actual use | Asked staff what they already use, standardized from there |
| The Hollow Bench | Rollout works, apprenticeship loop breaks | Juniors ask AI before their manager | Rebuilt review as teaching, not just checking |
| The Unfunded Hour | Tools arrive, learning time never does | Willing people who say they haven't had time | Put learning hours on the calendar, protected like client work |
| The Enthusiasm Flood | Leadership pushes everything at once | Several initiatives live, engagement falling | Sequenced one owned workflow at a time |
The Perpetual Pilot
The pilot never gets a production owner or a deadline, so it demos forever. It plays well in a slide. It gets applause in a leadership meeting. Then it waits for someone to make it real. Nobody is against it. Nobody is responsible for it either, which is worse.
Early sign: no named owner and no date by week 4. This is the classic Stage 0-to-Stage 1 trap on the Delivery Model Ladder: enthusiasm with no organizational commitment behind it. Survivors did one unglamorous thing: assigned a single accountable operator and defined the outcome before writing another line of the build: one deliverable, a cost-to-deliver target, a date, and a decision about where the freed hours go next.
The Invisible Win
Hours genuinely drop. People finish work faster. And then nothing about the business changes, because pricing, staffing, and capacity plans stay exactly as they were. The saved time evaporates into slack instead of margin. Nobody can point to where it went.
The mechanism is worth naming precisely: the freed time doesn't sit unclaimed, waiting for someone to point it somewhere. It gets absorbed by the tasks around it. Checking a report starts taking longer than it used to, remaining work stretches to fill the same allotted hours, and the cost moves one step downstream instead of disappearing.
Early sign: someone reports "we're saving 15 hours a week" and the org chart hasn't moved an inch. Second early sign, subtler: the task next to the one you sped up is now slower. This is the evaporation pattern behind cost to deliver, the first of the Four Numbers: a cheaper deliverable doesn't help the P&L until the freed capacity is aimed somewhere that reaches margin or retention. On the Ladder, this is Stage 1 in its purest form: one number moved, the other three didn't, and the firm calls it progress. Survivors decided in advance, before the AI shipped, exactly where reclaimed capacity would go, measured the task before and after at the task level, and shipped the reallocation together with the tool. Wait until you notice a gain to claim it, and the adjacent tasks have already stretched to fill the gap.
The Shadow Rollout
Individuals adopt AI on their own, quietly, because it's useful and nobody stopped them. Usage climbs. Results are inconsistent, because no two people do the work the same way. Nothing about it is repeatable at the firm level.
And a meaningful share of that use isn't just unofficial, it's concealed on purpose: 42% of Gen Z workers use AI without their employer's knowledge, and the concealment is rational, because employees described as getting help from AI are rated lazier and less competent than employees getting the same help from a person. The compounding cost: the same silence that let someone skip asking permission also lets them skip reporting when the tool got something wrong.
Early sign: usage is high, but ask five people how they use it and you get five different answers, and nobody volunteers which tool did the work. That's improvisation with a subscription attached.
The cause is usually more boring than defiance. Across the firms I've observed, people using AI outside the sanctioned path are rarely trying to break a rule. They don't know what the rule is, and they can't tell who to ask. A firm with no stated position on AI has, in practice, stated one: figure it out yourself and don't mention it. That produces the two behaviors a firm least wants, avoidance and concealment, from people who would have complied with a policy if one existed.
The firms that recovered didn't ban the tools. Two moves, in order. First, they wrote down the boundary, and it can be short: these tools are fine, this part of the work is fine, this client data never leaves the building. A permissive boundary that exists beats a strict one nobody can find, because the point isn't restriction, it's giving people a sanctioned place to experiment inside their real work instead of beside it. Second, they picked one workflow, standardized it end to end, and rolled it out as a single repeatable process before starting the next one. The best place to find that first workflow is inside the shadow rollout itself: somewhere in the firm, one person has already built the thing, and the highest-return move available is usually not a new tool but taking one person's working process and giving it to the twelve other people who do the same job. Standards are also what make the gains compound: update the proposal skill or a delivery rule once, and everyone works the new way instantly. Shared workflows are how the whole company gets faster and delivers consistently better results. Improvisation can't do that.
The Unowned Adoption
The system works at launch. Everyone is trained. Usage is strong in month one. Then it decays quietly because no one watches it after the launch meeting ends. This differs from the Shadow Rollout: here the workflow was standardized correctly, but ownership stopped at go-live, as if adoption were a one-time event rather than an ongoing responsibility.
Early sign: usage falls three consecutive weeks and nobody notices. Adoption curves don't hold themselves up. The firms that kept their gains treated adoption as an operating metric, not a launch event, and measured it weekly, in the same rhythm as revenue. This is the habit-building, culture-changing part of the work. Without it, regression isn't a risk. It's the default.
The Quality Debt Spiral
Output volume goes up, because AI makes producing work faster. Review capacity doesn't rise with it, because reviewing is still a human, sequential task. Rework climbs. Client-facing errors slip through. The speed gain gets eaten by the cost of catching mistakes afterward. That is the mechanism that quietly drags down retention, the customer-side number in the Four Numbers framework.
Early sign: reviewers become the bottleneck within a month, and the backlog grows instead of shrinking. There's a worse variant one step further down: work with no review step at all, because nobody officially knows it's AI-assisted, so nobody built a gate for it. That variant fails in five recurring shapes: inconsistent style and voice across one deliverable, wrong facts stated with flat confidence, wrong conclusions drawn from correct data, hallucinated statistics that trace back to nothing, and client data leaving the building through a tool with no agreement covering it. The first four are rework costs. The fifth is a trust cost and possibly a legal one, the only failure in this taxonomy that a better review step alone can't recover.
This mode hits hardest at the Stage 1-to-2 transition on the Ladder. The visible costs arrive from the client side: more complaints, worse reviews, and human hours burned on corrections. All of it is a sign that the processes were never updated for the new capabilities. Survivors redesigned review as a structural part of the workflow, not a checkpoint bolted onto the end.
The Wrong-Number Dashboard
The firm builds an AI dashboard: prompts run, hours claimed as saved, seats adopted, usage streaks. None of it is a lie. None of it appears in the monthly financials. Activity metrics feel like proof of progress while measuring something adjacent to what matters. Two dashboards that never reference each other is itself a warning sign.
Early sign: the AI dashboard has no line that shows up in the P&L, and nobody in the finance meeting has ever seen it. Survivors stopped keeping AI reporting separate and brought AI results into the same meeting as financials, measured against cost to deliver, retention, margin, and revenue per person. The test for any metric is whether it feeds one of the Four Numbers. Unit cost per deliverable tells you whether Stage 1 is real. Rework rate and renewals tell you whether Stage 2 is arriving. Revenue per person tells you whether Stage 3 is in sight. Built that way, the dashboard doubles as a Ladder diagnostic. A metric that can't be traced to one of the four is an activity metric, however good it looks.
The Headcount Reflex
Demand for the firm's services grows. The instinct, trained by years of running the business the old way, is to hire. Nobody asks whether the redesigned workflow was supposed to absorb that growth without adding headcount. That was often the economic case for building it.
Early sign: job postings appear for roles the new workflow was meant to absorb. This is the clearest Stage 3 blocker on the Ladder: revenue is supposed to decouple from headcount there, and hiring reflexively undoes it. Survivors made hiring approval conditional: someone had to show, with numbers, that the workflow genuinely couldn't absorb the extra work. Detaching headcount from demand is the Stage 3 signature, but the transition has to start long before Stage 3. The discipline is simple to state: if AI can absorb the work, don't hire for it. Hire for the more complex work that creates outsized value for clients.
The two modes below are the newest additions. The pattern behind each is visible in published data, but the case bank behind them is thinner than behind the first seven, and their recovery paths are earlier drafts.
The Owner Ceiling
The firm's official AI adoption runs at whatever pace the owner personally sets, and that pace tracks the owner's cohort, not the team's. Boomer-owned small businesses adopt AI at 10.3%; Millennial-owned businesses at 22.1%, while 73% to 85% of workers across every generation already bring their own AI tools to work regardless of what's sanctioned. The distance between those two numbers is the size of the Shadow Rollout running underneath: this mode is that one's origin story. The firm hasn't decided against AI; the owner's personal ceiling became the firm's official one, and everything below it went unsupervised.
Early sign: official adoption sits far below actual use, which you find out by asking staff what tools they already use, not whether they use any. This mode lives at Stage 0 on the Ladder while pretending the firm never started. Survivors started from candor: assumed the answer was yes before asking, inventoried what was already running, and standardized from the workflows that were already quietly working instead of launching a sanctioned tool from zero.
The Hollow Bench
The rollout works, and that's the trap: this is the only mode that strikes after adoption succeeds. AI lifts the least experienced workers most, roughly 35% for novices while experts see flat gains, so juniors start producing senior-shaped output. Meanwhile the apprenticeship loop that turns juniors into seniors quietly re-routes: nearly half of Gen Z ask ChatGPT before asking their manager, and 45% say it knows them better than their boss. The questions that used to reach seniors, the ones apprenticeship actually runs on, stop arriving. The firm keeps shipping, but the pipeline that produces the next generation of judgment is hollowing out, and nothing on this quarter's P&L shows it.
Early sign: juniors ship work nobody taught them to judge, and seniors notice they haven't been asked a hard question in months. This is a post-Stage-2 failure: the workflow redesign succeeded and the decay is in the people system around it. Survivors rebuilt review as teaching rather than just checking, making the senior's edit a conversation the junior sits in, so judgment still transfers even when production doesn't require it.
The two modes below are newer still, and their evidence base is different in kind. They come from the pattern I see repeated across firms rather than from published research, and I'm naming them because the repetition is hard to ignore, not because someone has measured them. Treat the recovery paths as first drafts.
The Unfunded Hour
The firm bought the tools. The people want to use them. Nothing happens, because nobody gave anyone the hours. This is the mode operators consistently misread as resistance, and it almost never is. Ask why adoption is flat and the answer that comes back most often isn't "I don't think this works." It's "I know it would help, I haven't had time to sit down with it."
That answer is not an excuse, it's an accurate description of the arithmetic. Every tool that saves an hour a week costs some number of hours to learn first, and those hours have to come out of a week that was already full before AI existed. The learning curve lands on the people with the least slack, because the people with the least slack are the ones doing the work worth automating. So the tool becomes one more item on a to-do list that was already too long, and it stays there. The constraint on AI adoption inside most firms is not budget. It is bandwidth, and the two are easy to confuse because only one of them shows up in a spending review.
The executive version of this mistake has a consistent shape: leadership supplies two of the three ingredients and skips the third. Enthusiasm, yes. Tools and licenses, yes. Protected time to learn and experiment, no. The third one is the whole thing. Early sign: seats bought far exceed seats used, and the people not using them describe wanting to rather than doubting it. On the Ladder this is a Stage 0 that never becomes a Stage 1, and it is invisible in the P&L because nothing was ever attempted. Survivors put learning on the calendar and defended it like client work: named hours, taken off delivery capacity on purpose, with someone accountable for the fact that they happened. There's a link worth drawing to the Invisible Win here. The first honest place to aim reclaimed capacity is at learning the next thing. A firm that spends its first efficiency gain buying itself the time to compound the second one has solved two modes with one decision.
The Enthusiasm Flood
This is the Owner Ceiling inverted, and it's the mode nobody warns operators about, because it's caused by doing what everyone told them to do. Leadership becomes convinced, correctly, that this matters. Then it arrives all at once: several tools, several new processes, a mandate to redesign workflows, and a standing expectation that everyone is now also learning AI. The intent is good. The effect on a delivery team is change fatigue, and the rational response to a flood of half-specified initiatives is to disengage from all of them and wait to see which one is real.
What makes this one hard to see from the top is that it produces exactly the symptom that reads as skepticism, so leadership's instinct is to push harder, which is the one intervention that reliably makes it worse. The resistance was manufactured by the enthusiasm, not by doubt about the technology. Early sign: more than a couple of AI initiatives are live at once, none of them has a single named owner, and engagement drops as the announcements increase. A second sign, from the staff side: people describe AI as something happening to them rather than something they're doing.
This mode sits at Stage 0 and Stage 1, and its specific damage is that it blocks the Stage 1-to-2 decision. That decision requires pointing at one workflow and rebuilding it, and a firm running six simultaneous initiatives has nothing to point at. Survivors sequenced. One workflow, one owner, standardized and finished before the next one starts — the same discipline that exits the Shadow Rollout, applied to leadership's own appetite rather than to the staff's improvisation. It's worth holding this mode and the Owner Ceiling together, because they look like opposites and are the same error: in both, the owner's personal relationship to AI substitutes for a delivery-model decision. Too little and too much both leave the firm at Stage 0.
How do I know my firm's AI rollout is failing?
Check four things: is there a named owner with a date, has anyone been given protected time to learn the thing, has anyone measured adoption this week, and does any AI-related number appear in the same report as revenue and margin. A no on any means you're inside one of the eleven modes above, visible or not. The Four Numbers page goes deeper on which metric to watch per failure.
What did the firms that survived do differently?
Every survivor in the case bank treated the failure mode as an operating problem, not a technology problem. Ownership, cadence, and measurement fixed what better prompting never could. None of the eleven exits required a bigger model or a new vendor. Someone accountable. A number in the P&L conversation. Hours on a calendar. A decision made before the tool shipped rather than after.
How does the Production Gap connect to the Delivery Model Ladder and the Four Numbers?
The Delivery Model Ladder shows where a firm sits, from Stage 0 to Stage 3 AI-native delivery. The Four Numbers show what to measure while it moves. The Production Gap explains why firms get stuck: the Owner Ceiling, the Enthusiasm Flood, the Unfunded Hour, and the Perpetual Pilot strike at Stage 0 and 1, the Quality Debt Spiral at the Stage 1-to-2 transition, the Headcount Reflex at the Stage 3 threshold, and the Hollow Bench after Stage 2, once the rollout has already worked.
Rod Amora writes about how service firms rebuild delivery around AI: what works in production, what fails, and how the firm changes when the model sticks. Drawn from ten years working with service-based businesses (since 2017) and a vantage point built by watching hundreds of firms adopt AI through a franchise network's delivery data spanning 4,000+ businesses and 150+ franchise units and growing.
FAQ
Is the Production Gap a fixed list, or will it grow?
It's a living document. As of July 2026 it holds eleven modes, drawn from a growing case bank. New ones get added once a pattern repeats. The evidence behind them is not uniform, and I would rather say so: the first seven have the deepest case bank, the Owner Ceiling and the Hollow Bench are backed by published data with thinner cases behind them, and the Unfunded Hour and the Enthusiasm Flood rest on repeated observation rather than published research.
Which failure mode is most common early in an AI rollout?
The Perpetual Pilot and the Invisible Win account for most early stalls I have observed. Both precede any decision about ownership.
My team isn't using the tools we bought. Is that resistance?
Usually not. Ask them, and the most common answer is that they want to and have not had time, which is the Unfunded Hour rather than skepticism. The two look identical from a usage dashboard and need opposite responses: resistance needs a reason, the Unfunded Hour needs hours. Pushing harder on the second one produces the Enthusiasm Flood, which does create real resistance.
Can a firm be in more than one failure mode at once?
Yes, commonly. No named owner often means no adoption tracking either. The same missing accountability causes both.
Does the Production Gap apply outside consulting and professional services?
The pattern base comes mainly from consulting and professional-services firms, observed through a franchise network's delivery data. The case bank for other verticals is still thin.
How is a failure mode different from a bad AI tool?
Almost none of the eleven modes are caused by the tool itself. They're caused by what happens, or doesn't happen, around it: ownership, review capacity, measurement cadence, hiring decisions.
What's the fastest check for the Wrong-Number Dashboard mode?
Put your last AI usage report next to your last monthly financials. If no line from the first appears in the second, you are in it.
Where should I look next?
The Delivery Model Ladder shows where your firm sits. The Four Numbers show which metric each mode damages. I write about these patterns in the newsletter as the case bank grows.