Framework · Rod Amora ·
The Production Gap
The demo worked, but getting the team to depend on it every day is another job. Here are eleven places that job can get stuck, and what to try next.
The demo can work, people can like it, and later nobody is using it in the work clients actually receive. The Production Gap names eleven ways an AI project can stall between the demo and daily work.
I built the list after watching hundreds of service firms adopt AI through a franchise network's delivery data. The model often worked, but the same decisions were missing, who owned the rollout, how the work would be checked, what to measure, or where the free capacity would go.
As of August 2026, the list has eleven failure modes. Each one includes an early warning sign and a practical next move. Some have a deeper case history than others, and I note that difference below.
The list connects each failure to a stage of The Delivery Model Ladder. Use it to name what is happening before a quiet stall turns into an abandoned project.
Why do AI projects stall before daily use?
Many pilots do not end with a clear rejection. A proof of concept works, people like it, and then it waits. Nobody gives it an owner, a date, or a place inside a real client workflow.
The most common stalls I have seen happen at Stage 0 or Stage 1 of the Ladder. At Stage 0, people use AI on their own. At Stage 1, the firm adds AI to an old workflow without redesigning the work.
Broad survey data shows the same lack of business change. A 2026 NBER survey of nearly 6,000 executives found that 89% reported no labor-productivity gain from AI over the previous three years. The group reported a realized gain of 0.29% and expected 1.4% over the next three years.
An April 2026 Census Bureau working paper found that 64% of firms using AI made no organizational changes and 66% only added AI to existing tasks. That is the shape of Stage 1: the tool is present, but the firm still runs the same way.
What are the eleven failure modes?
This table gives the short version. The sections below explain what each mode looks like and what firms did next.
| Failure mode | What it looks like | Early warning sign | Practical next move |
|---|---|---|---|
| The Perpetual Pilot | Demos continue, but nothing ships | No owner or decision date by week four | Name an owner, outcome, and date |
| The Invisible Win | Work gets faster, but the business stays the same | Hours saved have no named use | Decide where the free capacity will go |
| The Shadow Rollout | People use AI privately in different ways | No shared workflow or clear boundary | Write the boundary and standardize one workflow |
| The Unowned Adoption | Use is strong at launch, then fades | Usage falls for three weeks without a response | Give someone ownership after launch |
| The Quality Debt Spiral | Output rises faster than review capacity | Review becomes the new bottleneck | Build checks into the workflow |
| The Wrong-Number Dashboard | The firm tracks activity instead of results | AI measures never reach the financial review | Connect AI work to the Four Numbers |
| The Headcount Reflex | Demand grows and the firm hires as before | Open roles match work the new process should absorb | Test capacity before approving the hire |
| The Owner Ceiling | Official use moves at the owner's pace | Actual use is much higher than reported use | Ask what staff already use |
| The Hollow Bench | Production improves, but learning breaks | Junior work ships without teaching or review | Use review to teach judgment |
| The Unfunded Hour | Tools arrive without time to learn them | Willing people say they have not had time | Protect learning time like client work |
| The Enthusiasm Flood | Leadership starts too many changes at once | Initiatives rise while engagement falls | Finish one owned workflow before starting another |
The Perpetual Pilot
The demo works, but nobody is responsible for putting it into daily work. It keeps returning to leadership meetings without reaching a client workflow.
Early sign: there is no named owner or decision date by week four.
Next move: choose one deliverable, one owner, one target, and one date. Decide where any free capacity will go before building more.
The Invisible Win
AI makes a task faster, but pricing, staffing, and capacity plans stay the same. Nearby work fills the open time, so nobody can show where the gain went.
Early sign: someone reports fifteen hours saved each week, but those hours have no named use. The task beside the faster one may also begin taking longer.
Next move: use the gate in the Four Numbers. Ask what share of the free capacity now has a named job. Assign it to more delivery, better delivery, learning, or a new offer before the rollout ships.
This is a Stage 1 failure. Cost may fall, but retention, price, and margin stay flat.
The Shadow Rollout
People use AI on their own because it helps. The firm gets many private methods instead of one process it can repeat.
Some use is hidden on purpose. One survey found that 42% of Gen Z workers used AI without their employer's knowledge. A separate study found that workers described as receiving AI help were rated as lazier and less capable than workers receiving the same help from a person.
Early sign: five people give five different answers when asked how they use AI. Nobody knows which tool touched the work or what data it received.
Next move: write a short boundary. Say which tools are allowed, what work they may touch, and which data must stay inside the firm. Then take one useful private workflow, standardize it, and share it with everyone doing that job.
The Unowned Adoption
The workflow is ready and training goes well. Use is strong in the first month, then falls because nobody owns what happens after launch.
Early sign: usage falls for three weeks and nobody responds.
Next move: give one person ownership of adoption after launch. Review use each week, ask why it changed, and fix the part of the workflow people stopped using.
The Quality Debt Spiral
AI increases output, but the people checking it cannot keep up, so the work starts waiting for review. Rework grows, and mistakes begin reaching clients.
Early sign: the review queue grows within a month of the output increase. A more serious version has no review step because the work was never marked as AI-assisted.
The failures often include mixed voice, wrong facts, bad conclusions from correct data, invented statistics, and client data sent to an unapproved tool. The last one needs a data boundary as well as a review step.
Next move: put checks inside the workflow instead of adding one final review at the end. Define what must be checked, who checks it, and what happens when the work fails.
The Wrong-Number Dashboard
The firm tracks prompts, logins, seats, and estimated hours saved. These numbers show activity, but they do not show whether delivery or the business changed.
Early sign: the AI report has no number that appears in the monthly financial review.
Next move: connect the rollout to cost to acquire, cost to deliver, retention, or price. Also track what share of the free capacity has a named use. The Four Numbers explains how to read those measures together.
The Headcount Reflex
Demand grows and the firm hires the same role it would have hired before the workflow changed. Nobody checks whether the new process can absorb the work.
Early sign: a job post describes work the AI-supported workflow was built to handle.
Next move: test the new capacity before approving the hire. If the workflow cannot absorb the demand, record where it failed. Hire for work that still needs people, especially judgment and client-facing work.
This matters near Stage 3 of the Ladder, where revenue can grow without headcount rising at the same rate. It is different from cutting people before delivery has been rebuilt, a risk covered in the AI layoffs article.
The evidence is thinner for the next two modes than for the first seven. Published research supports the wider pattern, but the case history is smaller.
The Owner Ceiling
The firm's official adoption moves at the owner's pace while employees use AI privately.
These numbers come from different studies, so you cannot subtract one from the other and call that your adoption gap. They do give you a reason to ask what is happening in your own firm. AI adoption was reported at 10.3% for boomer-owned small businesses and 22.1% for millennial-owned businesses. Another study found that 73% to 85% of workers across generations brought their own AI tools to work.
Early sign: official use is far below what staff describe in private. The firm acts like adoption never started even though tools are already inside the work.
Next move: ask which tools people already use and what work they touch. Start with the private workflows that are useful, then bring them inside the firm's boundary and review process.
The Hollow Bench
The rollout works, but the way junior people learn begins to break. AI helps them produce work that looks more experienced, while fewer questions reach senior people.
One study found gains of roughly 35% for less experienced workers while expert gains stayed flat. Another report found that nearly half of Gen Z asked ChatGPT before asking a manager.
Early sign: junior work ships without anyone teaching the junior how to judge it. Senior people also notice that fewer difficult questions reach them.
Next move: turn review into teaching. Let the junior see the senior's changes, ask why they were made, and update the rule used for the next piece of work.
The final two modes come from patterns I have observed repeatedly rather than published research. Their recovery paths should be treated as practical first moves, not measured standards.
The Unfunded Hour
You buy the tools, but nobody gets time to learn them. From the owner’s side that can look like resistance, even when people want to try.
A tool that may save an hour each week still takes time to learn. Those first hours have to come from a week that was already full.
Early sign: the firm has many unused seats, and the people not using them say they have not had time.
Next move: put named learning hours on the calendar and remove them from delivery capacity on purpose. Give someone responsibility for making sure that time happens. Then keep checking, because enablement is a loop, not a roadmap, and the seats nobody looks at are the ones that go quiet. The AI Readiness Assessment checks for this pre-adoption stall.
The Enthusiasm Flood
Leadership starts several tools and workflow changes at once. The team cannot tell which initiative is real, so people wait or disengage.
More pressure can make the problem worse. What began as too much work starts to look like resistance from the top.
Early sign: several AI initiatives are active, few have a clear owner, and engagement falls as announcements rise.
Next move: choose one workflow, name one owner, and finish the change before starting the next one. The Owner Ceiling moves too slowly. The Enthusiasm Flood moves too fast. Both let the owner's pace replace the pace the work can support.
How do I know whether my rollout is in trouble?
Check four things:
- Does the rollout have a named owner and decision date?
- Did the people doing the work receive protected learning time?
- Did someone review use and problems this week?
- Does the rollout connect to a number reviewed beside revenue and margin?
A “no” does not prove failure, but it gives you a specific risk to investigate. The table above shows the likely mode and first response.
What did the firms that recovered do differently?
The firms that recovered changed how they ran the work, they named an owner, set a date, protected learning time, built review into the workflow, and chose a business number to check.
Those are decisions the owner and managers have to make, a larger model or another vendor cannot make them for you.
How does this connect to the other Proofwork frameworks?
The Delivery Model Ladder shows where the firm sits. The Four Numbers show what changed. The Production Gap shows why movement between the stages can stop.
Most early failures appear at Stage 0 or Stage 1. The Quality Debt Spiral becomes more likely during the move to Stage 2. The Hollow Bench appears after the rollout works. The Headcount Reflex blocks the move toward Stage 3.
I built this framework from working with service-based businesses since 2017 and watching hundreds of firms adopt AI through a franchise network's delivery data covering 4,000+ businesses and 150+ franchise units and growing.
FAQ
Why do most AI projects fail?
In the failures I keep seeing, the model often works, but nobody owns what happens next, people have no time to learn, review is missing, or nobody decided what the saved time is for. The eleven Production Gap modes help you see those missing decisions in the work.
Is the Production Gap a fixed list?
No. As of August 2026, it has eleven modes from a growing case bank. The first seven have the deepest case history. The Owner Ceiling and Hollow Bench have published research but fewer cases. The Unfunded Hour and Enthusiasm Flood come from repeated observation rather than published research.
Which failure mode appears most often early in a rollout?
The Perpetual Pilot and Invisible Win are the most common early stalls I have observed. In the first, nobody owns the move from demo to daily work. In the second, the work gets faster but nobody decides what the free capacity should do.
My team isn't using the tools we bought. Is that resistance?
Before calling it resistance, ask when they were supposed to learn. If someone wants to use the tool but every hour is already promised to client work, that is the Unfunded Hour pattern, and you need to make room in the week rather than ask them to try harder.
Can a firm have more than one failure mode at once?
Yes. One missing owner can create a Perpetual Pilot and an Unowned Adoption at the same time. Missing review can also turn a Shadow Rollout into a Quality Debt Spiral.
Does the Production Gap apply outside consulting and professional services?
The examples come mainly from consulting and professional-services firms observed through a franchise network. The case bank for other industries is still thin. Use the questions to examine your own workflow; the examples do not establish results across industries.
How is a failure mode different from a bad AI tool?
A bad tool cannot do the task well enough. A Production Gap problem can happen even when the tool works, because nobody owns the workflow, review cannot keep up, people have no time to learn, or the firm never checks the business result.
What's the fastest check for the Wrong-Number Dashboard?
Put the latest AI usage report beside the latest monthly financial review. If no measure connects to cost to acquire, cost to deliver, retention, price, or the use of free capacity, the dashboard is tracking activity rather than change.
Where should I look next?
Use the Delivery Model Ladder to see where your firm stands, then the Four Numbers to choose what to measure. If you want help deciding what to do next, the AI Readiness Assessment turns your answers into a short plan.