AI Projects Don't Fail at Your Size. They Go Underground.

An AI project can fail while work continues in personal ChatGPT accounts. Version history, source trails, and staff interviews show where the work went.

An empty workstation holds a monitor and a separate laptop, with a bag of files on a stool.
The work can keep arriving after the approved tool has lost its place.

An AI project can stop working in your firm while the team keeps using AI, and from your side of the business it can look as though the rollout succeeded.

The proposals arrive faster, people talk about the tools, and the client still gets the work. What you don't see is that the approved system lost its place and someone went back to a personal ChatGPT account to finish the job.

I saw this at a firm that introduced an internal AI tool for proposals and reports. It was supposed to replace people copying client notes into personal accounts, but it kept losing the client context that made the proposal useful.

With two hours left before a Thursday proposal deadline, a consultant went back to ChatGPT. Later, in a staff interview, the consultant explained it plainly, “I had two hours left. I used what I knew would work.”

Now, you can read that as an employee refusing to follow the new process. But the client deadline hadn't moved, the approved tool wasn't giving them what they needed, and the old path had worked before.

That doesn't settle whether the personal account was safe to use. It does tell you what problem the employee was solving, and it wasn't the problem management thought the rollout had already solved.

The proposal's version history showed the change. The first draft had the internal tool's generic structure, while the final version changed voice and included language from the team's old ChatGPT prompt.

By the next week, that prompt was being shared again. The team hadn't stopped wanting AI, some of the work had simply moved back to the route they knew.

And the owner had reasons to believe the new system was working. He had seen it demonstrated in Teams, proposals were arriving faster, and people were openly talking about using AI.

Those signals told him something true, the work was AI-assisted. They didn't tell him which tool had produced the version that reached the client.

That's the failure I want you to look for after a rollout disappoints. A canceled project can leave no AI in the workflow, but it can also leave shadow AI, work happening in tools the firm hasn't approved or can't see.

In this case the personal use existed before the rollout. The new tool didn't create it, the rollout made it easier to mistake the existing workaround for a working company system.

The cost is more than a license nobody uses. Client information is going into a different account, the prompt shaping the answer isn't the shared company prompt, and managers may be checking the result under different assumptions.

I have seen one manager check every AI-written sentence while another assumes the approved system has already done the checking. If they don't even know which system made the final version, they aren't starting from the same picture of the work.

I have also seen fabricated statistics used with client financial data. My records stop there, I don't have a recorded discovery time, client response or account of how that case was fixed, so I won't turn it into a more complete story than it is.

You don't need a public incident to have a problem worth correcting. When the source and prompt stay in one person's account, a correction made there may never reach the next person preparing the same kind of work.

The firm keeps producing, but it isn't keeping a shared account of what went wrong and what should change. So the next proposal can start with the same weakness, even though somebody already learned the lesson.

This is why I would ask which tool produced this version, rather than asking whether the team uses AI. The second question can get an enthusiastic yes while leaving the failed rollout completely hidden.

Now let me separate that from the big failure numbers, because they don't tell you how often this happens at a firm your size. When people ask why AI projects fail, they often get a statistic about a different kind of project, or a forecast presented as something that has already happened.

Fortune's August 2025 coverage of MIT NANDA reported that the vast majority of the enterprise pilots it described were stalled, with little to no measurable impact on profit and loss. That's the coverage behind the widely repeated 95 percent headline.

I'm relying on Fortune's account here. The report's methodology has been contested, and its headline is not a measured probability that your proposal tool will fail, much less that your staff will return to ChatGPT.

In 2024, Gartner predicted that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025. The same release discussed deployment approaches costing $5 million to $20 million, which gives you an idea of the scale of change it was talking about.

And Gartner's June 2025 forecast put cancellation of agentic AI projects above 40 percent by the end of 2027. Those are dated forecasts, not a count of what subsequently happened.

The causes in the research can still help you. RAND's 2024 report, based on interviews with 65 experienced data scientists and engineers, describes teams solving the wrong problem, lacking suitable data or the systems needed to use it, choosing technology for its own sake and asking AI to do work beyond its ability.

That last part matters, because the answer isn't always better management or another connection between tools. Sometimes the task itself is a poor fit, and no amount of pressure on the team changes that.

But the interviews don't give you a measured failure rate for small service firms. RAND's much-quoted “more than 80 percent” line is introduced as an estimate from elsewhere, not a rate calculated from those 65 interviews.

I would use this research to ask better questions about the work, not to assign your firm a set of odds. In the proposal case, we could name the failure much more precisely, the tool lost client context and the workflow still required a client-ready proposal by Thursday.

I've seen other paths into the same kind of problem too. Leadership buys something the team hasn't had time to learn, or everybody wants the tool but it can't fit the way the work is done, or employees want to improve the process and the company won't fund the tools and time it would take.

Money, learning time, people and agreement about the job all matter. If you only fund the software, the team still has to find the rest of those things somewhere in the working week.

But I don't want the title to leave you thinking every failed project continues in secret. I also saw a finance workflow where the approved tool was too slow, failed and stayed dead.

That work involved payroll and client cash data. Access was limited, uploads were monitored, and the finance manager required a source trail for every conclusion, so staff couldn't quietly complete the same work in a personal AI account.

That's the counter-case, and it's an important one. In the proposal workflow, there was a working route around the approved tool, while the finance workflow had boundaries that prevented that route.

So after a rollout fails, inspect what actually happened rather than assuming either that AI stopped or that everyone went underground. You can start with one recurring task and follow the latest finished version back through its sources.

Ask which tool produced it, what the person used that tool to do, what client or company information went in and where the output went afterward. Then name who checked it and what source or rule they checked it against.

Version histories, source records and a conversation with the person doing the work can give you a much clearer picture than the count of active licenses. Explain that you're trying to understand the working process, because if the conversation feels like a hunt for someone to blame, you make the next workaround harder to see.

Then decide whether the use should be approved, changed or stopped. Approval needs a clear data boundary, a reviewer who can judge the result and a shared way to record corrections, not just a decision that the output looked good this time.

If it needs changing, give people protected time to learn the replacement. The Production Gap calls this the Unfunded Hour, the learning work a company expects while leaving the person's delivery commitments unchanged.

And if a use exposes restricted information or cannot be checked safely, stop that exposure rather than waiting for a nicer tool. Keep the work manual or use another safe fallback while you prepare the replacement, because protecting a deadline can't make an unsafe route acceptable.

Where the issue is fit rather than immediate risk, have the workable replacement ready for the next deadline before you remove the route the team relies on. Otherwise you've recreated the choice that sent the consultant back to ChatGPT in the first place.

Give the person who found the workaround a say in the redesign, and give them credit for what they learned. Don't quietly make them the unpaid trainer and support desk for the entire firm, with responsibility for every future mistake and no authority to change the system.

This check assumes you can inspect the workflow and name someone who can judge the output. If you can't, keep that work manual until you can, and start by finding out which tool produced the last version your client received.

Topics: