You Bought Automation and Installed a Guess
AI works best when you give it a small piece of information work, not the whole responsibility. Start with the real source, a clear check, and one person who owns the result before you give the system more to do.
Why do AI projects fail? Usually, you give a guessing system the wrong kind of job.
Sometimes the job needed the same answer every time, and ordinary software would have done it better. Other times it needed a person to make an exception, protect a relationship, or approve something you couldn't undo. The AI can be useful too, but not when you ask it to do too much at once and can't see where it went wrong.
Start with one small piece of information work, then add more only when it holds up on the real work.
A clean summary can be a damaged record
I saw this in an AI tool we built for our own proposals and reports. It was meant to replace people copying client notes into personal accounts, and it could produce a clean transcript, a written record of the conversation, and a meeting summary much faster than doing the work by hand.
One consultant needed to send information to a client that was missing from the summary, so he had to rewatch the whole meeting.
A week later, the client had questions about something they had discussed and had ideas about, but those details were not in the summary. The consultant had to rewatch the entire meeting again, which wasted a lot of time and delayed the reply. That was not the client experience we wanted.
The problem was that clean wasn't the same as complete. It lost topics from the meeting, important facts and comments, and some hard data, numbers, because it was trying to simplify the transcript and summary.
We didn't blame the consultant or abandon AI. We switched to a higher-quality transcription service, changed the template the agent had to use, and created an eval, an evaluation based on one template, to judge how well the agent would write new summaries. After those changes, the missing details and repeated client questions stopped.
That kind of failure is hard to spot when someone shows you the tool. The page looks shorter, the writing looks smoother, and nobody has to read through the whole conversation again. You only notice the missing number when somebody needs it for the proposal, and by then the person doing the work has to go back through the recording or the notes to find what disappeared.
A summary that leaves out the reason a client made a decision isn't a better summary. It's a shorter record with a hole in it.
The person who goes back to the old way isn't necessarily against using AI. They're trying to get the work out, and if the new way keeps dropping details, going back is safer.
You own the bill. You chose the tool, chose the job, and gave the team a result they had to trust without testing it against the original material.
Buying the tool didn't change the job
The easy explanation is that the team didn't use the tool. They had paid accounts, there was training, somebody showed them how it worked, and the work still looked much like it did before.
That's half right, the team needs the right information, permission to use it, connections to the other software, and a tool that fits the job. But saying the team didn't use the tool skips the decision you made about the work. You left the job, the source material, who approves it and who checks it unchanged, and hoped buying the tool would change the result.
In his February 17, 2026 speech, Federal Reserve Governor Michael Barr said that making AI useful will likely require “fundamental changes in business practices and organization.” He was talking about changing how the company works, training people again, having managers work out better ways to do the job, and paying for trial and error. That's what buying a tool leaves out.
Your team already has client work waiting, payroll to run, support requests to finish, and deadlines that don't move because a new system is learning. When the new way fails, the old one is sitting there, paid for and familiar, so going back gets the work done while hiding what the failed first try cost.
That's why a clean demo, where someone shows you the tool working, doesn't tell you enough. It doesn't prove the tool can keep the details your work needs, follow your approval steps, or produce work someone can use without doing it over.
Sometimes the work goes back to a personal account and happens out of sight, as I wrote in AI Projects Don't Fail at Your Size. They Go Underground.. Sometimes the tool is cancelled. Counting paid accounts won't tell you which happened.
Give AI information work, not the whole responsibility
Some work needs software that follows fixed rules, giving the same answer or taking the same action each time you give it the same information. A rule, formula, template, or ordinary tool that repeats the steps for you is usually the better fit. Doing a calculation, checking that a form has the required details, naming a file, or moving a record from one known place to another doesn't get better just because AI can describe it in a friendly sentence.
A user in an accounting discussion compared Claude, an AI tool, with VBA macros, small programs that repeat set steps: “Claude will not repeat a process exactly the same way everytime. VBA macros, however, will.” That's one user's experience, not a study comparing the two. But it asks the right question, does this job benefit from different answers, or does it need the same result again?
Then there's information work, and AI can help make that information useful. Finding information, writing summaries, spotting useful patterns, exploring ideas, picking out the records you need, sorting information, and studying numbers can all be useful starting points when the original material is available and the system has the right tools.
The words at the end matter. An AI system working with numbers without access to the right records or tools to do the math is still guessing about the numbers. A summary without the full source can be a damaged record. A tool taught to sort records using the wrong examples can send an item to the wrong place.
Then there's work where a person has to make the call. Someone needs to decide whether to make an exception, whether a price protects the business, whether a client needs a difficult conversation, or whether work is safe to send. AI can prepare the information for that decision, but it doesn't take care of the relationship or answer for what happens.
The right tool can still be doing the wrong part of the job
Research shows that AI can help with one part of a job and get another part wrong. The study run with BCG and published in Organization Science tested GPT-4 with 758 BCG consultants on 18 realistic consulting tasks.
On the tasks the AI could handle well, people using it completed 12.2 percent more subtasks and worked 25.1 percent faster. Beyond what it could handle well, they were 19 percent less likely to get the answer right than people without AI.
Those numbers come from selected tasks and highly skilled consultants early in their careers, using the version of GPT-4 available at the time. They don't tell you how often AI will get things right in your finance work, your customer service team, or your next proposal.
They tell you something more useful. “AI helps” and “AI belongs on this part of the job” are different decisions.
The stories from people using these tools point at the same problem, although they don't tell us how common it is. An agency owner wrote, “If your manual process is a mess, automation just makes a mess happen 10x faster.” That's what one user reported, not a measured result for the product.
If a wrong answer forces somebody to check every result, find the missing source, fix the number, and explain the mistake to the next person, the tool didn't save the work. It changed who has to do it.
Start with one small piece of the work
Pick one job your team does regularly before you pick the AI tool. Write down what finished means, from what comes in to what the client, manager, or next team actually receives.
Then split that job. Which part needs the same answer each time? Which part is finding or sorting information? Which part needs someone to decide what's good enough or what happens next?
Choose one small AI step with a clear start and finish. Don't ask it to write a summary, decide what matters, update the system, contact the client, and close the job all at once. If one part fails, you need to know which part failed.
Use the real material, not a clean example made to show the tool at its best. If the job depends on topics discussed in a meeting, test a messy written record of one, with interruptions, side comments, numbers, and decisions made out loud. If the job depends on company records, give it the records it will actually have, with the access and tools it will actually use.
Give one person responsibility for checking the result against the original material and what the business needs. That person isn't there to read every sentence forever, but somebody has to decide when the step is good enough to move forward and what happens when it's wrong.
In our case, the eval checked whether a new summary kept the information the template required. You don't need our exact setup, but you do need a repeatable check that compares the output with what the job actually needs.
Follow the work to the end. Did the finished work pass its checks sooner? Did checking take longer? Did it come back for more fixes? Did the client or the next team receive something they could use? A task finishing inside the tool isn't the same as the business finishing its work.
Write down how the work gets done now, try the small change, and compare the finished work. Measure how much work passes its checks, the time spent checking and fixing it, and whether the business got the benefit you wanted.
Then give AI more to do only when the first step works well on real work. Add the next part when you understand the information it needs, the tool it uses, the rule it follows, the person who checks it, and how it can fail.
The work gets safer as you get clearer about what AI should and shouldn't do. It gets dangerous when you hide all those decisions inside one impressive set of instructions.
Fix the AI step, replace the tool, or stop using it
A failed trial doesn't always mean the whole idea was wrong. It means you have to look at the job before you buy another tool.
Try to fix the AI step when its job is clear and limited, you can supply the missing material or access, one person can judge the result, and you can say what good work looks like in one sentence.
Replace it with software that follows fixed rules when the job needs the same answer or action each time and a rule, formula, template, or ordinary tool can repeat the work without guessing.
Keep the decision with a person when the work means making an exception, protecting a relationship, giving approval, or answering for something the system can't take responsibility for. AI can prepare the material and suggest what to do, but the business still needs somebody whose name sits beside the decision.
Stop the AI step when it keeps dropping important information, needs a person to check all the work, makes an existing mess happen faster, or leaves nobody able to say whether the final work is correct. Don't keep paying for a guess because the demo was impressive.
Take one finished piece of work and follow it back through the original material, the tool, the instructions, the person who checked it, and what happened in the end. Record whether the important details survived, whether review took longer, and whether the client got a timely answer. That will tell you more than the number of paid accounts you bought.
Where this advice stops
This only works when the business has a job it does regularly, access to the original material, a person who can judge the result, and enough time to run a small test. If nobody can say what a correct result looks like, the information can't be used safely, or nobody understands how the work gets done, keep doing it by hand or use software that follows fixed rules while you fix what's missing.
The research doesn't prove that every AI failure comes from giving a guessing tool a job that needs fixed rules. The proposal example shows one failure and one repair, not a universal law or a measured time-saving result. The user reports are individual stories and the BCG study used selected consulting tasks.
A cheaper tool is not progress when the job needed a rule. A polished guess is not finished work.


