Your AI Is Just a Better Google Search
If AI can't reach your company information, you keep carrying the work into and out of a chat. Connect one workflow safely, write its rules and check whether the whole job improves.
If your AI can't reach the information your team works with, you spend a lot of your time explaining the company to it before it can help you.
You copy in the client notes, explain what happened in the meeting, find the old proposal and tell it how the firm likes things written. Then you take the answer back into the tools where the work actually happens.
That can still be useful, I don't want to dismiss it. But in the firms I've watched, the larger change starts when AI can work with the right company information and follow the firm's rules, so people no longer have to carry everything into and out of the chat by hand.
Without that connection, a lot of what you've bought is a better Google search. It helps someone find or prepare an answer, while the job around that answer stays much the same.
I want to walk you through a public example, because it shows the work that tends to disappear from the AI success story. In July 2026, Replit published The Self-Driving Company, describing how it had put agents into work across the business.
The number that gets your attention is 2.9 times as much code from the same group of authors. Replit also says its trends for code reversions and production incidents stayed flat, and that its hardest support tickets, those escalated to humans, closed 60 percent faster.
Those are Replit's own figures. The post doesn't establish an independently audited productivity result, and writing more code is not the same thing as creating proportionally more value for the company.
So I wouldn't put 2.9 times anything into your plan. What I would look at is what Replit says it did before the increase, because that is the part a service-firm owner can learn from without pretending to run an AI software company.
First, it built a controlled place for agents to work. It put rules around what they could access, recorded their actions and protected the network, and only then gave them access to the systems its people used, including GitHub, Slack, Notion and Zendesk.
That's an important order. Connecting the company doesn't mean giving an agent every password and hoping the instructions are enough, it means deciding what information it needs and what actions it is allowed to take.
Then the support team gave the agent skills for investigating problems and following its standard playbooks. The data team also made the company's data easier for the agent to understand, so it could work with what the numbers meant and how they related to each other.
Basically, the firm supplied the context and the rules that a good employee would need. The model didn't discover Replit's operating procedures by being clever, people did the work of making those procedures available.
I've watched hundreds of service firms adopt AI through a franchise network's delivery data, and the same part matters there. When agents can bring together the relevant files, emails, calendar entries, CRM information and meeting transcripts, they can help with work that a disconnected chat struggles to do.
A follow-up can start from what the meeting actually agreed. A proposal can use the firm's previous work and the client's history, and a project update can draw from the documents and messages where the project is being run.
Those are possible workflows, not promises that you can connect the accounts and leave them alone. Someone still needs to decide which source to trust when two records disagree, what the agent should do if information is missing and who checks the result.
That is where people's jobs can begin to change around the tool. They spend less time gathering the same material and more time deciding what to do with it, but the gathering has to work well enough for them to trust it first.
If your rollout hasn't reached that point, I wouldn't start by blaming the team for only asking simple questions. They may be using the only part of the tool you made practical, while the company information and the authority to act are still somewhere else.
You own the work needed to close that gap. Your people can help write the instructions and test them, but they need time to do it and someone who can make decisions when the current process doesn't fit.
There is a temptation to use the failure statistics as proof that connection is the whole answer. It isn't, and the famous MIT NANDA headline doesn't establish that every stalled project failed for the same reason.
Fortune's August 2025 coverage described a learning and integration gap in the report, with tools that didn't adapt to how firms worked. That's secondary reporting on a report with contested methodology, not a measured explanation for every company that bought a chatbot and saw little return.
The more useful question is what you can actually see inside your rollout. In Thomson Reuters' 2026 survey, organization-wide AI use rose from 22 to 40 percent, but only 18 percent of respondents said they knew their organization was tracking the return on its AI tools.
That doesn't prove the rest weren't measuring, because respondents may not know. It does show why you shouldn't assume that more use comes with a clear account of what improved.
And when the agent can act in your systems, seeing what it does becomes even more important. Sinch's May 2026 research release says 74 percent of surveyed enterprises had rolled back or shut down a live AI customer-communications agent after a governance failure.
That was a survey of 2,527 senior decision-makers at organizations with at least 1,000 employees, so it isn't the failure rate for a consultancy your size. It also measures whether a company had pulled back an agent, not the percentage of all agents that failed.
Sinch reported an even higher rate, 81 percent, among organizations with mature controls, and suggested better monitoring was helping them find failures others missed. That's the company's interpretation of the pattern, it doesn't prove that controls caused the difference.
But I agree with the operating point, being able to notice and stop a bad run is useful. A quiet system isn't necessarily a good one, it may just be a system that isn't showing you enough.
So what would I copy at your size? Start with one internal piece of work where the source material already exists, rather than trying to connect every department at once.
Meeting follow-ups are one possibility. You already have notes or transcripts, somebody already turns them into tasks, and you can compare what the agent proposes with what the meeting actually agreed.
Before you let it change the project tool, decide which meetings it may read, which projects it may write to and which proposed actions a person must approve. Keep a record of the source and what the agent did, so a wrong date can be traced and corrected rather than becoming a mystery later.
Then write down the procedure the person currently follows, including the parts that require a decision. If a task has no owner, if the transcript doesn't give a date, if two people disagree, the agent needs a rule for stopping or asking, not encouragement to fill the gap.
This is the documentation work firms tend to skip because their experienced people already know how to handle those situations. But if the rules live only in their heads, neither an agent nor the next hire can reliably follow them.
And the documentation isn't done when it reads well. Give it to someone who does the work, try it on real examples and see where they need to add information, because that is part of the process you've been relying on without writing it down.
For customer service, my recommendation is to keep the first agent in an internal support role. Let it research the account, gather the history and draft the reply, while a person stays responsible for what goes to the client.
Replit's 60 percent figure is specifically about tickets escalated to humans, which makes that part of its story relevant. Its post also says the agent can offer a response or escalate, so I'm not claiming Replit kept a person sending every reply, the stricter starting boundary is my advice for your firm.
Once that internal process works, you can decide whether a narrower set of replies is safe to automate. That decision needs evidence from the work and a way to catch mistakes, not just the fact that another company calls itself self-driving.
I would also be careful about copying Replit's decision to build its own tools. It says it canceled a seven-figure software contract because an internal application was better, and that may be a good decision for a company whose own product is software for building software.
It doesn't follow that your accounting practice should replace its main systems with applications an agent made. Buying a suitable product also buys you some of the maintenance and support, and if you build, that responsibility still has to belong to someone.
My starting preference for a service firm is to use the tools that already fit, connect them safely and find out what problem remains. Build when you can name the gap, judge whether the result is right and support it afterward, rather than because generating the first version has become easy.
There is a measurement trap here too. METR's early-2025 developer trial found people believed AI had made them faster even though their work took longer in that experiment.
METR now warns that the finding no longer represents current tool impact, and its February 2026 update describes evidence of improvement alongside problems measuring it. So I use the old trial as a warning about trusting impressions, not as proof that today's tools slow your team down.
For your own rollout, follow the work to an acceptable result. How long did it take, how much checking was needed, what had to be redone and what did the tool and the setup cost?
Replit's report is more useful when it talks about reversions, incidents and review time than when it only talks about code volume. Copy that habit of putting quality beside output, while remembering its chosen measures don't capture every kind of quality either.
And decide what any freed time is supposed to make room for. Otherwise you can get a better-connected tool and still lose the benefit in the work around it.
There is usually a person involved in helping this change spread. In firms I've watched, someone starts improving workflows, creating agent skills, sharing instructions and helping less technical colleagues use the tools, often before anyone gives them a formal role.
I call that person an AI enabler. They aren't necessarily a new hire, and noticing them is only the first step, they need time and support if you're going to ask them to help the rest of the team.
Replit describes people seeing engineers use the agent in Slack and trying it themselves. That gives you a more useful thing to encourage than a contest over who sends the most prompts, let people see a piece of work improve and help them try it on their own work.
You can plan an initial round of this over a quarter, but the scope has to fit what you're starting with. If the source information is missing or nobody can judge the output, connecting more systems will not fix that for you.
So take one internal workflow to the person who owns it, write down what the agent needs to know, what it may do and where it must stop, then test whether the whole job gets better. Expand after you can show that, not after the first answer looks impressive.


