Your AI Is Just a Better Google Search
Replit says its engineers ship 2.9x the code with quality flat. MIT says 95% of AI pilots deliver nothing. The difference is plumbing: connected systems and written rules, and both are available to a firm your size.
On July 16, 2026, Replit published a post called The Self-Driving Company. It claims a stable cohort of Replit engineers now produces 2.9x as much code as before, with code rollbacks and production incidents flat, and support closing its hardest tickets 60% faster. In August 2025, MIT researchers reported that 95% of generative AI pilots deliver no measurable impact on the P&L.
One company triples output while most companies get nothing, in the same year, with the same models. The gap between those two numbers has little to do with model quality. The difference is groundwork: agents connected to the systems the firm runs on, given its written rules, and allowed to act. An AI that can only answer questions is a better Google search, and that is where most firms stop.
What is Replit actually claiming?
Stripped to its checkable claims, the post says this. Lines of code contributed grew 5.8x from January to June 2026, or 2.9x after holding the group of authors constant. An agent reviews code changes and calls in a second human only when the change looks risky, saving 30% of human review time. Rollbacks and incidents stayed flat while code volume tripled. Support closes escalated tickets 60% faster. They canceled a seven-figure SaaS contract because an internal agent-built tool beat it.
Be careful with these numbers. They are self-reported. Nobody outside Replit has audited them, and the post shows percentages without the raw counts behind them. One of the few serious reviews of the post put it well: "Those are meaningful claims. They are not yet a measured productivity result."
The direction is harder to dismiss. Replit's agents work across the whole company: incident investigation, code review, business intelligence, support triage, sales research. And Replit is one of several firms reporting the same pattern. Anthropic says about 80% of its merged code is now written by its own model. Cloudflare reports 93% of its R&D staff now use its internal agent stack, less than a year after it launched. The useful question is what these companies did first.
The paragraph everyone skips
The least quoted part of Replit's post is the one that explains the results. Before any agent did useful work, Replit built plumbing: rules for what each agent could touch, a log of everything agents did, and a locked-down network. Only after that did agents get access to GitHub, GCP, Linear, Notion, Slack, and ZenDesk. The post spends one paragraph on this and pages on the output charts.

The same order shows up inside every function that worked. The support team wrote its playbooks into agent skills before the agent touched a ticket. The data team built a map of the company's data so the agent knew which numbers to trust and how they connect. In every case the agent started from written rules and connected data. The 2.9x came after.
The TECHi review captures the distinction: "Agency without access control is a demo. Agency with bounded credentials, recorded actions and escalation can become an operating process."
Cloudflare tells the same story. It built the internal infrastructure first, then watched adoption climb to 93% of R&D in under a year. The failure studies describe companies doing the reverse: deploy an agent, discover it can see nothing and follows no rules, roll it back.
Without the plumbing, you bought a better search box
I've watched hundreds of service firms adopt AI through a franchise network's delivery data, and the pattern repeats. The firms that see real gains are the ones where AI can reach everything: files, email, calendar, CRM, meeting transcripts, the tools the team works in. Agents that can join data across those sources produce results that surprise their owners, and roles start to reorganize around them. Below that threshold, AI answers questions, and that is the whole product. You bought a better Google search.
This explains the failure statistics better than "AI doesn't work." The MIT researchers found the 95% of stalled pilots failed on integration and organizational readiness, not on model quality. A 2026 survey of 2,527 senior decision-makers found 74% of enterprises had rolled back or shut down a customer-facing AI agent after deployment. In professional services, Thomson Reuters found org-wide AI use almost doubled to 40% in a single year while only 18% of firms track whether it returns anything. Firms are buying chat and calling it adoption.
One stat from that rollback survey is worth sitting with. Companies with the most oversight rolled back their agents more often than average, 81%. Oversight did not prevent their failures. It made the failures visible, which is the point. The plumbing exists so you can see what your agents do, catch what goes wrong, and keep going. The firms with no plumbing fail blind.
I wrote about the firm-level version of this in the productivity paradox piece: firms adopt the tools, change nothing about delivery, and the promised hours evaporate. Connection is the delivery change.
The service-firm version of the checklist
You are not going to build custom sandboxes and a locked-down network, and you don't need to. The expensive parts of Replit's stack are the parts that only matter at their scale. The transferable parts cost discipline, and they map one to one. That checklist is the mechanics of one specific move: the jump from Stage 1, Enhanced, where tools are sanctioned and individuals get faster, to Stage 2, Augmented, where the workflow itself gets rebuilt.
| Replit did | A $1M–$20M service firm does |
|---|---|
| Gave agents access to GitHub, GCP, Linear, Notion, Slack, ZenDesk | Connect Drive, Gmail, calendar, CRM, the project tool, meeting transcripts |
| Wrote support playbooks into agent skills | Write down the rules, SOPs, and best practices the firm runs on |
| Built a map of its data warehouse | Get operational data where agents can reach and join it |
| Let adoption spread visibly through Slack | Name an evangelist and invest in enablement |
| Proved it in engineering, then expanded | Pick one area, measure results, expand |
The second row is the one firms resist. Agents follow written rules. If your procedures live in the heads of your two best people, an agent has nothing to work from, and neither does your next hire. Writing down how the firm operates was worth doing before AI existed. Now it has a compounding return, because everything you document becomes something agents can execute.
The order matters as much as the list. Connection and documentation come first. The evangelist and the measurement come before expansion. Every firm I've seen skip to the end, handing out chatbot licenses and announcing an AI initiative, lands in MIT's 95%.
Start inside, and keep humans facing clients
Pick the first area with a bias for internal operations. It is the easy win. Your firm already produces the raw material every day: call recordings, emails, Slack or Google Chat messages, meeting transcripts, notes, proposals, SMS or WhatsApp messages. Once that material is captured and connected, it becomes what agents work from. Follow-ups draft themselves from the transcript. Proposals start from the last twenty the firm sent. Project status comes from the messages and documents that already exist, without anyone writing a report.
Customer service deserves its own rule: run the agent as an internal support role, and keep humans in front of clients. The agent researches the account, drafts the response, and pulls the history. A person sends it.

The evidence for that boundary is direct. The 74% rollback figure above is specifically customer-facing agents. And Replit, the flagship case for agents everywhere, kept the same boundary in support: its agent follows playbooks and escalates the hard tickets to a human. The 60% number describes an agent making humans faster on exactly those escalated tickets, with the human still in the seat.
Copy Replit's metric types, too, before you copy anything else. Time to close an escalated ticket. Rework rates on internal output. Volume up with rework flat proves the model works. Self-reported hours saved prove nothing. The next section explains why.
What not to copy from Replit
Two parts of the post transfer badly to a service firm.
The first is build-don't-buy. Replit canceled a seven-figure SaaS contract because its internal tool won, and the post presents this as a general shift. For them it probably is. The MIT study found the opposite for everyone else: purchased AI solutions succeeded about 67% of the time, internal builds about a third as often. Replit is an AI company with engineers in every seat. A service firm earns the right to build after the bought tools stop being the bottleneck, and most firms are years from that point.
The second is the multipliers. Lines of code is a metric an agent can inflate without producing value. And people's sense of their own speed runs well ahead of the stopwatch. In a July 2025 randomized trial by METR, experienced developers using AI tools took 19% longer to complete their tasks while believing the AI had made them 20% faster. A 39-point gap between feeling and fact, in the profession with the most AI practice. That gap is why "measure results" appears in the checklist. Your team's enthusiasm tells you adoption is happening. It tells you nothing about results.
Copy the flat numbers instead. Replit's most convincing charts show rollbacks and incidents holding steady while output tripled. Quality held while volume grew. Ask for the same evidence from your own rollout.
The first new role is already on your payroll
When the groundwork is in place, the first visible change is a person, and you already employ them. In the firms I've watched get this right, someone starts helping colleagues improve their workflows. They create skills, share recipes, refactor processes, and prototype small tools that remove repetitive work. Call them AI enablers. They support adoption and show the ropes to the less tech-inclined, and nobody assigned them the job.
Replit's adoption spread the same way. The rest of the company saw engineers tagging the agent in Slack and tried it themselves. Adoption traveled by visibility, one team watching another.

Compare that with Meta, which tied AI usage to performance reviews and reportedly got employees running idle models to climb an internal token leaderboard before it was shut down.
Replit's summary of the human side holds for service firms: "People don't feel like they've been automated. They feel like they've been promoted." In a service firm the promotion is concrete. The person who spent Friday afternoons assembling status reports now directs the agent that assembles them, and spends the afternoon on the judgment work clients pay for. The enabler is the first person through that door. The rest of the org chart follows.
Read the post again, differently
Replit's post will be quoted for its multipliers for months. Read it instead as a sequence: connect the systems, write down the rules, bound what agents can do, make the work visible, measure, expand. The multipliers belong to Replit. The sequence is available to a twelve-person consultancy, and the boring work in it pays twice, because documentation and connection improve delivery on their own before any agent touches them.
The groundwork fits in a quarter. Connect the tools agents will need: email, calendar, files, CRM, meeting transcripts. Write down the three SOPs your delivery depends on. Pick one internal area and define its before-and-after metric. Name the enabler you already have; you thought of the person while reading this paragraph.
If you want the next piece in this series when it ships, the newsletter is below.


