Document the Process First, or the Agent Scales the Chaos
Your managers run the firm from memory. Until that knowledge is written down where a system can read it, the agent you build will run your current process faster, broken parts included.
The first AI agent you build will impress everyone in the room and change nothing about how your firm runs. That happens when your managers cannot describe, in writing, what they monitor every week, what they check, and what warns them a project is going wrong. An agent with no written job does the only thing available to it: it runs your current process faster, broken parts included.
Across the 150+ franchise units, roughly 80% are still running chatbots, 15% have moved to rigid workflows that run on triggers, and 5% are running autonomous agents defined by conditions. The split is pulled from franchise delivery data, not an estimate. What separates those groups is whether the work got written down before the tool arrived.
Asana's 2023 Anatomy of Work index found knowledge workers spend 58% of the day on work about work, and those workers estimate they could get 4.9 hours a week back with better processes.
The prototype everyone loved and nobody used
The first time we tried to put AI into a delivery operation I was running, I could not get the managers to describe how their work ran.
I asked what they do to monitor how their consultants work, what they check, how they track tasks and projects, and what warns them when something is going right or wrong. The room went quiet, not because they did not know the answers. They ran that operation every day, competently, from memory. It had never left their heads.
Teams are not used to taking things out of their heads and onto paper. Getting it into a document an AI agent can read is harder still.
We built the agent anyway. Everybody was impressed with the prototype, but nobody knew what to do with it. We had an expensive, powerful tool that changed nothing about how the firm earned money.
The Production Gap is the distance between a pilot that works and a system that changes the economics. A demo that impresses the room and never touches delivery is the most common version I see.
Nobody in that room was bad at their job
The AI model was not the failure. It worked in the demo, and it worked on the day we shipped it. The firm never handed it a job written well enough to run.
The managers were not resisting AI either. Nobody had ever asked them to write their process down for another person, let alone for a machine. A firm that never wrote its delivery process down is carrying real debt. The problem was a gap in the written process, not a verdict about the people.
Where does your real process live?
The real process lives in managers' memories, chat threads, and the exceptions nobody logged.
Your CRM and ticketing system show how the process is supposed to run. The path work actually takes includes the workarounds, the side spreadsheet, the approval that happens in a Slack DM, and the house rules someone made after a bad month. Reworked made this point in April 2026: one engineer told them that an agent trained only on the official process in the ticketing platform is "learning a fantasy". An agent given nothing at all stays a demo.
Documenting the process does not mean building a library of formal procedures that covers only the cases someone remembered to write down. It means writing down how the manager runs the work: how you monitor the people doing it, what you check before it goes out, how you track tasks and projects, and which signals tell you an engagement is healthy or in trouble. Michael Bennett's research on choosing between working explanations found the same pattern in a different domain: the explanation that leaves more possible cases open is the one that holds up when a case appears that nobody wrote down.
What changed when the knowledge was on paper
Once that knowledge was in a document the system could read, the agent started doing real work.
Triggers began producing reports. Those reports warned managers when to step into a project or give a consultant feedback. The model and engineers stayed the same. Only the inputs changed.
Twelve conditions run in the network now, all pulled from how management actually operates, and each is a test: overdue tasks, churn signals from customer meetings, the quality of WhatsApp responses, upsell triggers, and overdelivery signals, meaning something worth making that adds to what the client already bought. Each test flags when somebody needs to act to keep an engagement healthy. None came from a vendor. They all came out of a manager's head during the write-down.
In the chatbot version of the same firm, a manager opens a window and types "draft me a check-in email for this client." The email is fine, but a person still had to notice the client needed one. Noticing is the expensive part. That is the difference between the 80% and the 5%, and I wrote about the rest of it in what an AI employee actually is.

On the Delivery Model Ladder, writing the work down gets a firm off Stage 0 and makes Stage 2 possible. You cannot rebuild a process around AI while it exists only as something three people remember. The AI readiness assessment measures whether the process is written down. That check comes before the changes to who does what that I covered in Fewer Managers, More Managing.
You will not get it right the first time
This is the part that kills AI projects, and it has nothing to do with the technology.
Most software comes ready to install: you configure it and expect it to work. Agents need a different loop. You put one into use, get everyone involved while it is still unreliable, and then read the conversations, errors, and list of things it could not do. From that, you update three things: the software, the rubric (a written scorecard for its work), and the documentation, before repeating the loop.

Documentation is the leg people forget. Your first write-down will be wrong in places because everyone leaves out the step they stopped noticing years ago. Real work finds those gaps. The output-validator workflow covers the rubric side of the loop: who reviews what while the system is still learning.
The hard part is keeping the group committed to that loop. Getting through the weeks when the output is still mediocre is harder than building the system. The work demands a lot, and it more than pays for itself.
What should you document before buying an AI agent?
If your managers went quiet the first time, asking again the same way gets the same silence.
What changed for us was how we asked: Document the process first. Draw the work first, then give an agent the job. We wrote the questions down in advance instead of saying, "Describe your workflow," and somebody sat with each manager and wrote while they talked instead of assigning it as homework. We also allowed the first pass to be wrong, which got people started.
Do this with yourself or your ops lead. Answer five questions in a document an agent could read later:
- What do you check every day and every week when you monitor how your people work?
- How do you track tasks and projects, and where does that live?
- What triggers warn you that something is going right or wrong?
- For one critical deliverable, how do you know it is good enough before the client sees it?
- If one of those checks were done for you, where would the freed hours go?
Question four starts your rubric. Question five decides whether the hours turn into capacity or quietly disappear.
Be honest about the size of this. It is not one formal procedure by Friday. The first pass takes a few sessions, and the loop above runs for a quarter before the output is what the business needs.
Where this stops
This assumes somebody owns the delivery process end to end: an ops lead, a delivery manager, or an owner who still runs delivery. If nobody in the firm can answer these questions because nobody owns the work, the documentation problem is the second problem.
Writing it down once and stopping produces a binder. It only pays off if the firm commits to the loop of putting it to work, watching what happens, and updating it.
Large enterprises use process-mining platforms for a version of this. At $1M to $20M, you do not buy Celonis, a process-mining platform. You interview your own managers and write.
Documenting the process does not put you in the 5% running autonomous agents defined by conditions. It makes getting there possible. That is a smaller claim, and it is the one the evidence supports.


