Constraint Whack-a-Mole: What AI Adoption Feels Like

If AI makes ten reports and your team can safely check five, more drafts won't get more work to the client. Start by managing the whole queue.

A reviewer checks a report beside a waiting stack and a nearly empty outgoing tray.
A faster draft still has to get through the person who can clear it for the client.

You can make your team faster with AI and end up with work taking longer to reach the client, because the people making it are no longer the people setting the pace.

The drafts arrive faster, but somebody still has to check them, approve the promises and decide what can leave. If that part of the job hasn't changed, the extra work waits there.

What I have seen work is to keep following that pile, find what is holding the work up, decide whether the step is needed, and improve it if it is. Then you look again, because once that part works better, something else becomes the limit.

That is what I mean by constraint whack-a-mole. The constraint is simply the part of the job that limits how much useful work gets through, and using AI doesn't remove the need to manage it.

I have watched implementations get dismissed because the first results were poor while everything around the tool stayed the same. Other teams stayed curious about where the work was getting stuck and changed the functions and workflows around it.

I'm not saying a flat result proves the tool is good. It tells you to look at the whole job before deciding, because a faster first draft and a faster service are different results.

Think about ten reports as an example, not a measured case. AI can help write all ten this week, but the partner can safely check only five.

If you start all ten, the reports don't become delivered work just because the writing is finished. Five are waiting for a person whose week was already full, and some of that work may need updating before the person gets to it.

The team can be doing good work throughout this. The junior produced what you asked for, the partner is checking carefully, and the client still gets the report late because you started more than the whole process could finish.

So I would not begin by telling the reviewer to hurry up. You set the amount of work going into a process that depends on their time, and that is something you can change.

There is research from software delivery that makes this distinction visible. Google's 2024 DORA report, based on nearly 3,000 working professionals, associated a 25 percent increase in AI adoption with about 2.1 percent higher individual productivity.

The same increase was associated with about 1.5 percent lower delivery throughput and 7.2 percent lower delivery stability. In plain words, the individual measure improved while the measures of getting software out reliably got worse.

These are associations in software teams, not proof that AI caused a slowdown in service firms. DORA suggested that AI might be encouraging larger batches of code, which are harder to move safely, and explicitly treated that explanation as a hypothesis.

The service-firm version I want you to check is simpler. Have you increased how much work people can make while leaving review, approval and client decisions sized for the old amount?

There is a useful difference between work waiting to start and work waiting after someone has made it. A request on a list still needs attention, and a client waiting for it can have a real cost, so I wouldn't call that waiting free.

But you haven't yet spent the production time on a version that may go stale. If priorities change, you can move the request before someone has written the report with numbers that are no longer current.

That is why I prefer to hold work before production when review has no room for it. It is easier to choose the next thing carefully than to keep refreshing a pile of finished drafts nobody can release.

The first question at the stuck step is whether it needs to exist at all. A status report nobody uses may be something you can stop making rather than something that needs a better AI prompt.

Be careful with that question around controls. A step can protect the client without appearing in the final deliverable, so not seeing it in the report is not a reason to delete it.

Ask who uses the step, what decision depends on it and what would go wrong without it. If it is needed, work on making it easier to complete without losing its purpose.

For repeated work, I would make the queue visible in four states. There is work waiting to start, work being made, work being checked, and work that has passed the check and can go to the client.

The important part is that being made and being checked share one limit. Moving a draft from production to review doesn't create room for another draft, passing the check does.

Suppose you set that shared limit at five reports and three are still open from yesterday. There are two slots available, not five new starts because the calendar changed.

A limit on open work is different from the number you finish in a week. You need to watch both, because five open reports that take a day to clear behave very differently from five that take a month.

The person who owns the job from request to client sets and holds that limit. The reviewer tells them how much checking the team can safely handle, but the reviewer should not also have to chase everyone who starts work outside the agreement.

If the owner makes an exception, I want a reason, a named person responsible for it and a date it ends. Otherwise the exceptional workload becomes the normal workload without anyone choosing that.

Counting reports works only when the reports require roughly similar effort. If a custom analysis takes four times as long to check as the weekly report, counting both as one item hides the work, so give the custom job more of the available room or track checking hours instead.

I would begin by watching two full cycles of repeated work, often around ten working days for a weekly process. That is a starting window, not a guarantee that ten days captures every kind of variation in your firm.

Write down what arrived, what passed the final check and what remains open. Also look at how long the oldest item has waited, because a steady total can hide the same difficult file being passed over again and again.

My warning line is a queue that grows for two weeks running, by roughly 10 percent of a normal week's output each week. If you normally finish 50 similar items, that would mean five extra items left waiting one week and another five the next.

That is my rule of thumb, not a threshold measured across the franchise network. Its purpose is to make you act on a growing pile before everyone gets used to working around it, and a missed client deadline deserves attention before the threshold is reached.

Before raising the limit, I would want two steady cycles in which work finishes in the expected time. The oldest item must remain within the review window and within what you promised the client.

I would also check that rework hasn't risen, serious defects aren't reaching clients and quality is staying in the range you agreed. Faster completion bought by skipping checks is not extra capacity.

Then ask the reviewer what it took to get there. If they repaired the reports after dinner or worked through the weekend, those clean delivery numbers are borrowing time from a person, they are not showing what the process can sustain.

When those conditions hold, add a little room, one slot for comparable work, and measure again. I would not double the amount coming in after two good weeks and assume the next part of the job can absorb it.

Now, limiting starts protects the reviewer, but it doesn't by itself make the checking better. For that, you need to move some of the checks into the work instead of leaving everything for one person at the end.

Through a franchise network's delivery data, I have seen review move from people to agents as the checks improved. Meeting summaries went from human review to agent checking, research review began matching quotes to references automatically, and data analysis moved away from heavy manual checking toward agent reviewers.

The checks were specific, claims against online sources, results against the client's own data, recurring mistakes from earlier runs, and the familiar ways AI gets language or logic wrong. When the delivery team found a bad piece, the implementation team changed the workflow and the delivery team tested the change.

Across many projects, one week of output contained more than 300 finished pieces. After tuning, roughly 95 percent fewer pieces were flagged for incorrect information or data compared with the same system when it first went live.

That result concerns flagged quality problems. It does not establish a 95 percent reduction in all errors, and the comparison did not measure how many human review hours, people or additional delivered pieces were involved.

I want that limit clear because this is exactly where an owner can read a quality improvement as a staffing saving. The evidence says the checking got better at the problems being tracked, you would need another measurement to decide what capacity that released.

And better automatic checks still leave decisions that belong with people. A proposal can have correct totals, complete sections and acceptable standard terms while making a promise your firm should not make.

AI can prepare those checks and flag the unusual parts, but a named person owns the promise. They decide whether the scope is realistic, whether to accept the margin and who is responsible for the part nobody has properly assigned.

Some cases also need to bypass the ordinary queue. In one case in the network's delivery data, a major client's service complaint was tagged as a billing question and sat in a low-priority queue for 11 days.

The consultant learned about it when the client contacted them directly. The client had asked to cancel and later left, though I cannot say the delay was the only cause.

Within 24 hours the rule changed so cancellation language, threats to leave and messages from strategic accounts went around the normal sorting and reached a person. The routine flow stayed automatic, but the dangerous cases had a different route.

That is the kind of distinction I would make in your review work. Repeated low-risk work can use automatic checks with samples and alerts, more consequential work needs a person looking at the important parts, and high-risk promises stay with a named senior person who signs them off.

Before changing the number of juniors per senior, count the senior minutes those different kinds of work need. Also make room for juniors to see the difficult decisions and hear why they were made, because routing every hard case straight past them can remove the work through which they learn.

Once review gets faster, expect the next wait to become visible. Reports may now clear internally but sit with a client, support cases may reach the right queue but wait for the one person allowed to resolve them, proposals may be ready while the owner has no time to decide.

Follow the work to that point and ask the same questions again. You are managing a service from beginning to end, so a local improvement is useful only if the rest of the job can use it.

This approach fits repeated work where you can measure checking and someone owns the whole job. One-off work still needs judgment, and a short rush needs real checking time booked, a hard cap and an end date.

Pick the workflow where completed drafts are waiting now, and sit with the person who clears them. Before starting more, find out how much open work they can safely handle and what is keeping the oldest item there.

Topics: