The Four Stages of Reviewing AI Work
Adding AI to delivery makes most firms slower, because production got cheap and checking did not. Reviewing AI work moves through four stages, and most firms are stuck in the second one, paying for the volume and getting none of the gain.
Your firm produces more work than it did a year ago and moves slower. Production got cheap. Checking did not. Atlassian's 2026 survey of 12,000 knowledge workers and 170 Fortune 1000 executives found 89% of those executives say AI increased the speed of work while only 6% can point to organization-wide return. The extra output "backs up at reviews, approvals, and other human-judgment gates."
Reviewing AI work passes through four stages: marvel, where nothing gets checked; panic, where everything gets read twice; rubric, where AI does most of the checking against a written standard; and maturity, where less review ships more work. Most firms are stuck in the second.
The third stage takes weeks to months of tuning. Across many projects in a franchise network's delivery data, one week of output above 300 artifacts came back with roughly 95% fewer marked as carrying incorrect information, compared against the same system before tuning. On the Delivery Model Ladder, this is the Stage 1 to Stage 2 jump.
Why does more AI output make a firm slower?
You can see the slowdown in a normal week. The same two reports keep coming back around. A partner reads line by line because one wrong number once reached a client. Two people check the same document because nobody is sure who owns it, and neither wants to be the one who let it through.
None of that appears in a status meeting. It appears as senior people staying late and a delivery calendar that stopped improving. The hours AI freed are going into verification, which is why the firm's profit and loss has not moved. I've written about that leak before in The Extra Time AI Buys You Is Already Gone.
The worst version is quieter. When people use AI without saying so, nobody knows the work needs a different kind of check, so no check gets built for it. That failure is one of the eleven in The Production Gap, and it is the one this article is about: the amount of work climbs, review time stays flat, and rework eats the gain.
Why doesn't hiring another reviewer fix it?
Your team is not over-checking out of nervousness. After a bad piece of work reaches a client, reading everything is the only tool they have. They are using it correctly. The tool is just the wrong size for the amount of work.
The queue math explains why a second reviewer disappoints. Average wait time runs roughly as 1 divided by (1 minus utilization). Here, utilization means how much of a reviewer's time is already booked. When a reviewer is booked 50% of the time, waiting has already doubled against an idle system. Past about 80%, wait time rises almost straight up. A senior reviewer who is 90% booked has almost no room left. They are on the steep part of that curve, so one more body moves them only a short way down it while volume keeps climbing.
Hire a reviewer and you buy a few months. The bottleneck moves somewhere else, which is the pattern I described in Constraint Whack-a-Mole.
What are the four stages of reviewing AI work?
Across the firms I've watched through a franchise network's delivery data, review passes through four stages, and they arrive in order. Most firms I see are sitting in the second one and think it is permanent.
| Stage | What checking looks like | What it costs | How you know you're here |
|---|---|---|---|
| 1. Marvel | Output goes out barely checked. Worst where AI use is unofficial. | Errors reach clients. | Nobody can say who or what checked the last deliverable. |
| 2. Panic | After something goes wrong, everything is read line by line, sometimes twice. | The hours the automation freed. | Your most expensive people spend their day reading. |
| 3. Rubric | AI does most of the checking against a written standard. | Setup and weeks of tuning. | Review is a step in the workflow, not someone's evening. |
| 4. Maturity | Better rubrics, stronger models, history to compare against. | Model spend. | Less review, more work shipped, quality holding. |
Stage 2 is where Stage 1 firms stall on the Ladder. Stage 3 is what the Stage 2 rebuild looks like when someone finishes it. The two glossary terms that do the work here are eval, which scores output after it exists, and guardrail, which blocks it before. A tuned system uses both.
This is a sequence I've watched, not a scorecard. There is no percentage attached to each stage and no promise that a firm reaches the fourth.
Why does it get slower before it gets faster?
A new AI system makes a lot of mistakes at the start. Review load goes up, not down. I've watched teams operate slower for a stretch while the thing they were promised would speed them up was still wrong most days.
That stretch is where firms quit. It looks exactly like proof that AI does not work for their kind of work, and the conclusion arrives when the evidence is at its worst. The teams that persisted got the gain, after a few weeks to a few months of tuning.
The slow period comes with no schedule attached, and the firms that stopped were not careless. They saw a real signal and drew a reasonable conclusion from it. The signal was just early.
What does tuning actually involve?
The checking layer needs several parts working together. Four parts do most of the work: agents that fact-check claims against online sources, agents that cross-check assertions against the client's own data, checks against error patterns you have already seen, and guardrails for common AI language and logic failures on repetitive tasks.
The loop that improves them is cheaper than it sounds. Operations mark a piece of work as defective, which takes seconds. The implementation team gets that feedback immediately, works out the fix, and ships a new version. Operations test again, and the loop goes around.
Marking a defect takes seconds. Reviewing every piece takes the week. That gap is why you start: review time falls with each iteration, and the loop pays for itself while it runs. A firm in stage 2 already pays for the costly half and gets nothing back.
Operations sees the defects, and implementation fixes them, so both have to take part. Feedback has to reach the people making the fix immediately. A defect logged in a spreadsheet that someone reviews monthly is not a loop.
Be honest with yourself about what automated checking can do. In the published research, automated checkers are much better at passing good work than catching bad work. One study of AI-run consistency checking found the model "can successfully retrieve more than 95% consistent summaries, while performing rather poorly on identifying all the inconsistent ones". On a test for made-up facts, the best-performing model reached 58.5% accuracy at separating factual from fabricated summaries. The authors call that "barely above chance."

The 95% result at the top came from a tuned system with several checking layers and a human loop around it, measured against that same system before tuning. A single model scoring your work leaves you near a coin flip, which is what a single AI reviewer sells.
Rubrics are where most of the quality comes from, and they are unglamorous. Two findings should shape how you write them. Adding rubric items buys little: scaling the number of rubric items, whether written by people or by a model, "yields only marginal gains". A model generating a rubric without access to reference answers produces a noticeably worse one, while rubrics generated with those references perform about as well as human-written ones. Someone who knows the work has to supply the standard either way. The best use of an hour in this project is having your best reviewer write down what they check for.
Do you have to tell clients that AI did the work?
You do not have to advertise it. You must never lie about it. What matters more than either is that the validation is real.
The evidence on disclosure is uncomfortable. Thirteen controlled experiments published in 2025 found that people who disclose AI use are trusted less than people who do not, including for analyst-type work. Framing the disclosure differently did not prevent it. Making it voluntary or mandatory did not prevent it.
Two findings decide the question. Being exposed costs more trust than disclosing. And in a 2026 study of 370 marketplace users, reviews labeled AI-assisted were judged more favorably than reviews labeled AI-generated. That was consumer reviews rather than professional work, so use it as a clue, not a measurement. You can only describe your work the better way if you kept a record of who checked what.
The courts had to answer this faster than anyone. As of August 2026, a tracker of standing orders, local rules, and decisions on AI use holds 823 entries, and only 163 of them require disclosure or verification. The Fifth Circuit declined to adopt an AI-certification rule at all, after attorneys pointed out that existing rules already make them responsible for the accuracy of what they file. The question that stuck was not whether AI was used. It was who verified this.

That matches the position I put in the one-page AI policy: client terms decide disclosure, and when a client asks, the answer is the truth. Build the checking layer for your reputation rather than for compliance. A label with nothing behind it is worse than no label.
When do you stop tuning?
There is no finish line. As models improve, you take some guardrails out, because the model stops making the mistake the guardrail was catching. That better model costs more per unit of work.
The endpoint is a measured tradeoff. Pay more for a capable model that errs less, or pay less for the model and carry more human review. Both are costs of delivery. Every firm lands somewhere different, and the only way to find your spot is to measure both sides on your own work.
One warning from The Four Numbers. When tuning frees up review time, decide where that time goes before it appears. Otherwise the work around it absorbs the time and the margin never moves, which is the same evaporation that catches firms at Stage 1.
When does this not work?
This assumes the work is repeatable enough to write a rubric against. A one-off strategy engagement with no precedent has nothing to check against yet. The first move there is to write down how the work gets done. Build a checker once the work is defined.
It also assumes someone owns the implementation side. The loop needs a person who can ship a new version when operations flag a defect. A firm without anyone in that seat stays in stage 2, whatever its intentions.
Most professional-services deliverables still cannot be checked by a script. That is why this is a layered system with people inside it rather than an automated pipeline, and why the human review step survives at stage 4. It gets smaller and better aimed.
Pick the deliverable your firm produces most often. Find out whether anyone can say who or what checked the last one. That answer tells you which stage you are in. The AI readiness assessment asks the same question about review capacity alongside the rest of the checks.
FAQ
How long does it take to get from stage 2 to stage 4?
A few weeks to a few months of tuning, in the delivery data I've reviewed. The technology does not set the timing. Feedback speed does: how fast defect feedback reaches the person who can fix it.
Can we skip straight to stage 3 with a new AI system?
You can start there by building the checking layer alongside the workflow instead of after it. Most firms do not, because stage 1 feels fine until it doesn't, and the cost of stage 2 only becomes visible once volume climbs.
Is one AI checking another AI reliable?
Not on its own. Automated checkers pass good work far more reliably than they catch bad work, and the best model on one test for made-up facts scored barely above chance. Reliability comes from layering several checks with different jobs and keeping people in the loop on what gets flagged.
Who should write the rubric?
Whoever currently catches the most errors. A model can help draft it, but the research shows rubric quality falls off when the model has no expert reference to work from. The standard matters more than the number of criteria.
Should we tell clients we use AI?
Do not advertise it, and never deny it. Answer honestly when asked, put the position in your engagement terms, and make sure the human verification you describe actually happens.
Does this mean we stop reviewing work?
No. Review gets smaller and more targeted as the checking layer matures. The senior read moves from every document to the ones the system flags and the ones where judgment is the actual product.
What if we tried this and it made us slower?
That is the expected first result. The question is whether defect rates are falling between versions. If they are, keep going. If nothing improves after several iterations, the feedback loop is broken somewhere, usually between operations and implementation.
For more on what I see in the delivery data, subscribe to the newsletter.


