# Human in the loop

Human in the loop means placing a person at a defined point in an AI process to review, approve, correct, or take over the work.

Source: https://rodamora.com/glossary/human-in-the-loop
Updated: 2026-08-09

---

Human in the loop means placing a person at a defined point in an AI process to review, approve, correct, or take over the work. The point is chosen before the process runs. A person at every step is slower and can be less safe because the reviewer stops looking closely. In a 2023 Radiology study, very experienced readers rated 82.3% of mammograms correctly when the AI suggestion was right and 45.5% when it was wrong. The reviewer was present. Presence alone did not prevent the error. A useful loop puts judgment where it changes the result, gives that person authority to stop the work, and leaves cheap steps alone.

## What a real loop looks like in delivery

Take a proposal that an [AI workflow](/glossary/ai-workflow) assembles from a call transcript. Six steps may sit between the end of the call and the document reaching the client. The pricing and scope decision needs a person because a wrong number costs money and the judgment may not be in the transcript. Formatting does not need the same gate. Filing does not either.

That is a designed loop. Someone chose the step, named the person, and gave them time and authority to say no. The other steps can run.

The failed version looks similar on a process diagram. A person appears at the end with an approve button. They face an overloaded queue and cannot tell which item contains the costly mistake. They move through it quickly. The firm carries the model's error rate and pays for a review that changes little.

The difference is not the picture. It is whether the firm decided which step a person's judgment can change.

## Why review becomes a rubber stamp

The main failure is automation bias. People can favor a machine suggestion over their own judgment, even when the evidence in front of them disagrees.

Experience helps less than many firms expect. In the Radiology study, 27 radiologists read 50 mammograms with what they believed was an AI assistant. When the system suggested the wrong category, correct ratings fell from 79.7% to 19.8% among inexperienced readers. They fell from 82.3% to 45.5% among the most experienced. Senior readers held up better. They still got fewer than half right when the machine was wrong. [Experience is not a substitute for a well-placed gate](https://rodamora.com/blog/age-predicts-who-adopts-ai-skill-predicts-who-benefits).

Training alone does not solve it. A 2010 review in Human Factors found automation bias in naive and expert participants. It concluded that the bias “cannot be prevented by training or instructions” and can affect people and teams. The review also found complacency under multiple-task load, when other work competes for the operator's attention. That is the normal working condition for a service-firm reviewer.

A memo that says “review carefully” is not a control. It asks a busy person to supply a behavior the research does not show will hold.

There is a second cost. When the model produces and the person only approves, the work can become a poorer job for a skilled operator. That is the [output-validator workflow](https://rodamora.com/blog/ai-anxiety-and-the-output-validator-workflow). Put the gate where judgment changes the work, not at the end where the only choice is approve or reject.

## Where should the person go?

Two tests help place the gate.

First, can you check the output, and can you undo it? Those properties of the task decide how much autonomy it can carry. Work that is easy to check and reverse may not need a gate. Work that is hard to check or hard to reverse needs a real one. This is the rule behind [delegate the inputs, own the outputs](https://rodamora.com/blog/delegate-the-inputs-own-the-outputs). It is about your workflow, not a model's reputation.

Second, does the person's judgment change the answer, or only confirm it? Scoping, client context, pricing, and promises to a client usually need judgment. Formatting, filing, retrieval, and first-draft assembly often do not. Put the person where they add information.

Name an individual, not a team. “Reviewed by the delivery team” gives no one authority. A named person with time to look and power to reject is a control. Record that decision in the [one-page AI policy](https://rodamora.com/blog/the-ai-policy-template-that-fits-on-one-page).

## What does it cost to run?

The obvious cost is reviewer time. The larger operating issue is that [review stays human and sequential](https://rodamora.com/blog/the-four-stages-of-reviewing-ai-work) while production gets faster.

AI can raise output quickly. Review capacity stays flat. The queue grows, and the reviewer can become the bottleneck within about a month. That is one form of the [Production Gap](/production-gap). Firms often loosen the review when the queue grows. The quality problem appears later in work a client sees.

Measurement is another cost. Without [a repeatable test that scores output against a written standard](/glossary/eval), you do not know what share of bad outputs the gate catches. Seed known errors into the queue without warning and count the catches. A gate that approves nearly everything may be telling you that the review is weak.

Keep the reviewer sharp too. Someone who sees only machine-produced work can lose the feel for what good looks like. That judgment is the skill the control relies on.

Give the reviewer the context needed to make a decision. They need the source record, the task's acceptance check, and a clear way to send work back. A button that only says approve hides the decision and makes later review of misses harder.

## Where this sits on the Delivery Model Ladder

Choosing where a person sits is Stage 2, Augmented, on the [Delivery Model Ladder](/delivery-model-ladder). Stage 2 asks what the team no longer does by hand and what the redesigned review now checks. A firm that adds AI but leaves its old approval at the end remains at Stage 1.

A working loop also enables Stage 3. An agent can take a defined slice of delivery only when someone can verify the result. The loop is a precondition for [AI-native](/glossary/ai-native) delivery, not a temporary scaffold to remove.

Across the 150+ franchise units in the delivery data I track, roughly 80% still run chatbots, 15% run rigid workflows on triggers, and 5% run autonomous agents. Those numbers describe that network. The step that separates the second group from the third is whether checking has been designed.

An April 2026 Census Bureau working paper found that 66% of AI-using firms rely on AI only to augment existing tasks. That is a broad measure, not a test of review quality. It does fit the operating pattern: most firms are adding a person and a tool to existing work before they redesign the loop.

## Quick answers

**Is human in the loop the same as human oversight?** They overlap. Oversight often means someone is accountable at a governance level. In the loop means a named person at a named step who can stop the work. The operational version changes the result.

**Should the loop shrink over time?** It can, with evidence. If a step runs for months and a measured catch rate stays near zero, move from every item to a sample. If the catch rate is real, keep the gate. Do not loosen it only because the queue grew.

**What is the cheapest thing I can do this week?** Choose one AI-assisted deliverable. Write down which step carries the costly mistake, who owns that step, and how many errors they caught last month. The third answer tells you whether the gate is being measured.
