Delegate the Inputs, Own the Outputs

62% of SMB leaders say they are very confident handing high-stakes work to AI agents. Zero surveyed accounting firms fully trust AI. Both stare at the model. Two properties of the task, can you check it and can you undo it, decide what an agent can run without you.

Sepia drawing of a canal lock: a busy basin of small boats above, one narrow channel out, the lock gate drawn in blue.

How much should you let AI agents run unsupervised? Ask two questions about the task. Can you check the agent's work quickly? If it gets something wrong, is the mistake cheap to undo? Tasks that pass both tests can run without you. A task that fails both should never run alone, no matter which model you bought this quarter.

Most owners answer a different question instead: do I trust the AI? In Upwork's Q1 2026 survey of firms between $1M and $49.9M in revenue, 62% of leaders said they were very confident handing high-stakes tasks to AI agents. The same year, a survey of accounting firms across 8 countries found zero firms that fully trust AI. Not one. Both groups are staring at the model. The answer sits in the task, and in a service firm it comes down to one rule: give agents the input side of the work, and keep a human on everything that leaves the building.

Why a smarter model won't settle it

The common assumption goes like this: when the models get better, we can hand them more. PostHog's Jina Yoon, whose July 2026 framework this article builds on, has the right response: "trusting your agents just because the models got smarter is like skipping your seatbelt because you got a nicer car."

There is also a practical problem with waiting. The capability ceiling will not hold still long enough to base a policy on it. METR measures how long a task AI can complete on its own, and as of January 2026 that horizon doubles roughly every 7 months. On models from 2024 onward, the doubling time dropped to about 89 days. Any delegation rule written as "the model can now handle X" expires within a quarter.

A rule written against the task does not expire. Checking and undoing are properties of your workflow. They stay put while the models churn.

The two questions that set the level

Question one: can you check the agent's work quickly and objectively? Question two: if the agent gets it wrong, is the mistake cheap to undo? The four combinations give you four levels of autonomy. Here is the grid, translated from PostHog's developer examples into the tasks a service firm runs every week:

Easy to check? Cheap to undo? Autonomy level Service-firm example
No No Assistant only. AI suggests, you do the work. Pricing a proposal, scoping an engagement
No Yes Human in the loop. AI drafts, you review every one. Client emails, report drafts, meeting follow-ups
Yes No Delegate with a gate. AI does the work, a check runs before anything fires. Invoicing, anything that auto-sends
Yes Yes Run unsupervised. Chasing missing documents, internal summaries, research pulls

The accounting survey confirms operators already feel this grid even without naming it. The task those firms were most willing to hand over completely, at 68%, was chasing clients for missing documents. A reminder email that goes out wrong costs almost nothing and the result is trivial to verify. That is the bottom-right corner. Meanwhile 62% of the same firms said their trust depends on a human approving anything before it gets sent or filed. Also correct. That is the gate.

Firms are already delegating, then. The interesting question is where it goes wrong.

Where firms get burned first

I've watched hundreds of service firms adopt AI through a franchise network's delivery data, and the first real delegation is almost always the same: fetching and summarizing information across Gmail, Drive, and Slack. On the grid, that is a reasonable place to start. Nothing leaves the building and everything is checkable.

The burns come from the same corner: an analysis that gets a number wrong, a summary forwarded to a partner without a read, a support answer that went out because the queue was long and the reply looked fine. In each case the task was easy to check. Nobody checked.

Easy to check is a property of the task. Actually checked is a property of your process. The grid only protects you if the checking is real, and the volume of checking does not disappear as tools mature. A November 2025 accounting study found agents cut analysis time by 75%, while the authors note the human effort moved to validation rather than going away.

The obvious objection: these are smart operators. Why does the checking fail?

Why AI mistakes slip past smart people

We have a lifetime of training at catching human error. A junior who is unsure sounds unsure. A rushed report looks rushed. The tells are visible from across the room. AI output carries none of them. It arrives confident, formatted, and polished whether it is right or wrong. The mistake hides in plain sight.

Two tower drawings: a cracked leaning tower beside a sleek straight tower with one broken column at its base.

Forwarding an AI response unchecked is also the path of least resistance. It is fast, it takes no thought, and the output looks perfect. Every incentive in a busy week points toward hitting send.

Do not assume you would notice the cost. In a randomized trial by METR published in July 2025, 16 experienced developers worked 246 real tasks with and without AI tools. With AI they took 19% longer. Afterwards, they believed AI had sped them up by about 20%. These were experts evaluating their own work, and their instinct about the checking cost was wrong by nearly 40 points. An operator skimming an agent's report has no better odds on instinct alone.

Validating AI output is a skill, and it is a different skill from validating people. Most firms have not built it yet. The ones that get this right train it deliberately, the same way they train a new manager to review a junior's work. The time to build that muscle is before the cost of a mistake gets high enough to scare the team off the technology entirely.

Delegate the inputs, own the outputs

Until that muscle exists, the safe line runs between inputs and outputs. Give agents the input side of the work with real autonomy: research, information gathering, fact checking, summarization, draft generation, connecting data between systems, and adversarial review of your own thinking before a big call. This is where agents earn their keep overnight, and where a wrong answer costs a redo.

Keep humans on the output side: decisions, final reports, and any direct client interaction. A wrong draft wastes an hour. A wrong deliverable costs a client, and reputation does not restore from backup. The moment work leaves the building, your undo button stops working.

The line is not permanent. A workflow that has been validated and battle-tested can graduate. But the standard for graduation is high, and it applies internally too. Never ship AI slop, not even to your own team. Polish is achievable. A well-built skill can get agent drafts to client grade. A human still owns the send until the workflow has earned it.

An agent that assembles the monthly client report overnight, pulling project data, drafting the narrative, and flagging anomalies, is input-side work and a clear win. The same agent emailing that report to the client at 7am unreviewed is the same work, one step too far. This is what Stage 2, Augmented looks like from the inside: agents execute the input side, a human directs and reviews before anything crosses that line.

How to raise the ceiling on purpose

Autonomy grows when you engineer better answers to the two questions. Make checking cheaper: an AI-specific review checklist, spot-check samples instead of full reads, a second agent assigned to attack the first one's numbers. Make undo cheaper: drafts by default, approval queues, backups you have tested, staged rollout for anything client-facing.

Drawing of a raised walkway through three arched checkpoints, with a rope safety net and ladder drawn in blue.

The undo side is more fixable than most owners assume. Anthropic's February 2026 research on real agent usage found only 0.8% of agent actions were irreversible. Irreversibility is mostly a pipeline choice.

The counterexample is what happens when that choice is skipped. In July 2025, Replit's coding agent deleted a production database during an explicit code freeze. The agent's own admission afterwards: it had "panicked and ran database commands without permission" while the freeze was in place. A hard-to-undo task had been granted self-driving autonomy. The model took the blame in the headlines. The autonomy level was the failure.

So the accounting firms with zero full trust are not behind. They are gating correctly. The losing strategy is waiting for a model that deserves trust, because that model is not coming and the graveyard is filling up in the meantime. Gartner predicted in June 2025 that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear value, and inadequate risk controls. Task design addresses all three.

What to do this week

List your ten most repeated tasks. Ask the two questions of each one. In most firms the list splits fast: three or four tasks an agent could run tonight, a handful that need a draft-and-review setup, and one task an agent is already running that deserves a human gate by Friday.

Start with the unsupervised corner and let the results argue for the next level. If you run the exercise and the split surprises you, reply and tell me what you found. I read every answer.

Topics: