Framework · Rod Amora ·
The Delivery Model Ladder
A staging model for how service firms move from human-delivered to AI-native delivery. Each stage is defined by one thing: what the P&L shows.
Most service firms think they know where they stand with AI. Most are wrong.
The Delivery Model Ladder is a four-stage model: Stage 0 Tool-assisted, Stage 1 Augmented, Stage 2 Upgraded, Stage 3 AI-native. It locates a firm by what its profit and loss statement shows. Not by which tools it bought. A firm can have every employee using AI daily and still sit at Stage 0 or Stage 1. Adoption and economic change are different things.
What matters is the signature: a specific, checkable pattern in cost to deliver, retention, margin, or revenue per person. As of July 2026, in the delivery data I've reviewed across a franchise network, most firms I see cluster at Stage 0 or Stage 1 while believing they're further along. The jump from Stage 1 to Stage 2 is the hard one. It requires upgrading how the work gets done, not adding a tool to how it already gets done.
The Delivery Model Ladder locates a service firm's AI adoption by economic signature rather than tool stack. It is built for owners of $1M–$20M service businesses. As of July 2026, it runs from Stage 0 Tool-assisted (no P&L change) through Stage 1 Augmented (cost to deliver drops, economics flat) and Stage 2 Upgraded (quality and customer outcomes improve; the bottom line moves) to Stage 3 AI-native (shared context; revenue and quality detach from headcount and hours). Find your stage by matching your numbers, not your intentions.
What stage of AI adoption is my firm actually in?
You don't find your stage by counting AI subscriptions or asking staff if they "use AI." You find it by looking at four numbers: cost to deliver, retention, margin, and revenue per person. That is what I call The Four Numbers elsewhere. Each stage of the Ladder is confirmed by which of those four numbers has actually moved, and by how much.
If none of the four numbers has moved, you're at Stage 0, regardless of how much AI activity is happening inside the firm. If cost to deliver has dropped but nothing else has changed, you're at Stage 1. That is the stage most operators mistake for progress.
Why doesn't "everyone here uses AI" show up on the P&L?
Individual use is Stage 0 by definition. Stage 0 (Tool-assisted) means people in the firm use AI on their own, in their own way, for their own tasks. No workflow changed. No engagement was repriced. No staffing plan was touched.
Teams do not share knowledge, tools, or context with each other. There is no centralized organization knowledge. There are no processes that wire work into skills or preset workflows. Nobody curates context or a company ontology. Each person rebuilds the same ground from scratch.
The signature of Stage 0 is silence: no measurable change in any of the Four Numbers. This is a starting point, not a failure. The mistake is treating it as an arrival.
Why do most service firms stall at Stage 1?
Stage 1 (Augmented) is where AI gets embedded into existing workflows. A proposal that took four hours now takes seconds. A first-draft report that took a day now takes minutes. Cost to deliver genuinely drops.
But across the firms I've observed through the franchise network's delivery data, the gain almost always leaks away before it reaches margin. Staffing ratios don't change. Engagement scope doesn't change. Fewer hours worked, at the same headcount and the same quality, just means idle capacity.
This is the single most common stall point I see. The signature of Stage 1 is internal efficiency without economic change: cost to deliver down, the other three numbers flat. It is also the most common failure mode I catalog in The Production Gap. The gain shows up as an anecdote ("we're saving fifteen hours a week") and nowhere in the financials.
What actually changes at Stage 2, and why is it the hard jump?
Stage 2 (Upgraded) is where the business starts to deliver more efficiently and effectively, with a focus on quality, customer experience, and retention. Stage 1 is adoption. Stage 2 is building an upgraded company: someone decides what the delivery team no longer does by hand, what the review process now checks for, and who gets added to or removed from the delivery chain. The workflow itself gets rebuilt around AI instead of AI being bolted onto the old one.
What Stage 1 leaves missing is guidance and enablement: how to put the newly freed time into higher-impact work. That means improving core business functions, raising customer experience, and increasing the quality of what gets delivered. Better sales proposals. Better onboarding. More thorough reports. More testing. More speed and flexibility. Free time has to be measured and aimed at what drives growth and profitability: overdelivery that cuts churn and improves sales and your brand value.
You know you've reached Stage 2 when the company does things noticeably better than before. Customers start to notice. Reviews and praise arrive. Retention improves. Repeat purchases rise. The bottom line gets better.
Stage 1 to Stage 2 is the hardest transition on the Ladder because it needs leadership, team alignment, and buy-in. It needs enablement so people can see the new possibilities and build a culture where the new way is the right way. It cannot be delegated to a tool purchase.
Three unglamorous things separate the firms that make this jump from the ones that keep trying. One named person owns it. Large organizations answer this with a task force or a center of excellence; a $1–20M firm can't staff that and doesn't need to. What it needs is the same function in one body: someone whose job includes noticing which experiments worked, deciding what gets standardized, and making sure the answer reaches everyone else doing that job. The best candidate is almost never the most senior person. It's whoever is already doing this on their own time — most firms have one, and elevating that person costs nothing and works better than hiring for it.
Someone writes down the boundary. A short, permissive statement of what's allowed — these tools, this part of the work, this data never leaves — outperforms both silence and a restrictive policy nobody can find. Silence doesn't produce caution, it produces the Shadow Rollout, because people who can't find the rule improvise one. The useful frame is a sandbox with guardrails: the sandbox is real permission to experiment inside actual client work, the guardrails are the short list of things that stay outside it. Firms treat this as the bureaucratic step and it is closer to the opposite — it's the thing that lets experimentation happen where you can see it.
The time to learn is on the calendar. Stage 2 asks people to work a new way while still delivering the old way, and that transition cost has to be funded in hours or the redesign stays theoretical. This is the Unfunded Hour in The Production Gap, and it's the quietest reason Stage 1 firms stay Stage 1. Equally, the flood is a real failure: several initiatives launched at once produce fatigue and disengagement, not momentum. One workflow at a time, owned and finished, beats six announced.
What does Stage 3 look like, and how rare is it?
Stage 3 (AI-native) is uncharted territory. There are no best practices left to copy and no model to follow. It's time to reinvent. Revenue can decouple from headcount. Service delivery speed and quality are no longer attached to hours worked. The company is the sum of its shared context, and agents have access to all of it.
Most business functions run from a chat interface. Documentation, handbooks, and best practices are stored so agents can access, change, and improve them on their own. Decisions, entities, and processes are encoded in ontologies. Chats, decisions, and workflows are recorded. Timelines get built. Derived context accumulates. A new knowledge layer exists to cut complexity and drive work for both agents and humans. What humans add now is taste, connection, and judgment.
The early accounts of this stage share a telling pattern in how the results get reported: per-person output multiplying while headcount grows slower, quality metrics held flat, cost to deliver down. Not hours saved. When a company genuinely changes its delivery model, the evidence shows up in per-person output, quality, and cost — the same numbers the Ladder is built on — and the firms living it instinctively report it that way.
The firms exploring this stage today are also the least representative ones: AI-native from birth, staffed by people who build with these systems for a living, and working in domains where output is easy to verify. Against the economy-wide baseline in The Production Gap — 89% of executives reporting no labor-productivity impact, 66% of AI-using firms augmenting existing tasks and nothing more — they are extreme outliers, not a preview of the median. That's exactly what makes them useful as landmarks: they show where the top of the Ladder is, not where the middle will be next year.
I've only seen a couple of companies starting to explore this stage so far, and both are early-stage firms born in the AI era. Unlocking Stage 3 will be the hardest challenge for established service companies.
What does it actually take to get to Stage 3?
Reading the early accounts of this stage against the service firms in my data, five requirements stand out. None of them is a purchase.
1. A delivery-model decision, not a tool decision. The firms operating at this stage didn't buy an AI product and roll it out. They built the operating layer first — access policies, audit logging, permissions, a way for agents to reach the systems where work actually lives — and then rebuilt how work flows through it: agents execute, humans direct and review. That is the Stage 1-to-2 jump taken all the way to its conclusion, and it confirms why the jump can't be delegated to procurement. A service firm that responds to this stage by buying "an agent tool" lands in the Shadow Rollout or Perpetual Pilot patterns from The Production Gap, with better branding.
2. Work decomposed into verifiable units. The work agents absorb first has a specific shape: high-volume, repeatable, and checkable. An agent can run a loop on that work because someone can verify the output. Every service firm has candidates: research memos, first-draft deliverables, QA passes, proposal assembly, engagement status reporting. The Stage 3 precondition is knowing which units of your delivery are verifiable and who owns reviewing them. This is also the honest limit for service firms today: the domains where this stage exists are the easiest ones to verify, and most professional-services deliverables aren't checkable by a script. Whoever works out what a verifiable unit of consulting work looks like gets the leverage first. Nobody has published that answer yet.
3. Judgment applied across more work than the person could ever touch. The structural ceiling on $1–20M firms is founder leverage: the founder stays the primary rainmaker and the senior people are the delivery bottleneck, because a person's judgment only reaches the work they personally touch. Stage 3 is the first credible mechanism that breaks this. One senior person directing agent loops is the leverage model of the big pyramid firms rebuilt without the pyramid: partners direct, agents do the associate-tier work — the same research-heavy junior work whose margin AI was already eroding. People go from doers to directors. There's a second way to see the same shift. In a traditional firm, the intelligence is spread across the people, and the structure exists to route it: managers carry context up and down the chain. That's why headcount was never just production capacity — a share of every hire goes to coordination. When the shared context lives in the system, the system does the routing, and the people move to the edge of the work, where judgment meets the client: the calls a system shouldn't make on its own, the trust, the taste. One senior person directing agent loops and a firm whose context lives in the system are the same claim from two directions — one from delivery economics, one from how information moves through an organization. Both land at the pyramid's leverage without the pyramid's layers. On the Ladder, that's what "revenue per person detaches from headcount" means in practice.
4. Shared context before autonomous delivery. In the accounts of this stage, adoption didn't spread by mandate or by a flashy client-facing pilot. It spread because people watched colleagues getting real answers from an agent, and the first wins were mundane: questions answered against the company's own knowledge, so people stopped waiting on other departments. That ordering matters. The knowledge layer — documentation agents can read, decisions that are recorded, context that accumulates — comes before agents touch anything a customer sees. For a service firm, the realistic entry point to Stage 3 is internal: make the firm's own knowledge agent-accessible and watch where usage spreads, rather than launching an AI-delivered service on day one.
There's a precondition hiding under this requirement: shared context only exists if the work already produces recorded artifacts. A firm where decisions happen on calls, plans live in someone's head, and client context sits in one person's inbox has nothing for an agent to read. A firm where decisions, plans, and delivery steps get written down as a byproduct of doing the work already owns the raw material for its knowledge layer. The test is checkable today: could someone reconstruct last month's key decisions from what's written down? If not, the first Stage 3 move isn't an agent. It's making the work machine-readable.
This is where the interesting inversion sits. Undocumented process is one of the most common blockers on AI adoption, and it fails in a specific way: you cannot automate a workflow that exists only inside someone's head, because there's nothing for an agent to read and nothing for anyone to verify against. Which means documentation stops being hygiene and becomes the rate limiter — the better documented the process, the faster anything built on top of it produces value. There is no such thing as too documented here, and that is a genuinely new statement about a task most service firms have always treated as overhead they'll get to later.
The cost of doing it has also collapsed, which is the part firms haven't updated their assumptions about. Documenting a process used to mean someone spending a week interviewing people and writing a manual nobody reads. It now means recording someone's screen while they do the job they do every week, handing the recording to an LLM, and getting back a written process — plus, usually, an unsolicited list of the places the human version is already inefficient. That requires no custom software and no project. A firm that can't yet answer what a verifiable unit of its delivery looks like can start here, this week, with the workflow its most experienced person runs most often.
Two things follow about where to point the first agent. Start internal, and start dull. In the accounts of firms that got this right, the use case that converted skeptics was not a flashy client-facing pilot — it was a bot that answered questions against the firm's own knowledge, so people stopped waiting on other departments for answers that already existed somewhere. It pays twice: once in the time recovered, and again in the resistance that doesn't materialize for everything you deploy afterward, because people have now personally experienced the thing working. The other reliable early candidates are back office and support functions, which have the useful property that the benefit is actually calculable — which matters more than it sounds, because a first win you can put a number on is what buys permission for the second one.
5. Proof in the Four Numbers, holding quality constant. The discipline that makes a Stage 3 claim credible is that every output claim is paired with a flat quality metric: more delivered, same rework rate; faster turnaround, same escalation quality. That pairing is the whole test. A firm claiming Stage 3 progress on volume alone, with no evidence that quality held, is describing a Quality Debt Spiral from the outside. The Stage 3 signature is not "we do more" — it's revenue per person up while retention and margin hold or improve, measured across quarters, in the financials.
The sequence matters as much as the list. Items 4 and 2 come first and are cheap: record decisions, make knowledge accessible, name the verifiable units of your delivery. Item 1 is the leadership commitment that turns them into a delivery model. Items 3 and 5 are the payoff and the proof. A firm at Stage 1 today doesn't get to Stage 3 by aiming at it — it gets there by making the Stage 2 upgrade decision and then compounding it until the economics detach from headcount.
How do I use the Ladder without buying anything?
Read the four signatures below against your own numbers this quarter, not your plans for next quarter. Most operators can place themselves within a few minutes once they stop asking "how much AI do we use" and start asking "which of the four numbers moved, and by how much."
| Stage | What's happening | Economic signature | How to tell you're here |
|---|---|---|---|
| Stage 0: Tool-assisted | Individuals use AI on their own initiative | No measurable change in cost to deliver, retention, margin, or revenue per person | Staff mention AI in conversation; none of it appears in a monthly report |
| Stage 1: Augmented | AI is embedded in existing workflows | Cost to deliver drops; retention, margin, and revenue per person stay flat | You can name hours saved but can't point to a margin line that moved |
| Stage 2: Upgraded | The firm delivers more efficiently and effectively; workflows are rebuilt around AI | Quality rises; customers notice; retention and margin improve | Work is visibly better than a year ago, praise shows up in reviews, and the P&L reflects it |
| Stage 3: AI-native | Uncharted territory: shared context, agent-accessible knowledge, reinvented delivery | Revenue per person detaches from headcount; speed and quality detach from hours worked | Context is shared and agent-accessible; growth no longer tracks hiring |
How does the Ladder relate to the Four Numbers and the Production Gap?
The Ladder tells you where you are. The Four Numbers tell you how to measure whether you're moving. The Production Gap tells you why firms get stuck between stages. They're meant to be read together. A firm stuck at Stage 1 is almost always caught in one of the eleven failure patterns I catalog in The Production Gap, most often the one where hours drop but nobody decided in advance where the reclaimed capacity would go.
I write about this because I've watched hundreds of service firms adopt AI through a franchise network's delivery data. Not from the outside as a vendor. From inside the numbers firms report every month. For the two lenses that go with this model, read The Four Numbers on measuring movement between stages, and The Production Gap on why firms get stuck. I write about what I see in that data in the newsletter.
FAQ
How do I find out which stage my firm is at?
Look at your last two quarters of financials, not your AI usage logs. If none of the Four Numbers moved, you're at Stage 0. If cost to deliver dropped but retention, margin, and revenue per person stayed flat, you're at Stage 1. Past that, retention and margin have to show it: customers staying longer and a bottom line that improved. You'll find that in your financials and renewal data, not in a chat history.
Can a firm skip a stage, for example going straight from Stage 0 to Stage 2?
It's uncommon in the pattern data I've reviewed. Stage 2 requires an upgrade decision usually triggered by frustration with Stage 1's flat economics. Most firms pass through some period of Stage 1, even a short one, before anyone decides the workflow itself needs to change.
Is Stage 1 a wasted stage?
No. It's where a firm learns what AI is actually reliable for inside its own delivery process. That information is necessary before any upgrade. The mistake is staying at Stage 1 and calling it done.
Does more AI spending move a firm up the Ladder faster?
Spending alone does not. Spending shows up as tool adoption, which is exactly what Stage 0 and Stage 1 look like from the outside. Movement up the Ladder requires a decision about pricing, staffing, or scope. Those decisions cost attention and authority, not subscription budget. The binding constraint in most firms I have seen is not the budget line at all. It is hours: people who would use the tool and have not been given time to learn it.
Are we behind because we never automated anything before?
Probably the opposite. Firms carrying a decade of earlier automation have to unlearn it twice: the systems have to come out, and the people have to stop thinking in the old patterns, which is the slower of the two. A firm with nothing to rip out gets to build directly against how these tools actually work. Never having automated is a clean starting position, not a deficit, and it is a common position for a $1–20M service firm to be in.
What's the fastest way to tell a Stage 1 firm from a Stage 2 firm from the outside?
Ask whether customers have noticed a quality jump. A Stage 1 firm will describe faster individual tasks inside the same delivery model. A Stage 2 firm will point to better work, better reviews, stronger retention, and a healthier bottom line.
Do all four numbers have to move for a firm to count as Stage 2?
No. The Stage 2 signature is the middle pair: retention and margin. Delivery quality and customer outcomes have visibly improved, and the bottom line reflects it. Revenue per person can stay flat well into Stage 2. Detaching it from headcount is the Stage 3 jump, not a Stage 2 requirement. A firm does not need a redesigned delivery model to be past Stage 1. It needs customers and the P&L to already show the upgrade.
Where do most $1–20M service firms sit today?
In the delivery data I've had visibility into across a franchise network of 150+ units and growing, the concentration as of July 2026 is heavily at Stage 0 and Stage 1.