Framework · Rod Amora ·

The Four Numbers

Every AI implementation in a service firm succeeds or fails on four numbers. Hours saved isn't one of them.

AI doesn't save a service firm time. It changes four numbers: cost to deliver, retention, margin, and revenue per person. Every AI implementation I've watched succeed or fail comes down to which of those four it actually moved. Time saved is a proxy. It shows up on a dashboard and nowhere on a P&L.

Cost to deliver is what it costs to produce one unit of the service. Retention is whether customers stay and come back. Margin is what the firm keeps. Revenue per person is how much the business earns per employee. These four work the same whether the firm bills hours, sells subscriptions, or delivers one-off projects. AI can move any of them. Most firms let it move none of them, because they measure hours instead.

The Four Numbers is the scoreboard for AI adoption in a service firm: cost to deliver, retention, margin, and revenue per person. Each stage of The Delivery Model Ladder is confirmed by one or two of these numbers moving, and each works on any billing model: hourly, fixed fee, subscription, or one-off project. It's built for owners of $1–20M service firms who want to know if an AI rollout changed the business. As of July 2026, firms that name which number they're moving before they build are the ones that see it move.

What should a service firm measure when it adopts AI?

Hours saved is the number every AI vendor pitches. It also shows up least often in a firm's financials. It is a proxy metric: easy to estimate, disconnected from margin.

A service firm should measure movement in four numbers instead:

  • Cost to deliver: what it costs to produce one unit of the service (an engagement, a project, a month of a subscription). Labor, tools, and rework included.
  • Retention: whether customers stay, renew, or come back for the next project, meaning higher LTV. For one-off project shops, read this as repeat purchase and referral.
  • Margin: what the firm keeps after the cost of delivery.
  • Revenue per person: total revenue divided by total headcount.

These four sit on the P&L, and none of them assumes a billing model. A firm on hourly rates, a firm on retainers, and a firm selling fixed-scope projects can all put the same four lines in the same monthly review. For where a firm sits on the broader adoption curve, from tool-assisted individuals to AI-native delivery, see The Delivery Model Ladder. The Four Numbers is how you check whether each stage of that ladder actually happened.

Which number moves at each stage of the Ladder?

Each stage of the Ladder has a signature in these four numbers. That is what makes them the scoreboard rather than another dashboard.

Ladder stage What moves What stays flat
Stage 0: Tool-assisted Nothing All four
Stage 1: Augmented Cost to deliver drops, margin increases Retention, revenue per person
Stage 2: Upgraded Retention and margin improve Revenue per person, mostly
Stage 3: AI-native Revenue per person detaches from headcount  

Stage 1 is the trap stage, and the table shows why: one number moves and three don't. Cost to deliver drops because AI genuinely compresses the work. But if retention and margin don't follow, the gain went to slack, not to the business. A firm that reports "we're saving fifteen hours a week" while margin sits still is a Stage 1 firm reading the wrong scoreboard.

Stage 2 is confirmed by the customer-facing pair: retention improves because the work got visibly better, and margin improves because the efficiency finally reached the bottom line. Stage 3 is confirmed by the last number: revenue grows without headcount growing with it.

Why don't AI time savings show up in profit?

Time saved has to land in one of the four numbers. In most firms, it doesn't. This is the evaporation pattern: cost to deliver drops on paper, but nobody decides where the freed capacity goes. The hours just go. Not to more delivery, not to better delivery, not to lower cost. They go to slack.

Across the firms I've watched adopt AI, the most common failure isn't the AI not working. It's the AI working exactly as advertised, cutting hours on a task, with no mechanism to catch the saved time before it evaporates. A senior person finishes a first draft in 40 minutes instead of three hours. If nothing else changes, the firm has the same costs, the same revenue, and a quieter afternoon. Margin didn't move because nothing about the business moved.

The Production Gap names this as a distinct failure mode: the gain is real and invisible at the same time, because nobody decided in advance where the reclaimed time would go.

How does AI move cost to deliver?

By compressing the labor inside each unit of delivery: research, first drafts, formatting, status reporting, routine analysis, internal handoffs. This is the number that moves first and moves easiest, because it doesn't require a structural decision. It's the Stage 1 number.

The measurement mistake is tracking hours saved as an estimate (a survey answer, a vendor claim, a guess from a team lead) instead of tracking the actual cost of producing one engagement, one project, one month of service. Cost to deliver is a real unit cost. Pick the unit, price its inputs, and watch the number quarterly. If you can't say what a deliverable cost you last quarter versus this quarter, you can't tell Stage 1 from Stage 0.

How does AI move retention?

Through quality the customer can feel. When freed capacity gets aimed at the work instead of evaporating (more thorough deliverables, faster turnaround, better onboarding, more proactive communication), customers notice. They stay longer, renew more, buy again, and refer. This is the Stage 2 number, and it's the one that tells you the upgrade was real rather than internal.

It can also go the other way, and firms measure this wrong by not measuring it at all. Unreviewed AI output creates quality debt: work that looks finished but isn't. Ship enough of it and retention falls while every internal metric says the rollout is going great. A firm tracking AI adoption without tracking retention is flying blind on the number that matters most to its future revenue.

Retention lags. It takes quarters to confirm, not weeks. That's appropriate for what it measures: stages are quarters-long transitions, and this number is the proof the transition happened.

How does AI move margin?

Margin is where cost to deliver and retention meet the P&L. When delivery gets cheaper and the firm holds or raises what it charges, margin expands. When better work keeps customers longer, margin expands again, because retained revenue costs less than replaced revenue.

The threat to margin is pricing. Once the cost of delivering an outcome visibly drops, the pricing model that assumed the old cost comes under pressure. Clients notice. Competitors reprice first. Or the firm's own hourly math stops working. Hourly billing implicitly prices the old cost of production. When AI compresses that cost, hourly billing hands the entire gain to the client as a fee cut the firm didn't choose. Firms that protect margin repackage the outcome (fixed fee, subscription, outcome-based) before the market forces the question. The billing model is a means. Margin is the number that tells you whether it worked.

How does AI move revenue per person?

Last, and only after everything else. Revenue per person detaches from headcount when the delivery model itself changes: when growth stops requiring proportional hiring because agents carry a real share of production, and the firm's knowledge is structured so that work doesn't bottleneck on individual people.

There's a structural reason this number can detach at all. In a growing firm, headcount was never just production capacity. It was coordination capacity: past a certain size, a share of every hire goes to routing information — relaying context between people, keeping teams aligned, carrying decisions up and down. When the firm's context is shared, recorded, and machine-readable, the system carries that routing instead of people. Growth stops requiring proportional hiring for two reasons, not one: agents carry production, and coordination no longer scales with headcount. That's why this number confirms Stage 3 of The Delivery Model Ladder and nothing earlier — it only moves when both halves change.

This is the Stage 3 number, and it's uncharted territory. In the delivery data I've reviewed, I've only seen a couple of early-stage firms where this number is starting to detach, both born in the AI era. For an established firm, watching revenue per person is mostly a way to keep an honest baseline: if it hasn't moved in years, the firm is still buying growth with headcount, whatever the AI stack looks like.

What do firms measure instead of the four numbers, and why does it fail?

What firms typically track What it actually tells you What to watch instead
Hours saved (estimated) An input estimate, not a financial change Cost to deliver: the unit cost of an engagement, project, or service month
AI adoption rate / logins Usage, not output quality Retention: do customers stay, renew, and come back
"We gave the team a tool" Activity, not economics Margin: what the firm keeps after delivery cost
Headcount and revenue growth, separately Two curves, no relationship Revenue per person: whether growth still requires hiring

This is the same trap named as a failure mode in The Production Gap: a firm builds a dashboard full of activity metrics that never appears in the monthly financial review. If a number doesn't show up in the same meeting as revenue and margin, it isn't one of the four numbers.

There's a reason firms land on the left column so consistently, and it isn't laziness. The honest state of the market is that most organizations aren't only failing to measure AI's return — they don't yet know what measuring it would look like. The frameworks everyone learned to evaluate an investment with were built for a purchase that does a defined thing at a defined cost, and AI doesn't fit that shape: it arrives as a capability that touches many tasks unevenly, and the value shows up somewhere other than where the spending did. So firms measure what's easy to count, which is activity, and end up with a number that is precise, true, and about nothing.

That is the whole reason this set exists, and why it's four lines from a P&L rather than a new dashboard. The four numbers are not a better way to measure AI. They're a refusal to measure AI at all. You measure the business, before and after, and let the technology be an explanation rather than a subject. That sidesteps the framework problem entirely, because a service firm already knows what a unit of its delivery costs, whether customers renewed, what it kept, and what it earned per person. Nobody has to invent a methodology to read those.

Which number should a firm move first?

Cost to deliver. It's the number AI touches without requiring a structural decision, and it moves within a quarter. That makes it the entry point, and also the least meaningful number on its own.

The real decision is what happens next. A firm that only moves cost to deliver is standing at Stage 1 with the evaporation clock running. In the delivery data I've reviewed across a franchise network of 150+ units and growing, the firms that shift their margin are the ones that aim the freed capacity at something within two or three quarters: better deliverables, tighter onboarding, overdelivery that shows up in retention. Name the second number before the first one finishes moving.


The Four Numbers is the scoreboard behind The Delivery Model Ladder and The Production Gap. Read the Ladder to find out where a firm actually sits. Read the Production Gap to see the specific ways the gain disappears before it reaches these four numbers. For more on how service firms rebuild delivery around AI, subscribe to the newsletter.

FAQ

What are the four numbers, in one line each?

Cost to deliver is what it costs to produce one unit of the service. Retention is whether customers stay, renew, and come back. Margin is what the firm keeps after delivery costs. Revenue per person is total revenue divided by headcount.

Do the four numbers work for firms that don't bill hours?

Yes, that is the point of this set. A subscription firm reads cost to deliver per service-month and retention as renewal. A project shop reads cost to deliver per project and retention as repeat purchase and referral. Margin and revenue per person read the same everywhere.

Why doesn't "hours saved" count as one of the four numbers?

Because it's an input estimate, not an output that appears on a financial statement. A firm can report large hours-saved numbers every quarter while all four numbers stay flat. Hours saved only matters once it lands in cost to deliver, and cost to deliver only matters once margin or retention follows.

Which number confirms which stage of the Delivery Model Ladder?

Cost to deliver dropping confirms Stage 1. Retention and margin improving together confirm Stage 2. Revenue per person detaching from headcount confirms Stage 3. If none of the four has moved, the firm is at Stage 0, whatever its AI usage looks like.

Can AI make one of the four numbers worse?

Yes. Retention is the number most at risk, because unreviewed AI output can degrade quality while every internal adoption metric improves. Margin can also worsen if a firm adds AI spend without changing anything about how work is delivered or priced.

How long does it take for a number to move after an AI rollout?

Cost to deliver can move within a single quarter. Margin follows only if the firm makes a decision about pricing or capacity. Retention takes two or more quarters to confirm, because customers respond to sustained quality, not a single good deliverable. Revenue per person moves on the scale of years.

Our AI spending has no ROI framework behind it. Where do we start?

You don't need a new framework, which is the point of this set. Most firms stall here because they're trying to build a methodology for measuring AI, and that's a genuinely hard problem nobody has solved. Measure the business instead: pick one deliverable, know what it costs you now, and watch that number plus retention, margin, and revenue per person across quarters. The AI becomes an explanation for movement in numbers you already track, rather than a thing requiring its own measurement discipline.

What's the fastest way to tell if an AI project is going to evaporate?

Ask which of the four numbers it's supposed to move, before the project starts. If nobody can name one, or the answer is "we'll save time," that is the early-warning sign.