# Eval

An eval is a repeatable test that scores AI output against a written standard, so a firm can track quality over time.

Source: https://rodamora.com/glossary/eval
Updated: 2026-08-09

---

An eval is a repeatable test that scores AI output against a written standard. The point is to see whether the work is good enough before clients or staff rely on it. A firm can track tool use and hours saved without knowing whether the deliverable is correct. In a survey of nearly 6,000 senior executives across the US, UK, Germany, and Australia, 89% reported no effect on labor productivity from AI over the prior three years. The survey does not prove that evals would change that result. It does show why a baseline matters. Start with 20 real deliverables and five checks your team agrees on. Run the same test next month. That small record gives you a way to spot quality loss before a client reports it.

## What an eval looks like in a service firm

For a service firm, an eval can be a grading checklist. You do not need a benchmark or a new platform to start.

Take 20 proposals made with AI help. Write down what makes one correct at your firm: the approved price is used, the scope matches the call, every claim has support, an owner is named, and the client details are right. Score each proposal against those checks. The result is a number you can compare with next month's number.

That is a complete first eval. The work may take an afternoon. More importantly, it forces the people who own delivery to agree on what good means.

The standard must be **written down**. Otherwise two reviewers can apply different rules. The cases must be **yours**, because generic test data does not show your real failure modes. The test must be **repeatable**, so a change in the score tells you something. A one-time audit gives you a snapshot. An eval lets you watch the work.

Keep the test tied to one decision. If a proposal fails the price check, the owner knows what to fix. If the score blends speed, tone, completeness, and retention into one number, the next action is unclear. A good first eval is narrow enough that a reviewer can explain a failure in one sentence.

Hours saved does not answer that question. It measures effort, and it is often an estimate. [The Four Numbers](/four-numbers) explains why usage and saved time are weak stand-ins for business results. An eval measures the thing the client receives.

## Why a feeling is not a measurement

Most firms first notice quality through a complaint. That signal arrives after the work has reached someone who matters.

[Quality can also slide slowly](https://rodamora.com/blog/cheap-mistakes-go-to-ai-expensive-ones-wait). Output volume rises, reviewers work faster, and small errors turn into rework or a missed renewal weeks later. If nobody recorded the earlier standard, the firm cannot tell how far quality moved.

A score also separates a real gain from a lucky month. Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027 because of rising costs, unclear value, or weak risk controls. The prediction is not about your firm. It is a reminder that expansion without a starting measure leaves little to compare at the end.

Evals matter more when an agent handles volume. A person can review every item while the queue is small. [Human in the loop](/glossary/human-in-the-loop) review uses attention on each item, so it becomes a limit as production grows. An eval can run across the set and send failed cases to a person.

## Who should grade the output?

Use the simplest grader that can answer the check.

**A person with a checklist** is usually the best first grader. It is accurate and slow. Writing the checklist also exposes places where senior people never agreed on what correct means.

**A rule** works for a narrow fact. It can check that a price matches the price list, that each figure appears in the source record, or that a required section exists. Rules are cheap and reliable inside that boundary.

**Another model** can help with judgment a rule cannot express, such as whether the tone fits a client or a conclusion follows from evidence. A 2023 study of LLM judges found more than 80% agreement with human preferences, about the same level of agreement reported between humans. The same study found position, verbosity, and self-enhancement bias, along with limited reasoning ability.

Treat a model grader as a tool that needs testing. Give it a sample of outputs already rated by people. Check it again on a schedule. Otherwise the firm has attached a number to another unmeasured opinion.

## What does an eval cost?

The first eval may take an afternoon and a difficult conversation. Three senior people may need to agree in writing on what makes a deliverable correct. That decision remains useful even if the firm later stops using AI.

The recurring cost is ownership. Someone must run the test, watch the score, and act when it moves. Without a named owner, an eval is run once and then becomes a slide in a deck.

The standard also needs maintenance. A price list changes. A client's requirements change. Delivery changes. Review the checks when the work changes, rather than waiting for a surprising score.

Record the cases that failed and the reason for each failure. That record helps the owner see whether the same mistake returns after a prompt or tool change. It also gives the team a small set of real examples for training and for the next version of the test.

People also need time to read the result. In the delivery data I have reviewed, [training that sticks takes 8 to 10 hours per person over about 3 weeks](https://rodamora.com/blog/your-team-isnt-resisting-ai-nobody-gave-them-the-hours). A score cannot guide a decision if the team does not know what to change.

## Where this sits on the Delivery Model Ladder

An eval sits at the hinge between Stage 2 and Stage 3 on the [Delivery Model Ladder](/delivery-model-ladder). Stage 3, [AI-native delivery](/glossary/ai-native), needs work that can be divided into pieces and checked. The eval supplies that check.

This is why [an AI employee](https://rodamora.com/blog/what-is-an-ai-employee) needs a validation method alongside scope, access, and a trigger. Those first three describe what the system may do. Validation tells you whether the result can pass.

The delivery data shows how uncommon that condition is. Across the 150+ franchise units I track, roughly 80% still run chatbots, 15% run rigid workflows on triggers, and 5% run autonomous agents defined by conditions. The numbers describe that network, not the market as a whole. The smaller group has a written answer to what a correct output looks like.

## Quick answers

**Do I need an eval before I start using AI?** No. Start the work, then write the eval when output first reaches a client. Treat it as a requirement before removing a human check or expanding the workflow.

**What should the first eval measure?** Measure one failure that would matter: a wrong number, an unsupported claim, or a missing disclosure. Overall quality is too broad for a useful first test.

**How does an eval affect what an agent can run alone?** Work should be [checked and undoable before you grant more autonomy](https://rodamora.com/blog/delegate-the-inputs-own-the-outputs). An eval makes the first part concrete. The answer can change as the task changes, even if the model stays the same.

**Does an eval appear in the Four Numbers?** Its early signal is quality. A quality drop can later appear as clients not renewing, which is [the ROI question firms often answer too late](https://rodamora.com/blog/what-roi-does-ai-actually-produce-in-a-service-firm). The eval gives you a measure before retention moves.
