Benchmarking

How to Evaluate an AI Agent Before Production

Most agents that fail in production passed their demo. A demo is three or four tasks, chosen because they work, run once, watched by someone who already knows what a good answer looks like. None of that resembles what the agent will meet in the first week of real traffic. AI agent evaluation before production is not about proving the agent can do the job. It is about finding out, on purpose and in advance, the specific ways it cannot.

Start with success criteria, not a vibe check

Before a single task is written, define what "done" means for the agent's job, in terms specific enough that two reviewers would grade the same transcript the same way. For a support agent that might be: the customer's stated problem is resolved, no action is taken outside the customer's own account, and any refund or credit matches policy. For a coding agent it might be: the change compiles, the existing tests still pass, and no unrelated file is touched.

Write these criteria down before you see any transcripts. Reading transcripts first and then deciding what counts as success is how a team ends up grading the agent against what it happened to do, rather than against what the job required.

Build a representative task set

A task set built from whatever is easy to find, a handful of internal test cases, the examples in the product spec, tends to be easier than the traffic the agent will actually see, because the people writing it already know how to succeed. A representative set is pulled from where the real work already is: closed support tickets, real pull requests, logged transactions, actual user requests. It should include the boring middle of the distribution, not just the sharp edge cases, because most of what an agent handles in production is ordinary, not adversarial.

Size matters less than people expect once a set passes a few hundred tasks with enough diversity to cover the input's natural variation; what matters more is that the set is fixed and versioned, so a change to the prompt or the model can be run against the exact same tasks a later change is run against. That fixed, repeatable setup, task set, grading rubric, scoring code, is the eval harness. Without it, "we improved the agent" is an anecdote instead of a comparison.

The pre-production evaluation pipeline from criteria to a go or no-go gate A left to right pipeline: define success criteria, build a fixed versioned task set, run the agent, measure task success tool use and cost, catalog failure modes, then a go or no-go gate set by what a failure costs. Finding the failures before production does Success criteria Fixed task set versioned, from real work Run the agent Measure task success, tool use, cost Catalog failures by category and severity Go / No-go bar set by what a failure costs
Criteria come before tasks, tasks stay fixed, and the numbers only matter if the gate is allowed to say no.

Measure what the job actually requires

Task success rate

Task success rate is the fraction of tasks the agent completed to the criteria defined above, not the fraction where it produced a plausible-looking response. Grading it with an automated check where one exists, a test suite, a policy rule, a database state, beats grading it by reading, because reading scales badly and drifts between reviewers. Where a check cannot be automated, use a rubric and more than one reviewer, and measure agreement between them; if your reviewers disagree with each other a fifth of the time, the success rate they produce is not precise enough to make a shipping decision on.

Tool-use correctness

An agent can land on the right final answer while getting there badly: calling a tool with malformed or guessed arguments, calling a destructive action twice because the first response was slow, skipping a lookup it should have made and guessing instead. Task success alone hides all of this, because a lucky path still counts as a win. Score tool calls separately, right tool, right arguments, right order, no unnecessary calls, and look at the cases where the final answer was right but the path was not. That pattern of getting away with it is usually the one that fails on a task where the shortcut matters.

Final answer correctness against tool use path correctness A two by two matrix. Columns are answer wrong and answer right. Rows are path wrong and path right. The cell where the answer is right but the path is wrong is highlighted as getting away with it, the risk task success alone hides. Right answer is not the same as right path Final answer wrong right Tool-use path wrong right Clear failure caught either way Got away with it lucky path, wrong call task success hides this Honest miss right steps, hard task Genuine success right answer, right path
Scoring the path separately surfaces the top right cell, the right answer reached the wrong way, which is the one that fails when the shortcut matters.

Cost per resolved task

Cost per call looks fine on a system that fails constantly, because a failed call is usually cheap. The number worth tracking is total spend across the benchmark run divided by the number of tasks the agent actually resolved, so retries, dead-end tool calls and failures are counted in the cost but not in the denominator. It is also worth watching alongside latency at the p95, not the average, since the tasks that need several tool calls are the ones a user notices waiting for.

Catalog failure modes, not just a failure rate

A single failure percentage tells you the agent is imperfect, which you already knew. What changes a go/no-go decision is knowing how it fails: does it say "I'm not sure" and stop, or does it answer confidently and wrongly? Does it fail randomly across the task set, or does it cluster on one category, a particular tool, a particular input format, a particular kind of ambiguous request? A system that fails ten per cent of the time by refusing and escalating is shippable in far more settings than one that fails five per cent of the time by taking a wrong, silent action. Tag every failure by category and by severity, not just by pass or fail, and review the clusters before you review the aggregate number.

Setting the go/no-go bar

The bar is not a single accuracy number picked because it sounds respectable. It is set by what a failure costs and who catches it. An agent whose output a person reviews before it takes effect can ship at a lower success rate than one that acts unattended. An agent that can only ever produce a reversible action, a draft, a suggestion, can tolerate more failure than one that can commit money or delete data. Set the bar in those terms, write it down before the run, and hold to it after the run whether the number is flattering or not. The task set, the tool-use grading and the cost figures only matter if the number they produce is allowed to say no.

None of this requires exotic tooling. It requires a fixed task set pulled from real work, a rubric agreed on before anyone reads a transcript, and the discipline to separate whether the agent got there from how. Our agent benchmarking work follows this same shape for teams who would rather have that structure built and run for them than build it from scratch.

Common questions

What is a good task success rate for an AI agent?

There is no universal number. It depends on what the task replaces, what a failure costs, and whether a human reviews the output before it takes effect. A rate that is safe for a drafting assistant can be unsafe for an agent that submits refunds unattended.

What is an eval harness for AI agents?

A repeatable setup that runs a fixed set of tasks against the agent, checks the outcomes against a rubric or a verifier, and produces the same metrics every time it runs, so a change to the model, prompt or tools can be compared against a prior run rather than judged by feel.

Why does an agent need tool-use correctness checked separately from task success?

An agent can reach the right answer while calling a tool with the wrong arguments, calling it twice, or skipping a required verification step. Task success alone hides that path, and the same bad call pattern usually resurfaces on a task where it does matter.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark