Model assurance

One accuracy number is the wrong shape for an agent

A classifier has one job, so one number describes it. An agent has four or five, and the number everyone quotes is an average across all of them.

What the average is averaging

Between a question arriving and an answer leaving, a typical agent decides what the user wants, searches for supporting material, chooses whether to call a tool, calls it, reads the result, and writes a reply in a shape something downstream has to parse. Six opportunities to be wrong, each with a different owner and a different fix.

Six stages collapsed into one average Six agent stages, understand, retrieve, choose tool, call tool, read result and write reply, each able to fail on its own, all feed into a single aggregate score that cannot say which stage moved. Six chances to be wrong, one number Understand can fail Retrieve can fail Choose tool can fail Call tool can fail Read result can fail Write reply can fail Aggregate 0.82 which stage moved?
Each stage has a different owner and a different fix. The average tells you none of them.

An aggregate score moves when any of them moves. It goes up when retrieval improves and down when the model becomes more cautious, and in a normal week it does both at once. The number changes, nobody can say why, and the meeting becomes an argument about anecdotes.

Slices, not a total

The fix is not a better metric. It is refusing to report one number.

Build the case set with named slices, decided by what the business actually cares about rather than by what is convenient to collect. For a support agent that might be billing questions, account access, product how-to, messages containing two requests, questions the documentation cannot answer, and adversarial input. Report each separately and report how many cases sit in each.

The moment you do this, the shape of the problem appears. A system at 0.82 overall is usually not uniformly mediocre. It is 0.94 on the three common slices and 0.42 on one, and the 0.42 is the slice that generates the complaint.

Per slice table with the total at the bottom Three common slices clear the 0.80 threshold at 0.94, 0.93 and 0.95, one edge slice fails at 0.42, and the total of 0.82 sits at the bottom as the least informative line. The total is the least informative line Slice Score Threshold Verdict Billing questions 0.94 0.80 clears Account access 0.93 0.80 clears Product how-to 0.95 0.80 clears Documentation cannot answer 0.42 0.80 fails Total 0.82 hides the failing slice
Read per slice, the shape appears at once: strong everywhere but the one slice that draws the complaint.

An aggregate that rises while a slice falls is normal

This is the part teams find surprising. Changes are rarely uniform improvements. A prompt revision that fixes twenty failures and breaks six successes moves the total up, and the six stay broken. If those six were in the slice that matters most, the release made the product worse and the dashboard said it got better.

Reporting per slice makes that visible in the same table that reports the win. No extra work, and it changes what the team argues about.

How to count a partial answer

Most disagreements about scoring are really disagreements about what counts as success. Settle it before the run, in writing, with the people who own the decision.

A three way scale is usually enough: the case passed, the case failed, or the answer was partially right and needs a human. Score a partial as half. What matters far more than the exact weight is that the same rule applies on every run, so a movement in the number is a movement in the system rather than in the grading.

What good looks like

A table with one row per slice, the case count, the pass count, the score, the threshold agreed beforehand, and whether it cleared. A total row at the bottom that everybody has agreed to treat as the least informative line in the table.

That is what a benchmark should hand you, and it is what our benchmarking engagement produces. The scorer that does it is open: r4agent reports per slice by default and treats the total as an afterthought, because that is the right way round.

Common questions

What is a good accuracy score for an AI agent?

There is no useful general answer, which is the point. A score only means something against a fixed case set and a threshold agreed in advance for each slice of that set. A system at 0.92 overall can be unfit for production if the slice it fails is the one carrying regulatory or financial consequence.

Should I report one overall accuracy figure to stakeholders?

Report it, but never on its own. The overall figure is an average across several different subsystems and it moves for reasons nobody can attribute. Put it at the bottom of a per slice table so it is read as a summary rather than as the finding.

How many test cases does an agent benchmark need?

Fewer than teams expect. A hundred cases argued over properly with domain experts will surface every ambiguity in the problem statement. Ten thousand cases labelled quickly will hide them and give a precise measurement of the wrong thing.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark