A classifier has one job, so one number describes it. An agent has four or five, and the number everyone quotes is an average across all of them.
Between a question arriving and an answer leaving, a typical agent decides what the user wants, searches for supporting material, chooses whether to call a tool, calls it, reads the result, and writes a reply in a shape something downstream has to parse. Six opportunities to be wrong, each with a different owner and a different fix.
An aggregate score moves when any of them moves. It goes up when retrieval improves and down when the model becomes more cautious, and in a normal week it does both at once. The number changes, nobody can say why, and the meeting becomes an argument about anecdotes.
The fix is not a better metric. It is refusing to report one number.
Build the case set with named slices, decided by what the business actually cares about rather than by what is convenient to collect. For a support agent that might be billing questions, account access, product how-to, messages containing two requests, questions the documentation cannot answer, and adversarial input. Report each separately and report how many cases sit in each.
The moment you do this, the shape of the problem appears. A system at 0.82 overall is usually not uniformly mediocre. It is 0.94 on the three common slices and 0.42 on one, and the 0.42 is the slice that generates the complaint.
This is the part teams find surprising. Changes are rarely uniform improvements. A prompt revision that fixes twenty failures and breaks six successes moves the total up, and the six stay broken. If those six were in the slice that matters most, the release made the product worse and the dashboard said it got better.
Reporting per slice makes that visible in the same table that reports the win. No extra work, and it changes what the team argues about.
Most disagreements about scoring are really disagreements about what counts as success. Settle it before the run, in writing, with the people who own the decision.
A three way scale is usually enough: the case passed, the case failed, or the answer was partially right and needs a human. Score a partial as half. What matters far more than the exact weight is that the same rule applies on every run, so a movement in the number is a movement in the system rather than in the grading.
A table with one row per slice, the case count, the pass count, the score, the threshold agreed beforehand, and whether it cleared. A total row at the bottom that everybody has agreed to treat as the least informative line in the table.
That is what a benchmark should hand you, and it is what our benchmarking engagement produces. The scorer that does it is open: r4agent reports per slice by default and treats the total as an afterthought, because that is the right way round.
There is no useful general answer, which is the point. A score only means something against a fixed case set and a threshold agreed in advance for each slice of that set. A system at 0.92 overall can be unfit for production if the slice it fails is the one carrying regulatory or financial consequence.
Report it, but never on its own. The overall figure is an average across several different subsystems and it moves for reasons nobody can attribute. Put it at the bottom of a per slice table so it is read as a summary rather than as the finding.
Fewer than teams expect. A hundred cases argued over properly with domain experts will surface every ambiguity in the problem statement. Ten thousand cases labelled quickly will hide them and give a precise measurement of the wrong thing.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark