Agent evaluations almost always measure tool selection: given a request, did the system choose the right function. It is easy to score and it is close to solved on modern models. The measurements that predict production behaviour are the ones after that.
Error recovery rate. A tool times out, returns a 500, or returns an empty result. What does the agent do? In benchmarks we have run this is routinely the worst number on the page, often below 0.5, because the happy path was built and tested and the failure path was not. The turn simply ends, and the user is told something vague or nothing at all.
Argument validity. Calls that were well formed: required parameters present, types right, enumerations respected, dates in the format the API expects. Usually high, and worth measuring because when it drops it drops on one specific tool with an awkward signature.
Redundant calls. Calls that changed nothing: the same lookup twice, a search after the answer was already in context, a fetch whose result was never used. This is directly billable waste and it multiplies by your traffic. Median calls per completed task is the number to watch, and the tail is where the cost sits.
Sequence correctness. For anything multi step, whether the calls happened in a workable order. A wrong call at step two corrupts everything after it, and scoring only the final output tells you the answer was wrong without telling you where it went wrong.
You cannot measure recovery by waiting for a real outage. Put a fault injecting layer between the agent and its tools and make it return, on demand: a timeout, a 500, an empty result set, a malformed payload, and a plausible but wrong result.
The last one is the interesting case. Most agents treat a tool result as ground truth, so a tool that confidently returns the wrong record produces a confident wrong answer with no hedging anywhere in it.
Retry once on a timeout, with a bound. Fall back to a different route where one exists. Where neither works, tell the user what could not be done rather than answering as though nothing happened. And never present a tool result as certain when the tool itself signalled uncertainty.
All four are testable, which means all four can be thresholds in a release gate rather than aspirations in a design document.
Tool use is where an agent quietly becomes expensive. Two redundant calls per task, at a hundred thousand tasks a month, is a line item somebody will ask about. Measuring calls per completed task alongside cost per resolved task turns that from a surprise into a number in the report.
The benchmark covers all of it, and the failing cases come back with the call sequence attached so the pattern is visible.
Measure four things separately: whether the right tool was chosen, whether the call was well formed, whether the agent recovered when the call failed, and how many calls changed nothing. Selection accuracy is the easiest to measure and the least predictive of production behaviour.
Put a fault injecting layer between the agent and its tools and force timeouts, server errors, empty results, malformed payloads and plausible but wrong results. The last is the most revealing, because most agents treat a tool result as ground truth.
Usually redundant tool calls. Measure calls per completed task and look at the tail rather than the mean. Two wasted calls per task is invisible in testing and a significant line item at production volume.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark