Applied ML

Your agent picked the right tool. Did it recover when the tool failed?

Agent evaluations almost always measure tool selection: given a request, did the system choose the right function. It is easy to score and it is close to solved on modern models. The measurements that predict production behaviour are the ones after that.

Four numbers, roughly in order of how much they will hurt you

Error recovery rate. A tool times out, returns a 500, or returns an empty result. What does the agent do? In benchmarks we have run this is routinely the worst number on the page, often below 0.5, because the happy path was built and tested and the failure path was not. The turn simply ends, and the user is told something vague or nothing at all.

Argument validity. Calls that were well formed: required parameters present, types right, enumerations respected, dates in the format the API expects. Usually high, and worth measuring because when it drops it drops on one specific tool with an awkward signature.

Redundant calls. Calls that changed nothing: the same lookup twice, a search after the answer was already in context, a fetch whose result was never used. This is directly billable waste and it multiplies by your traffic. Median calls per completed task is the number to watch, and the tail is where the cost sits.

Sequence correctness. For anything multi step, whether the calls happened in a workable order. A wrong call at step two corrupts everything after it, and scoring only the final output tells you the answer was wrong without telling you where it went wrong.

Tool use scores after selection accuracy A bar chart of three tool use rates against a nine tenths target. Argument validity is high, sequence correctness is middling, and error recovery is far below the target, which is routinely the worst number even when tool selection looks solved. Selection is solved. These are not. 1.0 0.0 0.9 target 0.97 Argument validity 0.70 Sequence correctness 0.45 Error recovery
The happy path is built and tested; the failure path is not. Error recovery is routinely the worst number on the page.

Test the failures deliberately

You cannot measure recovery by waiting for a real outage. Put a fault injecting layer between the agent and its tools and make it return, on demand: a timeout, a 500, an empty result set, a malformed payload, and a plausible but wrong result.

The last one is the interesting case. Most agents treat a tool result as ground truth, so a tool that confidently returns the wrong record produces a confident wrong answer with no hedging anywhere in it.

What good behaviour looks like

Retry once on a timeout, with a bound. Fall back to a different route where one exists. Where neither works, tell the user what could not be done rather than answering as though nothing happened. And never present a tool result as certain when the tool itself signalled uncertainty.

What a robust agent does when a tool call fails A fork after a failed tool call. The fragile path ends the turn with a vague answer. The robust path retries once within a bound, falls back to another route if one exists, and otherwise tells the user plainly what could not be done. After the call fails Tool returns an error timeout, 500, empty, malformed, or plausible but wrong fragile robust The turn just ends a vague answer, or nothing, as though the call had worked Retry once, within a bound Fall back to another route Else tell the user what failed Never present a tool result as certain when the tool itself signalled uncertainty.
All four behaviours are testable, so they can be thresholds in a release gate rather than aspirations in a design document.

All four are testable, which means all four can be thresholds in a release gate rather than aspirations in a design document.

Cost lives here too

Tool use is where an agent quietly becomes expensive. Two redundant calls per task, at a hundred thousand tasks a month, is a line item somebody will ask about. Measuring calls per completed task alongside cost per resolved task turns that from a surprise into a number in the report.

The benchmark covers all of it, and the failing cases come back with the call sequence attached so the pattern is visible.

Common questions

How do you evaluate tool use in an AI agent?

Measure four things separately: whether the right tool was chosen, whether the call was well formed, whether the agent recovered when the call failed, and how many calls changed nothing. Selection accuracy is the easiest to measure and the least predictive of production behaviour.

How do I test what my agent does when a tool fails?

Put a fault injecting layer between the agent and its tools and force timeouts, server errors, empty results, malformed payloads and plausible but wrong results. The last is the most revealing, because most agents treat a tool result as ground truth.

Why is my agent more expensive than expected?

Usually redundant tool calls. Measure calls per completed task and look at the tail rather than the mean. Two wasted calls per task is invisible in testing and a significant line item at production volume.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark