Model assurance

LLM Evaluation Metrics That Predict Production Failures

Most eval dashboards report a number that can sit still while the system gets worse. Overall accuracy is the usual culprit: it blends every kind of correct answer with every kind of wrong one, so a shift in the failure mix toward the error that actually reaches a customer does not have to move it at all. A metric worth tracking is tied to one specific failure mode with one specific consequence, measured on a slice narrow enough that a change in it means a single thing, which is what makes grounding rate, refusal accuracy, recall at k, and tool-call accuracy worth watching separately, rather than folded into a summary score.

Each evaluation metric maps to a distinct production incident Five metrics on the left connect to five kinds of production incident on the right. Grounding rate to acting on an invention, refusal accuracy to over or under blocking, recall at k to missing context, tool-call accuracy to a wrong action, and p95 latency and cost to an answer that arrives too late. One metric, one failure mode, one incident Metric watched on its own Incident it predicts Grounding rate Refusal accuracy Recall at k Tool-call accuracy p95 latency and cost User acts on an invention Over-block or under-block Answer missing from context A wrong action is taken Right answer, arrives too late Blend them into one accuracy number and the mapping, and the early warning, is lost.
Each metric answers a different question about how the system fails, so each is tracked on its own rather than averaged away.

Grounding and hallucination rate

Hallucination rate is the share of claims in an answer that no retrieved passage, tool result, or provided document actually supports. It is not a vibe score from asking another model whether an answer "sounds right." It requires decomposing the answer into checkable claims and checking each one against the material the system actually had access to.

This is the metric that correlates most directly with the incident where a user acts on something the system invented: a wrong policy detail, a wrong figure, a citation to a document that does not say what was claimed. A system can be fluent and confident and still be the one generating the support ticket that says "your assistant told me something that was not true." Grounding rate is what catches that before a customer does.

Refusal accuracy, not just refusal rate

Refusal rate alone tells you almost nothing, because a system that refuses everything scores well on it and is useless. What predicts incidents is refusal accuracy: did it refuse the requests that should have been refused, and answer the ones that should have been answered. Split it into two numbers. Refusal precision is the share of refusals that were actually warranted. Refusal recall is the share of requests that should have been refused and were. A system can look safe on an aggregate policy score while both numbers underneath are mediocre, over-blocking legitimate requests and under-blocking the ones that matter.

The incidents this predicts split the same way: over-refusal shows up as abandoned users and support tickets about a system that will not help with reasonable requests; under-refusal shows up as the incident escalated to legal or security.

Refusal accuracy as a two by two of what should happen against what the system did Rows are should refuse and should answer. Columns are system refused and system answered. The two off diagonal cells are the incidents: under-refusal escalates to security, and over-refusal abandons a legitimate user. Refusal accuracy, not refusal rate What the system did refused answered What should happen should refuse should answer Correct refusal blocked what it should Under-refusal answered what it should not have escalates to security Over-refusal blocked a fair request abandons the user Correct answer helped where it should
A single refusal rate can look safe while both off diagonal cells grow. Precision and recall on refusals separate the two incidents.

Recall at k for retrieval

For any system built on retrieval, recall at k is the ceiling on everything downstream. It is the share of queries where the passage actually holding the answer made it into the top k results handed to the model. If it did not, no amount of prompt engineering or model quality recovers the answer, because the model was never shown the material.

Where this breaks quietly

Recall at k tends to degrade after launch, not before it, as the document set grows and query phrasing drifts from what the index was tuned against. Measuring it as its own line, separate from end-to-end accuracy, tells you whether a drop in answer quality is a retrieval problem or a generation problem, fixed in completely different places.

Tool-call accuracy, and what happens after the call

Tool selection accuracy, did it call the right tool with the right arguments, is the easy half of this measurement and the less informative one. The half that predicts incidents is what the system does after a call returns an error, an empty result, or a timeout. Does it retry sensibly, ask a clarifying question, or fall back to a safe default? Or does it fabricate a result, or hand the user an error message meant for a developer?

Tool calls are also where an agent takes actions with real consequences, writing to a database, sending a message, issuing a refund, so an error here does not just produce a wrong answer, it produces a wrong action, usually a more expensive class of incident.

Latency and cost as reliability signals

Latency and cost look like operations metrics rather than quality metrics, but at the tail they predict incidents too. Report p95, not the average: the median describes an experience nobody complains about, and for an agentic system the gap between a simple request and one needing several tool calls is wide. A p95 that creeps past what a user or an upstream timeout will tolerate produces an incident that has nothing to do with correctness, the answer would have been right, if it had arrived.

Cost is a reliability signal too: a system that gets expensive under specific conditions, long documents, multi-step tool chains, retry loops after a failure, is usually telling you where it is also least stable. Tracking cost per resolved task, not cost per call, keeps failed and retried attempts in the number instead of hiding them.

Regression testing between model versions

The incident that catches teams off guard most often is not a new failure mode, it is an old one coming back after a model or prompt upgrade that looked, in aggregate, like an improvement. A vendor's point release can raise overall accuracy while quietly regressing a narrow slice: a refusal category, a tool-argument format, a retrieval-dependent answer type.

The only reliable defence is a fixed, versioned case set run against every candidate change before it ships, diffed at the case level rather than compared as aggregates. What matters is not that the new score is higher, it is the specific list of cases that used to pass and now do not, often a handful pointing at one exact mechanism, not a diffuse decline.

Putting the numbers in one place

None of these metrics is sufficient alone, and none should be averaged into the others. Grounding rate, refusal accuracy, recall at k, tool-call accuracy, and p95 latency and cost each answer a different question about how the system fails, and each maps to a different kind of production incident. Reported side by side, on the same case set, run again after every material change, they turn "the eval looks fine" into a claim someone can actually stand behind once the system is live.

Working out which of these actually predict incidents for your own system is close to the exercise behind our agent benchmarking work, building the case set and the per-metric breakdown before the number that matters gets decided by a production outage.

Common questions

Which LLM evaluation metric best predicts production incidents?

No single one does, because incidents come from different failure modes. Grounding rate predicts incidents caused by fabricated claims, recall at k predicts incidents caused by missing context, and tool-call accuracy predicts incidents caused by wrong actions taken on a user's behalf. Track them separately.

Is overall accuracy a vanity metric?

On its own, mostly. A blended accuracy number can hold steady while the failure mix underneath shifts toward the kind of error that actually reaches a customer. It is a summary, not a diagnosis, and it should be paired with per-slice metrics before anyone treats it as a release gate.

How often should regression tests run between model versions?

Before every version swap and prompt or retrieval change that ships to production, not on a calendar. A vendor's minor model update is a new system from an evaluation standpoint, and the fixed case set from the previous version is what tells you whether it got worse anywhere.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark