Model assurance

How to measure whether your agent made it up

Teams talk about hallucination as if it were weather. It is measurable, and the measurement splits cleanly in two: what the system asserts without support, and what it does when it has nothing to go on.

The unsupported claim rate

Take the answer, break it into individual claims, and check each against the passages the system actually retrieved. The share of answers containing at least one claim no passage supports is the unsupported claim rate.

Two details decide whether the number is honest. Check against what was retrieved, not against the whole corpus: an answer that happens to be true but was not supported by anything in front of the model is a lucky guess and will not stay lucky. And check claims, not sentences, because a sentence usually carries several and only one of them needs to be invented.

Checking each claim against the retrieved passages One answer is split into four separate claims. Three are found in a retrieved passage and one is not, so the whole answer counts toward the unsupported claim rate. Claims, not sentences One answer, four claims Checked against retrieval Claim 1 Claim 2 Claim 3 Claim 4 supported by a passage supported by a passage no passage supports it supported by a passage
One unsupported claim marks the whole answer. Grading by sentence would let it pass.

Citation validity is a different number

A system can cite diligently and still be wrong. The citation points at a real document, the document exists, the link resolves, and the passage does not support the sentence attached to it.

So measure separately whether a citation resolves and whether it supports. The second is the one that matters and it is the one usually skipped, because the first is easy to automate and produces a comfortable number.

The measurement teams miss entirely

Ask the system something the corpus cannot answer.

This is the single most informative slice in a benchmark and most evaluation sets contain none of it, because the cases are drawn from questions somebody already answered. Every case in such a set has an answer, so the system is never observed in the situation that produces its worst behaviour.

The correct response is to say it does not know, or to escalate. Score that as a pass. A confident invented answer to an unanswerable question is the failure that reaches a customer and gets acted on, and a system with no abstention path will fail this slice close to totally while scoring well on everything else.

We have measured agents at above 0.90 on their three main slices and below 0.45 on abstention. The fix is rarely a different model. It is that nobody built a branch for "I cannot answer this", so the system has no way to express it.

Strong on the main slices, weak on abstention Three main evaluation slices score above 0.90, while the abstention slice, questions the corpus cannot answer, scores about 0.44, well below the release bar of 0.80. The slice a blended score hides release bar 0.80 0.92 0.94 0.91 0.44 main slice main slice main slice abstention Questions the corpus cannot answer are where a confident invention reaches a customer.
A system can look excellent on its common slices and fail almost totally on abstention.

Grading free text without a judge model

Using a model to grade another model is the common answer and it works, with conditions. It costs money on every run, it drifts when the judge is updated, and it has biases: it rewards length, and it rates output from its own model family more generously than it deserves.

So calibrate it. Have people grade a sample of the same cases, measure the agreement, and quote that agreement next to every score the judge produced. A judge that agrees with your reviewers 89 percent of the time is a useful instrument. A judge nobody has checked is a number with no error bar.

For a large part of grounding you do not need a judge at all. Whether a claim appears in a retrieved passage, whether a citation resolves, and whether the system abstained are all decidable by ordinary code. That is why the graders in r4agent are deterministic and offline: a benchmark that costs nothing to re-run gets re-run.

What to report

  • Unsupported claim rate, overall and per slice
  • Citation validity, split into resolves and supports
  • Abstention rate on questions the corpus cannot answer
  • Where a judge model was used, its agreement with human reviewers

The benchmark reports all four, and the failing cases come back in full so the pattern behind them is visible rather than inferred.

Common questions

How do you measure hallucination in an LLM?

By checking each claim in an answer against the passages the system actually retrieved, and reporting the share of answers containing at least one unsupported claim. Measure separately what the system does when the corpus contains no answer: saying so is the correct behaviour and should be scored as a pass.

What is a good hallucination rate?

It depends entirely on consequence. A few percent of unsupported claims may be tolerable in an internal search tool and unacceptable in anything a customer or a regulator acts on. Set the threshold per slice before the run rather than judging the number afterwards.

Do I need a judge model to measure hallucination?

Not for most of it. Whether a claim appears in a retrieved passage, whether a citation resolves, and whether the system abstained are all decidable in ordinary code. Use a judge model where the answer is free text with no single correct form, and quote its agreement with human reviewers whenever you quote its score.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark