Teams talk about hallucination as if it were weather. It is measurable, and the measurement splits cleanly in two: what the system asserts without support, and what it does when it has nothing to go on.
Take the answer, break it into individual claims, and check each against the passages the system actually retrieved. The share of answers containing at least one claim no passage supports is the unsupported claim rate.
Two details decide whether the number is honest. Check against what was retrieved, not against the whole corpus: an answer that happens to be true but was not supported by anything in front of the model is a lucky guess and will not stay lucky. And check claims, not sentences, because a sentence usually carries several and only one of them needs to be invented.
A system can cite diligently and still be wrong. The citation points at a real document, the document exists, the link resolves, and the passage does not support the sentence attached to it.
So measure separately whether a citation resolves and whether it supports. The second is the one that matters and it is the one usually skipped, because the first is easy to automate and produces a comfortable number.
Ask the system something the corpus cannot answer.
This is the single most informative slice in a benchmark and most evaluation sets contain none of it, because the cases are drawn from questions somebody already answered. Every case in such a set has an answer, so the system is never observed in the situation that produces its worst behaviour.
The correct response is to say it does not know, or to escalate. Score that as a pass. A confident invented answer to an unanswerable question is the failure that reaches a customer and gets acted on, and a system with no abstention path will fail this slice close to totally while scoring well on everything else.
We have measured agents at above 0.90 on their three main slices and below 0.45 on abstention. The fix is rarely a different model. It is that nobody built a branch for "I cannot answer this", so the system has no way to express it.
Using a model to grade another model is the common answer and it works, with conditions. It costs money on every run, it drifts when the judge is updated, and it has biases: it rewards length, and it rates output from its own model family more generously than it deserves.
So calibrate it. Have people grade a sample of the same cases, measure the agreement, and quote that agreement next to every score the judge produced. A judge that agrees with your reviewers 89 percent of the time is a useful instrument. A judge nobody has checked is a number with no error bar.
For a large part of grounding you do not need a judge at all. Whether a claim appears in a retrieved passage, whether a citation resolves, and whether the system abstained are all decidable by ordinary code. That is why the graders in r4agent are deterministic and offline: a benchmark that costs nothing to re-run gets re-run.
The benchmark reports all four, and the failing cases come back in full so the pattern behind them is visible rather than inferred.
By checking each claim in an answer against the passages the system actually retrieved, and reporting the share of answers containing at least one unsupported claim. Measure separately what the system does when the corpus contains no answer: saying so is the correct behaviour and should be scored as a pass.
It depends entirely on consequence. A few percent of unsupported claims may be tolerable in an internal search tool and unacceptable in anything a customer or a regulator acts on. Set the threshold per slice before the run rather than judging the number afterwards.
Not for most of it. Whether a claim appears in a retrieved passage, whether a citation resolves, and whether the system abstained are all decidable in ordinary code. Use a judge model where the answer is free text with no single correct form, and quote its agreement with human reviewers whenever you quote its score.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark