We have argued before that retrieval is the ceiling. This is the practical follow up: how to measure that ceiling, so you know how much of your failure rate is even addressable by prompting.
Answerable at all. The share of your cases where the answer exists anywhere in the corpus. Establish this first, by hand, with the people who wrote the documentation. If it is 0.93, then seven per cent of your users will be failed by a content gap and no amount of model work touches it. Teams routinely spend a quarter tuning against a ceiling nobody measured.
Recall at k. Of the answerable cases, the share where the supporting passage appears in the top k results actually passed to the model. Use the k your system really uses. Recall at 20 is a comfortable number and irrelevant if you paste five passages into the prompt.
Mean reciprocal rank. Where in those results the right passage landed. Recall at 5 treats first place and fifth place identically; models do not. A passage at rank five competes with four distractors for attention, and the answer degrades before recall records anything.
With these three you can split every failing case into one of three buckets, which is the whole point of measuring retrieval separately:
In our experience the third bucket is the smallest and receives most of the attention, because it is the one that feels like AI work.
Fixed size chunking cuts a table in half, separates a heading from the paragraph it governs, and splits the sentence carrying the qualifier from the sentence carrying the claim. Retrieval then returns a fragment that is genuinely about the topic and does not contain the answer, which scores as a hit on a naive metric and fails the user.
This is why recall has to be judged against the passage that actually supports the answer, chosen by a person, rather than against whatever document the answer came from.
For each case, record the passage a knowledgeable person would point to. That mapping is slow to build and it is the most durable asset in the whole evaluation: it survives changing the embedding model, the vector store, the chunker and the language model, and it lets you compare any of those in an afternoon.
Our benchmark reports retrieval separately from generation for exactly this reason, so the report says which component to work on rather than that the system scored 0.78.
The share of answerable questions where the passage supporting the answer appears in the top k retrieved results that are actually passed to the model. Measure it at the k your system really uses, since a good recall at 20 means nothing if only five passages reach the prompt.
Usually because recall was measured against the right document rather than the right passage. Fixed size chunking often returns a fragment that is on topic and does not contain the answer, which scores as a hit and fails the user.
Measure which bucket your failures fall into before deciding. Failures split into the answer not existing in the corpus, existing but not reaching the prompt, and reaching the prompt but being mishandled. Only the third is a model problem, and it is usually the smallest of the three.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark