Applied ML

Recall at k is the ceiling on every answer your agent gives

We have argued before that retrieval is the ceiling. This is the practical follow up: how to measure that ceiling, so you know how much of your failure rate is even addressable by prompting.

Three numbers, in order of usefulness

Answerable at all. The share of your cases where the answer exists anywhere in the corpus. Establish this first, by hand, with the people who wrote the documentation. If it is 0.93, then seven per cent of your users will be failed by a content gap and no amount of model work touches it. Teams routinely spend a quarter tuning against a ceiling nobody measured.

Retrieval sets the ceiling on the final answer rate Answerable cases sit at 0.93, recall at k reaches 0.78, and final correct answers land at 0.72. The recall at k line is a ceiling that prompting cannot cross. Retrieval sets the ceiling Answerable in the corpus 0.93 Passage reaches the prompt, recall at k 0.78 Final correct answers 0.72 no prompting crosses this line
The final answer rate cannot rise above the passages that actually reach the prompt.

Recall at k. Of the answerable cases, the share where the supporting passage appears in the top k results actually passed to the model. Use the k your system really uses. Recall at 20 is a comfortable number and irrelevant if you paste five passages into the prompt.

Mean reciprocal rank. Where in those results the right passage landed. Recall at 5 treats first place and fifth place identically; models do not. A passage at rank five competes with four distractors for attention, and the answer degrades before recall records anything.

Attribute failures before you fix them

With these three you can split every failing case into one of three buckets, which is the whole point of measuring retrieval separately:

  • The answer was not in the corpus. A content problem. Write the missing page.
  • The answer was in the corpus but did not reach the prompt. A retrieval problem. Chunking, embeddings, query construction, reranking.
  • The answer reached the prompt and the model still got it wrong. A generation problem, and the only bucket where prompting helps.
Failing cases split into three buckets A bar of failing cases splits into content gaps, retrieval misses and generation errors. The generation bucket is the smallest, yet it receives most of the attention. Where a failure actually happened every failing case, by cause content 45% retrieval 40% gen 15% Not in the corpus a content problem write the missing page In corpus, not retrieved a retrieval problem chunking, embeddings, reranking Reached prompt, mishandled a generation problem the only bucket prompting fixes The model bucket is usually the smallest, and it receives most of the attention.
Measuring retrieval separately is what lets you sort failures into the bucket that owns the fix.

In our experience the third bucket is the smallest and receives most of the attention, because it is the one that feels like AI work.

Chunking is usually where the damage is

Fixed size chunking cuts a table in half, separates a heading from the paragraph it governs, and splits the sentence carrying the qualifier from the sentence carrying the claim. Retrieval then returns a fragment that is genuinely about the topic and does not contain the answer, which scores as a hit on a naive metric and fails the user.

This is why recall has to be judged against the passage that actually supports the answer, chosen by a person, rather than against whatever document the answer came from.

Build the retrieval set once, keep it

For each case, record the passage a knowledgeable person would point to. That mapping is slow to build and it is the most durable asset in the whole evaluation: it survives changing the embedding model, the vector store, the chunker and the language model, and it lets you compare any of those in an afternoon.

Our benchmark reports retrieval separately from generation for exactly this reason, so the report says which component to work on rather than that the system scored 0.78.

Common questions

What is recall at k in a RAG system?

The share of answerable questions where the passage supporting the answer appears in the top k retrieved results that are actually passed to the model. Measure it at the k your system really uses, since a good recall at 20 means nothing if only five passages reach the prompt.

Why is my RAG system wrong even though retrieval looks fine?

Usually because recall was measured against the right document rather than the right passage. Fixed size chunking often returns a fragment that is on topic and does not contain the answer, which scores as a hit and fails the user.

Should I improve the model or the retrieval first?

Measure which bucket your failures fall into before deciding. Failures split into the answer not existing in the corpus, existing but not reaching the prompt, and reaching the prompt but being mishandled. Only the third is a model problem, and it is usually the smallest of the three.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark