Applied ML

Retrieval is the ceiling. The model is rarely the problem

When a document question answering system gives a poor answer, the first instinct is almost always to change the model or rewrite the prompt. In our experience that is the right diagnosis perhaps one time in five.

Check whether the answer was in the prompt at all

Before anything else, take the failing question, run the retrieval step alone, and read what came back. Not the generated answer, the retrieved passages.

Most of the time the passage containing the answer is not among them. At that point the generation step is being asked to produce something it was never given, and it will do one of two things: say it does not know, which is the honest failure, or assemble something plausible from what it does have, which is the failure people complain about.

No amount of prompt engineering fixes an absent document. Retrieval sets a hard ceiling on the whole system, and everything downstream operates under it.

Retrieval recall sets a hard ceiling on the whole system A bar showing that when retrieval recall is seventy per cent, thirty per cent of questions can never be answered because the passage never reached the prompt, and all prompt and model work happens below that ceiling. Retrieval is the ceiling 100% 70% 0% 30% never retrieved 70% of answers reachable recall ceiling Above the line No prompt and no model can recover an answer that never arrived. Below the line Prompt work and model choice operate here, always under the ceiling.
Measure retrieval recall first. It is the hard ceiling every downstream step works beneath.

Chunking is where most of the damage happens

Splitting documents into fixed-size pieces is the default because every tutorial does it. It is also the single decision that costs most systems their accuracy.

Fixed-size splitting cuts wherever the character count runs out. That lands in the middle of tables, between a heading and the paragraph it governs, and between a definition and the sentence that qualifies it. The chunk that ends up in the index is then a fragment whose meaning depended on text that is now in a different chunk.

Two corrections do most of the work. Split on the document's own structure first, using headings, sections, and list boundaries, and only fall back to size limits inside a section that is too long. And overlap the windows, so a sentence sitting on a boundary appears in full in at least one chunk. Overlap costs storage, which is cheap, and it removes an entire class of failure.

Carry the heading path into the chunk text as well. A paragraph that reads "this must be filed within thirty days" is not retrievable on its own, because it contains none of the words a user would search with. Prefixed with its section title, it is.

Dense embeddings are bad at the things your users type

Vector search generalises well across phrasing, which is why it works so well in demonstrations. It generalises poorly across exact tokens, which is what people actually search for in a working system.

Part numbers, error codes, policy identifiers, surnames, dates, statute references: these are the queries that matter most in enterprise search and they are precisely where embeddings are weakest, because the embedding of a code is close to the embedding of a similar-looking code that means something else entirely.

The fix is not to abandon vectors. It is to run keyword search alongside them and merge the results. A hybrid of the two, combined by rank rather than by raw score, is substantially better than either alone and is not difficult to build. Raw scores from two different systems are not on a common scale and should not be added.

Measure recall before you measure anything else

Systems get tuned on answer quality because answer quality is what users see. It is the wrong first metric, because it confounds retrieval and generation.

Build a small set of questions with the passage that answers each one identified by hand. Then measure, for the retrieval step alone, how often that passage appears in the top results. That number is your ceiling. If it is seventy per cent, then thirty per cent of your questions cannot be answered correctly no matter what happens next, and every hour spent on the prompt is spent under that constraint.

It also tells you where to spend. A system with high recall and poor answers has a generation problem worth working on. A system with low recall has a retrieval problem, and the prompt is not involved.

Retrieve widely, then rank carefully

There is a tension between fetching enough to be sure the answer is present and fetching few enough that the answer is not buried. Trying to resolve it with a single search is unnecessary.

Fetch generously, then re-rank. A cross-encoder scores each candidate against the query directly rather than comparing precomputed vectors, which is far more accurate and far too slow to run over a whole corpus. Run over twenty or fifty candidates it is fast enough, and it reliably moves the right passage to the top.

This two-stage shape, cheap and broad then expensive and narrow, is the standard architecture in search for good reasons, and it transfers directly.

Two-stage retrieval, cheap and broad then expensive and narrow A three-stage flow: fetch fifty candidates cheaply with vector and keyword search, re-rank them with a cross-encoder against the query, then read ten well-ranked passages with the best placed where it will be used. Retrieve widely, then rank carefully cheap and broad expensive and narrow Fetch generously ~50 candidates vector plus keyword search, merged by rank Re-rank a cross-encoder scores each candidate against the query directly accurate, run on few Read 10 ranked passages best placed where it will be used
A broad, cheap fetch guarantees the answer is present, then a slow, accurate re-rank moves it to the top.

Position in the prompt is not neutral

Passages placed in the middle of a long context get used less than passages at the start or the end. This is well documented and it is stable enough to design around.

So the ordering of retrieved passages matters, and stuffing more of them in is not free. Ten well-ranked passages usually beat forty, both because the noise dilutes attention and because the best passage is more likely to sit where it will be read.

If a passage is worth including, it is worth putting somewhere it will be used.

Freshness is a correctness property

The last failure is organisational rather than technical. Indexes are built once during the project and then drift.

A retrieval system over a policy library, a product catalogue or a knowledge base is answering from a snapshot. When the underlying document is revised and the index is not rebuilt, the system does not degrade gracefully; it confidently returns the superseded version, and confidently is the operative word.

Track the source document's modification time alongside each chunk, rebuild on change rather than on a monthly cron, and surface the document date in the answer so a reader can judge for themselves. That last one costs nothing and prevents the most expensive kind of mistake.

Back to Insights

Tell us what you are trying to automate

We will say plainly whether machine learning is the right tool for it.

Start a conversation