Applied ML

Checking a protocol for contradictions is not a retrieval problem

We build an agent that reads a draft clinical trial protocol and reports where it contradicts itself. The first version that actually worked was much less impressive than the first version that looked like it worked, and the distance between those two is most of what the project taught us.

It looks like retrieval, and it is not

A protocol runs to a hundred pages or more, so the obvious architecture is the one everybody reaches for. Chunk it, embed it, retrieve against a question, let the model answer. That design answers questions about a document. Finding contradictions is a different shape of problem.

Retrieval ranks passages by similarity to a query. A contradiction is two passages that are similar in topic and different in content. The primary endpoint is change from baseline at week 12 and efficacy will be assessed at week 16 sit close together in embedding space and close to almost any query about endpoints. Retrieval will return one of them. Nothing in the ranking makes it return both, and a model cannot detect a disagreement when it has only been shown one side.

This is the same ceiling we have written about in a different setting: what limits these systems is usually what reaches the prompt rather than what the model does with it. Retrieval is the ceiling. Contradiction finding makes it sharper, because the requirement is not a relevant passage but a specific pair, and the pair is defined by a relationship the retriever does not represent.

Retrieval returns one side, a contradiction needs both On the left, retrieval ranks by similarity and returns only the week 12 passage while the week 16 passage is left behind. On the right, both passages are held together so the conflict between week 12 and week 16 is visible. A contradiction needs both halves at once What retrieval returns query: which week is the endpoint? Endpoint: change at week 12 returned, top ranked Efficacy assessed at week 16 not returned similarity ranks one side, not the pair What the check needs Endpoint: change at week 12 Efficacy assessed at week 16 conflict found week 12 vs week 16 both halves held together at once
The two passages are similar in topic and different in content, the one shape retrieval will not surface as a pair.

Extract first, compare second

What worked was abandoning the document as the unit of work.

The agent parses the protocol into a structured representation before it checks anything: objectives, endpoints, estimand attributes, eligibility criteria, the visit schedule, assessment timings, discontinuation rules, analysis populations. Every extracted item carries the section and the span it came from.

The checks then run over that structure rather than over prose. Does each endpoint named in the design section appear in the objectives. Does every analysis population have a definition. Does each assessment in the schedule have a matching timing in the section that specifies it. Does the data collection rule for discontinued participants permit the data the estimand strategy depends on.

The division of labour is the whole design. Extraction is fuzzy and needs a language model, because protocols express the same fact in prose, in tables and in footnotes, and no parser survives that. Comparison is exact and needs no model at all. Once week 12 and week 16 are two typed values on the same field, noticing they differ is not a reasoning task.

Extract with a model, compare in code Protocol prose, tables and footnotes are turned into typed facts with section anchors by a language model, then compared by plain deterministic code that flags mismatches. Extract with a model, compare in code Protocol prose, tables, footnotes extraction Typed facts each with its section span Compare values flag the mismatches comparison language model, fuzzy plain code, exact Comparison in code is reproducible and auditable, and it does not shift when the model is updated.
The division of labour is the design: the model handles what is fuzzy, code handles what must be exact.

Putting comparison in code rather than in a prompt bought three things. It is reproducible, which matters when the output has to be defended. It is auditable, because each check is a written statement someone can read and argue with. And it does not quietly change behaviour when the model is updated.

Precision is the metric, not accuracy

We built the evaluation before we built the agent, for the reasons we have set out elsewhere. The cases were real protocols with findings already known and confirmed, each one recorded with both of its source locations.

Choosing the metric took longer than building the harness. The instinct is to measure how many known issues the agent finds. That number is easy to move and describes the wrong property.

A reviewer opens a findings report once. If the first five entries are noise, the agent has spent its credibility and the rest go unread, whether or not they are correct. Recall does not survive that. What matters is the proportion of reported findings that are real, weighted towards the top of the list.

So the agent is tuned to be quiet. It drops findings it cannot anchor on both sides. It drops findings where extraction confidence on either half is low, because a comparison between two facts is only as good as the weaker of the two. A run reporting eight real findings is worth more than one reporting thirty of which twelve are real, and the second scores better on almost any accuracy measure you would write down first.

The constraint that made it usable

One rule did more for adoption than any modelling change. Every finding has to cite both halves, quoted, with section references. If the agent cannot produce both, the finding does not ship.

That is a hard constraint rather than a formatting preference. It removes an entire class of output that models generate readily: the confident general observation. The safety section appears inconsistent with the schedule cannot be checked and cannot be acted on. Two quotes and two section numbers can be confirmed in fifteen seconds, and dismissed just as fast when wrong.

Being easy to dismiss turned out to matter as much as being right. Reviewers trust a tool they can overrule quickly.

What it deliberately does not do

It does not judge whether the science is sound. It holds no view on whether an endpoint is the right endpoint, whether a sample size assumption is plausible, or whether the design answers a question worth asking. Those are the parts of review that need a person, and they are not the parts that consume the hours.

It does not claim completeness. It is a second reader. A clean report means the checks it ran found nothing, not that the protocol is consistent. Presenting it as more than that would be untrue, and in a regulated setting it would be a claim someone has to defend later.

Every run records the model version, the extracted structure, the checks executed and the spans they read. A qualified person reviews the output before it is used. That pattern is not specific to this agent, and it is the reason a tool like this can sit inside a validated process at all.

Where the leverage was

Not in the model. It was in deciding what the model was for. Reading a hundred pages and holding them in mind is the part that looked hard, and it became tractable as soon as the task was represented as typed facts with anchors instead of text with questions asked of it. Everything after that was ordinary engineering.

The agent this describes is one of several we run in clinical research, and the pattern generalises further than we expected. Any review task where the answer depends on two places in a long document agreeing has the same shape, and almost none of them are well served by asking a model to read it.

Back to Insights

Tell us what you are trying to automate

We will say plainly whether machine learning is the right tool for it.

Start a conversation