We build an agent that reads a draft clinical trial protocol and reports where it contradicts itself. The first version that actually worked was much less impressive than the first version that looked like it worked, and the distance between those two is most of what the project taught us.
A protocol runs to a hundred pages or more, so the obvious architecture is the one everybody reaches for. Chunk it, embed it, retrieve against a question, let the model answer. That design answers questions about a document. Finding contradictions is a different shape of problem.
Retrieval ranks passages by similarity to a query. A contradiction is two passages that are similar in topic and different in content. The primary endpoint is change from baseline at week 12 and efficacy will be assessed at week 16 sit close together in embedding space and close to almost any query about endpoints. Retrieval will return one of them. Nothing in the ranking makes it return both, and a model cannot detect a disagreement when it has only been shown one side.
This is the same ceiling we have written about in a different setting: what limits these systems is usually what reaches the prompt rather than what the model does with it. Retrieval is the ceiling. Contradiction finding makes it sharper, because the requirement is not a relevant passage but a specific pair, and the pair is defined by a relationship the retriever does not represent.
What worked was abandoning the document as the unit of work.
The agent parses the protocol into a structured representation before it checks anything: objectives, endpoints, estimand attributes, eligibility criteria, the visit schedule, assessment timings, discontinuation rules, analysis populations. Every extracted item carries the section and the span it came from.
The checks then run over that structure rather than over prose. Does each endpoint named in the design section appear in the objectives. Does every analysis population have a definition. Does each assessment in the schedule have a matching timing in the section that specifies it. Does the data collection rule for discontinued participants permit the data the estimand strategy depends on.
The division of labour is the whole design. Extraction is fuzzy and needs a language model, because protocols express the same fact in prose, in tables and in footnotes, and no parser survives that. Comparison is exact and needs no model at all. Once week 12 and week 16 are two typed values on the same field, noticing they differ is not a reasoning task.
Putting comparison in code rather than in a prompt bought three things. It is reproducible, which matters when the output has to be defended. It is auditable, because each check is a written statement someone can read and argue with. And it does not quietly change behaviour when the model is updated.
We built the evaluation before we built the agent, for the reasons we have set out elsewhere. The cases were real protocols with findings already known and confirmed, each one recorded with both of its source locations.
Choosing the metric took longer than building the harness. The instinct is to measure how many known issues the agent finds. That number is easy to move and describes the wrong property.
A reviewer opens a findings report once. If the first five entries are noise, the agent has spent its credibility and the rest go unread, whether or not they are correct. Recall does not survive that. What matters is the proportion of reported findings that are real, weighted towards the top of the list.
So the agent is tuned to be quiet. It drops findings it cannot anchor on both sides. It drops findings where extraction confidence on either half is low, because a comparison between two facts is only as good as the weaker of the two. A run reporting eight real findings is worth more than one reporting thirty of which twelve are real, and the second scores better on almost any accuracy measure you would write down first.
One rule did more for adoption than any modelling change. Every finding has to cite both halves, quoted, with section references. If the agent cannot produce both, the finding does not ship.
That is a hard constraint rather than a formatting preference. It removes an entire class of output that models generate readily: the confident general observation. The safety section appears inconsistent with the schedule cannot be checked and cannot be acted on. Two quotes and two section numbers can be confirmed in fifteen seconds, and dismissed just as fast when wrong.
Being easy to dismiss turned out to matter as much as being right. Reviewers trust a tool they can overrule quickly.
It does not judge whether the science is sound. It holds no view on whether an endpoint is the right endpoint, whether a sample size assumption is plausible, or whether the design answers a question worth asking. Those are the parts of review that need a person, and they are not the parts that consume the hours.
It does not claim completeness. It is a second reader. A clean report means the checks it ran found nothing, not that the protocol is consistent. Presenting it as more than that would be untrue, and in a regulated setting it would be a claim someone has to defend later.
Every run records the model version, the extracted structure, the checks executed and the spans they read. A qualified person reviews the output before it is used. That pattern is not specific to this agent, and it is the reason a tool like this can sit inside a validated process at all.
Not in the model. It was in deciding what the model was for. Reading a hundred pages and holding them in mind is the part that looked hard, and it became tractable as soon as the task was represented as typed facts with anchors instead of text with questions asked of it. Everything after that was ordinary engineering.
The agent this describes is one of several we run in clinical research, and the pattern generalises further than we expected. Any review task where the answer depends on two places in a long document agreeing has the same shape, and almost none of them are well served by asking a model to read it.
We will say plainly whether machine learning is the right tool for it.
Start a conversation