Model assurance

Two models, one blind spot: correlated errors in AI-vs-AI checking

Independent review works for a simple reason: two independent minds rarely make the same mistake in the same place. When a second reviewer signs off, the value is not that someone said yes, it is that someone who could have reached the answer by a different route reached the same one. Their agreement carries information precisely because their errors are uncorrelated. If one of them was going to slip, the other probably would not slip in the identical spot.

AI-vs-AI checking is built to look like that arrangement. You have one model produce an answer, then you have a second model grade it, and the second model's approval feels like a fresh pair of eyes. The trouble is that the second pair of eyes was often trained on the same data, built on the same architecture, and shaped by the same assumptions as the first. When that is true, the two models are not independent reviewers. They are closer to one reviewer consulted twice.

Correlated errors defeat AI-vs-AI checking Top: a producer model and a sibling checker share a blind spot, so they agree on a wrong answer. Bottom: an independent checker breaks the correlation. Agreement is not independence SHARED FAMILY Producer generates answer Sibling checker same training Comparator: agree shared blind spot ! false pass INDEPENDENT CHECK Producer generates answer Independent check other family / code / human Comparator: disagree correlation broken caught
Two models that share a blind spot agree on the same mistake. Independence is what catches it.

Shared training means shared blind spots

The mechanism is worth stating plainly. A model's errors are not random noise scattered evenly across every possible input. They cluster. A model gets a particular kind of arithmetic wrong, misreads a particular kind of ambiguous instruction, or repeats a particular factual error that was common in its training data. Those clusters are its blind spots, and they come largely from what it learned and how it was built.

Now put two models from the same family, or even two models from the same era trained on largely overlapping web-scale data, on either side of a check. Their blind spots overlap heavily. When the producer walks into one of those blind spots and generates a wrong answer, the checker walks into the same blind spot and sees nothing wrong. It approves. The check does not fail loudly. It passes quietly, on the exact input where you most needed it to catch something.

A comparison-based check is built to detect disagreement. That is its whole design. It flags cases where the two models diverge and lets through the cases where they agree. When errors are correlated, the dangerous cases are the ones where the models agree with each other and are both wrong, and those are exactly the cases the check is structurally unable to see.

The rule extends further than people think

Most teams already know the first version of the rule: do not let a model grade its own work. A model asked to score its own answer tends to rate it highly, because the same process that produced the answer also produced the belief that the answer is good. That much is widely accepted.

The part that gets missed is that the rule extends to siblings. Do not let a sibling model grade it either. Same family, same era, same failure patterns. The checker does not need to be the identical model to inherit the identical blind spot. It only needs to have learned the world in roughly the same way. Two different sizes of the same model line, or two models trained on the same public corpus a few months apart, are close enough that their agreement tells you much less than it appears to.

How to restore independence

The fix is to put something genuinely different on at least one side of the check. There are three practical options. Use a model from a different family, trained by a different team on different data with a different architecture, so its blind spots sit in different places. Use a deterministic check, real code that verifies the property you care about directly rather than judging it. Or use a human, at least on a sample.

Where the output is checkable, prefer the deterministic option. If the answer is a number, compute it. If it is a piece of JSON, validate it against the schema. If it is a claim about a document, retrieve the passage and compare. Deterministic checks do not have blind spots that correlate with a language model, because they are not language models. They have a different and much smaller failure surface, and when they pass, the pass means something. Save the model-based judgment for the cases where no deterministic check exists.

More models is not more independence

A tempting response is to add models. If one checker might share a blind spot, surely five checkers voting is safer. It is not, unless the five fail differently from one another. Five models drawn from the same family and the same training era share most of their blind spots, so when the producer hits one of them, all five checkers hit it too. They vote the same wrong way, in unison, and the tally comes back unanimous.

That unanimity is worse than a single checker, because it manufactures false confidence. A lone checker that approves a wrong answer is one opinion. Five that approve it look like overwhelming consensus, and a consensus is much harder to argue against in a review meeting. What matters is not the count of models. It is the count of independent failure modes among them, which is often one no matter how many models you line up.

More models multiply agreement, not independence Five similar models all vote to agree on a wrong answer because they share one blind spot, producing false confidence. A single genuinely independent check disagrees and catches the mistake. More models multiply agreement, not independence FIVE MODELS, ONE TRAINING FAMILY Model 1 Model 2 Model 3 Model 4 Model 5 Five agree, one shared blind spot: confident and wrong ! ONE GENUINELY INDEPENDENT CHECK Different family, deterministic code, or a human disagrees, so it can catch the shared mistake caught
Five similar models share one blind spot, so five votes are still one opinion. Real independence is a check that can disagree.

When AI-as-judge is legitimate

None of this means model-based judging is never valid. It has a real place. The right place is genuinely open-ended output, the kind where no deterministic check can exist because there is no single correct answer to compute against: the tone of a summary, the helpfulness of an explanation, the relative quality of two drafts. For those, a model judge is often the only scalable option, and a reasonable one.

The conditions are that independence is engineered in rather than assumed, so the judge comes from a different family or works alongside a human on a sample, and that its agreement is never treated as proof. A model judge produces a signal. A signal can be tracked, trended, and spot-checked against human judgment. It should not be the thing that stands between a wrong answer and a customer, because on the inputs where the producer was most confidently wrong, a correlated judge is most likely to agree.

Working out where a check is genuinely independent and where it only looks that way is a large part of our agent benchmarking work, because a benchmark that grades a model with its own relatives measures agreement, not correctness, and the difference is the whole point.

Common questions

Is using one model to grade another a valid form of independent review?

Only when the two are actually independent. If the checker shares training data, architecture, or assumptions with the producer, their errors correlate, so the checker tends to approve the same wrong answer the producer generated. The setup looks like independent review but behaves like one system checking itself.

Does adding more models make the check more reliable?

Not by itself. Five models that share a blind spot vote the same wrong way and produce a confident consensus on a mistake. More models add reliability only when they fail differently from each other. Count independent failure modes, not the number of models.

When is using an AI model as a judge actually appropriate?

For genuinely open-ended output where no deterministic check exists, with independence engineered in through a different model family or a human on one side, and with the judge's agreement treated as a signal rather than as proof. Prefer a deterministic check wherever the output is checkable.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark