Model assurance

Run it five times: the measurement almost everyone skips

Nearly every evaluation runs each case once. It is the natural thing to do and it makes an entire category of production failure invisible.

What one run per case cannot see

Generation is not deterministic in most deployed configurations. Temperature is above zero, retrieval can tie and break the tie differently, tool results change, and model endpoints are updated underneath you.

So a case that passes is not a case that passes. It is a case that passed once. Run the suite five times and the picture changes: some cases pass five times out of five, and some pass three.

The same cases scored across five repeated runs A grid of five cases by five runs of an unchanged system. Most cases pass every run, but two cases flip between pass and fail across runs, so their verdict is a coin toss that a single run would hide. One run cannot see a coin toss Run 1 Run 2 Run 3 Run 4 Run 5 Verdict Case 1 Case 2 Case 3 Case 4 Case 5 5 / 5 stable 5 / 5 stable 3 / 5 flaky 5 / 5 stable 4 / 5 flaky pass fail two of five cases disagree with themselves
Run the suite five times and the flaky cases surface. A single run reports each as a clean pass or fail.

The measurement is run to run disagreement, the share of cases that did not return the same verdict every time. We regularly see 5 to 10 per cent on systems whose owners believe them to be deterministic.

Why it matters more than the average suggests

An average absorbs it. A case passing three times in five contributes 0.6 and disappears into a slice score, where it is indistinguishable from a case that half worked every time. Those two are completely different problems: one is a capability gap, the other is a coin toss.

From the user's side the coin toss is worse. A system that is consistently wrong gets a workaround and a support article. A system that is right most of the time and occasionally not cannot be trusted or worked around, and it generates the support ticket that begins "it worked yesterday".

It also invalidates your comparisons

This is the practical sting. If run to run variation is eight per cent, then a change that moves the score by three points has told you nothing. Teams ship prompt revisions on movements smaller than their own noise floor every day.

Measure the noise floor first. Run the suite several times against an unchanged system and see how much the number moves on its own. That figure is the smallest difference any future comparison can honestly claim, and it belongs next to every result you report.

A change smaller than the noise floor claims nothing A scale of score movement in points, with the first eight points shaded as the run to run noise floor. A three point change sits inside the floor and means nothing, while a change past eight points is a claim you can defend. The noise floor is the smallest honest claim 0 8 pts 12 pts run to run noise floor = 8 points a 3 point change inside the floor, tells you nothing past the floor a result you can defend
If the suite moves eight points on its own, a three point improvement is noise wearing the costume of a result.

What to do about it

Set temperature to zero for the parts that should be deterministic, and accept that this does not make the system deterministic end to end. Pin model versions where the provider allows it. Make tie breaking in retrieval explicit rather than incidental. Then re-measure, because the remaining variation is the honest number.

Where variation is inherent, report it rather than hiding it: quote the score with the run to run disagreement beside it. r4agent takes a repeats argument for this and reports which specific cases were unstable, since the identity of a flaky case is more actionable than the rate.

Common questions

How many times should I run an LLM evaluation?

At least five for anything you intend to make a decision on. One run per case cannot distinguish a case that reliably passes from one that passes three times in five, and those are different problems with different fixes.

What is run to run disagreement?

The share of cases that did not return the same verdict across repeated runs of an unchanged system. It is the noise floor of your evaluation and the smallest difference any comparison between two versions can honestly claim.

Does setting temperature to zero make an agent deterministic?

It removes one source of variation, not all of them. Retrieval ties, changing tool results and silently updated model endpoints all remain, which is why the honest approach is to measure the residual variation rather than assume it is gone.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark