Nearly every evaluation runs each case once. It is the natural thing to do and it makes an entire category of production failure invisible.
Generation is not deterministic in most deployed configurations. Temperature is above zero, retrieval can tie and break the tie differently, tool results change, and model endpoints are updated underneath you.
So a case that passes is not a case that passes. It is a case that passed once. Run the suite five times and the picture changes: some cases pass five times out of five, and some pass three.
The measurement is run to run disagreement, the share of cases that did not return the same verdict every time. We regularly see 5 to 10 per cent on systems whose owners believe them to be deterministic.
An average absorbs it. A case passing three times in five contributes 0.6 and disappears into a slice score, where it is indistinguishable from a case that half worked every time. Those two are completely different problems: one is a capability gap, the other is a coin toss.
From the user's side the coin toss is worse. A system that is consistently wrong gets a workaround and a support article. A system that is right most of the time and occasionally not cannot be trusted or worked around, and it generates the support ticket that begins "it worked yesterday".
This is the practical sting. If run to run variation is eight per cent, then a change that moves the score by three points has told you nothing. Teams ship prompt revisions on movements smaller than their own noise floor every day.
Measure the noise floor first. Run the suite several times against an unchanged system and see how much the number moves on its own. That figure is the smallest difference any future comparison can honestly claim, and it belongs next to every result you report.
Set temperature to zero for the parts that should be deterministic, and accept that this does not make the system deterministic end to end. Pin model versions where the provider allows it. Make tie breaking in retrieval explicit rather than incidental. Then re-measure, because the remaining variation is the honest number.
Where variation is inherent, report it rather than hiding it: quote the score with the run to run disagreement beside it. r4agent takes a repeats argument for this and reports which specific cases were unstable, since the identity of a flaky case is more actionable than the rate.
At least five for anything you intend to make a decision on. One run per case cannot distinguish a case that reliably passes from one that passes three times in five, and those are different problems with different fixes.
The share of cases that did not return the same verdict across repeated runs of an unchanged system. It is the noise floor of your evaluation and the smallest difference any comparison between two versions can honestly claim.
It removes one source of variation, not all of them. Retrieval ties, changing tool results and silently updated model endpoints all remain, which is why the honest approach is to measure the residual variation rather than assume it is gone.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark