Model assurance

An acceptable error rate when one rare failure is the whole risk

Ask what error rate is acceptable for a system and the usual answer is a single number: ninety-five percent accurate, a one percent error rate, something that fits on a slide. That number is an average across everything the system does, and for a lot of systems that is fine. It stops being fine the moment one class of error carries almost all of the consequence. A medical triage assistant that is right ninety-nine times out of a hundred sounds excellent until you notice that the hundredth case is the one where it sends someone with a real emergency home. A payments agent that is accurate on average still ruins its quarter if the errors it does make are all wrong transfers. The average is describing the common case, and the common case was never the thing you were worried about.

The mistake is treating error as a single quantity to be minimized. Error is not one thing. A wrong answer that a user notices and shrugs off is not the same event as a wrong answer that triggers an irreversible action in a regulated process, and averaging them into one rate throws away the only distinction that matters. When the consequence is concentrated in a rare failure, the mean tells you almost nothing about your actual exposure. You can drive the mean down for months by improving the easy, high-volume cases and never touch the cell that holds the risk.

So the first move is to stop sorting errors by how often they happen and start sorting them by what they cost. Frequency and consequence are different axes, and it is the combination that decides how much attention an error deserves. A frequent mistake with no real cost is an annoyance you can tune later. A rare mistake with a severe cost is the thing that should set your bar, even though it almost never shows up in the logs. Laying the two axes against each other makes the point obvious.

Consequence versus frequency risk grid A two by two grid. The rare but severe cell is highlighted as the one that should set the acceptable error rate. Set the bar on the severe tail Consequence → Frequency → low high low high Rare + severe This sets your threshold. Give it its own budget, oversight, and a way to catch and reverse the failure. Frequent + severe Usually caught early. Fix or ship nothing. Rare + minor Log it and move on. Frequent + minor Tune for quality, not for safety.
Average accuracy hides the one cell that carries the risk. Threshold on it directly.

A threshold per failure mode, not one for the system

Once errors are separated by consequence, the idea of a single acceptable rate falls apart in a useful way. Each failure mode gets its own threshold. A one percent error rate might be perfectly acceptable for a mode where the worst case is a user rephrasing their question, and completely unacceptable for a mode where the worst case is a missed diagnosis or a payment sent to the wrong account. These are not the same number wearing different labels. They are different commitments, and writing them down separately forces the conversation that a blended number lets everyone avoid: how bad is the bad case here, and how often can we tolerate it.

This is where an error budget belongs, tied to consequence rather than spread evenly. The frequent-and-minor cell can carry a generous budget because spending it costs little. The rare-and-severe cell gets a tiny budget, and more than that, it gets the oversight and containment that the other cells do not need. That is where a human reviews before the action commits, where a second check runs, where the system is built to hold rather than proceed when it is unsure. Concentrating effort there is only possible once you have stopped pretending the cells are interchangeable.

A separate acceptable-error threshold for each failure mode Three failure modes, each with its own acceptable-error bar and threshold mark set by consequence. A single dashed system-average line crosses all three, showing that one blended number fits none of them. One threshold per failure mode tighter looser acceptable rate → system average (3.3%) fits none of them Rephrase needed low consequence 8% Wrong citation erodes trust 1.2% Missed safety-critical case severe, near zero 0.1%
Each mode gets a budget set by its consequence. The average line crosses all three and describes none.

The evaluation problem the average hides

Setting a strict threshold on the rare-and-severe cell creates an immediate measurement problem, and it is the part most teams underestimate. Rare cases are rare in your data too. If you sample production traffic at random and build a test set from it, the severe cell will contain a handful of examples, or none. You can report a beautiful measured error rate for that cell that is really just noise, because it rests on almost no observations. The one number you most need to trust is the one your naturally sampled data is least able to support.

The fix is to over-sample the tail on purpose. Go and find the rare severe cases, pull them from history, from incident reports, from the near-misses that never became incidents, and build a test set that is deliberately dense where production is sparse. Then push further with adversarial cases designed to trip the specific failure, and with synthetic cases that construct situations you have not seen enough of yet. This is the only way to get enough examples in the cell to measure anything.

Synthetic and adversarial evidence comes with a limit you have to state plainly. It can show that a failure is possible and that a given check does or does not catch it. It cannot tell you how often the failure happens in production, because you made the cases up. Use synthetic data to prove the failure exists and to prove your defence works, and keep the frequency estimate anchored to real traffic. Confusing the two, quoting a synthetic pass rate as if it were a production rate, is how a system looks safe on paper and fails in the field.

The plausible-but-wrong output is the real danger

The failure that gets through is almost never the obvious one. An output that is garbled or clearly off gets caught, by the user or by a simple sanity check. The dangerous output is the one that is plausible and wrong: the confident wrong dose, the transfer to a real account that is the wrong real account, the summary that reads correctly and misses the one clause that changes the decision. It passes a shallow check precisely because it looks like a good answer.

That means a general correctness check will not save you. A check has to be designed for the specific severe failure you are worried about. If the risk is a wrong dose, the check reasons about dose ranges, not about whether the reply is fluent. If the risk is a payment to the wrong party, the check verifies the party against an independent source, not against the model's own confidence. The check is as specialized as the failure it exists to catch, and building it starts from naming that failure exactly.

Detection and containment, not just a low rate

A low error rate on the severe cell is worth less than it looks if a failure, when it does happen, runs all the way to consequence with nothing in its path. The goal is not only that the wrong output is rare. It is that the wrong output is catchable and reversible. A wrong transfer that can be held for review and undone is a different risk from the same wrong transfer executed instantly. A wrong clinical suggestion that a clinician sees and signs off is a different risk from one that acts on its own.

So the severe cell needs two things the average never asked for: a check tuned to the specific failure, and a containment path that makes the failure recoverable when the check misses. Rare, caught, and reversible is a real safety posture. Rare on its own is a hope. Building the case set that tells the two apart, per failure mode and weighted by consequence, is close to the work behind our agent benchmarking work, where the number that matters gets decided before a production failure decides it for you.

Common questions

Why is a single averaged error rate a poor target for a high-stakes system?

Because it blends errors that cost nothing with errors that cause the whole harm. A model can hit 99 percent accuracy while every point of the missing 1 percent lands on the one failure mode that hurts a patient, moves money the wrong way, or breaks a regulated decision. The average looks healthy while the risk is entirely unaddressed. You have to set the bar on the specific severe failure, not on the blended number.

How do you set an acceptable error rate per failure mode?

Group errors by consequence rather than by frequency, then assign a threshold to each group. A one percent rate can be fine for a low-consequence mistake and completely unacceptable for a severe one, so the severe group gets a far stricter budget, plus oversight and a way to contain and reverse the error when it happens. The frequent-but-harmless group can absorb a much looser number.

Why does a random test set miss rare severe failures?

Rare-severe cases are, by definition, rare in naturally sampled traffic, so a random sample barely contains them and your measured rate on that cell is built from a handful of examples or none. You have to over-sample the tail on purpose and add adversarial and synthetic cases to probe it, while being explicit that synthetic evidence shows a failure can happen, not how often it happens in production.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark