Model assurance

The most useful thing an evaluation reports is the list of cases that used to pass

Ask most teams what their evaluation reports and they will tell you a score. The score is the least useful output. The useful one is the list of cases that passed last week and fail now.

Improvements are not uniform

A prompt revision, a retrieval change, a fine tune, a model upgrade: each tends to fix a group of failures and break a smaller group of successes. If the fixed group is larger, the aggregate goes up and the release is called a success.

The broken group is still broken. If it contains something that matters more than average, and it often does because the unusual cases are the fragile ones, then the release is a regression wearing an improvement's clothes and the dashboard is applauding.

Diff the results, not the totals

The mechanism is trivial. Every run writes a dated file with one row per case and its verdict. A second script compares two files and prints two lists: cases that went from pass to fail, and cases that went from fail to pass.

That is an afternoon of work. It changes the review conversation from "the score went from 0.86 to 0.89" to "eleven cases were fixed, four broke, and one of the four is in the billing slice", which is a conversation that can actually reach a decision.

A rising aggregate score next to the case level diff underneath it Left: the blended score rises from 0.86 to 0.89. Right: a diff of the two runs lists eleven cases newly passing and four newly failing, one of the four flagged critical in the billing slice. The score rose, and something broke Aggregate score 0.86 0.89 reads as a success the total says nothing broke Case level diff of the two runs Now passing 11 cases fixed Now failing 4 cases broke 1 is critical billing slice, blocks release
The aggregate rose, but the diff names four broken cases, one critical in billing, which is what a release decision actually turns on.

Sort the broken list by severity, not by count

Four newly failing cases is not a number that means anything on its own. Four newly failing cases where one is marked critical is a release blocker.

This is why cases carry a severity when they are written rather than being triaged when they break: at the moment a case breaks, everyone is looking at a release date and nobody is neutral about how serious it is.

Watch for the cases that quietly disappear

A diff should also report cases present in the old run and missing from the new one. Cases go missing when somebody renames an id, edits a suite while debugging, or removes a case that had become inconvenient.

None of that is misconduct. It is what happens when the same set of examples is the only thing anyone looks at, and it is exactly how an evaluation stops measuring what it was built to measure. A diff that reports disappearances makes it visible in the same afternoon rather than a quarter later.

Keep a set nobody debugs against

The last defence is a small set used only for release decisions, held by somebody who is not optimising against it. Everything else drifts towards the cases the team has been staring at; this one does not, and it is the only number in the process that stays honest by construction.

Our benchmark hands over the diff tool with the suite, and the open harness reports newly failing cases sorted by severity, plus anything that went missing between runs.

Common questions

How do I know if a change to my AI system made it worse?

Compare results case by case rather than comparing scores. An aggregate can rise while specific cases break, and the list of cases that used to pass and now fail is the only view that shows it.

What is regression testing for an LLM application?

Running a fixed case suite before and after a change and diffing the per case verdicts, so newly failing cases are named rather than averaged into a score. Sorting the newly failing list by severity turns it into a release decision.

Why do evaluation sets stop being useful over time?

Two reasons. The cases stop resembling production as users and inputs change, and the set leaks into development as prompts get tuned until those specific cases pass. Refreshing on a schedule and keeping a small set nobody debugs against are the defences.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark