Ask most teams what their evaluation reports and they will tell you a score. The score is the least useful output. The useful one is the list of cases that passed last week and fail now.
A prompt revision, a retrieval change, a fine tune, a model upgrade: each tends to fix a group of failures and break a smaller group of successes. If the fixed group is larger, the aggregate goes up and the release is called a success.
The broken group is still broken. If it contains something that matters more than average, and it often does because the unusual cases are the fragile ones, then the release is a regression wearing an improvement's clothes and the dashboard is applauding.
The mechanism is trivial. Every run writes a dated file with one row per case and its verdict. A second script compares two files and prints two lists: cases that went from pass to fail, and cases that went from fail to pass.
That is an afternoon of work. It changes the review conversation from "the score went from 0.86 to 0.89" to "eleven cases were fixed, four broke, and one of the four is in the billing slice", which is a conversation that can actually reach a decision.
Four newly failing cases is not a number that means anything on its own. Four newly failing cases where one is marked critical is a release blocker.
This is why cases carry a severity when they are written rather than being triaged when they break: at the moment a case breaks, everyone is looking at a release date and nobody is neutral about how serious it is.
A diff should also report cases present in the old run and missing from the new one. Cases go missing when somebody renames an id, edits a suite while debugging, or removes a case that had become inconvenient.
None of that is misconduct. It is what happens when the same set of examples is the only thing anyone looks at, and it is exactly how an evaluation stops measuring what it was built to measure. A diff that reports disappearances makes it visible in the same afternoon rather than a quarter later.
The last defence is a small set used only for release decisions, held by somebody who is not optimising against it. Everything else drifts towards the cases the team has been staring at; this one does not, and it is the only number in the process that stays honest by construction.
Our benchmark hands over the diff tool with the suite, and the open harness reports newly failing cases sorted by severity, plus anything that went missing between runs.
Compare results case by case rather than comparing scores. An aggregate can rise while specific cases break, and the list of cases that used to pass and now fail is the only view that shows it.
Running a fixed case suite before and after a change and diffing the per case verdicts, so newly failing cases are named rather than averaged into a score. Sorting the newly failing list by severity turns it into a release decision.
Two reasons. The cases stop resembling production as users and inputs change, and the set leaks into development as prompts get tuned until those specific cases pass. Refreshing on a schedule and keeping a small set nobody debugs against are the defences.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark