Model assurance

Build the evaluation before you build the model

Most machine learning projects that stall do not stall on modelling. They stall because nobody can say whether the last change made things better, and after a few weeks of that the team is arguing from anecdotes.

Without a harness, every change is an opinion

A team without a fixed evaluation set reasons from whatever they last looked at. Someone tries a different prompt, checks five examples, and the five look better. The change ships. Two weeks later a different person tries five different examples and concludes the system got worse.

Both observations are real. Neither is evidence, because neither was measured against the same thing. The purpose of an evaluation harness is not sophistication, it is having a number that means the same thing on Tuesday as it did on Monday.

This is why it has to come first. A harness built after the model exists is built from cases the model already handles, because those are the cases the team has been looking at. It measures agreement with the current behaviour rather than agreement with what is correct.

Building the evaluation first versus building it after the model Two lanes. Building the eval first measures the model against a defined notion of correct. Building it after the model measures agreement with behaviour the model already has. Order decides what the number means Evaluation built first Define correct with the domain Fixed case set versioned Build model against truth Evaluation built after Build model Cases it already passes Measured against its own behaviour Cases drawn from what the model already does only confirm it agrees with itself.
Built first, the harness scores the model against a defined answer. Built afterward, it scores the model against itself.

A hundred cases you argued about beats ten thousand you scraped

The instinct is to make the evaluation set large. Size is the least important property it has.

What matters is that each case has an answer somebody is prepared to defend. That means the set has to be assembled with the people who know the domain, and it means the disagreements found while assembling it are the most valuable output of the whole exercise. Two experts labelling the same case differently is not noise to be resolved by majority vote; it usually means the task itself is underspecified, and the model was never going to do better than the specification.

A hundred cases, argued over properly, will surface every ambiguity in the problem statement. Ten thousand cases labelled quickly will hide them and give you a precise measurement of the wrong thing.

Hold out the cases that are hard for a reason

A random sample of production traffic is dominated by the easy majority. A model that handles the common case and fails everything unusual can score above ninety per cent on it, and that number will be quoted in a steering meeting.

Build the set deliberately instead, with named slices. Cases the current system gets wrong. Cases where the input is malformed, truncated or in the wrong language. Cases that sit on a decision boundary the business actually cares about. Cases that are adversarial, in the plain sense that a person trying to get an unreasonable answer would send them.

Report per slice as well as overall. An aggregate that moves up while a slice moves down is the normal situation, not the exception, and only the per-slice view shows it.

A random sample hides weak slices that a deliberate set exposes Left: a random production sample is dominated by an easy majority and reports one high pass rate. Right: named slices are each measured on their own, and the hard slices show much lower pass rates. One number hides the slices that fail Random sample easy majority dominates 0.92 overall pass rate the hard cases are the thin sliver, and they hide Named slices, each scored Common case 0.97 Currently wrong 0.41 Malformed input 0.55 Decision boundary 0.62 Adversarial 0.38
A single blended rate looks healthy. The same cases split into named slices show where the system is actually weak.

The score is a tripwire, not a summary

The most useful thing an evaluation reports is not the average. It is the list of cases that used to pass and now fail.

Model changes are rarely uniform improvements. A prompt revision, a retrieval change, a fine-tune: each tends to fix a group of failures and break a smaller group of successes. If the fixed group is larger, the aggregate goes up and everyone is pleased. The broken group is still broken, and if it contains something that matters more than average, the release is a regression dressed as an improvement.

Diff the results, not just the totals. It takes an afternoon to build and it changes what the team argues about.

Evaluation sets go stale, quietly

Two failure modes, both slow enough to miss.

The set stops resembling production. Inputs shift as users learn what the system does, a new customer segment arrives with different phrasing, an upstream form changes its fields. The evaluation keeps reporting a stable number about a distribution that no longer exists.

The set leaks into development. Cases get used for debugging, prompts get tuned until those cases pass, and the measurement quietly becomes a training signal. This is not misconduct, it is what happens when the same hundred examples are the only ones anyone looks at.

The defence for both is refresh on a schedule. Sample new cases from production regularly, label them the same way, retire the oldest. And keep a small set that is never used for debugging, only for release decisions, held by someone who is not optimising against it.

Judging generated output is a measurement problem too

Where the output is free text, exact match does not work and human review does not scale. Using a model to judge is the common answer and it is workable, with a caveat.

The judge is a model, so it has its own failure modes. It rewards length. It prefers output that resembles its own style, which means it will rate a system built on the same family more generously. It is inconsistent near the boundary between adjacent scores.

Treat the judge as an instrument that needs calibrating. Have people grade a sample of the cases the judge graded, measure the agreement, and quote the agreement whenever you quote the score. A judge that agrees with your reviewers most of the time is a useful proxy. A judge nobody has checked is a number with no error bar and an unexamined bias.

What this looks like on day one

A file of cases in version control, each with an input, an expected answer, and a note on why that answer is right. A script that runs the current system over all of them and writes results to a dated file. A second script that diffs two result files and prints what changed in each direction.

That is a day of work and it is enough to stop the arguing. Everything else is refinement.

Back to Insights

Tell us what you are trying to automate

We will say plainly whether machine learning is the right tool for it.

Start a conversation