Most machine learning projects that stall do not stall on modelling. They stall because nobody can say whether the last change made things better, and after a few weeks of that the team is arguing from anecdotes.
A team without a fixed evaluation set reasons from whatever they last looked at. Someone tries a different prompt, checks five examples, and the five look better. The change ships. Two weeks later a different person tries five different examples and concludes the system got worse.
Both observations are real. Neither is evidence, because neither was measured against the same thing. The purpose of an evaluation harness is not sophistication, it is having a number that means the same thing on Tuesday as it did on Monday.
This is why it has to come first. A harness built after the model exists is built from cases the model already handles, because those are the cases the team has been looking at. It measures agreement with the current behaviour rather than agreement with what is correct.
The instinct is to make the evaluation set large. Size is the least important property it has.
What matters is that each case has an answer somebody is prepared to defend. That means the set has to be assembled with the people who know the domain, and it means the disagreements found while assembling it are the most valuable output of the whole exercise. Two experts labelling the same case differently is not noise to be resolved by majority vote; it usually means the task itself is underspecified, and the model was never going to do better than the specification.
A hundred cases, argued over properly, will surface every ambiguity in the problem statement. Ten thousand cases labelled quickly will hide them and give you a precise measurement of the wrong thing.
A random sample of production traffic is dominated by the easy majority. A model that handles the common case and fails everything unusual can score above ninety per cent on it, and that number will be quoted in a steering meeting.
Build the set deliberately instead, with named slices. Cases the current system gets wrong. Cases where the input is malformed, truncated or in the wrong language. Cases that sit on a decision boundary the business actually cares about. Cases that are adversarial, in the plain sense that a person trying to get an unreasonable answer would send them.
Report per slice as well as overall. An aggregate that moves up while a slice moves down is the normal situation, not the exception, and only the per-slice view shows it.
The most useful thing an evaluation reports is not the average. It is the list of cases that used to pass and now fail.
Model changes are rarely uniform improvements. A prompt revision, a retrieval change, a fine-tune: each tends to fix a group of failures and break a smaller group of successes. If the fixed group is larger, the aggregate goes up and everyone is pleased. The broken group is still broken, and if it contains something that matters more than average, the release is a regression dressed as an improvement.
Diff the results, not just the totals. It takes an afternoon to build and it changes what the team argues about.
Two failure modes, both slow enough to miss.
The set stops resembling production. Inputs shift as users learn what the system does, a new customer segment arrives with different phrasing, an upstream form changes its fields. The evaluation keeps reporting a stable number about a distribution that no longer exists.
The set leaks into development. Cases get used for debugging, prompts get tuned until those cases pass, and the measurement quietly becomes a training signal. This is not misconduct, it is what happens when the same hundred examples are the only ones anyone looks at.
The defence for both is refresh on a schedule. Sample new cases from production regularly, label them the same way, retire the oldest. And keep a small set that is never used for debugging, only for release decisions, held by someone who is not optimising against it.
Where the output is free text, exact match does not work and human review does not scale. Using a model to judge is the common answer and it is workable, with a caveat.
The judge is a model, so it has its own failure modes. It rewards length. It prefers output that resembles its own style, which means it will rate a system built on the same family more generously. It is inconsistent near the boundary between adjacent scores.
Treat the judge as an instrument that needs calibrating. Have people grade a sample of the cases the judge graded, measure the agreement, and quote the agreement whenever you quote the score. A judge that agrees with your reviewers most of the time is a useful proxy. A judge nobody has checked is a number with no error bar and an unexamined bias.
A file of cases in version control, each with an input, an expected answer, and a note on why that answer is right. A script that runs the current system over all of them and writes results to a dated file. A second script that diffs two result files and prints what changed in each direction.
That is a day of work and it is enough to stop the arguing. Everything else is refinement.
We will say plainly whether machine learning is the right tool for it.
Start a conversation