Classic drift detection was built for a world of numeric predictions. A model outputs a probability, a price, a score, or a class, and you watch the distribution of those numbers over time. When the distribution shifts, a monitoring line bends, an alert fires, and someone investigates. The whole practice assumes the output is a value you can bin, average, and plot. That assumption is the problem, because the output of a generative or agentic system is not a value at all.
When an agent writes code, explains a decision, drafts a summary, or works through a derivation, its output is qualitative. There is no single number to watch. Two answers can be the same length, use the same vocabulary, and score identically on any surface statistic while one is correct and the other is subtly wrong. A dashboard line stays flat because there is no line that captures the difference. The drift is real, it is affecting users, and none of your existing monitoring is built to see it.
The uncomfortable part is that this can happen when nothing on your side moved. A commercial model can change its behavior without changing its name or its version label. The provider retunes a safety layer, adjusts how requests are routed, or ships new served weights behind the same identifier, and your outputs shift even though your prompts, your code, your retrieval, and your data are exactly what they were yesterday.
From where you sit, you are calling the same model you always called. There was no upgrade you approved, no migration you scheduled, no changelog you read. The API contract looks identical. But the thing on the other end of the call is not the same system it was last week, and the answers it returns to your users have moved with it. Treating a stable model name as a stable model is the assumption that lets this drift run unnoticed for weeks.
You cannot watch a distribution for this, so you have to manufacture something you can watch. The detectable form of qualitative drift is regression on a fixed, versioned set of evaluation cases. You assemble a set of inputs that represent the work the system really does, you record the expected behavior for each one, and you version that set so it does not quietly change underneath you. Then you run it on every event that could move behavior: every model swap, every prompt change, and every retrieval change.
The essential discipline is how you read the result. You do not compare the new run against the old run as aggregates, because an average is exactly the thing that hides this problem. You diff it case by case. The output that matters is a list, the specific cases that used to pass and now fail. That list points at a mechanism. Ten cases flipping, all of them involving a particular tool argument or a particular kind of edge input, tells you far more than a score that slipped from 0.91 to 0.89. The aggregate tells you something changed; the case-level diff tells you what.
Alongside pass and fail on the case set, a handful of qualitative signals tend to move before a visible incident, and each one is countable even though the output itself is not. Grounding rate, the share of claims an answer can actually support from its retrieved material, drops when the model starts filling gaps with invention. Refusal accuracy, whether the system refuses the right requests and answers the right ones, shifts when a provider retunes a safety layer.
Tool-call recovery, what the agent does after a call returns an error or an empty result, degrades in ways that never show up in a selection-accuracy number. Parse and format failure rate, the share of outputs that do not match the structure your downstream code expects, climbs quietly when a model starts formatting things slightly differently. And the human override or correction rate, how often a person edits or discards what the system produced, is often the earliest honest signal you have, because it measures what your users actually did with the output rather than what a metric said about it.
The hardest version of this drift to catch is a narrow-slice regression sitting underneath a stable or even improved aggregate. A provider's update can genuinely raise overall quality while making one specific thing worse. The blended number goes up, everyone reads that as good news, and the slice that regressed, a single refusal category, one tool-argument format, one type of retrieval-dependent answer, gets worse for the users who depend on it. If you only look at the total, the improvement actively conceals the damage.
This is why per-slice tracking is not a refinement, it is the point. You break the case set into the categories that matter for your system and you watch each one on its own line, because a regression that a total absorbs is a regression that reaches production. And it is why the same-model-name point carries so much weight here. The moment you accept that the same identifier can serve a different system, you stop trusting names and start trusting the case set. The name is a label the provider controls. The case set is evidence you control.
The teams that handle this well do the unglamorous work first. They build a fixed case set while the system is calm, they wire it into every change that could move behavior, and they track it slice by slice so a narrow regression cannot hide. The teams that handle it badly wait, and the bar for what counts as acceptable drift gets set for them, in public, by a production outage and the incident review that follows.
Working out which cases and which slices actually predict trouble for your own system is close to the exercise behind our agent benchmarking work, assembling the fixed case set and the per-slice breakdown while you still have the time to do it calmly, rather than reconstructing it under pressure after the drift you could not see has already reached someone.
A chart plots a number, and generative output is not a number. Code, explanations, and derivations can change in meaning while every summary statistic you track stays flat. The change lives in the content, not in a distribution, so a line on a graph never bends to show it.
Yes. A provider can update routing, safety layers, or the served weights behind the same model name, and your outputs move even though your prompts, code, and data did not. From your side it looks like the same model. From a behavior standpoint it is a new system, which is why a fixed case set run on every swap is the only reliable check.
Run a fixed, versioned set of evaluation cases on every model swap, prompt change, and retrieval change, and diff it case by case rather than as an average. Track qualitative signals that move early, grounding rate, refusal accuracy, tool-call recovery, parse failure rate, and the human override rate, and watch each narrow slice on its own so a regression cannot hide under a stable total.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark