Once an LLM feature is live, most of the numbers on the dashboard are there to reassure you, not to warn you. Requests per minute, tokens processed, uptime, and a rolling average accuracy all tend to look healthy right up to the moment a customer escalation lands on someone's desk. They are activity metrics. They tell you the system is running and being used, which is worth knowing for capacity and billing, but none of them has to move when the quality of the output starts to slide. A model can serve every request, inside its latency budget, at a steady average score, and still be quietly getting a growing share of answers wrong.
The distinction that matters in production is between a metric that reports activity and a metric that leads an incident. A leading signal moves before the support ticket, before the churned account, before the postmortem. It moves because it is measuring the thing that actually degrades first, rather than a summary that averages the degradation away. Evaluation before release is a separate discipline with its own signals, covered in the evaluation metrics piece; this is about what to watch once real traffic is flowing and you no longer have clean ground truth for every request.
If there is a person in the loop, whether an agent approving a drafted reply, an analyst accepting a generated summary, or a reviewer signing off on a classification, that person is running a small evaluation on every single item. When they accept the output, they are voting that it was good enough. When they edit it before sending, or reject it and start over, they are telling you it was not. The override rate and the edit or correction rate are, together, the most sensitive leading signal you have, because they capture human judgement on live traffic in real time.
When these rise, quality has already moved, and it has moved before any accuracy dashboard built on delayed labels can show it. A support team that used to send AI-drafted replies untouched and now rewrites one in three is giving you a cleaner quality reading than any offline score, and giving it to you days earlier. Instrument the accept, edit, and reject actions in whatever tool the humans already use, capture what they changed where you can, and treat a rising correction rate as the first alarm rather than a productivity statistic.
Human review does not cover every request, and in a high-volume system it covers a small fraction. For the rest you need signals that a machine can compute on live output without waiting for a label.
Any output with structure you can check deterministically should be checked on the way out: does the JSON parse, does the value fall in the allowed set, does the cited order number exist, does the total add up. The failed validation rate is the share that does not pass those checks. It predicts the integration incident, the one where a downstream system chokes on a malformed field or acts on a value that was never valid. This rate climbing is often the first sign of a silent model or prompt change upstream, and it points straight at the broken step.
Related but worth its own line is the rate at which output cannot even be read into the shape the next step expects. Where a validation failure is a wrong value in the right shape, a format failure is the wrong shape entirely. It leads the incident where an agent stalls or a pipeline drops requests, and it tends to spike sharply rather than drift, which makes it a good candidate for a hard threshold alert.
Refusal behaviour shifts on its own in production as traffic changes and vendors update models underneath you. Watch whether the system is refusing more of the requests it should answer, or answering more of the ones it should refuse, and treat movement in either direction as drift worth investigating. Over-refusal drift leads the incident where users quietly give up on a feature that has started saying no; under-refusal drift leads the one that reaches legal or security. Grounding rate, the share of claims in an answer supported by the material the system actually retrieved, is the automated proxy that leads the fabrication incident, and it can be sampled continuously on live traffic even when you cannot label the answer as a whole.
Latency and cost read like operations metrics, but at the tail they lead incidents too. Report p95, not the average, because the median describes an experience nobody complains about while the tail is where timeouts and abandoned sessions live. A p95 creeping toward an upstream timeout predicts an incident that has nothing to do with correctness: the answer would have been right if it had arrived. Cost per resolved task, rather than cost per call, tells you where the system is burning retries and multi-step loops, which is usually the same place it is least stable.
A single healthy number today means very little, because the failures that hurt hide inside a stable aggregate. Overall correction rate can sit at a comfortable eight percent while the rate for one customer segment, one document type, or one language has doubled and is dragging a specific set of users toward the exit. Monitor each of these signals as a trend over time and broken out by the slices that matter for your traffic, and alert on the movement within a slice rather than the comfort of the whole. The point of monitoring is to catch the change while it is still small and localised, which a snapshot of the average will never do.
The expensive version of every lesson here is the one learned after the first outage, when nobody can say when the correction rate started climbing because it was never being recorded. None of these signals can be backfilled: you cannot reconstruct last month's override decisions or validation failures if you were not capturing them at the time. Wire the accept, edit, and reject events, the validation and parse checks, the refusal and grounding samples, and the p95 latency and cost breakdowns into the system before it ships, so the first question after any wobble has an answer already sitting in the data.
Deciding which of these signals lead incidents for your particular system, and where the alert thresholds belong, is the same groundwork behind our agent benchmarking work: mapping each failure mode to the signal that moves first, so the number that matters is not decided for you by a production outage.
In a human-in-the-loop system it is the human override and edit rate. When the people reviewing or using the output start rejecting or rewriting more of it, quality has already moved, usually days before any accuracy dashboard reflects it. Those people are running an evaluation every time they accept or change a result, so instrument that decision.
Requests, tokens, and uptime measure that the system is running, not that it is right. They stay flat, or even rise, while the answers underneath drift toward the wrong ones. They are useful for capacity planning and useless as an early warning, because none of them move when the model starts producing confident, well-formed, incorrect output.
Pre-release evaluation runs a fixed case set against a candidate before it ships and tells you whether it is good enough to release. Production monitoring watches live traffic you cannot fully label and tells you whether a system that was good enough is still behaving. The first uses ground truth you control; the second leans on operational proxies such as override rate, failed validations, and refusal drift because ground truth arrives late or never.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark