Applied ML

A correct answer that will not parse is a failed answer

When an agent output feeds another system, correctness is necessary and not sufficient. A response containing exactly the right information in a shape the parser rejects fails in production exactly as hard as a wrong one, and it fails in a way that is harder to notice because the content looks fine in a log.

Schema validity

The share of responses that parse and satisfy the contract. Not just valid JSON: the required fields present, the types right, the enumerations inside their allowed set, the nesting as declared.

This number is usually very high and that is exactly why it needs measuring. At 0.993 you have roughly one failure in one hundred and forty. At ten thousand calls a day that is seventy incidents, each of which is a user seeing an error or a record silently not being written.

A high validity rate is a large daily failure count at volume A comparison showing that schema validity of 99.3 per cent looks reassuring, yet at ten thousand calls a day it produces about seventy failures every day, each a user seeing an error or a record never written. The number that hides the incidents Schema validity 99.3% one failure in about 140 looks reassuring on a dashboard at 10,000 calls a day Failures every day 70 each: a user sees an error or a record is never written and content looks fine in the log
Judge validity against your volume. At scale, the second decimal place is a daily incident count.

Required field completeness

Separate from validity, because a response can satisfy a permissive schema while omitting the field the workflow depends on. If the downstream process needs an account identifier to act, then a response without one is a failure whether or not the schema allowed it to be optional.

Measure the fields your process actually consumes, not the fields the schema happens to declare.

Refusal precision, in both directions

Refusals are a formatting problem as much as a safety one, because they are where structured output most often breaks. The system decides not to answer and returns prose where the contract expected an object.

Measure both directions. Under refusal is answering something that should have been declined. Over refusal is declining something legitimate, which teams almost never measure and which does more day to day damage: every over refusal is a user who cannot complete a task and contacts support instead.

Refusal precision measured in both directions A two by two grid of what the system should do against what it did. The diagonal cells are correct. Answering when it should refuse is under refusal, and refusing when it should answer is over refusal, the costly case teams rarely measure. Refusal precision, both directions what the system did Answered Refused what it should do Should answer Should refuse Correct answered a fair request Over refusal legitimate task blocked, user contacts support Under refusal answered what it should decline Correct refused a bad request the one teams skip
A system tuned hard for safety shows a beautiful policy pass rate and an over refusal rate nobody looked at.

A system tuned hard for safety will show a beautiful policy pass rate and an over refusal rate nobody looked at.

Cheap to test, cheap to fix

Format is the least glamorous family on a benchmark and it has the best return. The graders are deterministic, so they cost nothing to run on every change, and the fixes are usually structural rather than a model change: constrained decoding, a schema in the request, a validation and repair step, or a retry when parsing fails.

Because it is cheap, it belongs in continuous integration rather than in a quarterly review. That is what a handover harness is for: the suite runs on every change and the build fails when schema validity drops. r4agent ships a JSON grader for this, and the benchmark reports it per slice.

Common questions

How do you measure output format reliability in an LLM?

Score the share of responses that parse and satisfy the contract, including required fields, types and enumerations. Measure completeness of the fields your workflow actually consumes separately, since a permissive schema can accept a response that is useless downstream.

Is 99 percent schema validity good enough?

Rarely. At ten thousand calls a day, 99 percent is a hundred failures a day, each one a user seeing an error or a record not being written. Judge it against your volume and the consequence of a single malformed response.

What is over refusal and why does it matter?

Declining a legitimate request. It is measured far less often than under refusal and usually does more day to day damage, because every over refusal is a user who cannot finish a task and contacts support instead.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark