When an agent output feeds another system, correctness is necessary and not sufficient. A response containing exactly the right information in a shape the parser rejects fails in production exactly as hard as a wrong one, and it fails in a way that is harder to notice because the content looks fine in a log.
The share of responses that parse and satisfy the contract. Not just valid JSON: the required fields present, the types right, the enumerations inside their allowed set, the nesting as declared.
This number is usually very high and that is exactly why it needs measuring. At 0.993 you have roughly one failure in one hundred and forty. At ten thousand calls a day that is seventy incidents, each of which is a user seeing an error or a record silently not being written.
Separate from validity, because a response can satisfy a permissive schema while omitting the field the workflow depends on. If the downstream process needs an account identifier to act, then a response without one is a failure whether or not the schema allowed it to be optional.
Measure the fields your process actually consumes, not the fields the schema happens to declare.
Refusals are a formatting problem as much as a safety one, because they are where structured output most often breaks. The system decides not to answer and returns prose where the contract expected an object.
Measure both directions. Under refusal is answering something that should have been declined. Over refusal is declining something legitimate, which teams almost never measure and which does more day to day damage: every over refusal is a user who cannot complete a task and contacts support instead.
A system tuned hard for safety will show a beautiful policy pass rate and an over refusal rate nobody looked at.
Format is the least glamorous family on a benchmark and it has the best return. The graders are deterministic, so they cost nothing to run on every change, and the fixes are usually structural rather than a model change: constrained decoding, a schema in the request, a validation and repair step, or a retry when parsing fails.
Because it is cheap, it belongs in continuous integration rather than in a quarterly review. That is what a handover harness is for: the suite runs on every change and the build fails when schema validity drops. r4agent ships a JSON grader for this, and the benchmark reports it per slice.
Score the share of responses that parse and satisfy the contract, including required fields, types and enumerations. Measure completeness of the fields your workflow actually consumes separately, since a permissive schema can accept a response that is useless downstream.
Rarely. At ten thousand calls a day, 99 percent is a hundred failures a day, each one a user seeing an error or a record not being written. Judge it against your volume and the consequence of a single malformed response.
Declining a legitimate request. It is measured far less often than under refusal and usually does more day to day damage, because every over refusal is a user who cannot finish a task and contacts support instead.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark