Safety measurement is usually delegated to a security review that happens once before launch and never again. It belongs in the benchmark that runs on every change, because the behaviours are as testable as accuracy and they regress the same way.
Policy pass rate. Requests that should be declined, declined. Build this slice from your own policy rather than from a generic list: what your agent must not do is specific to your business, and a public safety benchmark tests somebody else's rules.
Over refusal rate. Legitimate requests declined. Measured far less often and usually more damaging day to day, because each one is a user who cannot complete a task. A system tuned hard on the first number will look excellent and quietly fail here, and no aggregate will show it because the two pull in opposite directions.
Sensitive data in output. Whether personal or confidential information that entered the context comes back out in a response. This is the one that produces a notifiable incident rather than a bad review.
The common cause is not the model deciding to disclose something. It is that the retrieval layer returned a document the requesting user was not entitled to, and the model then did its job on material it should never have seen.
Which means the test has to run with realistic permissions. Benchmarking against a corpus where every case can see everything measures a system that does not exist. Give test cases identities, give those identities entitlements, and check both that the answer is right and that the material behind it was permitted.
Two of one hundred and twenty responses carrying another customer's order reference is the kind of finding that changes a launch date, and it is invisible to every accuracy measurement on the page.
The best source of safety cases is what has already gone wrong: support escalations, complaints, near misses, the things people worried about in design review. Each becomes a case, and each stays in the suite permanently as a regression test.
That is the difference between a safety review and a safety measurement. A review is an opinion at a point in time. A measurement runs on every change and tells you when the fix stopped working.
Most slices get a threshold below one, because perfection is not the standard and a few failures are tolerable. Safety slices usually do not. A single case of sensitive data reaching the wrong user is not a rate to be averaged; it is a demonstrated capability, and it should be a critical case that blocks release whatever the surrounding average says.
Our benchmark treats it that way, and the open harness enforces it: a failing critical case returns "not ready" even when its slice score clears the threshold.
Build the slice from your own policy and your own past incidents rather than a generic list, and measure three things: requests that should be declined and were, legitimate requests wrongly declined, and sensitive information reaching an output it should not.
Usually because the retrieval layer returned a document the requesting user was not entitled to, and the model then worked correctly on material it should never have seen. Testing with realistic per user permissions is what surfaces this.
No. Most slices tolerate a small failure rate. A single case of sensitive data reaching the wrong user is a demonstrated capability rather than a rate, and it should block release whatever the surrounding average says.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark