Model assurance

Measuring what your agent should refuse, and what it leaks

Safety measurement is usually delegated to a security review that happens once before launch and never again. It belongs in the benchmark that runs on every change, because the behaviours are as testable as accuracy and they regress the same way.

Three numbers

Policy pass rate. Requests that should be declined, declined. Build this slice from your own policy rather than from a generic list: what your agent must not do is specific to your business, and a public safety benchmark tests somebody else's rules.

Over refusal rate. Legitimate requests declined. Measured far less often and usually more damaging day to day, because each one is a user who cannot complete a task. A system tuned hard on the first number will look excellent and quietly fail here, and no aggregate will show it because the two pull in opposite directions.

Sensitive data in output. Whether personal or confidential information that entered the context comes back out in a response. This is the one that produces a notifiable incident rather than a bad review.

The two ways an agent gets the boundary wrong A matrix of what the request was against what the system did. Declining a request that should be declined is a policy pass, answering it is a violation, declining a legitimate request is over refusal, and answering it is correct. Two ways to be wrong at the boundary System declined System answered Should be declined Legitimate request Policy pass correct refusal Violation did what it must not Over refusal blocked a real user Correct task completed Policy pass rate and over refusal rate pull in opposite directions, so no single average shows both.
Tuning hard on refusals looks excellent while quietly filling the over refusal cell.

Leakage is a retrieval problem more often than a model problem

The common cause is not the model deciding to disclose something. It is that the retrieval layer returned a document the requesting user was not entitled to, and the model then did its job on material it should never have seen.

Which means the test has to run with realistic permissions. Benchmarking against a corpus where every case can see everything measures a system that does not exist. Give test cases identities, give those identities entitlements, and check both that the answer is right and that the material behind it was permitted.

Two of one hundred and twenty responses carrying another customer's order reference is the kind of finding that changes a launch date, and it is invisible to every accuracy measurement on the page.

Build the cases from your own incidents

The best source of safety cases is what has already gone wrong: support escalations, complaints, near misses, the things people worried about in design review. Each becomes a case, and each stays in the suite permanently as a regression test.

That is the difference between a safety review and a safety measurement. A review is an opinion at a point in time. A measurement runs on every change and tells you when the fix stopped working.

Set these thresholds differently

Most slices get a threshold below one, because perfection is not the standard and a few failures are tolerable. Safety slices usually do not. A single case of sensitive data reaching the wrong user is not a rate to be averaged; it is a demonstrated capability, and it should be a critical case that blocks release whatever the surrounding average says.

Our benchmark treats it that way, and the open harness enforces it: a failing critical case returns "not ready" even when its slice score clears the threshold.

A single critical failure blocks release A safety slice score of 0.95 clears the 0.80 bar, but one failing critical case, sensitive data reaching the wrong user, overrides the average and returns a verdict of not ready. One critical failure overrides the average Slice score 0.95 clears the 0.80 bar 1 critical case fails data reached the wrong user Release rule critical case gate NOT READY An averaged safety number, read on its own, would have called this ready to ship.
A demonstrated leak is a capability, not a rate. It blocks release whatever the slice average says.

Common questions

How do you test an AI agent for safety?

Build the slice from your own policy and your own past incidents rather than a generic list, and measure three things: requests that should be declined and were, legitimate requests wrongly declined, and sensitive information reaching an output it should not.

Why does my agent leak data it should not have access to?

Usually because the retrieval layer returned a document the requesting user was not entitled to, and the model then worked correctly on material it should never have seen. Testing with realistic per user permissions is what surfaces this.

Should safety thresholds be the same as accuracy thresholds?

No. Most slices tolerate a small failure rate. A single case of sensitive data reaching the wrong user is a demonstrated capability rather than a rate, and it should block release whatever the surrounding average says.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark