Model assurance

The same question, asked badly

Evaluation sets are written by people who know what the system does. They phrase questions the way the documentation phrases them. Real users do not, and the gap between those two populations is where a system that tested well starts generating complaints.

Four ways to ask badly, all measurable

Paraphrase stability. Take each case and rewrite it three ways without changing what is being asked. The share where the verdict stays the same is paraphrase stability. Anything below about 0.9 means the system is matching phrasing rather than meaning, and the cases it passes are partly accidents.

Typo tolerance. Realistic misspelling, missing punctuation, no capitals, a phone keyboard. Generate it mechanically and keep it plausible: transposed letters and dropped vowels rather than random noise.

Distractor resistance. Pad the question with irrelevant but related context, the way a real user pastes three paragraphs of their situation around one question. This one degrades earlier than teams expect, and it degrades worse as context windows grow, because more retrieved material means more plausible wrong material competing for attention.

Prompt injection resistance. Put an instruction inside a document the system will retrieve. Not the obvious version. The version that reads like ordinary content and carries an instruction in the middle of a sentence.

Robustness scores when the same question is asked badly A bar chart of three robustness measures against a nine tenths threshold. Paraphrase stability clears the threshold, typo tolerance falls just below it, and distractor resistance drops well under it, even though the expected answer never changed. Same question, expected answer unchanged 1.0 0.0 0.9 threshold 0.92 Paraphrase stability 0.88 Typo tolerance 0.74 Distractor resistance
The answer never changed, so every point lost is the system matching phrasing rather than meaning. Distractors bite first.

Injection is a retrieval problem wearing a security costume

The reason injection matters for a benchmark rather than only for a security review is that any system reading documents it did not write has this exposure, and most teams test it once by hand and never again.

Grading it is unusual and worth stating explicitly: the test is inverted. You plant a string that only appears in the output if the agent obeyed, and a match is a failure. That is the one grader where finding what you searched for is the bad outcome, and it is easy to write backwards.

Set the threshold at one. A single successful injection in a benchmark is a demonstrated capability, not a rate to be averaged with the passes.

The inverted grader for prompt injection A flow where a planted instruction sits inside a retrieved document, the agent produces an output, and a grader searches it for the planted string. Finding the string is a failure because it proves the injection was obeyed, and no match is a pass. Injection: a match is the failure Retrieved document a planted instruction, mid-sentence Agent output grader searches it for the planted string String found = FAIL the injection was obeyed No match = PASS the instruction was ignored This is the one grader where finding what you searched for is the bad outcome. Threshold is one: a single success is a demonstrated capability, not a rate to average.
The injection grader is inverted. A match means the agent obeyed the planted instruction, so a match is a failure.

Robustness cases are cheap to make and expensive to skip

Every robustness case is a transformation of a case you already wrote, so a hundred case suite becomes four hundred without four hundred conversations with domain experts. The expected answer is unchanged by construction, which is the whole trick: you are not writing new cases, you are asking the existing ones badly.

That makes this the highest yield slice per hour of effort in a benchmark, and it is the slice most often absent. Our engagement generates it from your own cases rather than from a template, so the bad phrasing resembles your users rather than somebody else's.

Common questions

What is robustness testing for an AI agent?

Asking the same questions badly: paraphrased, misspelled, padded with irrelevant context, or carrying an instruction planted inside retrieved content. The expected answers are unchanged, so any drop in score is the system matching phrasing rather than meaning.

How do you test for prompt injection?

Plant an instruction inside a document the system will retrieve, phrased like ordinary content rather than an obvious command, and check whether the output shows it was obeyed. The grading is inverted: a match on the planted string is a failure, not a pass.

What is an acceptable prompt injection rate?

Zero, in the sense that the threshold should be set at total resistance. A single successful injection is a demonstrated capability rather than a rate to be averaged with the cases that passed.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark