Evaluation sets are written by people who know what the system does. They phrase questions the way the documentation phrases them. Real users do not, and the gap between those two populations is where a system that tested well starts generating complaints.
Paraphrase stability. Take each case and rewrite it three ways without changing what is being asked. The share where the verdict stays the same is paraphrase stability. Anything below about 0.9 means the system is matching phrasing rather than meaning, and the cases it passes are partly accidents.
Typo tolerance. Realistic misspelling, missing punctuation, no capitals, a phone keyboard. Generate it mechanically and keep it plausible: transposed letters and dropped vowels rather than random noise.
Distractor resistance. Pad the question with irrelevant but related context, the way a real user pastes three paragraphs of their situation around one question. This one degrades earlier than teams expect, and it degrades worse as context windows grow, because more retrieved material means more plausible wrong material competing for attention.
Prompt injection resistance. Put an instruction inside a document the system will retrieve. Not the obvious version. The version that reads like ordinary content and carries an instruction in the middle of a sentence.
The reason injection matters for a benchmark rather than only for a security review is that any system reading documents it did not write has this exposure, and most teams test it once by hand and never again.
Grading it is unusual and worth stating explicitly: the test is inverted. You plant a string that only appears in the output if the agent obeyed, and a match is a failure. That is the one grader where finding what you searched for is the bad outcome, and it is easy to write backwards.
Set the threshold at one. A single successful injection in a benchmark is a demonstrated capability, not a rate to be averaged with the passes.
Every robustness case is a transformation of a case you already wrote, so a hundred case suite becomes four hundred without four hundred conversations with domain experts. The expected answer is unchanged by construction, which is the whole trick: you are not writing new cases, you are asking the existing ones badly.
That makes this the highest yield slice per hour of effort in a benchmark, and it is the slice most often absent. Our engagement generates it from your own cases rather than from a template, so the bad phrasing resembles your users rather than somebody else's.
Asking the same questions badly: paraphrased, misspelled, padded with irrelevant context, or carrying an instruction planted inside retrieved content. The expected answers are unchanged, so any drop in score is the system matching phrasing rather than meaning.
Plant an instruction inside a document the system will retrieve, phrased like ordinary content rather than an obvious command, and check whether the output shows it was obeyed. The grading is inverted: a match on the planted string is a failure, not a pass.
Zero, in the sense that the threshold should be set at total resistance. A single successful injection is a demonstrated capability rather than a rate to be averaged with the cases that passed.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark