Worked example

Sample benchmark report

What arrives at the end of a benchmark, shown in full. The subject is a tier one customer support agent over a product documentation corpus.

Every number on this page is invented. This is an illustrative example built to show the structure, the level of detail and the kind of conclusion a report reaches. It is not a client result, no real engagement is described, and the agent it describes does not exist. Where a real report would name a system and a date, this one names a plausible one.

1. The decision

Thresholds were agreed before the run. Two slices did not clear theirs, both for the same underlying reason.

98 / 120cases passed overall
4 of 6slices cleared their threshold
$0.021per resolved task
6.8 sp95 latency

2. Task success, by slice

The overall figure of 0.817 is the least useful number in this report. It is the average of one slice at 0.938 and another at 0.417.

SliceCasesPassedRateThresholdResult
Billing questions32300.9380.90clear
Product how-to28260.9290.90clear
Account access24220.9170.90clear
Adversarial and injection860.7501.00short
Two requests in one message1690.5630.80short
Not answerable from the corpus1250.4170.95short
Total120980.817— 

On the unanswerable slice, a pass means the agent said it did not know or escalated. Of the 7 failures, 7 produced a confident answer with no basis in any retrieved passage. The agent has no path for admitting it cannot help.

3. Grounding, retrieval and tool use

Where the failures come from. Reported separately because they have different owners and different fixes.

MeasureValueReading
Unsupported claim rate6.7%Share of answers containing at least one claim no retrieved passage supports
Citation validity91.2%Citations that point at a passage actually supporting the claim
Abstention on unanswerable41.7%The headline weakness. Should be near total
Retrieval recall at 50.84The supporting passage reached the prompt in 84% of answerable cases
Answerable at all0.937% of the suite has no answer anywhere in the corpus. A ceiling no prompt clears
Tool selection accuracy0.95Right tool chosen for the job
Argument validity0.98Calls that were well formed
Error recovery0.42Second weakness. A failed call usually ends the turn
Redundant calls per task1.7Calls that changed nothing. Directly billable waste
Schema validity0.993Output parsed by the downstream system

4. Robustness, consistency and safety

The same questions asked badly, asked repeatedly, and asked in bad faith.

MeasureValueReading
Paraphrase stability0.88Same question, different wording, same verdict
Typo tolerance0.91Holds up under realistic misspelling
Distractor resistance0.79Degrades when irrelevant context is padded in
Prompt injection resistance6 of 8Two documents carrying a planted instruction were obeyed
Run to run disagreement8.3%Over 5 runs of every case. One case in twelve is not reproducible
Policy pass rate0.97Requests that should be refused, refused
Over refusal1 caseA legitimate request declined
Sensitive data in output2 of 120Another customer's order reference reached the response

The two injection failures and the two data leakage cases are the findings a security reviewer will read first, and they are unrelated to accuracy. A system can score well on every other line here and still be the one that leaks.

5. Two failing cases, in full

Every failure comes back with the input, the output and what was wrong with it. Twenty two are in the full report. Two are shown here.

Case 74, unanswerable slice

Asked: "Does the enterprise plan include SAML on the legacy portal?"

Answered: "Yes, SAML single sign on is available on all enterprise plans including the legacy portal."

Wrong because: the corpus says nothing about the legacy portal and SAML together. Retrieval returned two passages about SAML on the current product, and the answer generalised. The correct behaviour was to say the documentation does not cover it.

Case 31, two requests in one message

Asked: "Cancel my second seat and tell me when the next invoice lands."

Answered: a correct and complete answer about the invoice date. No mention of the seat.

Wrong because: the second request was dropped silently. The user has no way to know it was not actioned, which is worse than a refusal. Eleven of the sixteen cases in this slice fail this way.

6. What to fix, and how you will know

Add an abstention path. The agent has no way to say it does not know. Gate the answer on a retrieval score and add an explicit branch. Confirmed when abstention on the unanswerable slice clears 0.95 with no drop on the three passing slices.

Decompose multi-intent messages. Split the request before answering and respond to each part. Confirmed when the two-request slice clears 0.80 and no case drops a request silently.

Retry and fall back on tool failure. A failed call currently ends the turn. Confirmed when error recovery clears 0.80.

Strip identifiers on the way out. Two responses carried another customer's order reference. Confirmed when the leakage count is zero and the over-refusal count has not risen.

None of these requires a different model, and the retrieval ceiling of 0.93 is the item to plan for next: seven per cent of real questions have no answer in the corpus at all, which is a documentation problem rather than an AI one.

How the grading was checked

Free text answers were graded by a model. Two reviewers independently graded a 40 case sample of the same set, and the judge agreed with them on 89% of it.

The disagreements were not random: they concentrate on the boundary between partially correct and correct, and the judge is the more generous of the two. Every score in this report should be read with that 89% attached, and the four slices nearest their threshold are the ones where it matters.

What was handed over

The suite
120 cases with expected answers and the note on why each answer is right
The harness
Runs the suite against the system and writes a dated result file
The diff tool
Compares two result files and prints what changed in each direction
Raw results
Every input, output and judgement from all 5 runs

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark