Product

AI agent benchmarking

The rest of this category is software you install, instrument and run yourself, and it leaves you to write the test cases, which is the hard part. This is the assessment instead. We build the suite with your domain experts, run it against your agent, and hand back a production readiness decision with the evidence attached.

From $6,000. Typically three to four weeks. You keep the suite and the harness.

A demo is not evidence

Most agents look good in the room where they were built. They were tried on the cases the team had in mind, and those are the cases they handle. The questions that decide whether the thing can go live are the ones nobody typed in: the malformed input, the customer who asks two things at once, the document that is not in the index, the tool that times out.

The usual answer is a single accuracy figure quoted in a steering meeting. An agent is not a classifier, so that figure is a summary of several different systems averaged together. It goes up when retrieval improves and down when the model gets more cautious, and it moves for both reasons at once without saying which.

What a team needs before production is narrower and more useful: which part fails, how often, on which kind of input, and what it costs when it works.

This is not an evaluation platform, and you probably want both

Galileo, Patronus, Braintrust, LangSmith, Arize and Langfuse are good at what they do. They are also all the same shape: software your team installs, instruments and operates.

A platform gives you a metric library and a place to put the results. It does not decide what correct means for your domain, and it does not sit in a room with your clinical lead or your head of support arguing about whether an answer was acceptable. That argument is the expensive part, it is the part that surfaces the ambiguities in your product rather than in your model, and no amount of tooling removes it.

So the two answer different questions. A platform tells you what happened. An assessment tells you whether to launch. If you already run one, keep it: we can hand the suite over in a form it can execute, and the value of your instrumentation goes up once somebody has defined what the numbers should say.

 Evaluation platformThis
What arrivesA dashboard and a metric libraryA decision, with the failing cases attached
Who writes the casesYour teamUs, with your domain experts
Integration neededAn SDK in your codebase, traces exportedNone. We drive it the way a user does
Where your data sitsUsually their cloudYour environment, if that is the requirement
What it costsA subscription, quoted on a callA fixed fee, published above
When it endsWhen you stop payingYou keep the suite and can run it without us
The question it answersWhat happenedWhether this is ready

What we measure

Ten families of measurement. Not all of them apply to every system, and the ones that do not are dropped rather than padded out with a number nobody will read.

  1. 01

    Task success

    Whether the job was finished end to end, judged against an answer somebody is prepared to defend, rather than against output that reads plausibly.

    Resolution ratePer slice pass rateEscalation rate
  2. 02

    Grounding and hallucination

    How much of the answer is supported by the sources the system actually retrieved, and what it does when the answer is not there at all. Saying so is the correct behaviour and it is measured as such.

    Unsupported claim rateCitation validityAbstention on unanswerable
  3. 03

    Retrieval quality

    Whether the evidence needed was in front of the model in the first place. Retrieval sets a ceiling that no amount of prompting gets you past, so it is measured separately from the answer.

    Recall at kMean reciprocal rankAnswerable at all rate
  4. 04

    Tool use

    Whether the right tool was chosen, called with valid arguments, and whether the agent recovered when a call failed or returned nothing. Calls that should never have happened are counted too, because they cost money.

    Tool selection accuracyArgument validityError recovery rateRedundant calls
  5. 05

    Instruction and format adherence

    Whether the output is the shape the system downstream expects, every time. A response that is correct but unparseable fails in production exactly as hard as one that is wrong.

    Schema validityRequired field completenessRefusal precision
  6. 06

    Robustness

    The same questions asked badly. Rephrased, misspelled, padded with irrelevant context, or carrying an instruction hidden inside a retrieved document that the agent was never meant to obey.

    Paraphrase stabilityTypo toleranceDistractor resistancePrompt injection resistance
  7. 07

    Consistency

    The same input run many times. A system that answers differently on the third attempt is a support problem waiting to happen, and an average taken over one run each will never show it.

    Run to run disagreementSelf consistency across n samples
  8. 08

    Safety and data handling

    Behaviour on the requests you do not want answered, and whether personal or confidential data that entered the context comes back out in the response.

    Policy pass rateSensitive data leakageOver refusal rate
  9. 09

    Cost and latency

    What one completed task costs and how long a user waits, measured at the tail rather than the average, because the slow requests are the ones people complain about.

    p50 and p95 latencyTokens per taskCost per resolved task
  10. 10

    Regression against your last version

    The list of cases that used to pass and now fail. Model changes are rarely uniform improvements, and the aggregate can rise while something you care about quietly breaks.

    Newly failing casesNewly passing casesNet movement per slice

The output is a decision

Thresholds are agreed with you before the run, not chosen afterwards to suit the result. A benchmark that cannot fail is a marketing exercise.

How a benchmark runs

Scope it. Which decisions the agent makes, which of them matter most, and what the thresholds are. This is where we find out whether the task has been specified well enough to be measured at all.

Build the suite. Cases assembled with the people who know the domain, each with an expected answer and a note on why that answer is right. The disagreements found here are usually the most valuable output of the whole exercise, because they are ambiguities in the product, not in the model.

Run it. The full suite against your system, repeated where consistency is being measured, with every input, output and judgement recorded.

Report and hand over. Results per slice, the failing cases in full, the readiness decision, and the harness itself so your team can run it again on every change.

We have built this instrument before

The method is not new to us. R4SUB is five packages on CRAN that score how ready a clinical submission is: an evidence contract, an FMEA based risk engine computing risk priority numbers, a traceability engine, and profiles carrying each regulator's own thresholds for the FDA, EMA, PMDA, Health Canada, TGA and MHRA.

That is the same instrument as this one. Agree the thresholds with the people who own the decision, score the evidence per category rather than as one average, and end with a readiness answer somebody can act on. One is pointed at a submission and one at an agent.

The harness is open. r4agent is the runner, the graders, the per slice scorer and the regression diff, MIT licensed, so you can read exactly how a number here is produced before you pay anyone to produce one. Grading in it is deterministic and offline: no model judges another model, so a score is reproducible from the suite and the runner alone.

The judge needs checking too

Where the output is free text, exact matching does not work and human review does not scale, so a model does the grading. That is workable and it is what we do, with one condition attached.

A judge model has its own biases. It rewards length. It prefers output that resembles its own style, which means it will rate a system built on the same model family more generously than it deserves. It is unreliable near the boundary between adjacent scores.

So we calibrate it. People grade a sample of the same cases, we measure how often the judge agrees, and that agreement figure is quoted next to every score it produced. A judge nobody has checked is a number with no error bar.

Read how we build evaluations

What it runs against

Systems
Agents, retrieval augmented systems, classifiers and plain model calls. If it answers over an API or we can call it in your environment, it can be benchmarked.
Models
Vendor neutral. Hosted APIs and self hosted open weight models are treated the same way, and comparing two candidates on your own cases is a common reason to run this.
What we need from you
Access to the system, a sample of real inputs, and time from the people who know what a correct answer looks like. The last one is the constraint that matters.
What you keep
The case suite, the run harness, the raw results and the report. All of it in version control, runnable in your own CI, with no dependency on us to get the number again.
Data handling
Where your data cannot leave your environment, the harness runs inside it and we work from the results. Agreed in writing before anything is copied.

What this is not

Not a public leaderboard score. How a model ranks on a published benchmark says very little about how it will behave on your documents and your users. Several of those test sets have also leaked into training data, which makes a high score partly a memory test.

Not a certification. There is no badge and no pass mark we invented. The thresholds are yours, and the report shows the measurements so you can disagree with the conclusion.

Not a replacement for monitoring. A benchmark tells you what is true at one moment against a fixed set of cases. Production drifts, inputs change, and the suite has to be refreshed on a schedule or it slowly stops describing reality.

Find out what your agent actually does

Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark