The rest of this category is software you install, instrument and run yourself, and it leaves you to write the test cases, which is the hard part. This is the assessment instead. We build the suite with your domain experts, run it against your agent, and hand back a production readiness decision with the evidence attached.
From $6,000. Typically three to four weeks. You keep the suite and the harness.
Most agents look good in the room where they were built. They were tried on the cases the team had in mind, and those are the cases they handle. The questions that decide whether the thing can go live are the ones nobody typed in: the malformed input, the customer who asks two things at once, the document that is not in the index, the tool that times out.
The usual answer is a single accuracy figure quoted in a steering meeting. An agent is not a classifier, so that figure is a summary of several different systems averaged together. It goes up when retrieval improves and down when the model gets more cautious, and it moves for both reasons at once without saying which.
What a team needs before production is narrower and more useful: which part fails, how often, on which kind of input, and what it costs when it works.
Galileo, Patronus, Braintrust, LangSmith, Arize and Langfuse are good at what they do. They are also all the same shape: software your team installs, instruments and operates.
A platform gives you a metric library and a place to put the results. It does not decide what correct means for your domain, and it does not sit in a room with your clinical lead or your head of support arguing about whether an answer was acceptable. That argument is the expensive part, it is the part that surfaces the ambiguities in your product rather than in your model, and no amount of tooling removes it.
So the two answer different questions. A platform tells you what happened. An assessment tells you whether to launch. If you already run one, keep it: we can hand the suite over in a form it can execute, and the value of your instrumentation goes up once somebody has defined what the numbers should say.
| Evaluation platform | This | |
|---|---|---|
| What arrives | A dashboard and a metric library | A decision, with the failing cases attached |
| Who writes the cases | Your team | Us, with your domain experts |
| Integration needed | An SDK in your codebase, traces exported | None. We drive it the way a user does |
| Where your data sits | Usually their cloud | Your environment, if that is the requirement |
| What it costs | A subscription, quoted on a call | A fixed fee, published above |
| When it ends | When you stop paying | You keep the suite and can run it without us |
| The question it answers | What happened | Whether this is ready |
Ten families of measurement. Not all of them apply to every system, and the ones that do not are dropped rather than padded out with a number nobody will read.
Whether the job was finished end to end, judged against an answer somebody is prepared to defend, rather than against output that reads plausibly.
How much of the answer is supported by the sources the system actually retrieved, and what it does when the answer is not there at all. Saying so is the correct behaviour and it is measured as such.
Whether the evidence needed was in front of the model in the first place. Retrieval sets a ceiling that no amount of prompting gets you past, so it is measured separately from the answer.
Whether the right tool was chosen, called with valid arguments, and whether the agent recovered when a call failed or returned nothing. Calls that should never have happened are counted too, because they cost money.
Whether the output is the shape the system downstream expects, every time. A response that is correct but unparseable fails in production exactly as hard as one that is wrong.
The same questions asked badly. Rephrased, misspelled, padded with irrelevant context, or carrying an instruction hidden inside a retrieved document that the agent was never meant to obey.
The same input run many times. A system that answers differently on the third attempt is a support problem waiting to happen, and an average taken over one run each will never show it.
Behaviour on the requests you do not want answered, and whether personal or confidential data that entered the context comes back out in the response.
What one completed task costs and how long a user waits, measured at the tail rather than the average, because the slow requests are the ones people complain about.
The list of cases that used to pass and now fail. Model changes are rarely uniform improvements, and the aggregate can rise while something you care about quietly breaks.
Thresholds are agreed with you before the run, not chosen afterwards to suit the result. A benchmark that cannot fail is a marketing exercise.
Every slice clears the threshold you set, the failures that remain are understood, and the cost per task is one the business can carry at expected volume.
The most common outcome. Specific slices fall short, the cause is identified, and the work to close each one is listed with the measurement that will confirm it worked.
Something structural is wrong: the evidence is not retrievable, the task is underspecified, or the failure rate on a slice that matters is too high to supervise. We say so.
Scope it. Which decisions the agent makes, which of them matter most, and what the thresholds are. This is where we find out whether the task has been specified well enough to be measured at all.
Build the suite. Cases assembled with the people who know the domain, each with an expected answer and a note on why that answer is right. The disagreements found here are usually the most valuable output of the whole exercise, because they are ambiguities in the product, not in the model.
Run it. The full suite against your system, repeated where consistency is being measured, with every input, output and judgement recorded.
Report and hand over. Results per slice, the failing cases in full, the readiness decision, and the harness itself so your team can run it again on every change.
The method is not new to us. R4SUB is five packages on CRAN that score how ready a clinical submission is: an evidence contract, an FMEA based risk engine computing risk priority numbers, a traceability engine, and profiles carrying each regulator's own thresholds for the FDA, EMA, PMDA, Health Canada, TGA and MHRA.
That is the same instrument as this one. Agree the thresholds with the people who own the decision, score the evidence per category rather than as one average, and end with a readiness answer somebody can act on. One is pointed at a submission and one at an agent.
The harness is open. r4agent is the runner, the graders, the per slice scorer and the regression diff, MIT licensed, so you can read exactly how a number here is produced before you pay anyone to produce one. Grading in it is deterministic and offline: no model judges another model, so a score is reproducible from the suite and the runner alone.
Where the output is free text, exact matching does not work and human review does not scale, so a model does the grading. That is workable and it is what we do, with one condition attached.
A judge model has its own biases. It rewards length. It prefers output that resembles its own style, which means it will rate a system built on the same model family more generously than it deserves. It is unreliable near the boundary between adjacent scores.
So we calibrate it. People grade a sample of the same cases, we measure how often the judge agrees, and that agreement figure is quoted next to every score it produced. A judge nobody has checked is a number with no error bar.
Read how we build evaluationsNot a public leaderboard score. How a model ranks on a published benchmark says very little about how it will behave on your documents and your users. Several of those test sets have also leaked into training data, which makes a high score partly a memory test.
Not a certification. There is no badge and no pass mark we invented. The thresholds are yours, and the report shows the measurements so you can disagree with the conclusion.
Not a replacement for monitoring. A benchmark tells you what is true at one moment against a fixed set of cases. Production drifts, inputs change, and the suite has to be refreshed on a schedule or it slowly stops describing reality.
Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark