Two numbers get quoted about agent economics and both are misleading. Cost per call ignores that a failed call still costs money and delivers nothing. Average latency ignores that nobody complains about the median.
Take the total spend across a benchmark run and divide by the number of cases that actually succeeded. Failures, retries and redundant tool calls are in the numerator; only successes are in the denominator.
This is the number that answers the question a business is really asking, which is what it costs to get one piece of work done. It also has a useful property: improving accuracy improves it. A system that resolves eighty per cent of tasks at a given cost per call is a quarter more expensive per resolved task than one resolving a hundred per cent, before anyone touches a price list.
Report p50 and p95, and make decisions on p95. The median describes an experience nobody complains about. The ninety fifth percentile describes the one they escalate, and for an agent the gap between them is unusually wide because a task needing three tool calls takes several times longer than one needing none.
Where a request fans out into parallel calls, measure the whole task rather than the individual calls. A user waits for the last one.
Total tokens per task is worth breaking into input and output, because they price differently and they respond to different fixes. Input dominated cost is usually a retrieval problem: too many passages, passages too long, a system prompt nobody has pruned since the prototype. Output dominated cost is usually a verbosity problem and is often cheap to fix by asking for less.
Retrieved context is the line that grows quietly. Every improvement that adds a passage adds it to every request forever.
Cost and accuracy are usually reported by different people at different times, which is how a system gets shipped that is accurate and unaffordable, or cheap and unusable.
In the same table you can see the trade directly: a slice at 0.94 costing four cents a task, and a slice at 0.96 costing thirty, because it is the one that triggers the expensive tool chain. That is a product decision, and it can only be made when both numbers sit in the same row.
The benchmark reports cost and latency alongside every accuracy slice for exactly that reason.
Divide total spend across a run by the number of tasks actually completed, so failures, retries and redundant tool calls are counted in the cost but not in the denominator. Cost per call flatters a system that fails often.
Make decisions on p95. The median describes an experience nobody complains about. For agents the gap between median and tail is unusually wide, because a task requiring several tool calls takes several times longer than one requiring none.
Usually retrieved context. Every improvement that adds a passage to the prompt adds it to every request from then on. Splitting tokens per task into input and output shows immediately which side is growing.
From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.
Book a benchmark