Applied ML

Cost per resolved task is the only unit that means anything

Two numbers get quoted about agent economics and both are misleading. Cost per call ignores that a failed call still costs money and delivers nothing. Average latency ignores that nobody complains about the median.

Cost per resolved task

Take the total spend across a benchmark run and divide by the number of cases that actually succeeded. Failures, retries and redundant tool calls are in the numerator; only successes are in the denominator.

This is the number that answers the question a business is really asking, which is what it costs to get one piece of work done. It also has a useful property: improving accuracy improves it. A system that resolves eighty per cent of tasks at a given cost per call is a quarter more expensive per resolved task than one resolving a hundred per cent, before anyone touches a price list.

Cost per call versus cost per resolved task A hundred attempts cost the same whether or not they succeed, but only eighty resolve. Dividing spend by successes rather than by calls raises the true unit cost by a quarter. The denominator is the whole argument 100 attempts, every one billed 80 resolve, 20 fail or retry resolved failed or retried, still billed Cost per call $0.04 Cost per resolved task $0.05 a quarter more, same price list
Failures and retries stay in the spend but leave the denominator, so cost per resolved task is the honest unit.

Measure latency at the tail

Report p50 and p95, and make decisions on p95. The median describes an experience nobody complains about. The ninety fifth percentile describes the one they escalate, and for an agent the gap between them is unusually wide because a task needing three tool calls takes several times longer than one needing none.

Where a request fans out into parallel calls, measure the whole task rather than the individual calls. A user waits for the last one.

Task latency distribution with the median and the ninety fifth percentile marked A right skewed distribution of task times. The median sits low where nobody complains, and the long tail holds the slow multi tool tasks that a user escalates, marked at the ninety fifth percentile. The median is not the experience they escalate task completion time p50 nobody complains p95 the task they escalate multi tool tasks
For an agent the tail runs far to the right, so decisions belong on p95, the slow multi tool task, not the comfortable median.

Tokens per task, split

Total tokens per task is worth breaking into input and output, because they price differently and they respond to different fixes. Input dominated cost is usually a retrieval problem: too many passages, passages too long, a system prompt nobody has pruned since the prototype. Output dominated cost is usually a verbosity problem and is often cheap to fix by asking for less.

Retrieved context is the line that grows quietly. Every improvement that adds a passage adds it to every request forever.

Put it next to the accuracy table

Cost and accuracy are usually reported by different people at different times, which is how a system gets shipped that is accurate and unaffordable, or cheap and unusable.

In the same table you can see the trade directly: a slice at 0.94 costing four cents a task, and a slice at 0.96 costing thirty, because it is the one that triggers the expensive tool chain. That is a product decision, and it can only be made when both numbers sit in the same row.

The benchmark reports cost and latency alongside every accuracy slice for exactly that reason.

Common questions

How do you calculate the cost of an AI agent?

Divide total spend across a run by the number of tasks actually completed, so failures, retries and redundant tool calls are counted in the cost but not in the denominator. Cost per call flatters a system that fails often.

Should I measure average or p95 latency?

Make decisions on p95. The median describes an experience nobody complains about. For agents the gap between median and tail is unusually wide, because a task requiring several tool calls takes several times longer than one requiring none.

Why is my agent cost growing without more traffic?

Usually retrieved context. Every improvement that adds a passage to the prompt adds it to every request from then on. Splitting tokens per task into input and output shows immediately which side is growing.

Back to Insights See the benchmark

Find out what your agent actually does

From $6,000, typically three to four weeks. Tell us what it is meant to do and we will tell you how we would measure it.

Book a benchmark