Glossary · 4 minute read
What Is Cost per Task? The Metric That Governs AI Economics
Cost per task measures total AI spend divided by successfully completed outcomes, including retries, escalations, and human review time. It is the only view that supports model selection and design decisions, because a cheaper model that fails more often costs more per outcome than an expensive one that succeeds.
AI economics is discussed in cost per million tokens, which is a rate card rather than a measurement. The number that governs whether an AI system is worth operating is cost per completed task, and it frequently points in a different direction from the price list. This explainer covers how to compute and use it. It complements what is token accounting and llm cost per task benchmarking, and reflects FISTA Solutions' approach in AI enablement delivery.
What goes into the figure?
Total spend attributable to a task type, divided by tasks successfully completed. Spend includes every attempt, not only the successful one: retries, failed paths, retrieval and tool calls, and any human time spent reviewing or correcting output.
That last component is the one most often left out, and excluding it makes automation look considerably better than it is.
| Component | Usually counted | Should be counted |
|---|---|---|
| Successful model call | Yes | Yes |
| Failed attempts and retries | Rarely | Yes |
| Retrieval and tool calls | Sometimes | Yes |
| Human review time | Rarely | Yes |
| Escalated case handling | Rarely | Yes |
| Infrastructure | Sometimes | Yes |
Why does per-token pricing mislead?
Because it tells you nothing about consumption. A model priced at half the rate that requires three times the context, more few-shot examples, and two attempts to produce usable output costs more per outcome.
Rate comparisons are useful for estimating and useless for deciding. The decision requires running representative tasks through each candidate and measuring what they actually consume to reach a correct result.
How much does a cheaper model raise cost?
Frequently enough that it should be the default hypothesis to test. The mechanisms are consistent: longer prompts to compensate for weaker instruction following, more examples, higher retry rates, and more human correction downstream.
A model that is 40% cheaper per token and succeeds 15% less often is usually more expensive per task, and the additional human review time can double the gap.
What is the right comparison?
The human baseline for the same task, measured honestly. That means including supervision, error correction, and the time a person spends on the cases automation escalates — not just the nominal hourly rate.
Without that baseline a cost per task figure floats free of any decision. With it, the question becomes concrete: does this cost less than the alternative, and is the quality comparable.
Why measure per task type?
Because profiles differ by orders of magnitude. A document classification costs a fraction of a multi-step research task, and averaging them produces a number that describes no real workload.
Per-type measurement also reveals which tasks are worth automating. Frequently the distribution is skewed: a few task types account for most of the spend, and they are where optimisation belongs.
How does the metric change design?
It makes accuracy and cost the same conversation. Improving success rate reduces cost per task, which means retrieval quality, prompt clarity, and error handling become cost optimisations rather than quality investments competing with cost work.
That framing tends to produce better decisions than treating the two as opposed. See how to reduce ai costs.
What should you do first?
Pick your highest-volume AI task and compute cost per successful completion, including human review. Most teams find the number differs substantially from what they assumed, and the direction of the surprise usually determines what to work on next.
How does it apply to agents?
Sharply, because agent cost varies enormously per task. The distribution is typically skewed: most tasks complete in a few iterations and a small number consume many times the median. Reporting the mean hides that tail, and the tail is usually where both the cost and the failures concentrate.
The useful views are the median, the ninety-fifth percentile, and a list of the most expensive recent tasks with their traces attached. That last one regularly reveals a single failing tool or an unbounded loop that accounts for a disproportionate share of spend.
What about the cost of getting it wrong?
It belongs in the comparison even though it is harder to quantify. An automated task that is wrong 2% of the time carries the cost of those errors: rework, customer impact, and occasionally worse. A human baseline has its own error rate and its own cost, and comparing automation's price against a human's price while ignoring both error rates produces a misleading result in whichever direction the error rates differ.
Where errors are consequential, the honest comparison includes the expected cost of being wrong on each side, which frequently changes which option looks better.
How FISTA Solutions helps
FISTA Solutions measures cost per successful task including retries, tool calls, and human review, benchmarks candidate models on real workloads rather than rate cards, compares against an honestly measured human baseline, and treats accuracy improvements as cost optimisations, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 47% efficiency gains.
To understand what your AI actually costs to run, message FISTA on WhatsApp, or read llm cost per task benchmarking.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why is per-token pricing misleading?
Because it says nothing about how many tokens a task consumes. A model at half the price that needs three times the context and two attempts costs more per outcome. Price per token is a rate card, not an economic measurement.
02What belongs in the figure?
All attempts including failures, retrieval and tool costs, retries, and the human time spent reviewing or correcting output. That last item is frequently the largest and the one most often excluded, which flatters automation considerably.
03Can a cheaper model raise cost per task?
Routinely. If it needs longer prompts, more few-shot examples, more retries, or more human correction downstream, the total rises even as the per-token rate falls. This is the most common error in cost-driven model selection and it is rarely caught without per-task measurement.
04What is the right comparison?
The human baseline for the same task, measured honestly including supervision and error correction. Without it, a cost per task figure has no context, and nobody can say whether the automation is worth operating at all.
05Why measure per task type?
Because a document classification and a multi-step research task have profiles differing by orders of magnitude, and averaging them produces a number describing no real workload. Each task type needs its own figure before it can support any decision about design or model choice.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.