Cost · 5 minute read
AI Agent Cost by Use Case: What Actually Drives the Number
AI agent cost is driven by integration surface, approval and control requirements, evaluation depth, and the consequence of failure — not by the model. A support agent and a finance agent using the same model differ by an order of magnitude because those four factors differ, not because inference costs more.
Agent cost estimates go wrong because they anchor on the model. Two agents built on the same model, by the same team, can differ by an order of magnitude in total cost, and the difference is entirely in what they touch, what happens when they are wrong, and what has to be proven before they are allowed to run. This guide covers those drivers, drawing on FISTA Solutions' AI agents delivery. It complements what is cost per task and ai total cost of ownership.
Why is the model not the main cost?
Because inference is inexpensive relative to engineering. A production agent's cost concentrates in integration, control implementation, evaluation, and maintenance, and the per-token spend is frequently a minor line even at meaningful volume.
That is counter-intuitive because the model is the visible novelty. It is also the reason cost estimates built from token pricing are consistently wrong in the same direction.
| Driver | Effect on cost | Varies by |
|---|---|---|
| Systems integrated | Large | Count and API quality |
| Consequence of error | Large | Domain and action type |
| Approval requirements | Large | Regulatory and financial exposure |
| Evaluation depth | Moderate to large | Consequence and volume |
| Ongoing maintenance | Recurring | System count and change rate |
| Inference volume | Small to moderate | Usage and context size |
What drives integration cost?
The number of systems an agent must read from and act on, the quality of their interfaces, and whether user authorisation can be propagated through them.
An agent working against one well-documented API is a fraction of the cost of one working across six systems of varying age, where two have no usable API and one requires a service account that breaks per-user authorisation. That last problem in particular turns a straightforward build into an architecture exercise.
How does failure consequence affect cost?
It determines everything downstream. An agent drafting internal text needs light evaluation and no approval workflow. An agent initiating payments needs approval gates, audit trails, idempotency, calibrated confidence, and evidence that would satisfy a regulator.
That is not a configuration difference; it is a substantial engineering difference. The same capability with a different consequence profile is a different project.
How does evaluation scale with the use case?
With consequence and with breadth. A narrow agent doing one thing in one domain needs a modest evaluation set. A broad agent handling varied requests needs coverage across that variety, including the cases where it should abstain.
Building the evaluation set is human labelling work, and keeping it current as usage evolves is ongoing. Teams routinely budget the build and omit the maintenance. See what is continuous evaluation.
What does an approval workflow actually cost?
More than expected. It requires serialisable agent state so a run can pause and resume, routing to the right approver with context, notification and escalation, an audit record, and handling for approvals that never arrive.
That is real engineering, and it is required the moment an agent can do anything consequential. Use cases that need it should be estimated with it included rather than as an addition.
What is recurring?
Evaluation maintenance, prompt and model version management, monitoring and alerting, incident response, and the labelling that keeps evaluation representative. Integration maintenance too, since every system the agent touches changes independently.
Those recurring costs scale with the number of integrated systems and the rate at which they change, which is another reason integration surface matters more than it appears.
How do common use cases compare?
Internal drafting and summarisation sit at the low end: few integrations, low consequence, light evaluation. Support agents sit in the middle: several integrations, moderate consequence, substantial evaluation because the output reaches customers. Finance and operations agents that take action sit at the high end: many integrations, high consequence, approval workflows, and evidence requirements.
The ordering is stable across organisations, even though the absolute figures are not.
What makes an estimate credible?
Counting the systems, naming the consequence of error, listing the approvals required, and sizing the evaluation set. Those four answers predict cost better than any benchmark or comparison to another company's agent.
An estimate that cannot answer them is an estimate of the demonstration rather than of the system.
What should you do first?
Take your intended use case and write down those four answers. If the systems count is high or the consequence is serious, the cost is in the engineering rather than the model, and the estimate should be built from there.
What makes cost fall over time?
Reuse. The second agent in an organisation costs less than the first because the integration layer, the evaluation harness, the approval workflow, and the observability already exist. That compounding is the strongest argument for building the platform properly on the first use case rather than treating it as a one-off.
Organisations that build each agent as an isolated project pay the full cost every time and wonder why the programme does not become cheaper. The saving is architectural rather than incremental.
How FISTA Solutions helps
FISTA Solutions estimates agent programmes from integration surface, failure consequence, approval requirements, and evaluation depth rather than from model pricing, and builds the recurring evaluation and maintenance into the plan rather than discovering it later, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To estimate an agent programme on the drivers that actually determine cost, message FISTA on WhatsApp, or read what is cost per task.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why is the model not the main cost?
Because inference is cheap relative to engineering. A production agent's cost is dominated by integration work, control implementation, evaluation, and ongoing maintenance, and the per-token spend is frequently a minor line even at substantial volume.
02What drives integration cost?
The number of systems an agent must read from and act on, their API quality, and whether authorisation can be propagated through them. An agent touching one well-documented system is a fraction of the cost of one touching six legacy ones.
03How does failure consequence affect cost?
It determines how much evaluation, control, and human review the system needs. An agent drafting internal text needs light evaluation; one initiating payments needs approval workflows, audit trails, and evidence that would pass a regulator.
04What is recurring rather than one-off?
Evaluation maintenance, prompt and model version management, monitoring, incident response, and the labelling that keeps evaluation sets current. These are ongoing and are the costs most often omitted from initial estimates.
05How should a use case be estimated?
By its characteristics rather than by comparison to another agent. Count the systems, assess the consequence of error, determine what approvals are required, and size evaluation accordingly. Those four answers predict cost better than any benchmark.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.