Cost · 5 minute read
Digital FTE Cost: What an AI Agent Role Really Costs to Run
A Digital FTE's cost is the cost per completed task of an AI agent role: model inference, retrieval and tool calls, infrastructure, human oversight for approvals and exceptions, a share of the shared platform, and maintenance of the specification and evaluation suite. It is highest at launch, falls as autonomy rises, and must be compared at equal quality.
Digital FTEs are budgeted badly because they fit neither existing template. They are not software, where the cost is a license line, and they are not people, where the cost is salary plus overhead. They are capacity with a run cost, an oversight cost, and a platform behind them, and their economics change over their lifecycle in a way neither software nor headcount does. This guide sets out what a Digital FTE costs to run and how to compare it honestly. FISTA does not publish price points because run cost depends on your volume, model choice, and integration surface; the components below are what to compute. It expands on what is a Digital FTE and the Digital FTE economics whitepaper.
What is the unit of cost?
The unit is a completed task of a defined type at the required quality: one invoice matched, one ticket resolved, one claim triaged. Everything spent to reach acceptable completion belongs to the unit, including failed attempts, retries, and the human work that finishes what the agent escalated. Cost per model call is easy to measure and misleading, because it hides steps, retries, and human finishing.
What are the six cost components?
| Component | What it includes | Driver |
|---|---|---|
| Inference | Model tokens across all steps and retries | Steps per task × context size × model price |
| Retrieval and tools | Vector search, database and API calls, tool-gateway overhead | Tool calls per task |
| Infrastructure | Runtime compute, sandboxes, queues, storage, observability | Concurrency and data volume |
| Oversight | Human time on approvals, exceptions, and quality sampling for this role | Exception rate × handling time × loaded hourly cost |
| Platform share | Gateway, integration layer, evaluation tooling, security controls, allocated across the fleet | Fleet size |
| Maintenance | Spec updates, evaluation upkeep, model migrations, incident handling | Engineering hours per quarter ÷ task volume |
Why does oversight dominate early?
A new Digital FTE launches at the suggest level, where a person reviews every case, or at act with approval, where consequential steps pause for a human. Either way, human time is attached to most tasks. As evaluation evidence accumulates and the agent advances to act with sampling, oversight shrinks to a sample plus the exceptions, and the cost per task falls sharply. This curve is the single most important feature of Digital FTE economics, and plans that assume end-state costs on day one are wrong in the expensive direction. The autonomy model is described in human-in-the-loop AI explained.
| Phase | Cost per task | Why |
|---|---|---|
| Shadow mode | Highest, no savings | Full human cost plus agent cost plus evaluation build |
| Suggest | High | Every case touched by a person |
| Act with approval | Falling | Only consequential steps pause; exceptions shrink |
| Act with sampling | Lowest | Sample plus exceptions only |
| Maturity | Stable, watch for drift | Context creep and spec bloat raise cost quietly |
How does the platform share work?
The first Digital FTE pays for the gateway, the integration layer, the evaluation tooling, and the security controls. The second one reuses them. By the tenth, the platform share per agent is a fraction of what the first carried, which is why the economics of a fleet are better than the economics of a pilot and why the business case should be made at the portfolio level. The platform components are described in the LLM gateway architecture whitepaper and the Model Context Protocol for the enterprise whitepaper.
How should the comparison with alternatives be built?
Compare against the fully loaded alternative for the same volume and quality, over three years.
| Element | Human or outsourced path | Digital FTE |
|---|---|---|
| Direct cost | Loaded labor or per-transaction fee | Inference, tools, infrastructure |
| Management and tooling | Supervisor time, licenses, vendor management | Platform share, maintenance |
| Oversight and quality | QA sampling, service levels | Approvals, exceptions, sampling |
| Errors and rework | Measured error rate × cost to fix | Measured error rate × cost to fix |
| Delay | Backlog cost, late fees, churn | Usually lower; measure it |
Two rules keep it honest: measure the human path's error rate rather than assuming zero, and use a multi-year horizon so the front-loaded platform investment and the falling oversight cost are both visible. The full method is the AI agent unit economics whitepaper; the comparison with outsourcing specifically is in Digital FTE vs BPO outsourcing.
What levers reduce the cost?
- Specification quality: fewer steps, fewer retries, fewer escalations; the cheapest lever.
- Model routing: simple cases on efficient models, escalation on failure or low confidence.
- Context discipline: retrieve and include only what the task needs.
- Caching: repeated context and answers served from cache where safe.
- Step and token budgets: hard limits that prevent runaway runs.
- Autonomy earned by evidence: the lever that moves oversight, the largest early component.
- Fleet growth: platform share per agent falls.
How should Digital FTE spend be governed?
Set a budget per Digital FTE enforced at the gateway, show cost per task next to quality on one dashboard so cost is never cut blind, alert on step-count and retry spikes, attribute every component to an owner, and review cost per task against the alternative each quarter. The budgeting process for a fleet is in how to budget for Digital FTEs.
What are the common mistakes?
- Pricing calls, not tasks.
- Treating oversight as free.
- Assuming day-one economics are end-state economics.
- Comparing with a perfect human path.
- Optimizing cost without re-running evaluation, so quality falls unnoticed.
How does FISTA Solutions help?
FISTA Solutions builds the cost model with finance and operations teams as part of its AI enablement practice, instruments it through the gateway and evaluation pipeline, and delivers every Digital FTE as a governed AI agent with the telemetry the model needs. Our forward deployed engineers measure the agent's real profile in shadow mode so projections rest on observed data. FISTA has delivered 150+ projects for 50+ companies across 12+ countries with 47% average efficiency gains.
To build the cost model for one role, message FISTA on WhatsApp, or read Digital FTE vs human FTE for the comparison it feeds.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How much does a Digital FTE cost?
There is no single figure, because run cost depends on task volume, steps per task, model choice, integration complexity, and how much human oversight the autonomy level requires. The honest answer is a model with six components computed for your workload; this guide explains each component and how it changes over the agent's lifecycle.
02Which cost component is usually underestimated?
Human oversight. At launch a Digital FTE runs at the suggest or act-with-approval level, so every case or every consequential action touches a person. Plans that count only model prices miss this entirely and are surprised by the early cost; the component falls as evidence allows autonomy to rise.
03Does a Digital FTE get cheaper over time?
Usually, for three reasons: oversight cost falls as autonomy rises on evidence, specification improvements reduce steps and retries, and the platform share per agent falls as more Digital FTEs share the same gateway, integration layer, and evaluation tooling. Model price declines help but should not carry the business case.
04How do you compare a Digital FTE's cost with a person's?
Use the fully loaded cost of the human path for the same task volume and quality: labor, management, tooling, error and rework costs, and cost of delay, over a multi-year horizon. Compare it with the Digital FTE's full cost per task including oversight and maintenance. Equal quality, measured on both sides, is the condition for a fair comparison.
05What drives Digital FTE cost down fastest?
Tighter specifications that cut steps and escalations, routing simple cases to cheaper models with escalation on failure, caching repeated context, retrieving only what the task needs, step budgets that prevent runaway runs, and raising autonomy as the evaluation evidence supports it, which shrinks the oversight component.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.