Cost · 5 minute read
AI Agent Hosting Cost: Compute, State and Long-Running Work
Agent hosting cost is driven by isolated execution environments, state persistence, long-running task handling, and sandboxing rather than by inference spend. Agents pause for approvals and wait on external systems, which creates idle capacity that conventional request-based infrastructure neither anticipates nor handles economically.
Hosting agents is not hosting services. Agents run for minutes or hours, hold state across steps, may execute generated code, and can pause indefinitely waiting for a human to approve something. Infrastructure designed for stateless request-response handles none of that well. This guide covers what it actually requires, drawing on FISTA Solutions' AI agents delivery. It complements what is agent sandboxing and what is model serving.
Why does agent hosting differ?
Because the execution model differs fundamentally. A conventional request arrives, is handled, and completes in milliseconds with no retained state. An agent run may last minutes or hours, maintains working state across many steps, calls external systems that may be slow, and may stop to wait for a person.
Each of those requires infrastructure that request-response hosting does not provide, and the cost sits in providing them rather than in the model calls.
| Requirement | Conventional service | Agent |
|---|---|---|
| Execution duration | Milliseconds | Minutes to hours |
| State | Stateless | Persistent across steps |
| Code execution | Rare | Common, needs isolation |
| External waits | Bounded | Potentially indefinite |
| Concurrency cost | Low per request | High per run |
| Resumption after pause | Not needed | Required |
What does sandboxing cost?
Isolated ephemeral environments, created per task and discarded afterwards, with no credentials present, restricted network egress, and enforced resource limits.
Creating and tearing those down has real cost, and it is the price of allowing an agent to execute generated code at all. The alternative â running generated code in a shared environment with access to production credentials â is not a cost saving, it is an unmitigated vulnerability.
How do approval waits affect infrastructure?
They create idle capacity. An agent paused awaiting a human decision may wait hours or overnight, and holding an execution environment open for that period wastes resources entirely.
The pattern that works persists the agent's state, releases the environment, and resumes when the approval arrives. That requires serialisable state, which is an architectural requirement rather than an optimisation, and one that is far cheaper to design in than to retrofit. See what is a scratchpad in ai agents.
Why do concurrency limits matter more than capacity?
Because a modest number of concurrent agent runs consumes capacity that would serve a far larger number of conventional requests. Twenty concurrent agents each executing code and calling models is a substantial load.
Admission control â declining to start a new run when the system is saturated â is more effective than scaling, because scaling accelerator and execution capacity is slow. Declining to start is also better than starting many runs and finishing none.
What state must be persisted?
The agent's working conclusions, its progress against the goal, pending approvals, and enough context to resume coherently. Not the full conversation history, which can be summarised, and not raw tool outputs, which can be re-fetched or filtered.
Deciding what must survive a pause is a design question with cost implications, since storing everything makes resumption expensive and storing too little makes it impossible.
What is most often under-budgeted?
The infrastructure around the agent. Execution environments, state storage, orchestration, queuing, and observability together frequently exceed the model spend â particularly for agents that run long, execute code, or produce content-heavy traces.
Budgets built from token pricing miss all of it.
How should capacity be planned?
From concurrent runs rather than from request volume, and from the distribution of run durations rather than the mean. Agent workloads have long tails, and a system sized for the median will be saturated by the tail regularly.
What should you do first?
Measure your agents' run durations and concurrency. Those two numbers size the infrastructure honestly, and they are frequently different from what request-based reasoning would predict.
What about failure and retry?
Agent runs fail partway through more often than conventional requests, and restarting from the beginning wastes everything already done. Checkpointing progress so a failed run resumes rather than restarts is both a cost saving and a reliability improvement, and it depends on the same serialisable state that approval waits require.
That makes state design the single decision with the widest effect on agent infrastructure cost. Systems designed without it pay in wasted compute, in inability to pause for approval, and in restarts that consume the budget twice.
How does this scale with agent count?
Better than expected if the infrastructure is shared. Execution environments, state storage, orchestration, and observability serve every agent, so the marginal cost of an additional agent is mostly its own inference and integration.
Organisations that build each agent with its own infrastructure pay the platform cost repeatedly, which is the same compounding argument that applies to integration and evaluation â and the same reason the first agent should be built as platform rather than as a point solution.
How FISTA Solutions helps
FISTA Solutions builds agent infrastructure with isolated ephemeral execution environments, serialisable state that allows pause and resume across approval waits, admission control rather than reactive scaling, and capacity planned from concurrent runs and duration distribution, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To host agents without paying for idle environments, message FISTA on WhatsApp, or read what is agent sandboxing.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why does agent hosting differ from service hosting?
Because agents run for minutes or hours rather than milliseconds, hold state across steps, may execute code, and can pause waiting for a human or an external system. Conventional request-response infrastructure assumes none of those.
02What does sandboxing cost?
Isolated ephemeral environments per task, with no credentials, restricted network egress, and enforced resource limits. Creating and discarding those per run has real cost, and it is the price of letting an agent execute generated code safely.
03How do approval waits affect infrastructure?
They create idle time. An agent paused for a human decision may wait hours, and holding its execution environment open for that period wastes capacity. Persisting state and resuming later is the pattern, and it requires serialisable state.
04Why do concurrency limits matter?
Because a modest number of concurrent agent runs can consume capacity that would serve far more conventional requests. Admission control â declining to start a run when saturated â is more effective than scaling, which is slow.
05What is most often under-budgeted?
The infrastructure around the agent rather than the inference. Execution environments, state storage, orchestration, and observability together frequently exceed the model spend, particularly for agents that run long or execute code.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.