Playbook · 6 minute read
How to Scale an AI Agent to Millions of Requests
Scaling an agent means handling three things that pilot volume hides: provider rate limits and their failure behaviour, cost per task at volumes where small differences matter, and quality variance across a much wider input distribution than the pilot ever saw.
Agents that work at pilot volume fail at scale for reasons the pilot could not surface: provider limits, cost per task, and a much wider input distribution. This playbook covers each, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
When volume is growing towards levels where per-task economics matter, or when a system is already hitting limits and failing in ways that look like infrastructure problems but are not.
It is not worth doing pre-emptively for a system whose volume is not growing. Scaling work adds complexity, and complexity added for volume that never arrives is pure cost.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Find the real bottleneck | Usually provider limits, not compute |
| 2. Measure cost per task | At the volume you are heading for |
| 3. Route by difficulty | Keep the expensive path small |
| 4. Handle limits deliberately | Queue, shed, back off |
| 5. Widen the evaluation set | New inputs bring new failures |
| 6. Scale the monitoring | Sampling, drift, rates |
Step 1 â Find the real bottleneck
Instrument where requests wait and where they fail. In most agent systems the ceiling is provider rate limits rather than your own compute.
Provider quotas apply per account and per model and are shared across everything your organisation runs, which means another team's batch job can exhaust the capacity your interactive product depends on. That is an organisational problem before it is a technical one.
Establish the quota, the current headroom, and the process for raising it â which frequently takes weeks and should start before you need it.
Step 2 â Measure cost per task at target volume
Project the current cost per completed task to the volume you are heading for and check whether the business case survives.
This is the calculation that kills scaling projects, and it is better done early. A system costing a few cents more per task than it needs to is irrelevant in a pilot and decisive at a million a day.
If the projection does not work, the fix is architectural rather than operational: routing, context reduction, and caching. Those are easier to introduce before scale than after. See what is cost per task.
Step 3 â Route by difficulty
Send the straightforward majority down a cheaper path and reserve the expensive model for cases that need it.
At scale this is usually the difference between viable and not. Most request distributions are heavily weighted towards simple cases, and serving those with an expensive general model is the largest single waste in most systems.
Validate the routing per route rather than in aggregate, and monitor the proportion going each way. A routing scheme that drifts towards the expensive path as inputs change will quietly rebuild the cost problem.
Step 4 â Handle limits deliberately
Design what happens when capacity runs out: queue with an honest wait, shed load with a clear message, or degrade to a cheaper path.
The default â retry on failure â is the worst option, because retries during saturation increase load and turn a capacity problem into an outage. Exponential backoff with jitter, circuit breakers, and a queue with visible depth are the tools.
Decide which requests matter most and prioritise them explicitly. Under load, treating all requests equally means the important ones fail alongside the rest. See what is backpressure in ai pipelines.
Step 5 â Widen the evaluation set
Volume brings inputs the pilot never saw: malformed submissions, unexpected languages, very long documents, and deliberate probing.
Sample real production traffic and add the unusual cases to the evaluation set. This is continuous rather than a one-off, because the distribution keeps widening as the audience grows.
Watch specifically for inputs that cause expensive behaviour â long context, many tool calls, retries â because those are both a cost and a denial-of-service surface.
Step 6 â Scale the monitoring approach
Reviewing every interaction stops being possible early. Move to sampled review with stratification: a random sample plus every escalation, every low-confidence case, and every expensive trajectory.
Add drift detection on inputs so a change in what users are asking surfaces before it shows up as a quality complaint.
Alert on rates rather than individual events. At volume, individual failures are constant and meaningless; a change in the rate is the signal. See what is continuous evaluation.
What about multi-region and failover?
Both become relevant at scale, and both add complexity worth adding deliberately.
Multi-region helps latency for distributed users and raises data residency questions that need settling first. Failover to a second provider helps availability and requires the abstraction, prompts, and evaluation for that provider to be maintained rather than theoretical.
Decide which problem you are solving. Adding both because they sound like scale practices produces complexity with no specific benefit.
When should you consider self-hosting?
At high, steady, predictable volume where the per-token economics and the operational burden favour it.
Self-hosting trades a variable per-token cost for fixed infrastructure and an operations commitment. That is a good trade at sustained high volume and a poor one at spiky or growing volume, where capacity has to be provisioned for the peak.
Model the comparison on your actual pattern including engineering time, not on list prices. See AI agent hosting cost.
Who needs to be involved?
An engineer who can change the architecture, someone who owns the provider relationship and quotas, and someone accountable for both cost and quality at volume.
The quota owner is frequently overlooked and is on the critical path, because limit increases take time.
How long does it take?
Four to eight weeks for the architectural changes, with quota increases and provider negotiation running in parallel and sometimes longer.
What are the common failure modes?
Assuming your infrastructure is the ceiling. Projecting cost late. Retrying into saturation. Evaluation sets that stop growing. Reviewing everything until it becomes impossible. And adding multi-region without a specific reason.
How do you know it worked?
Volume served within provider limits, cost per task holding as volume grows, quality stable across a wider input distribution, and overload handled visibly rather than as timeouts.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Find out your current provider quota and how close you are to it. That number is usually lower and closer than teams expect, and raising it takes time you should start spending now.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: provider limits and quota headroom established before load becomes the constraint, cost per task projected to target volume before the architecture is fixed, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read what is rate limiting for LLM APIs.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What usually limits scale?
Provider rate limits rather than your own infrastructure. Quotas are per account and per model, they apply across your whole organisation, and hitting them produces errors that need handling rather than autoscaling.
02Why does cost per task matter more at scale?
Because small per-task differences multiply. A design that costs a few cents more per task is irrelevant at a thousand a day and decisive at a million, which is why architecture that ignores it fails the business case rather than the load test.
03What changes about quality?
The input distribution widens. Cases that never appeared in the pilot arrive regularly at volume, including malformed inputs, unusual languages, and adversarial attempts, and each is a new failure mode.
04How should overload be handled?
With deliberate queueing and load shedding rather than timeouts and retries that compound. A system that queues honestly and tells users what to expect degrades better than one that accepts everything and fails slowly.
05What monitoring is needed at volume?
Sampled quality review rather than full review, drift detection on inputs, cost per task tracked continuously, and alerting on rates rather than individual failures. Reviewing everything stops being possible early.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.