Playbook ┬╖ 6 minute read
How to Plan AI Capacity Before It Becomes a Problem
AI capacity planning differs from conventional planning because the binding constraint is usually a provider quota rather than your own infrastructure. Model demand including peaks, establish the real limits, secure headroom in advance, and decide what degrades first when capacity runs out.
AI capacity planning differs from conventional planning in one important way: the binding constraint is usually a provider quota you do not control rather than infrastructure you do. This playbook covers planning around that, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
Before a launch that will materially increase volume, before a seasonal peak, and periodically for any system whose usage is growing.
It is also worth doing immediately after any incident where requests failed under load, because that is evidence the current planning was wrong.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Model demand | Including the peak, not the average |
| 2. Establish real limits | Provider quotas, shared across teams |
| 3. Calculate headroom | Against the peak with margin |
| 4. Request increases early | Weeks, not days |
| 5. Decide degradation order | Before it is needed |
| 6. Monitor and revisit | Usage and limits both change |
Step 1 тАФ Model demand including the peak
Project requests per period, tokens per request, and the shape of the peak тАФ daily, weekly, and seasonal.
Averages mislead badly here. A system averaging comfortable volume can exceed its quota for an hour every morning, and the users affected experience it as broken rather than as briefly busy.
Include the growth trajectory and the events that cause spikes: campaigns, releases, month-end, and incidents in adjacent systems that push traffic your way.
Step 2 тАФ Establish the real limits
Find out your provider quotas per model, per minute and per day, and who else in the organisation consumes them.
The shared nature of quotas is the part that surprises teams. Another team's batch process can exhaust the capacity your interactive product depends on, and the first sign is failures in your system caused by something you do not control.
Establish visibility across the organisation, ideally through a gateway that attributes usage by team. Without it, capacity planning is planning for your own consumption in a shared pool. See what is rate limiting for llm apis.
Step 3 тАФ Calculate headroom against the peak
Compare projected peak consumption against the quota, with margin for error in the projection.
Plan for the peak plus a meaningful buffer. Systems planned to the average fail at the peak, and the peak is when the traffic matters most.
Where the margin is thin, that is the finding: either increase the quota, reduce consumption per request, or plan the degradation carefully. Discovering it during the peak is the expensive route to the same conclusion.
Step 4 тАФ Request increases early
Quota increases involve the provider's own capacity allocation and approval, and they take weeks rather than hours.
Request them as part of planning rather than in response to constraint. A request submitted when you are already failing means being constrained for however long the process takes.
Providers generally want usage forecasts to support the request, which is another reason to have modelled demand properly. A specific forecast gets a better response than a general request for more.
Step 5 тАФ Decide what degrades first
Rank request types by importance and decide in advance what happens when capacity is short: queue, shed, route to a cheaper model, or serve a cached response.
Without that decision, load sheds arbitrarily. Important requests fail alongside trivial ones, and the system's behaviour under stress is the opposite of what the business would choose.
Implement the priority rather than documenting it. A ranking that exists in a plan and not in code does nothing during the incident. See what is backpressure in ai pipelines.
Step 6 тАФ Monitor consumption and revisit
Track consumption against quota continuously, with alerting at a threshold that leaves time to act.
Alert at seventy or eighty per cent rather than at the limit. An alert that fires when you are already constrained is a notification of an incident rather than a warning.
Revisit the plan when usage patterns change, when a new team starts consuming the shared quota, and when the provider changes its limits тАФ which happens without much notice.
Is reserved capacity worth buying?
Where predictability matters and volume is steady, frequently yes. It converts a variable risk into a fixed cost and removes the shared-quota problem.
It suits systems with service commitments and steady load. It suits spiky or low-volume systems poorly, because you pay for the peak continuously.
Model it against your actual demand curve rather than against the peak alone. The answer depends heavily on how much of the time you would be using what you reserved.
What about multi-provider capacity?
Running a second provider as overflow gives capacity headroom and resilience at the cost of maintaining prompts and evaluation for both.
It works well for workloads where quality differences between providers are tolerable тАФ classification, extraction, routine generation тАФ and poorly where the primary provider's specific behaviour matters.
Test the overflow path regularly. Capacity that has never been used is capacity that may not work when needed.
Who needs to be involved?
An engineer who can measure consumption, whoever owns the provider relationship, and someone who can rank request importance from a business perspective.
The ranking needs business input. Engineering rankings of what matters tend to reflect what is technically interesting rather than what is commercially important.
How long does it take?
One to two weeks for the modelling and limit discovery, plus however long quota increases take тАФ which is the part to start first.
What are the common failure modes?
Planning to the average. Not knowing about shared consumption. Requesting quota increases late. No degradation order. Alerting at the limit. And buying reserved capacity without modelling the demand curve.
How do you know it worked?
Peaks served without failures, quota headroom visible and monitored, degradation behaving as designed when it happens, and no surprises from another team's consumption.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Find out your current provider quota and what proportion of it you use at your busiest hour. That single number tells you how urgent the rest of this is.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: quota headroom modelled against real peaks rather than averages, degradation order implemented in code rather than documented in a plan, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to forecast AI inference cost.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What actually limits AI capacity?
Provider rate and token quotas in most systems. Those are set per account, shared across every team using them, and cannot be autoscaled. Your own infrastructure is rarely the binding constraint for API-based systems.
02Why do quotas surprise teams?
Because they are shared. A batch job another team runs can exhaust the capacity your interactive product depends on, and neither team knows about the other until requests start failing.
03How far ahead should you request increases?
Weeks. Quota increase requests involve the provider's own capacity planning and approval, and they are not instant. Requesting one when you are already constrained means being constrained for the duration.
04What should degrade first?
Whatever matters least, decided in advance. Without that decision, load sheds randomly and important requests fail alongside trivial ones, which is the worst possible allocation.
05Is reserved capacity worth it?
Where predictability matters and volume is steady, frequently yes. It converts a variable risk into a fixed cost, which suits systems with a service commitment attached and suits spiky low-volume systems poorly.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.