FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook · 6 minute read

How to Plan AI Capacity Before It Becomes a Problem

AI capacity planning differs from conventional planning because the binding constraint is usually a provider quota rather than your own infrastructure. Model demand including peaks, establish the real limits, secure headroom in advance, and decide what degrades first when capacity runs out.

By FISTA Solutions· AI-Native Engineering Team·
How to Plan AI Capacity Before It Becomes a Problem article cover

AI capacity planning differs from conventional planning in one important way: the binding constraint is usually a provider quota you do not control rather than infrastructure you do. This playbook covers planning around that, drawing on FISTA Solutions' AI agents work.

When is this worth doing?

Before a launch that will materially increase volume, before a seasonal peak, and periodically for any system whose usage is growing.

It is also worth doing immediately after any incident where requests failed under load, because that is evidence the current planning was wrong.

What does the sequence look like?

StepPurpose
1. Model demandIncluding the peak, not the average
2. Establish real limitsProvider quotas, shared across teams
3. Calculate headroomAgainst the peak with margin
4. Request increases earlyWeeks, not days
5. Decide degradation orderBefore it is needed
6. Monitor and revisitUsage and limits both change

Step 1 — Model demand including the peak

Project requests per period, tokens per request, and the shape of the peak — daily, weekly, and seasonal.

Averages mislead badly here. A system averaging comfortable volume can exceed its quota for an hour every morning, and the users affected experience it as broken rather than as briefly busy.

Include the growth trajectory and the events that cause spikes: campaigns, releases, month-end, and incidents in adjacent systems that push traffic your way.

Step 2 — Establish the real limits

Find out your provider quotas per model, per minute and per day, and who else in the organisation consumes them.

The shared nature of quotas is the part that surprises teams. Another team's batch process can exhaust the capacity your interactive product depends on, and the first sign is failures in your system caused by something you do not control.

Establish visibility across the organisation, ideally through a gateway that attributes usage by team. Without it, capacity planning is planning for your own consumption in a shared pool. See what is rate limiting for llm apis.

Step 3 — Calculate headroom against the peak

Compare projected peak consumption against the quota, with margin for error in the projection.

Plan for the peak plus a meaningful buffer. Systems planned to the average fail at the peak, and the peak is when the traffic matters most.

Where the margin is thin, that is the finding: either increase the quota, reduce consumption per request, or plan the degradation carefully. Discovering it during the peak is the expensive route to the same conclusion.

Step 4 — Request increases early

Quota increases involve the provider's own capacity allocation and approval, and they take weeks rather than hours.

Request them as part of planning rather than in response to constraint. A request submitted when you are already failing means being constrained for however long the process takes.

Providers generally want usage forecasts to support the request, which is another reason to have modelled demand properly. A specific forecast gets a better response than a general request for more.

Step 5 — Decide what degrades first

Rank request types by importance and decide in advance what happens when capacity is short: queue, shed, route to a cheaper model, or serve a cached response.

Without that decision, load sheds arbitrarily. Important requests fail alongside trivial ones, and the system's behaviour under stress is the opposite of what the business would choose.

Implement the priority rather than documenting it. A ranking that exists in a plan and not in code does nothing during the incident. See what is backpressure in ai pipelines.

Step 6 — Monitor consumption and revisit

Track consumption against quota continuously, with alerting at a threshold that leaves time to act.

Alert at seventy or eighty per cent rather than at the limit. An alert that fires when you are already constrained is a notification of an incident rather than a warning.

Revisit the plan when usage patterns change, when a new team starts consuming the shared quota, and when the provider changes its limits — which happens without much notice.

Is reserved capacity worth buying?

Where predictability matters and volume is steady, frequently yes. It converts a variable risk into a fixed cost and removes the shared-quota problem.

It suits systems with service commitments and steady load. It suits spiky or low-volume systems poorly, because you pay for the peak continuously.

Model it against your actual demand curve rather than against the peak alone. The answer depends heavily on how much of the time you would be using what you reserved.

What about multi-provider capacity?

Running a second provider as overflow gives capacity headroom and resilience at the cost of maintaining prompts and evaluation for both.

It works well for workloads where quality differences between providers are tolerable — classification, extraction, routine generation — and poorly where the primary provider's specific behaviour matters.

Test the overflow path regularly. Capacity that has never been used is capacity that may not work when needed.

Who needs to be involved?

An engineer who can measure consumption, whoever owns the provider relationship, and someone who can rank request importance from a business perspective.

The ranking needs business input. Engineering rankings of what matters tend to reflect what is technically interesting rather than what is commercially important.

How long does it take?

One to two weeks for the modelling and limit discovery, plus however long quota increases take — which is the part to start first.

What are the common failure modes?

Planning to the average. Not knowing about shared consumption. Requesting quota increases late. No degradation order. Alerting at the limit. And buying reserved capacity without modelling the demand curve.

How do you know it worked?

Peaks served without failures, quota headroom visible and monitored, degradation behaving as designed when it happens, and no surprises from another team's consumption.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

Find out your current provider quota and what proportion of it you use at your busiest hour. That single number tells you how urgent the rest of this is.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: quota headroom modelled against real peaks rather than averages, degradation order implemented in code rather than documented in a plan, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read how to forecast AI inference cost.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What actually limits AI capacity?

Provider rate and token quotas in most systems. Those are set per account, shared across every team using them, and cannot be autoscaled. Your own infrastructure is rarely the binding constraint for API-based systems.

02Why do quotas surprise teams?

Because they are shared. A batch job another team runs can exhaust the capacity your interactive product depends on, and neither team knows about the other until requests start failing.

03How far ahead should you request increases?

Weeks. Quota increase requests involve the provider's own capacity planning and approval, and they are not instant. Requesting one when you are already constrained means being constrained for the duration.

04What should degrade first?

Whatever matters least, decided in advance. Without that decision, load sheds randomly and important requests fail alongside trivial ones, which is the worst possible allocation.

05Is reserved capacity worth it?

Where predictability matters and volume is steady, frequently yes. It converts a variable risk into a fixed cost, which suits systems with a service commitment attached and suits spiky low-volume systems poorly.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project