Checklist ┬╖ 5 minute read
AI Capacity Planning Checklist: Sizing Before You Scale
AI capacity planning is constrained by provider rate limits and per-call cost rather than by your own hardware. Model demand including peaks, confirm rate limits against that peak, define latency budgets, set cost ceilings, and test behaviour under throttling before it happens in production.
AI capacity planning has unusual constraints, because the ceiling is usually someone else's rate limit rather than your own hardware. This checklist covers sizing properly, drawn from FISTA Solutions' AI enablement operational work.
What actually constrains capacity?
Six constraints, in the order they typically bind.
| Constraint | How it shows up |
|---|---|
| Provider requests per minute | Throttling at peak |
| Provider tokens per minute | Throttling on long prompts |
| Cost ceiling | Budget exhausted mid-month |
| Latency budget | Interactive use becomes unusable |
| Downstream system limits | Tools rate limit before the model does |
| Human review capacity | Queue grows faster than it clears |
Demand modelling
Model the peak, and model growth.
- Current request volume measured per hour, not per day
- Peak-to-average ratio calculated from real traffic
- Seasonal and weekly patterns identified
- Growth projected for twelve months
- Token volume modelled, not only request count
- Volume per feature separated, since they scale differently
- A launch or campaign spike scenario included
Provider limits
The usual ceiling, and raising it takes time.
- Current rate limits documented for every provider and model
- Limits compared against projected peak, not average
- Process and lead time for raising limits understood
- Behaviour under throttling defined and tested
- Retry with backoff implemented so throttling does not compound
- A secondary provider available for overflow
- Limit headroom monitored with an alert before it binds
Latency budgets
Decide the acceptable response time per use case before choosing a model.
- A latency target set for each use case
- Time to first token targeted separately for streaming
- Targets measured at the ninety-ninth percentile
- Context size constrained to meet the target
- Model choice validated against the target under load
- Synchronous versus asynchronous decided per use case
- Degradation path defined when the target cannot be met
Cost ceilings
Affordability is a capacity limit. See LLM cost control checklist.
- Monthly cost ceiling agreed with the budget owner
- Cost per task known for every operation
- Projected cost at projected volume computed
- Spend rate alerts configured per hour
- A hard cap or throttle defined for runaway spend
- Cheaper routing options identified before the ceiling binds
- Behaviour when the ceiling is reached defined and communicated
Downstream and human capacity
The model is rarely the only constraint in the chain.
- Rate limits on every tool and downstream system documented
- Database and API capacity checked against agent call volume
- Human review capacity modelled against escalation rate
- Escalation queue growth monitored with an alert
- Approval turnaround measured against expectation
- Capacity for exception handling provisioned, not assumed
- Weekend and overnight coverage considered
Testing and headroom
Test the failure behaviour, not just the happy capacity.
- Load tested at projected peak, not current volume
- Behaviour under provider throttling observed in a test
- Queue depth and drain rate measured under load
- Graceful degradation verified rather than assumed
- Headroom target agreed, typically well above peak
- Batch workloads moved off the interactive peak
- Capacity reviewed whenever volume grows materially
What are the most common failures?
Planning against average volume. Discovering rate limits during a launch. Latency measured serially. No cost ceiling. And ignoring human review capacity, which is frequently the first thing to saturate.
Who should own this?
Engineering models the technical capacity; the business owner agrees the cost ceiling and the review staffing. Capacity planned without the cost and staffing conversation is incomplete.
How often should it run?
At launch, before any expected volume increase, and quarterly thereafter. Re-check provider limits after any provider or plan change, because they are not always carried over.
What evidence should it produce?
The demand model with peak figures, documented rate limits, load test results including throttling behaviour, and the agreed cost ceiling. That package makes scaling decisions defensible.
What if you cannot raise provider limits in time?
Queue non-urgent work, route overflow to a second provider, and degrade gracefully for the rest. All three need to exist before the peak arrives.
The better position is to request increases well ahead of need, since lead times are measured in days or weeks rather than minutes. See what is a fallback chain.
What should you do first?
Compare your peak hourly volume against your provider's rate limits. If the margin is thin, that is the constraint to address before anything else.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: capacity modelled against peak and projected growth rather than current average, with throttling behaviour tested before it occurs in production, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read LLM cost control checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What limits capacity in an AI system?
Usually the provider's rate limits on requests and tokens per minute, not your own infrastructure. Those limits are negotiable but not instantly, so they need checking against projected peak.
02Why model peak rather than average?
Because business systems are bursty. A daily average that fits comfortably within a rate limit can exceed it for an hour every morning, which is when users notice.
03What is a latency budget?
The maximum acceptable response time for a use case, decided in advance. It determines model choice, context size, and whether work can be done synchronously at all.
04How is cost a capacity constraint?
Because a system that could serve ten times the volume technically may not be affordable at that volume. The cost ceiling is a real limit and should be planned like any other.
05What should be batched?
Anything nobody is waiting for тАФ overnight classification, bulk enrichment, report generation. Batch pricing is cheaper and it moves load away from the interactive peak.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.