Checklist · 4 minute read
AI Capacity Model Template: Sizing With Numbers That Hold
A capacity model turns guesses into numbers you can defend. Start from request volume at peak, multiply by measured tokens per request, compare against provider rate limits, project cost, and state every assumption so the model can be corrected when reality differs from it.
A capacity model turns guesses into numbers you can defend in a budget conversation. This template covers building one, drawn from FISTA Solutions' AI enablement operational work.
What does the model compute?
Six outputs from a small set of measured inputs.
| Output | Derived from |
|---|---|
| Tokens per period | Volume times tokens per request |
| Rate limit headroom | Peak requests against the limit |
| Concurrency required | Peak rate times request duration |
| Cost per period | Tokens times price per model |
| Review hours required | Escalation rate times handling time |
| Growth runway | Current headroom against growth rate |
Demand inputs
Measured from real traffic, at peak.
- Request volume measured per hour over several weeks
- Peak hour identified and its ratio to average computed
- Weekly and seasonal patterns documented
- Growth rate derived from history, not from ambition
- Volume split by operation, since they differ
- Launch or campaign spikes modelled separately
- Source of each figure recorded
Token maths
Measured, because estimates are consistently low.
- Input tokens measured per operation on real requests
- Output tokens measured, including the tail
- Retrieved context tokens included
- System prompt and tool definition tokens included
- Conversation history growth modelled where relevant
- Agent trajectories counted as multiple calls
- Retries included in the total
Throughput and limits
Compare the arithmetic against what the provider allows.
- Provider requests-per-minute limit recorded
- Provider tokens-per-minute limit recorded
- Peak demand compared against both
- Headroom expressed as a multiple, not a percentage
- Lead time to raise limits recorded
- Secondary provider capacity considered
- Downstream system limits included
Cost projection
Per model, per operation, over the planning period.
- Current prices recorded per model with a date
- Cost computed per operation per period
- Caching discounts applied where they exist
- Batch pricing applied to batchable work
- Projection over twelve months with growth
- Sensitivity to the main assumptions shown
- Comparison against the budget ceiling
Human capacity
Frequently the first constraint to bind.
- Escalation rate measured or estimated with a stated basis
- Handling time per escalation measured
- Review rate for outputs requiring approval
- Hours required per period computed
- Staffing compared against the requirement
- Coverage across time zones and languages included
- Peak concentration accounted for, not just totals
Assumptions and review
Visible assumptions are correctable ones.
- Every assumption listed explicitly with its basis
- Confidence noted per assumption
- The assumptions most affecting the outcome identified
- Model compared against actuals quarterly
- Divergence investigated and assumptions corrected
- Model owner named
- Version history kept so past projections can be reviewed
What are the most common failures?
Modelling from average volume. Estimating tokens. Headroom expressed so thinly that normal variation breaches it. Human capacity omitted. And assumptions buried in formulas where nobody can check them.
Who should own this?
Engineering builds the model; the business owner supplies volume expectations and accepts the cost projection. A model built without business input projects the past rather than the plan.
How often should it run?
Built before launch, compared against actuals quarterly, and rebuilt when volume, routing, or pricing changes materially.
What evidence should it produce?
The model with its stated assumptions, measurement records behind the inputs, and the quarterly comparison against actuals.
What if you have no traffic to measure yet?
Model from a comparable process — the manual volume the system will handle — and state that clearly as the basis. Then re-measure as soon as real traffic exists.
A pre-launch model is a planning tool, not a commitment. The value comes from revisiting it with real numbers within the first month. See AI capacity planning checklist.
What should you do first?
Measure actual tokens per request for your main operation. It is usually well above what anyone estimated, and it changes the whole model.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: capacity modelled from measured tokens at peak volume with every assumption stated, and compared against actuals quarterly so it stays honest, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI capacity planning checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What are the model's inputs?
Request volume at peak, tokens per request measured from real traffic, concurrency, the model's throughput and rate limits, and the human review rate. Everything else derives from those.
02Why measure tokens rather than estimate?
Because estimates are consistently low. Retrieved context, system prompts, and conversation history all add tokens that a mental estimate omits, frequently by a large factor.
03What does the model output?
Required rate limit headroom, projected cost per period, concurrency needed, and human review hours. Those four numbers answer the questions capacity planning exists to answer.
04Why state assumptions?
Because they will be wrong, and a model whose assumptions are visible can be corrected. One with assumptions buried in formulas gets discarded when reality diverges.
05When should it be rebuilt?
When volume changes materially, when the model or routing changes, or when actual figures diverge from projection by enough to matter. Quarterly comparison against actuals keeps it honest.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.