FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Checklist · 4 minute read

AI Capacity Model Template: Sizing With Numbers That Hold

A capacity model turns guesses into numbers you can defend. Start from request volume at peak, multiply by measured tokens per request, compare against provider rate limits, project cost, and state every assumption so the model can be corrected when reality differs from it.

By FISTA Solutions· AI-Native Engineering Team·
AI Capacity Model Template: Sizing With Numbers That Hold article cover

A capacity model turns guesses into numbers you can defend in a budget conversation. This template covers building one, drawn from FISTA Solutions' AI enablement operational work.

What does the model compute?

Six outputs from a small set of measured inputs.

OutputDerived from
Tokens per periodVolume times tokens per request
Rate limit headroomPeak requests against the limit
Concurrency requiredPeak rate times request duration
Cost per periodTokens times price per model
Review hours requiredEscalation rate times handling time
Growth runwayCurrent headroom against growth rate

Demand inputs

Measured from real traffic, at peak.

  • Request volume measured per hour over several weeks
  • Peak hour identified and its ratio to average computed
  • Weekly and seasonal patterns documented
  • Growth rate derived from history, not from ambition
  • Volume split by operation, since they differ
  • Launch or campaign spikes modelled separately
  • Source of each figure recorded

Token maths

Measured, because estimates are consistently low.

  • Input tokens measured per operation on real requests
  • Output tokens measured, including the tail
  • Retrieved context tokens included
  • System prompt and tool definition tokens included
  • Conversation history growth modelled where relevant
  • Agent trajectories counted as multiple calls
  • Retries included in the total

Throughput and limits

Compare the arithmetic against what the provider allows.

  • Provider requests-per-minute limit recorded
  • Provider tokens-per-minute limit recorded
  • Peak demand compared against both
  • Headroom expressed as a multiple, not a percentage
  • Lead time to raise limits recorded
  • Secondary provider capacity considered
  • Downstream system limits included

Cost projection

Per model, per operation, over the planning period.

  • Current prices recorded per model with a date
  • Cost computed per operation per period
  • Caching discounts applied where they exist
  • Batch pricing applied to batchable work
  • Projection over twelve months with growth
  • Sensitivity to the main assumptions shown
  • Comparison against the budget ceiling

Human capacity

Frequently the first constraint to bind.

  • Escalation rate measured or estimated with a stated basis
  • Handling time per escalation measured
  • Review rate for outputs requiring approval
  • Hours required per period computed
  • Staffing compared against the requirement
  • Coverage across time zones and languages included
  • Peak concentration accounted for, not just totals

Assumptions and review

Visible assumptions are correctable ones.

  • Every assumption listed explicitly with its basis
  • Confidence noted per assumption
  • The assumptions most affecting the outcome identified
  • Model compared against actuals quarterly
  • Divergence investigated and assumptions corrected
  • Model owner named
  • Version history kept so past projections can be reviewed

What are the most common failures?

Modelling from average volume. Estimating tokens. Headroom expressed so thinly that normal variation breaches it. Human capacity omitted. And assumptions buried in formulas where nobody can check them.

Who should own this?

Engineering builds the model; the business owner supplies volume expectations and accepts the cost projection. A model built without business input projects the past rather than the plan.

How often should it run?

Built before launch, compared against actuals quarterly, and rebuilt when volume, routing, or pricing changes materially.

What evidence should it produce?

The model with its stated assumptions, measurement records behind the inputs, and the quarterly comparison against actuals.

What if you have no traffic to measure yet?

Model from a comparable process — the manual volume the system will handle — and state that clearly as the basis. Then re-measure as soon as real traffic exists.

A pre-launch model is a planning tool, not a commitment. The value comes from revisiting it with real numbers within the first month. See AI capacity planning checklist.

What should you do first?

Measure actual tokens per request for your main operation. It is usually well above what anyone estimated, and it changes the whole model.

How FISTA Solutions helps

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: capacity modelled from measured tokens at peak volume with every assumption stated, and compared against actuals quarterly so it stays honest, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI capacity planning checklist.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What are the model's inputs?

Request volume at peak, tokens per request measured from real traffic, concurrency, the model's throughput and rate limits, and the human review rate. Everything else derives from those.

02Why measure tokens rather than estimate?

Because estimates are consistently low. Retrieved context, system prompts, and conversation history all add tokens that a mental estimate omits, frequently by a large factor.

03What does the model output?

Required rate limit headroom, projected cost per period, concurrency needed, and human review hours. Those four numbers answer the questions capacity planning exists to answer.

04Why state assumptions?

Because they will be wrong, and a model whose assumptions are visible can be corrected. One with assumptions buried in formulas gets discarded when reality diverges.

05When should it be rebuilt?

When volume changes materially, when the model or routing changes, or when actual figures diverge from projection by enough to matter. Quarterly comparison against actuals keeps it honest.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project