Glossary · 5 minute read
What Is Throughput in AI Systems? Capacity Planning Explained
Throughput measures how much work an AI system completes per unit time, usually as tokens or requests per second. It trades directly against latency through batching: larger batches use hardware more efficiently and make individual requests wait longer. Capacity planning must account for context length, not just request counts.
Capacity planning for AI systems fails when it borrows assumptions from conventional web services, where requests are roughly uniform and scaling is fast. Neither holds here: request sizes vary by orders of magnitude, and accelerator capacity is slow and expensive to acquire. This explainer covers how throughput actually behaves. It complements what is batch inference and what is model serving, and reflects FISTA Solutions' approach in AI enablement delivery.
Why tokens rather than requests?
Because requests are not comparable units. A thousand short classifications and a thousand long document summaries have identical request counts and workloads differing by a large factor.
Tokens per second, split into input processing and output generation, describes what the system is actually doing. Request rate describes the traffic arriving, which is useful for a different purpose.
| Factor | Effect on throughput |
|---|---|
| Batch size | Higher batches raise throughput, raise latency |
| Context length | Longer contexts reduce concurrent capacity |
| Output length | Generation is slower than input processing |
| Model size | Directly reduces tokens per second |
| Quantization | Raises throughput, may affect quality |
| Speculative decoding | Raises per-request speed, less under load |
How does batching trade against latency?
Directly. Processing several requests together uses the hardware more efficiently, because the expensive part is loading model weights rather than the arithmetic for any one sequence. Larger batches mean better utilisation and longer waits for individual requests to be included.
Continuous batching, where arriving requests join a batch already in flight, improves this considerably and is standard in current serving frameworks. The configuration should still follow the workload: interactive interfaces want small batches, bulk processing wants large ones.
Why does context length affect capacity?
Because attention memory scales with sequence length, and that memory competes with the memory available for batching. A system whose average context doubles can serve fewer concurrent requests even at identical request volume.
This makes context discipline a capacity concern as well as a cost one. Trimming retrieved passages and summarising history improves throughput, which is an argument that lands with infrastructure teams as well as finance. See what is context engineering.
How does this work with hosted APIs?
Capacity appears as rate limits, usually on both requests and tokens per minute. The planning task becomes knowing your peak demand against those limits, requesting increases in advance of needing them, and handling limit responses properly.
Treating a rate limit response as an error rather than as backpressure is a common mistake. It should trigger backoff and queueing, not a failure surfaced to the user.
How should peaks be handled?
By queueing with backpressure and communicating the wait. A user who waits four seconds during a peak with a visible indication is better served than one who receives an error, and far better served than one who waits in silence.
Provisioning for peak is the alternative and it is expensive, because accelerator capacity idles between peaks. For most workloads a combination of modest headroom and graceful queueing is the better economic choice.
What should be monitored?
Tokens per second in and out, queue depth and wait time, batch sizes achieved, rate limit responses, and utilisation. Queue wait is the one that most often becomes the dominant latency term while every other metric looks healthy.
What should you do first?
Measure your peak-to-average ratio in tokens rather than requests. Most teams find it higher than expected, and that ratio is what determines whether the answer is more capacity, better queueing, or moving eligible work to batch.
How does agent work change the picture?
Agents multiply the request count per user action unpredictably. One user request may become fifteen model calls and twenty tool invocations, and the relationship between visible traffic and actual load stops being proportional. Capacity planned on user-facing request rates will be wrong by whatever the average iteration count turns out to be.
The planning input that matters is model calls per completed task, measured from traces, with its distribution rather than its mean. A workload where most tasks take three calls and some take forty needs headroom sized for the tail, not the median.
What about multi-tenant fairness?
Under contention, one tenant's heavy workload can consume capacity that others are waiting for, and without per-tenant limits the experience degrades for everyone because of one account's behaviour. Per-tenant rate limiting, applied in tokens rather than requests, is what keeps that contained.
The limits should be visible to the tenants they apply to. A customer who knows their ceiling can plan around it; one who experiences unexplained slowness raises a support case instead.
How do you know when to add capacity?
When queue wait time consistently exceeds the latency users tolerate, not when utilisation crosses a threshold. High utilisation with acceptable wait times is efficient rather than dangerous, and adding capacity because a gauge is amber is how AI infrastructure budgets grow without anyone experiencing an improvement.
How FISTA Solutions helps
FISTA Solutions plans AI capacity in tokens rather than requests, configures batching to match interactive and bulk workloads separately, treats rate limits as backpressure with queueing, monitors queue wait alongside latency, and moves eligible work to batch processing to flatten peaks, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To size AI capacity against traffic you actually get, message FISTA on WhatsApp, or read what is batch inference.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why is tokens per second the better measure?
Because requests vary enormously in size. A thousand short classifications and a thousand long document summaries are the same request count and completely different workloads. Token throughput describes the work; request throughput describes the traffic.
02How does batching affect this?
It raises aggregate throughput by keeping hardware busy and increases individual latency because requests wait for a batch. Continuous batching, where requests join an in-flight batch, improves the trade substantially and is standard in modern serving stacks.
03Why does context length matter for capacity?
Because attention memory grows with sequence length, so long contexts consume capacity disproportionately and reduce how many requests can be batched together. A system whose average context doubles loses throughput even at constant request volume.
04What about hosted APIs?
Capacity appears as rate limits on requests and tokens per minute. Planning means knowing your peak against those limits, requesting increases before you need them, and handling limit responses with backoff rather than treating them as errors.
05How should peaks be handled?
With queueing and backpressure rather than failure. A request that waits four seconds during a peak is better than one that errors, and communicating the wait is better than a silent delay. Provisioning for peak is the expensive alternative.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.