Cost · 5 minute read
LLM Gateway Cost: What a Gateway Buys and What It Costs
An LLM gateway centralises routing, caching, rate limiting, cost attribution, and policy across providers. Its savings come mainly from routing and caching; its cost is build and operational burden. Below a certain scale it is premature infrastructure, and above it the attribution alone justifies it.
An LLM gateway is a piece of infrastructure that solves several real problems at once and introduces one significant risk: it sits in the request path. Whether it is worth building depends on how many applications call models, how much is being spent, and whether anyone can currently say where that spend goes. This guide covers the calculation, drawing on FISTA Solutions' AI enablement work. It complements how to build an mcp gateway and what is token accounting.
What does a gateway buy?
A single place for routing between models, prompt and semantic caching, rate limiting and backpressure, cost attribution by feature and team, policy enforcement, and provider failover.
Each of those can be implemented in an application. The gateway's value is that they apply consistently across every application rather than in whichever team got round to it, which matters once there are several.
| Capability | Per-application | Via gateway |
|---|---|---|
| Model routing | Duplicated | Consistent |
| Prompt and semantic caching | Duplicated | Shared cache |
| Rate limiting and backpressure | Duplicated | Centrally managed |
| Cost attribution | Rarely implemented | Automatic |
| Policy and guardrails | Inconsistent | Uniform |
| Provider failover | Rarely implemented | Centrally handled |
Where do the savings come from?
Routing and caching, principally. Sending straightforward requests to a smaller model rather than the largest available, and caching repeated prefixes and semantically equivalent queries, are where the material reductions sit.
Both are available without a gateway. What the gateway changes is coverage: they apply to every application rather than to the one that implemented them, and a shared cache serves hits across applications that a per-application cache would miss.
Why is attribution underrated?
Because without it AI spend arrives as a single provider invoice that answers no useful question. Which feature is expensive, which customer is unprofitable, whether last week's change cost anything — none of it is visible.
A gateway that tags every request with feature, team, model, and customer turns spend into something manageable. Teams frequently find that benefit exceeds the routing savings, because it changes decisions rather than just reducing a number. See what is token accounting.
What does it cost to operate?
Build or licensing, plus the responsibility of sitting in the request path. A gateway outage is an outage of every AI feature in the organisation, which means it needs the reliability engineering that implies — redundancy, monitoring, capacity headroom, and an on-call rota.
It also needs maintenance as providers change interfaces, add models, and adjust pricing. That is ongoing work rather than a one-off build.
When is it premature?
With one or two applications and modest spend. At that point the gateway's capabilities can be implemented directly in the applications with less total effort, and the availability risk is not worth taking for benefits that are small in absolute terms.
The signals that it is time are several teams calling models independently, nobody able to attribute spend, and routing or caching opportunities that no single team has the incentive to implement.
What about vendor options?
Several exist, and the build-versus-buy question turns on the same thing it does elsewhere: whether the organisation can operate what it chooses. A bought gateway removes the build cost and adds a dependency in the request path; a built one is fitted to the estate and must be maintained.
For most organisations buying is the right answer initially, with the option to replace if the requirements diverge.
What should be measured?
Savings from routing and caching against build and operating cost, attribution coverage, gateway availability, and added latency. That last matters: a gateway adds a hop, and if it adds meaningful latency the user experience cost may exceed the savings.
What should you do first?
Try to answer which of your features is most expensive in AI spend. If you cannot, that gap is the strongest argument for a gateway, and it is a gap that grows more expensive to fix as the estate expands.
How does it interact with governance?
Usefully, because a gateway is a natural enforcement point. Policy on which models may be used for which data classes, which teams may access which capabilities, and what must be logged can be applied once rather than trusted to every application.
That makes the gateway an asset to the governance function as well as to engineering, and it frequently strengthens the business case more than the routing savings do — particularly in regulated organisations where per-application policy compliance is otherwise unverifiable.
What are the failure modes?
Becoming a bottleneck, organisationally rather than technically. A gateway team that must approve every new model or capability slows the teams it serves, and they route around it — which reintroduces the inconsistency the gateway existed to remove.
The arrangement that works treats the gateway as infrastructure with self-service access and policy guardrails, not as a gatekeeper with a request queue.
How FISTA Solutions helps
FISTA Solutions assesses whether a gateway is justified at current scale, implements routing and caching where the savings are, builds cost attribution by feature and team, and treats the gateway as request-path infrastructure with the reliability engineering that requires, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To find out where your AI spend goes and whether a gateway would pay, message FISTA on WhatsApp, or read what is token accounting.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What does a gateway actually buy?
A single point for model routing, prompt and semantic caching, rate limiting, cost attribution by team and feature, policy enforcement, and provider failover. Each is achievable per application; the gateway makes them consistent and centrally managed.
02Where do the savings come from?
Routing simple requests to smaller models and caching repeated prefixes and equivalent queries. Those two typically dominate, and both are available without a gateway — the gateway makes them apply everywhere rather than in whichever application implemented them.
03Why is attribution underrated?
Because without it AI spend arrives as a single provider invoice that answers no question. A gateway that tags every request with feature, team, and customer turns spend into something manageable rather than merely observable.
04What does it cost to operate?
Build or licensing, plus availability responsibility. It sits in the request path, which makes it a single point of failure requiring the reliability engineering that implies, and it needs updating as providers change their interfaces.
05At what scale is it justified?
When several applications call models, spend is large enough that attribution matters, and routing or caching would produce meaningful savings. One application with modest spend does not need one.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.