Comparison ¡ 5 minute read
LLM Gateway Comparison: What a Proxy Layer Should Give You
A gateway centralises what every AI application ends up needing: routing between providers, cost attribution, rate limiting, caching, and failover. Compare candidates on attribution granularity, failover behaviour under real conditions, added latency at the tail, and whether they become a single point of failure you cannot bypass.
A gateway centralises what every AI application needs and introduces a dependency every call passes through. This guide covers comparing them, drawing on FISTA Solutions' AI enablement operational work.
What should a gateway provide?
Six capabilities, in order of practical value.
| Capability | What to look for | Why it matters |
|---|---|---|
| Cost attribution | Per team, app, and user | The most-used feature |
| Routing | Rules by task and model | Cost and quality control |
| Failover | Automatic, fast, sensible fallback | Provider outages happen |
| Rate limiting | Per consumer, not global | Stops one app starving others |
| Caching | Prefix and semantic | Direct cost saving |
| Logging | Structured, exportable | Audit and debugging |
Why is attribution the main benefit?
Because without it nobody knows where the spend goes.
A gateway sitting on every call can tag by application, team, feature, and user, which turns an opaque provider invoice into a breakdown you can act on. Most organisations discover surprises the first time they see it.
Check the granularity available and whether it requires callers to supply tags. A gateway that attributes only by API key gives you less than one that accepts arbitrary dimensions. See LLM cost control checklist.
What should routing support?
Rules by task, with fallback and the ability to change without a deployment.
Routing simple operations to cheaper models is the largest cost lever available, and a gateway is the natural place to implement it. The useful capability is changing routes through configuration rather than code.
Check whether routing can consider request properties, whether it supports weighted rollout of a new model, and whether route changes are auditable. See the quiet rise of small models.
How should failover behave?
Detect quickly, fall back sensibly, and recover automatically.
Detection speed matters: a gateway that waits for a long timeout on every request during an outage makes things worse. Fallback quality matters too â a cheaper model that produces unusable output is not a fallback.
Test it by blocking a provider and observing. Many gateways handle a clean failure well and handle slow degradation badly, which is the more common real case. See what is a fallback chain.
What is the single-point-of-failure risk?
Real, and worth designing around.
Every model call passing through one component means its availability becomes your availability. A self-hosted gateway needs the same reliability engineering as any critical service; a managed one adds a dependency.
Keep a bypass path: applications able to call the provider directly if the gateway is unavailable, even in a degraded mode. That single design decision removes most of the risk.
When do you not need one?
One application, one provider, modest volume.
At that scale a gateway adds a hop and a dependency for capabilities you can implement in a few lines. Attribution, retries, and a timeout are not hard when there is one consumer.
The threshold is several applications, several providers, or a need to attribute cost across teams. Below that, the layer is premature. See the consolidation of AI tooling.
What about lock-in?
Moderate, and worth checking before adopting.
Routing rules, cost attribution history, and cached data all live in the gateway. Check whether configuration exports in a usable form and whether your applications' interface to the gateway resembles the provider's own.
A gateway presenting a provider-compatible interface is easy to remove; one with its own request format means changing every caller. See AI vendor offboarding checklist.
How do you run your own comparison?
Put it in front of a real workload and measure added latency at the tail. Then block a provider and observe failover behaviour, including how long detection takes.
Check whether the attribution it produces answers the cost questions you actually have. A breakdown by API key when you need one by feature is not useful.
What does switching cost later?
Depends on the interface. A gateway presenting a provider-compatible API can be removed by changing a base URL. One with a bespoke request format requires touching every caller.
Prefer the former, and keep routing configuration in version control rather than only in the gateway's own store.
What do people get wrong here?
Adopting one for a single application. No bypass path. Attribution granularity that does not match your questions. Failover untested. And a bespoke request format that makes the gateway hard to remove.
Build or buy the gateway itself?
For basic needs â credentials, retries, attribution â a thin internal layer is straightforward and gives full control.
Buying earns its place when you want caching, sophisticated routing, failover, and a management interface without building them. The decision is the same build-versus-buy question as everywhere else. See build vs buy AI agents.
Which should you choose?
Add a gateway when several applications or several providers are involved, or when cost attribution across teams matters. Compare on attribution granularity, failover behaviour, and added latency, and keep a bypass path so it does not become a single point of failure.
What should you do first?
Check whether you can currently break down model spend by feature. If not, that alone may justify the layer.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: gateways assessed on attribution granularity and tested failover behaviour, with a bypass path so the layer does not become a single point of failure, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To run this comparison against your own workload, message FISTA on WhatsApp, or read LLM cost control checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What does a gateway actually give you?
One place for provider credentials, routing rules, cost attribution, rate limits, caching, and failover â instead of each application implementing its own version inconsistently.
02When is it worth adding?
When several applications call models, when you route between providers, or when you need cost attribution across teams. A single application talking to one provider gains little.
03What should you check about failover?
Whether it detects provider failure quickly, whether the fallback model produces acceptable output, and whether failover is automatic or requires intervention. Behaviour varies substantially.
04What is the risk?
Becoming a single point of failure. Every model call routes through it, so its availability becomes your availability, and a bypass path is worth having.
05Does it add latency?
Some, and it should be measured rather than assumed negligible. A well-implemented gateway adds a few milliseconds; a poorly placed one adds a network hop across regions.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.