FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Web & Mobile ┬╖ 5 minute read

API Gateway: What It Should Do and What It Should Not

An API gateway should handle routing, authentication, rate limiting, and observability at the edge. Business logic, consumer-specific data transformation, and cross-service orchestration belong behind it, because a gateway that accumulates them becomes a monolith with a deployment bottleneck every team has to queue behind.

By FISTA Solutions┬╖ AI-Native Engineering Team┬╖
API Gateway: What It Should Do and What It Should Not article cover

API gateways solve genuine problems тАФ routing, authentication, rate limiting, observability тАФ and then accumulate responsibilities that do not belong at the edge. This guide covers where the line sits, drawing on FISTA Solutions' web and mobile work.

What belongs at the edge?

Cross-cutting concerns that every service would otherwise implement separately.

ConcernGateway or service?
TLS terminationGateway
RoutingGateway
AuthenticationGateway
Rate limitingGateway
Request logging and tracingGateway
AuthorisationUsually the service
Business logicService, always
Consumer-specific shapingBehind the gateway

Why keep business logic out?

Because logic in the gateway is logic every team must coordinate to change.

The pattern is gradual: a small transformation for one consumer, then a special case, then a rule about which requests go where. Within a year the gateway contains behaviour nobody can safely modify and every deployment is coordinated.

That reintroduces exactly the bottleneck a service architecture exists to remove. Keep it thin deliberately, and push back on the first special case rather than the tenth.

Where should authorisation live?

Behind the gateway, in the service that owns the resource.

The gateway can establish identity тАФ who is calling тАФ because that is uniform. Deciding what they may do requires knowing the resource and its rules, which the edge does not have without duplicating service logic.

The exception is coarse-grained access: rejecting a caller with no valid credentials, or blocking an entire route for a consumer class. That is cheap at the edge and does not require resource knowledge.

How should rate limiting be designed?

Per consumer and per endpoint, with the limits and remaining quota communicated in response headers.

Global limits mean one heavy consumer degrades service for everyone, which is the failure rate limiting exists to prevent. Per-consumer limits contain the impact.

Differentiate by endpoint cost too. An expensive aggregation endpoint and a cheap lookup should not share a limit, because a consumer hitting the expensive one heavily is a different problem.

What about consumer-specific response shaping?

It belongs behind the gateway, usually in a dedicated layer per consumer type.

Mobile clients and web clients frequently want different shapes of the same data. Handling that in the gateway produces consumer-specific logic at the edge; handling it in a layer owned by the team serving that consumer keeps ownership clear.

That is the backend-for-frontend pattern, and it works better than gateway transformations because the logic sits with the people who need it. See GraphQL federation guide.

What observability should it provide?

Request and response logging with correlation identifiers, latency per route and per consumer, error rates by status class, and traces that propagate downstream.

This is the cheapest observability available, because it is one place covering every service. A gateway that logs only aggregate traffic wastes that position.

Per-consumer breakdown matters. Aggregate error rates hide the case where one integrator is failing constantly while everyone else is fine.

How do you avoid the gateway becoming a bottleneck?

By making configuration declarative and owned by the service teams who need it, reviewed rather than implemented by a central team.

Gateways requiring a central team to make every routing change become a queue, and the queue becomes an argument for teams routing around the gateway entirely.

Self-service with guardrails is the workable arrangement: teams declare their routes and limits, the platform validates and applies them.

What are the common mistakes?

Business logic at the edge. Global rate limits. Authorisation duplicated in the gateway. Central team required for every change. And aggregate-only observability.

How do you test it?

Test routing, authentication rejection, rate limit behaviour at the boundary, and what happens when a downstream service is slow or unavailable.

The last is the one most often untested and the one that determines whether a single slow service degrades everything behind the gateway.

What does it cost to operate?

Modest for a managed gateway, higher for a self-operated one including the operational burden. The larger cost is the coordination overhead if it becomes a bottleneck.

At high volume the per-request cost of a managed gateway becomes material and worth modelling against the operational cost of running one yourself.

What should you measure?

Latency added by the gateway, error rates by route and consumer, rate limit rejections, and time from a team requesting a routing change to it being live.

How does this apply to AI services?

Directly, and with an extra benefit: a gateway in front of model provider calls gives you authentication, cost attribution per team, rate limiting, and logging in one place.

That is the practical form of the model gateway that governance and cost programmes both depend on. Building it once saves every team implementing partial versions. See LLM gateway cost.

When is this the wrong approach?

For a single service with one consumer, a gateway is infrastructure with no consumers of its own. It earns its cost with several services, several consumer types, or a need for uniform authentication and observability.

What should you do first?

List what your gateway currently does and mark which items are business logic. Anything in that column is a candidate for moving behind it.

How FISTA Solutions helps

FISTA Solutions builds and operates production systems through web and mobile, AI enablement, and staff augmentation: gateways kept thin with business logic behind them, per-consumer rate limiting and observability configured so one integrator cannot degrade the rest, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To scope this work, message FISTA on WhatsApp, or read authentication architecture guide.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What belongs in a gateway?

Routing, TLS termination, authentication, rate limiting, request and response logging, and basic request validation. Those are cross-cutting concerns every service would otherwise implement separately and inconsistently.

02What does not belong there?

Business logic, consumer-specific response shaping, and orchestration across services. Those accumulate until the gateway is a monolith every team must change, which reintroduces the bottleneck the architecture was meant to remove.

03Where should authorisation happen?

Usually behind the gateway, in the service that owns the resource. The gateway can establish who the caller is; deciding what they may do requires knowledge of the resource that the edge does not have.

04How should rate limiting work?

Per consumer and per endpoint rather than globally, with limits communicated in responses. Global limits mean one heavy consumer degrades everyone, which is the failure mode rate limiting exists to prevent.

05How do you keep it maintainable?

By keeping it thin and by making configuration declarative and owned by service teams rather than by a central gateway team. Gateways requiring a central team's involvement for every change become a queue.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.

Start a project