Glossary · 4 minute read
What Is Model Serving? Running Models in Production Explained
Model serving is the infrastructure layer that turns a model into a production service: handling requests, batching them for efficiency, managing memory and hardware, scaling with load, versioning deployments, and exposing metrics. Hosted APIs provide it as a service; self-hosting means operating it yourself, with the control and the burden that implies.
Model serving is where a working model becomes a dependable service, and it is the part most underestimated in project planning. The gap between a model producing correct output in a notebook and one serving traffic within a latency budget at acceptable cost is substantial. This explainer covers what a serving layer does and the decisions around it. It complements what is inference optimization and gpu cost for ai, and reflects FISTA Solutions' approach in AI enablement delivery.
What does a serving layer handle?
Request intake and queuing. Batching, so that the hardware processes several requests together rather than idling between them. Memory management for weights and attention caches. Backpressure when demand exceeds capacity. Routing between model versions. Streaming responses. And metrics detailed enough to diagnose problems.
Each of these is straightforward alone and non-trivial in combination, which is why dedicated serving frameworks exist rather than everyone writing a wrapper around a model call.
| Concern | Hosted API | Self-hosted |
|---|---|---|
| Operational burden | Provider's | Yours |
| Cost model | Per token | Per hour of capacity |
| Data residency control | Limited | Full |
| Model choice | Provider catalogue | Any open-weight model |
| Scaling | Automatic | Your capacity planning |
| Version control | Pinning where offered | Complete |
When does self-hosting make sense?
When data residency or sector regulation requires it. When sustained volume makes reserved capacity cheaper than per-token pricing. When a specific open-weight or fine-tuned model is needed that no provider offers. When latency requires the model close to the application.
Absent one of those, hosted APIs are the correct default. The operational cost of running inference infrastructure well — capacity, upgrades, monitoring, on-call — is routinely underestimated, and a preference for control does not by itself justify it.
Why are cold starts different here?
Because model weights are large. Loading tens of gigabytes onto accelerators takes minutes, not the milliseconds a container start takes. Scale-to-zero, which is ordinary practice for stateless services, means the first request after an idle period waits for that load.
The practical consequence is that serving infrastructure keeps warm capacity, and warm capacity costs money whether or not it is serving traffic. Utilisation, not peak capacity, is the number that determines whether self-hosting is economical.
What makes GPU autoscaling hard?
Acquisition latency, regional availability, warm-up time, and unit cost. A conventional service scales in seconds on abundant commodity instances; accelerator capacity may take minutes to acquire and may not be available at all in a given region at a given moment.
Scaling therefore reacts on a timescale slower than traffic changes. Capacity planning — provisioning for expected peak with a buffer — matters more than elasticity, which inverts the intuition most teams bring from web infrastructure.
How should versioning work?
By pinning explicitly and upgrading deliberately. Provider-side model updates change behaviour, sometimes subtly, and a system whose evaluation was run against one version has no guarantee against the next.
The pattern that works: pin the version, run the evaluation suite against a candidate before switching, deploy the change like a code change, and keep the ability to roll back. Treating a model version as configuration rather than as a given is the core discipline. See ai evaluation checklist.
What should be monitored?
Request latency including time to first token, which matters more than total time for streaming interfaces. Tokens per second. Queue depth and wait time. Batch sizes achieved. Memory utilisation. Error and timeout rates by cause. Cost per request.
Request-level latency alone hides the common failure where a system is nominally healthy while queue time has quietly become the dominant term.
How does batching affect latency?
It trades individual latency for aggregate throughput. Larger batches use hardware more efficiently and make each request wait longer for the batch to form. Continuous batching, where requests join an in-flight batch rather than waiting for a new one, substantially improves that trade and is standard in modern serving stacks.
The configuration should follow the workload: interactive interfaces want small batches and low wait, bulk processing wants the opposite.
What about multi-model serving?
Common and awkward. Serving several models on shared hardware saves cost when each is individually under-utilised, but memory is the binding constraint and swapping weights is slow. Approaches that share a base model across multiple fine-tuned adapters are considerably more efficient than running separate full models, where the use case allows it.
How does this connect to routing?
Directly. A serving layer that can route between a small fast model and a large capable one, based on task complexity, is often the largest available cost lever. The routing decision belongs above the serving layer, but the serving layer must support both models being available with predictable latency for it to work.
How FISTA Solutions helps
FISTA Solutions makes hosted-versus-self-hosted decisions on evidence about volume, residency, and latency, plans accelerator capacity rather than relying on elasticity, pins and evaluates model versions before upgrades, configures batching to the workload, and monitors token-level metrics, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To put a model into production reliably, message FISTA on WhatsApp, or read ai evaluation checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What does a serving layer actually do?
Receives requests, batches them to use hardware efficiently, manages model weights and attention caches in memory, handles queuing and backpressure, routes between model versions, and exposes metrics. It is the difference between a model that runs and a model that serves traffic reliably.
02When is self-hosting justified?
When data residency or regulation requires it, when volume makes the economics favourable at sustained scale, when a specific fine-tuned or open-weight model is needed, or when latency demands proximity. Preference for control alone rarely covers the operational cost.
03Why are cold starts a problem?
Because loading large model weights onto accelerators takes minutes, not milliseconds. Scale-to-zero, which is routine for stateless services, means the first request after idle waits for a full load, so serving usually keeps warm capacity and pays for it.
04What makes GPU autoscaling hard?
Slow instance acquisition, limited availability in some regions, long warm-up times, and high per-unit cost. Scaling reacts in minutes while traffic changes in seconds, so capacity planning matters far more than elasticity does for conventional stateless workloads.
05Why pin model versions?
Because provider-side model updates change behaviour, and a system evaluated against one version is not guaranteed against the next. Pinning plus a deliberate re-evaluation before moving is the difference between a controlled upgrade and an unexplained regression.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.