Trends ¡ 5 minute read
The Quiet Rise of Small Models in Production Systems
Small specialised models handle more production work than the public discussion suggests. Most tasks in a real system are narrow â classification, extraction, routing, formatting â and a small model tuned for them is faster, cheaper, and more predictable than a frontier model doing the same job.
Most production tasks are narrow, and small models keep appearing in real systems for reasons that rarely feature in the public discussion. This piece covers where they fit, drawing on FISTA Solutions' AI agents engineering work.
Where do small models fit?
The tasks where scope is genuinely narrow.
| Task | Why a small model suits it |
|---|---|
| Intent classification | Constrained output space |
| Field extraction | Clear definition of correct |
| Routing decisions | High volume, low complexity |
| Reformatting | Mechanical transformation |
| Safety screening | Fast, frequent, binary |
| Short summarisation | Bounded input and output |
Why does cost favour them so heavily?
Because per-call cost differs by an order of magnitude or more.
A classification running on every incoming message is a high-volume operation. Performed by a frontier model it can dominate an AI budget; performed by a small tuned model it is close to negligible.
That difference is what makes previously uneconomic automation viable. The question is not whether the large model is better â it usually is, marginally â but whether the margin justifies the multiple. See how to reduce AI costs.
What does latency buy?
Interactive features that would otherwise feel slow.
A small model responding in tens of milliseconds can sit in an interaction loop â suggesting as someone types, validating as they enter data, routing as a message arrives. A frontier model taking seconds cannot.
That opens a category of features rather than merely making existing ones cheaper, which is the more interesting consequence. See how to optimize AI latency.
Why is predictability better?
Because a narrow model on a narrow task has less room to do something unexpected.
A general model asked to classify may occasionally explain its reasoning, refuse, or produce a category nobody defined. A small model fine-tuned to emit one of eight labels emits one of eight labels.
For production systems, that constraint is a feature. It is also why the deterministic-scaffolding argument and the small-model argument point the same way. See the return of determinism.
How should routing work?
By task, with evaluation determining the assignment.
Run your evaluation suite against both a small and a large model. Tasks where the small model is adequate move to it; tasks where it is not stay. Re-run periodically, because both models change.
A fallback path matters: if the small model's confidence is low or its output fails validation, escalate to the larger one. That keeps the quality ceiling while capturing most of the saving. See how to build an agent evaluation harness.
When is fine-tuning worth it?
When the task is stable, narrow, and you have examples.
A few thousand well-labelled examples of a specific task frequently produce a small model that outperforms a much larger general one on that task. The training cost is modest and the inference saving is permanent.
The cost is maintenance: a tuned model is a versioned artefact needing retraining as the task drifts, plus evaluation to detect that drift. Teams should be honest about whether they will do that.
What does on-device enable?
Features that cannot send data anywhere.
Local inference removes network latency, works offline, incurs no per-call cost, and keeps data on the device. For privacy-sensitive applications that last property is decisive rather than convenient.
The constraints are real: model size, memory, battery, and the difficulty of updating a model already distributed. It suits narrow, stable tasks. See mobile app architecture guide.
What is the counter-argument?
The counter is that model routing adds complexity for a saving that may not matter at low volume, and that frontier model prices keep falling anyway. Both are fair. The threshold is volume: below a certain call rate, use one good model and spend the engineering time elsewhere.
What does this change for engineering teams?
It means the architecture has several models rather than one, with routing logic and per-route evaluation. That is more moving parts and it needs the abstraction layer to be real.
It also means fine-tuning becomes an operational capability: datasets, training runs, versioning, and drift detection.
What does this change for buyers?
It means asking vendors what runs where, because a vendor routing everything to a frontier model has a cost structure that will show up in your pricing.
And asking about on-device options where data residency matters.
What should leaders do about it now?
Look at your highest-volume AI operation and check which model serves it. That single route is usually where most of the spend is and where a small model would be adequate.
Then require evaluation before any routing change, because routing without measurement is guessing.
How does this apply to agents?
Directly. An agent's trajectory contains many routine steps â deciding which tool to call, formatting arguments, checking a result â that do not need frontier reasoning.
Routing those to a small model while reserving the large one for genuine decisions reduces both cost and latency substantially, which matters because agents make many calls per task. See AI agent production readiness checklist.
How will you know if this is happening?
Watch for one high-volume operation dominating model spend, for latency complaints on interactive features, and for tasks where output variance is causing downstream problems. Each is a small-model candidate.
How FISTA Solutions reads this
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: model routing decided by evaluation against real cases rather than by preference, with validation and escalation so the quality ceiling is preserved, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To discuss what this means for your roadmap, message FISTA on WhatsApp, or read how to reduce AI costs.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why use a small model at all?
Cost, latency, and predictability. For a narrow task, a small tuned model frequently matches a frontier model's accuracy while costing a fraction and responding several times faster.
02Which tasks suit them?
Classification, extraction, routing, formatting, and short structured generation â anything with a constrained output space and a clear definition of correct.
03How do you decide what goes where?
With evaluation. Run both models against your actual cases and see where the small one is adequate. That measurement is the whole method and it is why teams without evaluation cannot do this.
04Does fine-tuning help?
Substantially, for narrow tasks with enough examples. A small model tuned on a few thousand good examples often outperforms a much larger general model on that specific task.
05What about on-device?
Small models make local inference viable, which removes network latency, keeps data on the device, and eliminates per-call cost. That suits privacy-sensitive and latency-sensitive features.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.