Leadership · 4 minute read
Model Routing Explained for Executives
Model routing sends each task to the cheapest model that meets its quality threshold, with escalation to stronger models when confidence is low. It is the most effective cost control in most AI programs, it requires evaluation to set the thresholds, and it needs governance so routing changes are deliberate rather than silent.
Most AI programs pay frontier prices for work that does not need frontier capability, because the first prototype used the best available model and nobody revisited the choice. Model routing fixes that, and it is usually the largest single cost saving available. This explainer covers what it is, what it requires, and how to govern it.
What is model routing?
A rule that decides which model handles each request. A simple routing policy might send document classification to a small fast model, customer-facing answers to a mid-tier model, and complex analysis to a reasoning model, with escalation when the cheaper option signals it is struggling. The rule lives at a gateway, so agents request a capability rather than naming a provider.
The reasoning models explained for executives and small language models explained for executives pieces cover the two ends of the range being routed between.
Why does it matter so much for cost?
Because AI spend is dominated by volume on routine tasks.
| Task | Volume | Capability needed | Typical mistake |
|---|---|---|---|
| Document classification | Very high | Low | Frontier model at scale |
| Data extraction from standard forms | Very high | Low to moderate | Frontier model at scale |
| Routine customer answers with retrieval | High | Moderate | Frontier model |
| Complex analysis and planning | Low | High | Correctly uses strong model |
| Code generation on complex changes | Low to moderate | High | Correctly uses strong model |
The rows with very high volume and low capability requirements are where the money goes, and they are the easiest to fix. Companies implementing routing for the first time commonly find a substantial share of spend sitting in the top two rows. The AI agent unit economics whitepaper covers the cost model.
What does routing require?
Evaluation, first. You cannot route safely without knowing how each candidate model performs on each task type. The required numbers are pass rate, cost per task, and latency, measured on the company's own cases. Routing without evaluation is guesswork that degrades quality somewhere nobody is watching.
Fortunately the evaluation sets that support routing are the same ones that gate releases, so the investment serves twice. The AI evaluation explained for executives piece covers building them.
A gateway. Routing implemented inside each agent is unmanageable; implemented at a gateway it is central, auditable, and changeable without touching the agents. The AI gateways explained for executives piece covers the layer.
How does escalation work?
Route cheap first, escalate when needed. The cheaper model attempts the task; if it signals low confidence, fails a validation check, or the input matches an escalation rule, the request is retried on a stronger model. This preserves quality on the difficult tail while keeping the common case cheap.
The design decision is the escalation trigger: too sensitive and everything escalates, eliminating the saving; too insensitive and hard cases get weak answers. Tune it against the evaluation set and monitor the escalation rate in production.
Why is routing a governance artifact?
Because a routing change alters quality and cost simultaneously, often without anyone noticing. A well-intentioned change to save money can degrade customer-facing answers; a change to improve quality can multiply spend. Three controls:
- Business owners approve quality thresholds, since they own the outcome.
- Routing changes are logged and visible, not buried in configuration.
- The monthly review covers routing: current rules, escalation rates, cost per task, and any changes.
Silent routing changes are a common cause of the unexplained quality shifts that erode trust in an AI program. The AI operating rhythm for leadership teams guide describes where this fits in the review.
What else does routing enable?
Resilience. A gateway that can route between providers can fail over when one has an outage, which is otherwise a single point of failure for every agent.
Migration. When a provider deprecates a model or changes pricing, routing makes the switch a configuration change rather than a project, provided the evaluation set exists to validate the alternative. The how to design a model routing strategy guide covers the implementation.
Negotiating position. A company that can move volume between providers negotiates differently from one that cannot.
What should executives ask?
- What share of our model spend is on high-volume routine tasks, and what model handles them?
- Do we have pass rates per task type per model, or are we guessing?
- Is routing implemented at a gateway or inside individual agents?
- What is our escalation rate, and has it changed?
- Who approved the current quality thresholds?
How can FISTA Solutions help?
FISTA Solutions benchmarks models on clients' own cases, implements routing at a governed gateway with escalation and failover, and builds AI agents that request capability rather than naming providers, through its AI enablement practice. Since 2017, FISTA has delivered 150+ projects for 50+ companies across 12+ countries.
To find out what routing would save on your current workload, talk to FISTA on WhatsApp, or read the multi-model strategy whitepaper.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is model routing?
A rule that decides which model handles each request, typically sending routine tasks to cheaper, faster models and hard ones to stronger models, with escalation when the cheaper model signals low confidence or fails a check. It is implemented at a gateway so agents do not choose models themselves.
02Why is routing the main AI cost control?
Because cost is driven by volume on routine work, not by the occasional hard problem. Running high-volume classification or extraction on a frontier model can cost many times what a smaller model would, for an identical outcome. Routing captures that difference without reducing quality where quality matters.
03What do you need before implementing routing?
Evaluation per task type across candidate models: pass rate, cost per task, and latency. Without those numbers, routing is guesswork and will degrade quality somewhere invisible. The evaluation sets that support routing are the same ones that gate releases.
04How does escalation work in model routing?
The cheaper model attempts the task; if it signals low confidence, fails a validation check, or the case matches an escalation rule, the request is retried on a stronger model. Escalation preserves quality on the hard tail while keeping the common case cheap, at the cost of some added latency on escalated cases.
05Who should own routing rules?
Engineering implements them; the business owner of each agent approves the quality thresholds, because a routing change is a quality and cost decision. Changes should be logged and reviewed in the monthly review, since silent routing changes are a common cause of unexplained quality shifts.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.