Whitepaper ¡ 9 minute read
Model Routing and Cost Control: An Enterprise Whitepaper
Enterprise AI cost is controlled at the routing layer, not in application code. A gateway that routes each task to the right model tier, enforces context budgets, caches stable prefixes and repeated results, attributes spend to teams and use cases, and applies budgets with alerts turns unpredictable inference spend into a managed line with quality evidence per route.
AI spend behaves differently from other infrastructure spend. It is variable per request, sensitive to changes nobody flags as cost-relevant such as a longer prompt or a broader retrieval setting, and it arrives as a single provider invoice that resists attribution. Organisations that let applications call providers directly discover this at quarter end, and their first instinct, restricting usage, trades value for savings. The better answer is architectural: a routing and cost control layer that makes spend visible, attributable, and adjustable without changing application code. This whitepaper sets out that layer. It draws on FISTA Solutions' AI enablement practice and complements the LLM gateway architecture whitepaper and the AI total cost of ownership model whitepaper.
Where does enterprise AI spend actually go?
| Driver | Share of typical spend | Controllable by |
|---|---|---|
| Context tokens | Usually the majority | Retrieval selection, history bounds, compression |
| Model tier | Large multiplier | Routing policy per task |
| Output tokens | Moderate | Output schemas, length limits |
| Retries and failures | Underestimated | Validation, idempotency, backoff design |
| Agent step count | Grows fast in multi-step systems | Step bounds, state summarisation |
| Duplicate work | Invisible without caching | Prefix and semantic caching |
| Non-production usage | Often unbudgeted | Environment attribution, quotas |
The pattern that surprises teams is that the model choice, which is what everyone debates, is usually less consequential than the amount of context sent to it. A workload on an expensive model with tight context often costs less than the same workload on a cheap model with an unbounded window.
What does a routing policy contain?
A named route per task type, each specifying the model tier, fallback and escalation behaviour, context budget, output constraints, caching policy, and the evaluation set that governs changes to it. Routes are versioned and deployable independently of application code.
Routing inputs that matter in practice: task type, which is the primary determinant; input size and complexity; data classification, since some data may only go to specific models or deployments; tenant or customer tier where service levels differ; and current cost posture, allowing a controlled degradation to cheaper tiers under budget pressure rather than an outage.
The property that makes this valuable is that changing which model serves a task becomes a configuration change with evaluation evidence, rather than a code change across several applications. Routing mechanics are in what is an llm router.
How is escalation designed?
So that cheap paths handle the common case and expensive paths handle the hard one. A small model attempts the task; validation checks the output against a schema and business rules; a confidence or abstention signal indicates uncertainty; and failures escalate to a larger model, with the escalation logged so the rate is visible.
Two disciplines keep this honest. Bound the escalation chain, so a pathological input cannot traverse every tier at full cost. And monitor escalation rate as a quality signal: a rising rate means the small model's task has drifted, which is useful information that a single-tier architecture never surfaces.
Why is context discipline the largest lever?
Because tokens processed dominate cost and most requests carry material the task does not use. Retrieval configured to return ten passages when three suffice triples that portion of every request forever. Conversation history carried in full rather than summarised grows unbounded. Structured records passed whole when the task reads four fields waste the rest.
Tightening these usually improves quality at the same time, because irrelevant context degrades results, which makes context work the rare optimisation with no trade-off. The measurement that drives it is tokens per request broken down by context component, which most teams have never looked at. See the context engineering whitepaper and ai cost optimization checklist.
How much does caching save, and where does it break?
Prefix caching, where a stable leading block of instructions and shared context is reused across requests, saves meaningfully on workloads with long fixed preambles, which describes most enterprise assistants. It requires stable ordering, which is a reason to assemble context deterministically.
Semantic caching, where an equivalent prior question returns a stored answer, saves more but carries risk: the answer may be stale, or the new question may differ in a way that matters despite similarity. Safe practice scopes semantic caching to stable content domains, sets short time-to-live, keys on the user's entitlements so cached answers cannot cross permission boundaries, and excludes anything account-specific or time-sensitive. See what is prompt caching and what is a semantic cache.
How should cost be attributed?
At the gateway, on every call, with tags for team, use case, route, environment, and where relevant customer or tenant. Attribution reconstructed later from provider invoices is approximate and arrives too late to change behaviour.
Attribution enables the three things that actually control spend: budgets per team and use case with alerts before thresholds rather than after; unit economics per use case, meaning cost per ticket resolved or document processed, which is the number that justifies or kills a use case; and accountability, since a cost line with no owner is never reduced. Unit economics are in the AI agent unit economics whitepaper.
What should budgets and limits do?
Degrade gracefully rather than fail. A workload approaching its budget should route to cheaper tiers, reduce optional context, or queue non-urgent work before it stops serving. Hard caps belong on non-production environments and on individual request cost, to contain runaway loops.
Rate limiting deserves specific attention in agent systems, where a single triggering event can fan out into hundreds of model calls. Per-workflow concurrency limits and step bounds prevent a bug from becoming an invoice. See what is rate limiting for llm apis.
How is quality protected while cost is reduced?
By treating every cost change as a behaviour change. Routing a task to a cheaper model, tightening retrieval, shortening history, or enabling semantic caching all change outputs, and each must pass the route's evaluation set before deployment and be compared in canary against the incumbent.
The failure this prevents is common and expensive: a cost programme reports substantial savings, quality degrades gradually across several routes, and the business effect appears months later as rising escalations or falling conversion that nobody connects to the optimisation. Evaluation practice is in the AI evaluation and testing whitepaper.
What does the gateway need to enforce?
Model access by data classification, so regulated data cannot reach an unapproved model. Route resolution and version pinning, so provider updates are scheduled rather than surprising. Context budget limits. Output constraints. Caching policy with entitlement-aware keys. Attribution tagging. Budget and rate policy. Full logging of prompts, responses, versions, and cost, joined to a request identifier. And a kill switch per route.
That list is the argument against direct provider calls: none of it can be enforced in scattered application code, and all of it is required to run AI as a managed service rather than an expanding liability.
What does a cost reduction programme look like?
- Instrument first. Route every call through the gateway with attribution; change nothing else for two weeks and observe.
- Find the top five workloads by spend and break each into context, output, retries, and step count.
- Fix context on those workloads, measuring quality before and after.
- Evaluate tier changes per route with reference sets and canary deployment.
- Enable caching where content is stable and entitlements allow.
- Set budgets and unit economics per use case, with owners.
- Review monthly with both cost and quality in the same report.
Teams that run steps three to five without step one cannot tell whether savings came from the change or from a traffic shift.
What goes wrong?
Applications calling providers directly, which makes everything above impossible. Cost programmes with no quality measurement. Semantic caching keyed without entitlements, which leaks between users. Budget caps that fail workloads outright instead of degrading. Model pinning skipped, so provider updates change both cost and behaviour unannounced. And optimisation of the model tier while ignoring context, which is usually the larger number.
How does this interact with provider commitments?
Most enterprises eventually negotiate committed spend with one or more providers in exchange for discounts, and the routing layer changes that negotiation in the organisation's favour. With attribution in place, a buyer knows its workload mix, its growth curve, and how much of its volume is genuinely portable between providers, which are the three facts that determine leverage.
The architectural consequence is worth stating plainly: routes that can move between providers are worth more than routes that cannot. Keeping prompts, evaluation sets, and retrieval content in the organisation's own environment, and treating provider-specific features as route-level choices rather than application-level dependencies, preserves that portability. Teams that build deeply against one provider's proprietary orchestration features find at renewal that their negotiating position consists of asking politely.
Commitments also interact with routing policy in an unhelpful way if nobody notices: a discount tied to one provider can make the cost-optimal route for a task different from the quality-optimal one. That trade-off belongs in an explicit policy decision with the numbers visible, not in an engineer's judgement call inside a pull request.
What does a monthly cost and quality review look like?
One report, both dimensions. Spend by team, use case, and route against budget. Unit economics per use case, meaning cost per resolved ticket, processed document, or completed task, trended. Quality scores per route from the reference sets, with any regressions flagged. Escalation rates. Cache hit rates. And a short list of the largest workloads with their context breakdown.
Reviewing cost without quality produces optimisation that quietly degrades the product. Reviewing quality without cost produces systems nobody can afford to scale. The organisations that manage AI well put both on the same page in front of the same people.
How FISTA Solutions delivers this
FISTA Solutions builds the gateway and routing layer that makes AI spend attributable and adjustable, runs the per-route evaluations that keep cost reductions from becoming quality losses, and installs the budgets, caching, and unit economics reporting that keep it managed, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime and 47% efficiency gains where measured.
To bring AI spend under control without cutting value, message FISTA on WhatsApp, or read the LLM gateway architecture whitepaper.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why does AI cost become unmanageable?
Because applications call providers directly with no shared layer, so spend cannot be attributed to a team or use case, context grows unchecked, every workload uses whichever model was chosen at build time, and nobody sees the trend until the invoice arrives. The fix is architectural, not behavioural.
02What is the biggest lever on inference cost?
Context size. Most enterprise requests carry far more context than the task requires, and tokens processed dominate cost. Tightening retrieval selection, bounding history, and compressing what need not be verbatim typically cuts spend substantially while improving quality.
03How does caching help, and what are its limits?
Prefix caching avoids reprocessing stable instruction blocks and shared context on every request, and semantic caching returns a stored answer for repeated equivalent questions. The limits are freshness, since cached answers age, and correctness, since near-duplicate questions can differ in ways that matter.
04How should AI spend be attributed?
By team, use case, environment, and where relevant customer or tenant, captured at the gateway on every call rather than reconstructed from provider invoices. Without attribution, budgets cannot be set or enforced and no one owns a cost reduction.
05How do you avoid cutting quality while cutting cost?
Measure quality per route with a reference set and treat a routing change like any other behaviour change: evaluate, canary, compare, and revert if quality regresses. Cost reductions without per-route evaluation are quality changes nobody measured.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.