FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook ┬╖ 6 minute read

How to Run an AI Cost Reduction Program That Holds

Reducing AI running cost works when you measure cost per completed task rather than per token, find the small number of patterns that dominate spend, route easy work to cheaper paths, and assign ownership so the savings do not erode as features are added.

By FISTA Solutions┬╖ AI-Native Engineering Team┬╖
How to Run an AI Cost Reduction Program That Holds article cover

AI running costs drift upward quietly because nobody owns the number and every individual change is small. This playbook covers reducing them and, harder, holding the reduction, drawing on FISTA Solutions' AI enablement work.

When is this worth doing?

When AI spend has become material enough to notice, when unit economics are marginal, or when a system that worked at pilot volume is becoming expensive at production volume.

It is not worth doing on a system still being built. Optimising cost before the behaviour is settled produces work that gets undone, and early systems usually have quality problems worth more than the spend.

What does the sequence look like?

StepPurpose
1. Instrument per taskCost attributed to completed outcomes
2. Find the dominant patternsA few request types usually dominate
3. Cut context firstThe most common avoidable cost
4. Route by difficultyCheap path for the easy majority
5. Cache where inputs repeatExact first, semantic with care
6. Assign ownershipOr it erodes within two quarters

Step 1 тАФ Instrument cost per completed task

Attribute spend to outcomes rather than to calls. A task that took four model calls and two retries cost the sum of all of them, and that total is the number that matters.

Most systems cannot do this initially, because logging records calls rather than tasks. Adding a task identifier that threads through every call is a small change that makes the whole programme possible.

Without it, optimisation proceeds on intuition, and intuition in this area is reliably wrong about which patterns are expensive. See what is cost per task.

Step 2 тАФ Find the patterns that dominate

Group tasks by type and rank by total spend. Almost every system has a small number of patterns accounting for most of the cost.

They are rarely the ones people expect. A rare complex query that everyone worries about may be trivial in aggregate, while a high-volume routine request nobody thinks about consumes the majority.

That ranking determines where effort goes. Optimising a pattern representing two per cent of spend is a well-executed waste of a week.

Step 3 тАФ Cut context before changing models

Oversized context is the most common avoidable cost and the easiest to fix.

Systems accumulate context: retrieved documents nobody trimmed, conversation history nobody summarised, instructions that grew by accretion. Much of it does not improve the answer and all of it is billed on every call.

Measure the effect of trimming rather than assuming. Run the evaluation suite with reduced context and compare тАФ frequently quality holds or improves, because the model is no longer distracted by irrelevant material. See what is a context budget.

Step 4 тАФ Route by difficulty

Send the straightforward majority to a smaller cheaper model and escalate only the hard cases.

The routing decision has to be cheap and reasonably accurate. A classifier, a confidence threshold, or a rule based on input characteristics all work; using an expensive model to decide which model to use does not.

Validate the routing against your evaluation set, split by route. A routing scheme that sends ten per cent of hard cases down the cheap path has degraded quality for those users, and an aggregate score will hide it.

Step 5 тАФ Cache where inputs genuinely repeat

Exact-match caching is safe and helps where the same question arrives repeatedly, which in support and internal knowledge systems is more often than expected.

Semantic caching covers more and requires care: two similar questions can have different correct answers, particularly where the underlying data changes. Set similarity thresholds conservatively and invalidate on corpus changes.

Measure the hit rate. A cache with a two per cent hit rate is complexity without benefit, and knowing that early saves maintaining it. See what is a semantic cache.

Step 6 тАФ Assign ownership and a target

Name someone accountable for cost per task and give them a target treated as a budget rather than an aspiration.

Without that, costs return. Every new feature adds context or calls, each addition is individually reasonable, and nobody is looking at the aggregate until the bill arrives.

Put the number on a dashboard the team sees, and review it when features ship. Cost regressions caught at review are cheap; caught at quarter-end they are a project.

What about self-hosting?

It changes the cost shape rather than always reducing it. Self-hosted models trade per-token charges for infrastructure and operational burden, which pays off at high steady volume and not at low or spiky volume.

Run the comparison on your actual usage pattern, including the engineering time to operate it. Teams that model only the compute cost reach the wrong answer, usually optimistically. See AI agent hosting cost.

What should you not cut?

Evaluation, logging, and human review capacity.

Those are the first things to look expensive and the last things worth removing. A system that saved twenty per cent by cutting evaluation has removed its ability to detect the quality regression that the other optimisations will eventually cause.

If the budget genuinely requires cutting something, cut scope rather than controls.

Who needs to be involved?

An engineer who can instrument and change the system, someone who can judge whether quality held, and a named owner for the cost number.

Cost programmes run without quality judgement produce savings and complaints, usually in that order.

How long does it take?

Two to four weeks to instrument, analyse, and implement the main changes, then continuous ownership. The instrumentation is frequently the longest part and the most valuable.

What are the common failure modes?

Optimising without per-task measurement. Starting with model switching rather than context. Routing without validating by route. Caching without measuring hit rates. And finishing without an owner.

How do you know it worked?

Cost per task down meaningfully with evaluation scores held, the dominant patterns understood, and the number staying down two quarters later.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

Add a task identifier that threads through every model call, and run for a week. The resulting breakdown usually makes the first three optimisations obvious.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: cost attributed to completed tasks rather than calls, quality held against the evaluation suite as each optimisation lands, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read what is cost per task.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why measure cost per task rather than per token?

Because a cheaper model that needs three attempts costs more than an expensive one that succeeds first time. Per-token pricing is an input; cost per completed task is the number that determines whether the system is economic.

02Where does the money usually go?

Into a small number of high-volume patterns, oversized context, and retries. Most systems have a handful of request types that account for most of the spend, and they are rarely the ones people assume.

03What is routing by difficulty?

Sending straightforward cases to a smaller cheaper model and escalating only the hard ones to an expensive model. It preserves quality where it matters while removing most of the cost, provided the routing decision is itself cheap and accurate.

04Does caching help?

Where inputs genuinely repeat, substantially. Exact-match caching is safe and limited; semantic caching covers more and needs care, because two similar questions can have different correct answers in a changing domain.

05Why do savings erode?

Because new features add context and calls, and nobody is watching the aggregate. Without an owner and a per-task cost target treated as a budget, spend returns to its previous level within a couple of quarters.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.

Start a project