FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook ┬╖ 6 minute read

How to Forecast AI Inference Cost You Can Defend

A defensible AI cost forecast starts from measured cost per completed task, projects volume with a stated basis, models the drivers that change with scale, and presents a range with assumptions rather than a single number. Forecasts extrapolating token prices are reliably wrong.

By FISTA Solutions┬╖ AI-Native Engineering Team┬╖
How to Forecast AI Inference Cost You Can Defend article cover

AI cost forecasts are usually wrong because they extrapolate token prices rather than modelling task cost and volume. This playbook covers building one that survives contact with a real bill, drawing on FISTA Solutions' AI enablement work.

When is this worth doing?

Before a budget cycle, before a launch that changes volume materially, or when the current bill has diverged from what was expected.

It is not worth doing for a system whose spend is immaterial. The effort is justified when the number is large enough that being wrong matters.

What does the sequence look like?

StepPurpose
1. Measure cost per taskFrom real traffic, not calculation
2. Segment by request typeMix drives the average
3. Project volume with a basisNot a growth percentage
4. Model what changes at scaleMix, caching, tiers
5. Add the hidden costsRetries, failures, evaluation
6. Present a rangeWith assumptions attached

Step 1 тАФ Measure cost per completed task

Instrument spend attributed to completed tasks rather than to individual calls, from real production traffic.

This number already includes what calculations miss: retries, multi-step trajectories, failed attempts that consumed tokens, and the context that accumulated along the way.

A system without task-level attribution cannot forecast credibly. Adding a task identifier that threads through every call is a small change and the prerequisite for everything here. See what is cost per task.

Step 2 тАФ Segment by request type

Different request types cost very different amounts, and the mix determines the average.

A forecast built on a blended average breaks when the mix shifts тАФ which it does as the audience grows, as new features launch, and as users learn what the system can do.

Forecast per segment and combine, with an explicit assumption about how the mix changes. That makes the forecast both more accurate and easier to correct when reality diverges.

Step 3 тАФ Project volume with a stated basis

Tie volume to something real: user growth, a rollout schedule, campaign plans, seasonal patterns, or a known migration.

A growth percentage with no underlying reason is a guess wearing a number. When the actual figure diverges, nobody can say which assumption was wrong, and the next forecast is no better.

Where the basis is genuinely uncertain, say so and forecast scenarios rather than inventing precision.

Step 4 тАФ Model what changes at scale

Cache hit rates improve with volume as repeated queries become more common. Request mix shifts as the audience broadens. The proportion going down expensive paths changes. Provider pricing tiers may apply.

A forecast assuming today's characteristics hold at ten times the volume is usually optimistic on caching and pessimistic on mix, and the two do not cancel out reliably.

Model each explicitly with a stated assumption rather than folding them into a single factor.

Step 5 тАФ Add the costs people forget

Evaluation runs, shadow deployments, development and testing usage, retries, and the tokens consumed by tasks that failed and were abandoned.

Evaluation in particular can be material for a system with a large suite running on every change. It is worth forecasting separately so it is not mistaken for production growth.

Development usage is small per engineer and adds up across a team, and it is the line most often absent from a forecast that then looks wrong by exactly that amount.

Step 6 тАФ Present a range with assumptions

Give a low, expected, and high figure with the assumptions producing each, and name the two or three factors that would move it most.

Ranges are more useful than point estimates and considerably more defensible. A forecast that came in within its stated range was correct; one that missed a single number by fifteen per cent was wrong, even if it was better analysis.

Record the assumptions somewhere you will look again. Re-forecasting is much easier when you can see which assumption broke.

How often should you re-forecast?

When the mix changes, when a launch lands, or when actuals diverge from the range тАФ not on a calendar.

Monthly re-forecasting produces work without information when nothing has changed. Event-driven re-forecasting catches the changes that matter and skips the months where the answer is the same.

Set a divergence threshold that triggers a review automatically.

What if the forecast is unaffordable?

Then the architecture needs changing, and it is considerably cheaper to know now.

The levers are routing by difficulty, reducing context, caching, and in some cases self-hosting at sustained high volume. Each has a quality trade worth measuring rather than assuming.

A forecast that kills a design before it is built is the most valuable output this process produces, even though it does not feel like it at the time. See how to run an ai cost reduction program.

Who needs to be involved?

An engineer who can instrument and measure, and someone from finance who will use the number.

Finance involvement early prevents the common outcome where a technically sound forecast is presented in a form nobody can put in a budget.

How long does it take?

One to two weeks, mostly instrumentation if task-level attribution does not exist. The modelling itself is a day or two once the data is there.

What are the common failure modes?

Multiplying token price by volume. Blended averages. Growth percentages with no basis. Ignoring evaluation and development usage. Single-point estimates. And calendar-driven re-forecasting.

How do you know it worked?

Actuals landing within the stated range, divergences attributable to a named assumption, and finance able to use the number without translation.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

Instrument cost per completed task for one week. That measurement alone usually changes the forecast more than any amount of modelling.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: forecasts built from measured cost per completed task rather than token arithmetic, ranges presented with the assumptions that produce them, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read AI total cost of ownership.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why do AI cost forecasts go wrong?

Because they multiply a token price by an assumed volume. That ignores retries, failed tasks that consumed tokens, growing context, and the distribution of easy and hard requests, all of which dominate the actual bill.

02What should the unit be?

Cost per completed task, measured from real traffic rather than calculated from list prices. That number already includes retries, multi-step trajectories, accumulated context, and failed attempts that consumed tokens, all of which a per-token figure leaves out and all of which dominate a real bill.

03How do you project volume?

From a stated basis: user growth, campaign plans, seasonal patterns, or a known rollout schedule. A percentage growth assumption with no underlying reason is a guess presented as an analysis.

04What changes with scale?

Request mix, cache hit rates, the proportion going down expensive paths, and sometimes provider pricing tiers. A forecast assuming today's mix holds at ten times the volume is usually optimistic.

05How should uncertainty be presented?

As a range with the assumptions that produce each end, and the factors that would move it. A single number invites false precision and gets treated as a commitment.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.

Start a project