FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Whitepaper ¡ 9 minute read

AI FinOps: Governing the Cost of Intelligence

AI FinOps applies cost management practice to inference spend: attributing every call to a team and use case, expressing cost as unit economics rather than a monthly total, forecasting from usage drivers rather than trend lines, optimising in an order that protects quality, and running a monthly review where engineering and finance see cost and quality together.

By FISTA Solutions¡ AI-Native Engineering Team¡
AI FinOps: Governing the Cost of Intelligence article cover

Technology finance functions spent a decade learning to manage cloud spend, and most of that discipline transfers to AI. What does not transfer is the assumption that cost tracks provisioned resources. AI cost tracks behaviour: how much context a request carries, which model serves it, how many steps an agent takes, how often a cache hits. Those are engineering decisions made continuously, often without anyone realising they moved a cost line. This whitepaper adapts FinOps practice to that reality. It draws on FISTA Solutions' AI enablement work and complements the AI total cost of ownership model whitepaper and the model routing and cost control whitepaper.

What makes AI cost hard to govern?

PropertyConsequence
Variable per requestNo stable unit to budget against without measurement
Driven by context sizeCost changes when a retrieval setting changes
Scales with adoptionSuccess raises cost faster than headcount raises budget
Sensitive to model choiceOrder-of-magnitude differences between tiers
Agent step fan-outOne user action can trigger dozens of calls
Opaque invoicingProvider bills cannot be decomposed retroactively
Quality coupledEvery cost lever is also a quality lever

The last row is what distinguishes AI FinOps from cloud FinOps. Right-sizing a virtual machine does not change what the application does. Reducing context, changing model tier, or enabling semantic caching all change outputs, which means cost work without quality measurement is product change without testing.

What does visibility require?

Attribution captured at the point of the call, not reconstructed later. Every request passes through a gateway that tags it with team, use case, route, environment, and where relevant customer or tenant, and records tokens in and out, model and version, latency, cache status, and cost.

From that single stream everything else follows: spend by team and use case, unit economics, forecasting inputs, and the detail needed to diagnose a spike. Without it, the organisation has a provider invoice and a guess. This is the first investment in any AI FinOps practice and it typically pays for itself within a quarter simply by revealing which workloads consume the budget, which is rarely what teams expect.

Why unit economics rather than totals?

Because a rising total is meaningless without volume. Spend up forty percent with volume up eighty percent is an efficiency gain; spend up forty percent with flat volume is a regression. Only unit economics distinguish them.

The units that matter are business units, not technical ones: cost per resolved support ticket, per processed invoice, per generated draft, per completed agent task. Those numbers compare directly against the cost of the alternative, whether that is a person, an outsourcer, or the previous system, which is what makes them usable in a business case and defensible in a board pack.

Cost per token is an engineering diagnostic, not a business metric, and presenting it to finance invites the wrong conversation. See the AI agent unit economics whitepaper.

How is AI spend forecast?

Bottom-up from drivers. For each use case: expected volume, measured cost per unit, and any planned change to either. Add planned launches with estimated volumes and a wide uncertainty band. Subtract optimisations that are committed with owners and dates, not aspirational ones.

Trend extrapolation fails for AI because the curve is dominated by adoption steps and feature launches rather than by gradual growth. A forecast built on last quarter's slope will be wrong in the month a new use case reaches general availability.

Two practices improve accuracy materially: forecast a range rather than a point, and reforecast monthly rather than quarterly, since the drivers move faster than planning cycles.

What is the optimisation ladder?

In order of return and ascending quality risk:

  1. Context discipline. Tighten retrieval selection, bound conversation history, pass only the record fields the task uses, and compress what need not be verbatim. Usually the largest saving, and it typically improves quality because irrelevant context degrades results.
  2. Caching. Prefix caching for stable instruction blocks; semantic caching scoped carefully to shared, stable content with entitlement-aware keys.
  3. Routing. Send bounded tasks to smaller models with escalation, rather than serving everything from one tier.
  4. Output constraints. Schemas and length limits that stop generation producing more than the consumer uses.
  5. Step and retry discipline. Bound agent steps, summarise intermediate state, and fix retry storms, which are frequently a larger cost than anyone realises.
  6. Model tier changes. Last, because it carries the most quality risk and should follow per-route evaluation.

Teams reliably do this in reverse order, starting with the model because it is the most visible number, and discovering that context was the actual problem.

How is quality protected during cost work?

By treating each optimisation as a behaviour change subject to the same gates as any feature: run the route's evaluation set, compare against the incumbent, canary on a traffic share, watch the quality metrics as well as the cost ones, and revert on regression.

The failure this avoids is common enough to name: a cost programme reports substantial savings, quality declines gradually across several routes, and the business impact surfaces months later as rising escalations, falling conversion, or customer complaints that nobody connects back to the optimisation. Savings claimed without quality evidence are not savings; they are an unmeasured trade.

Who owns what?

RoleOwns
Platform engineeringGateway, attribution, caching, routing infrastructure
Use case ownerThat use case's unit economics and volume forecast
FinanceBudget, chargeback model, reporting into the P&L
Engineering leadershipOptimisation backlog and quality gates
ProductPricing where AI cost affects customer economics

The pattern that fails is a central AI cost owner with no authority over the teams whose decisions create the cost. Attribution plus ownership is what makes reduction happen; either alone does not.

Should cost be charged back?

Where usage is discretionary and the owner can influence it, yes. Chargeback changes behaviour in a way showback does not, because a team that sees its own budget consumed will tighten context and question volumes, while a team that sees a report will not.

A workable hybrid for shared platforms: the platform's fixed costs sit centrally as an enabling investment, while per-request inference cost is charged to the consuming team or product. Development and experimentation get a separate allocation so that cost pressure does not suppress the experimentation the organisation wants.

Where AI is embedded in a product sold to customers, the cost belongs in product margin rather than IT, and product management owns the pricing response.

What does the operating rhythm look like?

Weekly, engineering reviews anomalies: spikes, unusual cache miss rates, escalation rate changes, and any route whose cost per unit moved.

Monthly, a joint review with finance and use case owners covering spend against budget by team and use case, unit economics trended, quality scores per route alongside cost, the optimisation backlog with owners and dates, and the reforecast.

Quarterly, a strategic review: provider commitments and negotiation posture, concentration risk, which use cases justify their cost and which should be reconsidered, and the capacity plan for the coming period.

The single most important detail is that cost and quality appear in the same report, in front of the same people. Separating them guarantees that one is optimised at the other's expense.

What does maturity look like?

Initial. One provider invoice, no attribution, surprise at quarter end, and reactive usage restrictions.

Managed. Gateway attribution in place, spend visible by team, budgets set, anomalies detected within days.

Defined. Unit economics per use case, driver-based forecasting, an optimisation backlog with quality gates, and chargeback where appropriate.

Optimised. Cost per unit trending down while volume grows, quality stable or improving, provider strategy informed by measured portability, and AI cost discussed in business terms rather than technical ones.

What goes wrong?

Attribution deferred until spend becomes a problem, at which point the history cannot be recovered. Optimisation without evaluation. Cost reported without volume. Forecasts extrapolated from trend. Chargeback without the ability to influence usage. Hard budget caps that fail workloads rather than degrading them. And the organisational failure of putting AI cost with infrastructure finance while the decisions that drive it sit with product and engineering.

How does this change the business case for new use cases?

Once unit economics exist for deployed systems, new proposals can be assessed against evidence rather than vendor claims. A team proposing an agent for a new workflow can be asked for expected volume, an estimated cost per unit derived from a comparable existing route, and the cost of the current way of doing the work. That converts an argument about whether AI is worthwhile into an arithmetic comparison, which is a far better conversation.

It also disciplines the portfolio. Use cases whose cost per unit approaches the cost of the human alternative are not necessarily wrong, since speed, availability, and consistency have value, but they should be chosen deliberately rather than discovered later. Use cases whose cost per unit exceeds the alternative by a wide margin and whose quality advantage is unproven are candidates for stopping, and an organisation with unit economics can stop them without a political fight.

What should be reported to the board?

Three numbers and one trend, on a single slide. Total AI spend against budget. Cost per unit for the two or three largest use cases, trended. The measured business outcome those use cases produced, in the organisation's own operational metrics. And the trend in cost per unit over time, which is the clearest single indicator of whether the programme is maturing.

What does not belong there: token counts, model names, and comparisons of provider pricing. Boards are equipped to judge whether a cost per resolved ticket is falling and whether the resolution rate held; they are not equipped to judge whether a context window change was wise, and presenting it invites decisions at the wrong altitude.

How FISTA Solutions delivers this

FISTA Solutions installs the attribution, unit economics, forecasting, and optimisation practice that makes AI spend a managed line rather than a surprise, with quality measured alongside every cost change, through AI enablement, AI agents, and forward deployed engineers working with engineering and finance. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime and 47% efficiency gains where measured.

To make AI spend predictable and defensible, message FISTA on WhatsApp, or read the AI total cost of ownership model whitepaper.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What makes AI cost different from cloud cost?

It varies per request rather than per resource, responds to changes nobody tags as cost-relevant such as a longer prompt or broader retrieval, scales with adoption rather than infrastructure, and arrives as a provider invoice that cannot be decomposed into teams or use cases after the fact.

02What unit economics should be tracked?

Cost per unit of business work: per resolved ticket, processed document, generated draft, or completed task, plus cost per active user and per tenant where relevant. These numbers decide whether a use case is viable and are the only fair basis for comparing AI with the alternative.

03How should AI spend be forecast?

From usage drivers rather than historical trend: expected volume per use case multiplied by measured cost per unit, plus planned launches, minus committed optimisations. Trend-based forecasts fail because adoption curves and feature launches dominate the shape.

04In what order should cost be optimised?

Context first, since it is usually the largest share and tightening it often improves quality; then caching; then routing to appropriate model tiers; then model tier changes themselves. Start with the levers that carry the least quality risk.

05Should AI cost be charged back to business units?

Where usage is discretionary and owners can influence it, yes, because chargeback changes behaviour and showback rarely does. Where AI is embedded in a shared platform, a hybrid of platform cost centrally and usage charged back is usually workable.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project