Glossary · 5 minute read
What Is Token Accounting? Attributing AI Cost Explained
Token accounting attributes token consumption and cost to the features, teams, customers, and requests that caused it, rather than reporting a single aggregate. Without attribution, spend can only be observed, not managed, because nothing identifies which part of the system to change.
AI spend reaches finance as a single number from a provider, and that number can answer no question anyone wants to ask. Which feature is expensive, which customer is unprofitable, whether last week's prompt change cost anything — none of it is visible without attribution built into the system. This explainer covers what to instrument. It complements how to build an ai cost dashboard and what is cost per task, and reflects FISTA Solutions' approach in AI enablement delivery.
Why is the aggregate useless?
Because it reports an effect with no cause. Spend rose forty percent. That could be more users, a longer system prompt shipped on Tuesday, an agent loop retrying more often, a model change, or one enterprise customer discovering a feature and using it heavily.
Each of those requires a different response, and an aggregate figure distinguishes none of them. Teams without attribution typically respond by asking everyone to be careful, which has no measurable effect.
| Dimension | Answers |
|---|---|
| Feature or endpoint | Which functionality is expensive |
| Team or product owner | Who is accountable |
| Customer or tenant | Which accounts are unprofitable |
| Model and version | What a model change cost |
| Input vs output tokens | Which optimisation applies |
| Cached vs uncached | Whether caching is working |
What should be captured?
At each model call: the feature, the owning team, the customer where relevant, the model and version, the prompt version, input tokens split by cached and uncached, output tokens, and whether the overall task succeeded.
That set answers nearly every question that arises, and capturing it at the call site is straightforward. Retrofitting attribution into a system that logged only totals is considerably harder, which argues for doing it early.
Why separate input and output?
Because they are priced differently, with output typically costing several times more per token. A feature that processes long documents and returns short summaries has a cost profile dominated by input; one that generates long content is dominated by output.
The optimisations differ completely. Input-heavy workloads benefit from caching and context trimming; output-heavy ones benefit from brevity constraints and structured output. Without the split, teams optimise the wrong side.
What does cache tracking reveal?
Whether caching is actually working. Cached prefixes are billed at a substantially reduced rate, so hit rate is directly a cost metric. A prompt change that moves a variable element earlier in the context can destroy cache hits entirely, raising cost with no other visible symptom.
Tracking hit rate alongside token counts catches that within a day rather than at the invoice. See what is inference optimization.
Why per successful task?
Because per-call cost misleads. A cheaper model that requires three attempts and a retry costs more than an expensive one that succeeds immediately, and a per-call view shows the cheaper model winning.
Cost per successful task is the figure that supports a decision, and it requires knowing whether the task succeeded, which means task outcome has to be instrumented alongside spend.
How does this support chargeback?
By making it defensible. Internal chargeback for AI spend requires attribution people trust, and a model that allocates by headcount or by guess produces arguments rather than accountability. Per-team, per-feature attribution derived from actual calls is arguable only about methodology, not about the numbers.
What should you do first?
Try to answer which of your features is most expensive. If you cannot, that is the gap, and instrumenting feature-level attribution at each call site is usually a day of work that pays back within the first month of having it.
How does this fit with budgets and alerts?
Attribution makes budgets enforceable at the level where someone can act. A team-level budget with a burn-rate alert gives the owning team a signal while there is still room to respond, which an organisation-wide monthly figure never does.
The alerting that matters most is on rate of change rather than on absolute spend. A feature whose cost per request doubled overnight is a defect, even if its total is small, and it will not be found by a threshold set against the total.
What about agent workloads specifically?
They need per-step attribution, because an agent's cost is the sum of an unpredictable number of iterations. Knowing that a task cost eleven times the baseline is useful; knowing that it spent nine iterations retrying one failing tool is actionable, and only step-level accounting distinguishes them.
Agent cost also has a long tail: most tasks are cheap and a few are very expensive. Reporting the mean hides that entirely, so percentiles and a list of the most expensive recent tasks are the more useful views.
Who should see the numbers?
The engineers who can change them, with the aggregate rolling up to whoever owns the budget. Cost visibility confined to finance produces requests to reduce spend without information about how; visibility at the call site produces engineers who notice that a prompt change tripled a feature's cost, which is where the saving actually happens.
How FISTA Solutions helps
FISTA Solutions instruments token attribution at the call site across feature, team, customer, model, and prompt version, separates input, output, and cached tokens, tracks cost per successful task rather than per call, and builds chargeback on attribution people can verify, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 47% efficiency gains.
To find out where your AI spend actually goes, message FISTA on WhatsApp, or read how to build an ai cost dashboard.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why is an aggregate figure insufficient?
Because it tells you spend rose without indicating what to change. A 40% increase could be more users, a longer system prompt, a retry loop, or one customer's unusual usage, and each requires a different response. Attribution distinguishes them.
02What dimensions should be captured?
Feature or endpoint, team or product owner, customer or tenant where relevant, model and version, and request type. Together those answer nearly every cost question that arises, including the awkward one about which customers are unprofitable.
03Why separate input and output tokens?
Because they are priced differently, usually with output costing several times more. A feature generating long responses has a different cost profile from one processing long documents, and the optimisations that help each are completely different.
04What about cached input?
Cached prefixes are typically billed at a substantially reduced rate, so tracking cache hit rate alongside token counts shows whether caching is working. A falling hit rate after a prompt change is a cost regression with no other visible symptom.
05Why measure per successful task?
Because cost per call can fall while cost per outcome rises. A cheaper model that needs three attempts to succeed is more expensive overall than one that succeeds first time, and only per-task accounting makes that visible rather than appearing as a saving.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.