Checklist · 5 minute read
LLM Cost Control Checklist: Finding Spend That Does Nothing
Model spend is controllable once it is attributed. Work through attribution, then routing by task, then caching, then prompt size, then retry and abandonment behaviour. Most organisations find a large share of their bill going to one operation that a cheaper model would handle adequately.
Most AI cost problems come down to a few specific patterns, and all of them are visible once spend is attributed. This checklist works through them, drawn from FISTA Solutions' AI enablement operational work.
Where does the money actually go?
Six places, in the order worth checking.
| Area | Typical finding |
|---|---|
| Attribution | Nobody knows the split |
| Model routing | One operation on an expensive model |
| Caching | Long shared prefixes not cached |
| Prompt size | Context grown without review |
| Retries | Exponential retry on a failing provider |
| Abandonment | Cancelled requests still billing |
Attribution
Without this, every other item is a guess. See how to reduce AI costs.
- Every model call tagged with feature, workflow, and caller
- Input and output tokens recorded separately
- Cost per task computed and visible to the owning team
- Spend broken down by model and by operation
- Cost per user or per account available where relevant
- A dashboard the owning team actually looks at
- Month-on-month trend by operation, not just in total
Routing by task
The largest single lever in most systems. See the quiet rise of small models.
- The highest-volume operation identified
- A cheaper model evaluated against it on real cases
- Routing implemented with per-task model selection
- Escalation path when the cheap model's output fails validation
- Confidence or validation used to decide escalation, not guesswork
- Routing decisions re-evaluated after any model release
- Quality monitored per route, not only overall
Caching
Cheap where it applies and useless where it does not, so check which case you are in.
- Prompt prefix caching enabled where the provider supports it
- System prompts and fixed instructions placed at the start for cache reuse
- Semantic caching considered for repeated user questions
- Cache hit rate measured, not assumed
- Cache invalidation defined when underlying content changes
- Cached responses excluded from quality sampling or labelled
- Cache storage cost compared against the token saving
Prompt and context size
Input tokens bill on every call, so context that grew without review is a permanent tax.
- Average and maximum prompt size measured per operation
- Retrieved passage count tuned with measurement, not set defensively
- System prompts reviewed for accumulated instructions nobody needs
- Few-shot examples reduced to the minimum that maintains quality
- Conversation history truncated or summarised rather than appended forever
- Tool definitions trimmed of unused capabilities
- Output length constrained where long responses are not needed
Retries, loops, and abandonment
This is where runaway costs come from, and they arrive quickly.
- Retry limits set and enforced per request
- Backoff implemented so a failing provider is not hammered
- Agent loops have a maximum step count
- Cancellation propagates from client to provider call
- Abandoned request cost measured
- Timeouts set at every layer
- A circuit breaker stops repeated calls to a failing dependency
Budgets and alerting
Alerts on rate catch problems while they are still small.
- Spend rate alerts configured per hour, not per month
- Per-feature budget thresholds defined
- Anomaly detection on token volume
- A hard cap or kill switch for runaway spend
- Alert recipients are the team that can act, not just finance
- Provider rate limits understood and monitored
- Cost reviewed as part of the regular operational meeting
What are the most common failures?
Optimising before attributing. Serving high-volume classification with a frontier model. Retry without backoff. Context that grew and was never reviewed. And alerting on monthly totals, which is always too late.
Who should own this?
The team that owns the feature owns its cost, with a platform or finance function providing the attribution infrastructure. Cost owned centrally and spent locally produces no behaviour change.
How often should it run?
Attribution continuously, a review monthly, and a full pass through this checklist quarterly or after any significant volume change. Re-check routing after every provider model release.
What evidence should it produce?
A cost-per-task trend by operation, routing decisions with the evaluation that justified them, cache hit rates, and a record of alert thresholds. That set shows cost is managed rather than observed.
What about agents specifically?
Agents multiply everything here, because one task becomes many model calls. Step limits, routing within the trajectory, and cancellation matter more than in single-call systems.
The highest-return agent optimisation is usually routing routine steps â tool selection, argument formatting, result checking â to a small model while reserving the expensive one for genuine decisions. See AI agent production readiness checklist.
What should you do first?
Tag every model call with its feature and look at the split. The largest line is almost always a surprise, and it is where the work is.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: spend attributed per feature and workflow before any optimisation, and routing decisions justified by evaluation against real cases, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read how to reduce AI costs.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Where do you start?
Attribution. Until you know cost per feature, per workflow, and per user, every optimisation is guesswork. The instrumentation is straightforward and it almost always reveals a surprise.
02What usually dominates the bill?
A single high-volume, low-complexity operation â classification, routing, or extraction running on every item â served by a frontier model because that is what was available when it was built.
03How much does caching save?
Substantially where requests share a long common prefix, such as a system prompt and a fixed instruction set. Where every request is unique, it saves nothing and adds complexity.
04What is abandoned request waste?
Generation that continues after the user has navigated away, because cancellation was never propagated to the provider call. On consumer products this can be a meaningful share of spend.
05Why alert on rate rather than total?
Because a monthly total tells you about overspend after it happened. A rate alert catches a runaway loop or a misconfigured retry within minutes, which is when it is still cheap.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.