Checklist ¡ 4 minute read
AI Performance Tuning Checklist: Making Systems Feel Fast
AI latency is dominated by architectural decisions rather than by code efficiency. Optimise time to first token, reduce prompt size, cache shared prefixes, parallelise independent work, route simple tasks to faster models, and use streaming so perceived speed exceeds measured speed.
AI latency is dominated by architectural decisions rather than by code efficiency. This checklist covers what actually moves it, drawn from FISTA Solutions' AI agents engineering work.
Where does the time go?
Six contributors, in the order worth attacking.
| Contributor | Typical remedy |
|---|---|
| Input processing | Reduce prompt and context size |
| Model speed | Route to a smaller model |
| Sequential tool calls | Parallelise independent ones |
| Retrieval | Index and query tuning |
| Generation length | Constrain output |
| Perception | Stream and show progress |
Measurement
Measure the right thing before changing anything.
- Time to first token measured separately from total time
- Both measured at the ninety-ninth percentile
- Provider time separated from your own processing time
- Retrieval latency measured separately
- Tool call latency measured per tool
- Measurements taken under realistic concurrency
- A target agreed per use case
Prompt and context size
Affects latency as well as cost. See why context beats prompting.
- Current prompt size measured per operation
- Retrieved passage count tuned with measurement
- Accumulated system prompt instructions pruned
- Few-shot examples reduced to the minimum
- Conversation history summarised rather than appended
- Tool definitions trimmed of unused capabilities
- Latency re-measured after each reduction
Caching
Cheap where prompts share a long prefix.
- Prompt prefix caching enabled where supported
- Static content placed at the start of the prompt
- Cache hit rate measured
- Semantic caching considered for repeated questions
- Retrieval results cached where the corpus is stable
- Cache invalidation defined
- Latency improvement from caching quantified
Parallelism
The usual agent bottleneck.
- Independent tool calls identified
- Independent calls issued in parallel
- Retrieval and other preparation overlapped where possible
- Speculative work considered where it is cheap
- Concurrency limits set to avoid downstream overload
- Timeouts set so one slow call does not block everything
- Partial results used where the remainder is optional
Model and routing
Frequently the largest single improvement. See the quiet rise of small models.
- Latency compared across candidate models on your tasks
- Simple operations routed to faster models
- Quality verified on each route before switching
- Escalation path for cases the fast model handles poorly
- Output length constrained where long responses are not needed
- Streaming enabled wherever the interface supports it
- Routing re-evaluated after provider releases
Perceived speed
Frequently cheaper to improve than actual speed. See streaming UI patterns for AI apps.
- Streaming enabled so output starts quickly
- Progress shown during tool calls and long steps
- Immediate acknowledgement of the request
- Layout stable so growing content does not shift the page
- Long work moved to background jobs with notification
- Optimistic interface updates where safe
- Cancellation available and responsive
What are the most common failures?
Optimising total time rather than time to first token. Measuring averages. Sequential tool calls. Context grown without review. And ignoring perceived speed, which is usually the cheapest improvement available.
Who should own this?
The team running the system owns its latency targets and tuning. Targets should be agreed with the product owner, since the acceptable figure is a product decision.
How often should it run?
Measured continuously, reviewed monthly, and re-tuned after any change to prompts, retrieval, or model. Re-check after provider releases, since latency characteristics change.
What evidence should it produce?
Latency measurements at the tail before and after each change, cache hit rates, and routing decisions with the evaluation that justified them.
What if latency is inherent to the task?
Then move it off the interactive path. Work that genuinely takes a minute belongs in a background job with a notification, not behind a spinner.
That is a product decision as much as a technical one, and it usually produces a better experience than any amount of tuning. See how to optimize AI latency.
What should you do first?
Measure time to first token at the ninety-ninth percentile for your main operation. That number, rather than the average total time, is what users experience.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: time to first token optimised ahead of total generation time, with independent tool calls parallelised and routing decided by measurement, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read how to optimize AI latency.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What should be optimised first?
Time to first token. Users judge responsiveness by when something happens, so reducing the delay before output begins helps more than reducing total generation time.
02Does prompt size affect speed?
Yes. Processing a large input takes time before generation begins, which shows up directly as a slower first token. Trimming context helps latency and cost together.
03What slows agents down?
Sequential tool calls. An agent making five calls in sequence waits for all five; making independent ones in parallel can cut total time substantially.
04How much faster are smaller models?
Frequently several times, for tasks they handle adequately. Routing classification and extraction to a small model removes latency that no other optimisation could.
05Why measure at the tail?
Because the slow requests are what users notice and what times out. A good average with a poor ninety-ninth percentile produces a worse experience than the average suggests.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.