Playbook ┬╖ 6 minute read
How to Optimize AI Latency Without Losing Quality
Reducing AI latency starts with measuring the distribution rather than the average, attributing time across retrieval, context assembly, model calls, and post-processing, and then fixing whichever stage dominates. The model is frequently not the slowest part of the request, and assuming it is wastes the effort.
AI latency problems are usually not the model. Retrieval, context assembly, sequential tool calls, and post-processing collectively account for more time than teams expect. This playbook covers finding and fixing the real contributor, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
When users describe the system as slow, when latency is blocking an interactive use case, or when a system that felt responsive at pilot volume has degraded under load.
It is not worth doing for a batch system where nobody is waiting. Latency work has a quality cost, and spending it on a process that runs overnight is a poor trade.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Measure the distribution | Median, p95, p99, by request type |
| 2. Attribute time by stage | Retrieval, context, model, post-processing |
| 3. Fix the dominant stage | Not the one you assumed |
| 4. Stream early | Perception improves before total time does |
| 5. Parallelise independent work | Sequential independent calls are waste |
| 6. Verify quality held | Every change, against the suite |
Step 1 тАФ Measure the distribution by request type
Record latency at the median, ninety-fifth, and ninety-ninth percentiles, broken down by request type.
Averages mislead in both directions. A system with an acceptable average may have a category of requests that are consistently terrible, and users in that category experience the system as broken while the dashboard looks fine.
Break down by type because the fix differs. Simple lookups and multi-step agent tasks have different profiles, and averaging them together hides both.
Step 2 тАФ Attribute time across every stage
Instrument each stage: input processing, retrieval, context assembly, each model call, tool calls, and post-processing.
This is where assumptions break. Teams frequently discover that retrieval against a large index takes longer than the model call, or that a post-processing step nobody thought about adds half a second.
Without this attribution, optimisation targets whatever the team believes is slow, and the belief is usually about the model because that is the visible component. See what is distributed tracing for ai.
Step 3 тАФ Fix the dominant stage first
Work on whichever stage actually dominates, which the attribution now tells you.
Retrieval latency responds to index tuning, filtering before search, and reducing the number of candidates reranked. Context assembly responds to trimming and to caching assembled context. Model latency responds to smaller models for easy cases, shorter outputs, and provider selection.
The discipline is doing one thing at a time and measuring. Bundled changes make it impossible to know what helped, and the next regression becomes unattributable.
Step 4 тАФ Stream as early as you can
Time to first token is what users judge by. A response that begins in three hundred milliseconds and completes in four seconds feels considerably faster than one delivered complete at two.
Streaming requires the interface to handle partial content well, and it interacts with any post-processing that needs the complete output. Where post-processing is required, consider streaming an initial acknowledgement or a partial result while it completes.
Accessibility matters here: streamed content needs careful handling for assistive technology. See AI and ADA accessibility.
Step 5 тАФ Parallelise independent work
Independent retrieval calls, independent tool calls, and any precomputation not dependent on model output should run concurrently.
Sequential execution of independent steps is the most common avoidable latency in agent systems, and it usually exists because the code was written in the order someone thought about the problem.
Be careful with ordering assumptions. Parallelising steps that turn out to have a dependency produces intermittent wrong results, which is a worse outcome than slow ones.
Step 6 тАФ Verify quality after every change
Run the evaluation suite after each optimisation, not at the end.
Latency work reduces context, skips steps, and uses smaller models, and each of those can cost quality. A change that halved latency and reduced answer accuracy by ten per cent is usually a bad trade, and nobody will notice without measurement.
Where a trade is worth making, make it knowingly and record it. Undocumented quality reductions taken for speed resurface later as a mysterious decline.
What about perceived latency?
It responds to design as much as to engineering. Progress indication, partial results, and optimistic interfaces all change how long a wait feels.
The highest-return change in many products is not making the system faster but making the wait informative. A system showing what it is doing тАФ searching, reading, drafting тАФ is tolerated at latencies that would otherwise feel broken.
What about cold starts and provider variability?
Both produce tail latency that no amount of local optimisation fixes.
Provider response times vary, sometimes substantially, and a request that normally takes a second occasionally takes ten. Handle it: timeouts with a fallback path, and retries that do not compound the delay.
Measure your provider's latency distribution independently of your own. It is useful to know which part of your tail is yours. See what is a fallback chain.
Who needs to be involved?
An engineer who can instrument and change the system, and someone who can judge whether quality held.
Latency programmes run without quality judgement reliably produce a faster system that people trust less.
How long does it take?
One to three weeks including instrumentation, which is usually the longest part. The fixes themselves are frequently quick once the attribution exists.
What are the common failure modes?
Optimising the model because it is visible. Measuring averages. Bundling changes. Parallelising steps with hidden dependencies. And skipping quality verification.
How do you know it worked?
Tail latency down at the ninety-fifth and ninety-ninth percentiles, evaluation scores held, and users no longer describing the system as slow.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Instrument every stage of one request type and look at where the time actually goes. The answer is usually not what the team expected, and it redirects the whole effort.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: latency attributed across every stage before anything is optimised, quality verified against the evaluation suite after each change, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read what is throughput in AI systems.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why measure the distribution rather than the average?
Because users experience the tail. A system averaging two seconds with a slow tail at fifteen feels unreliable, and the average hides exactly the requests that damage trust. Track median, ninety-fifth, and ninety-ninth percentiles.
02What is usually the slowest part?
Frequently not the model. Retrieval against a large index, context assembly, sequential tool calls, and post-processing each add time, and in many systems they collectively exceed the model call.
03Does streaming actually help?
It changes perception substantially without changing total time. Time to first token is what users judge responsiveness by, and a system streaming at 300ms feels faster than one delivering complete at two seconds.
04What can be parallelised?
Independent retrieval calls, independent tool calls, and any pre-computation not dependent on the model output. Sequential execution of independent steps is the most common avoidable latency in agent systems.
05How do you avoid degrading quality?
By running the evaluation suite after each change. Latency optimisations frequently reduce context, skip steps, or use smaller models, and each of those can cost quality in ways an average latency number will never show.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.