FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Glossary ¡ 5 minute read

What Is Speculative Decoding? Faster Inference Explained

Speculative decoding accelerates text generation by having a small draft model propose several tokens which the larger target model verifies in a single forward pass. Accepted tokens are kept and rejected ones discarded, so the output matches what the target model would have produced alone, at lower latency.

By FISTA Solutions¡ AI-Native Engineering Team¡
What Is Speculative Decoding? Faster Inference Explained article cover

Speculative decoding is one of the few inference optimisations that improves latency without trading anything away in output quality, which makes it unusually attractive. It is also frequently misunderstood as an approximation, which it is not. This explainer covers how it works, when it helps, and what it costs. It complements what is inference optimization and how to reduce ai costs, and reflects FISTA Solutions' approach in AI enablement delivery.

How does it work?

A small, fast draft model generates several candidate tokens ahead. The large target model then processes all of them in a single forward pass, checking whether each matches what it would have produced. Matching tokens are accepted; at the first mismatch, the target model's own token is used and the rest are discarded.

The cycle repeats. When the draft model is often right, several tokens emerge per target-model pass rather than one.

AspectEffect
Output qualityIdentical to target model alone
Latency per requestLower when acceptance is high
Throughput under batchingLittle or no gain
MemoryHigher; two models resident
ComplexityModerate; usually handled by the serving stack
Best-suited textPredictable, formulaic, structured

Why does verifying several tokens cost so little?

Because single-token generation is bound by memory bandwidth rather than by arithmetic. Loading the model's weights for one forward pass is the expensive part; the computation for one token versus several candidates is comparatively cheap once the weights are loaded.

Speculative decoding exploits that imbalance. It converts a bandwidth-bound sequential process into one that gets more useful work from each weight load.

What determines the speedup?

The acceptance rate, which depends on how well the draft model predicts the target's output for this kind of text. Boilerplate, structured formats, code following conventional patterns, and repeated phrasing all accept well. Creative writing and highly specialised domain text accept poorly.

Real-world speedups vary widely by workload, which is why the honest answer to how much faster it is always begins with what kind of text.

When is it not worth it?

When acceptance is low, because rejected draft tokens are wasted computation. Below a workload-dependent threshold the overhead exceeds the benefit and the arrangement is slower than plain generation.

It also helps less under heavy batching. The technique exploits spare capacity in a bandwidth-bound process, and a server already saturated with concurrent requests has less spare capacity to exploit. Systems optimising for aggregate throughput rather than per-request latency may see little gain.

What are the operational costs?

Memory, primarily. Both models must be resident, and the draft model's footprint reduces what is available for batching and context. On constrained hardware that trade can be unfavourable.

There is also compatibility maintenance: draft and target must share a tokenizer and vocabulary, and updating the target model means revisiting the draft.

How does it relate to other optimisations?

It is complementary. Quantisation reduces memory and can speed both models. KV caching avoids recomputing attention over prior tokens and is orthogonal. Batching improves throughput at some cost to individual latency. Prompt caching avoids reprocessing shared prefixes entirely.

Speculative decoding addresses the sequential nature of generation specifically, which none of the others do. See what is inference optimization.

Should teams build it?

Rarely. Major serving frameworks and most hosted inference providers implement it already, and a correct implementation requires care around the acceptance criterion to preserve output equivalence. For nearly all teams it is a configuration choice — enable it, pick a draft model, measure — rather than an engineering project.

How should it be evaluated?

By measuring end-to-end latency on representative production traffic, not on a synthetic benchmark. Acceptance rate should be monitored as an operational metric, because it will differ by workload and can shift when prompts or use cases change.

Because output is unchanged, quality evaluations do not need re-running, which is one of the technique's practical advantages over approximating optimisations.

What about self-speculation?

Some approaches use the target model itself, with early layers or a partial pass generating drafts, avoiding a separate draft model and its memory cost. These are attractive where memory is the binding constraint, though the speedup is typically smaller than a well-matched separate draft model achieves.

Where does it fit in a latency budget?

Generation time is only part of what a user experiences. Retrieval, tool calls, network round trips, and rendering all contribute, and in many agent systems the model is not the dominant term. Measuring the full path before optimising generation prevents the common outcome of a substantial decoding improvement that users cannot perceive.

Where generation genuinely dominates — long outputs, streaming interfaces, high token counts per response — the improvement is felt directly, and those are the workloads worth prioritising.

How does it interact with streaming?

Well, and slightly counter-intuitively. Tokens emerge in bursts rather than at a steady rate, because several are accepted at once. Interfaces that assume even pacing may display output in a way that feels uneven, and smoothing the render rate on the client is a small change that makes the faster generation feel faster rather than merely jumpier.

How FISTA Solutions helps

FISTA Solutions evaluates inference optimisations against real client workloads, measures acceptance rates and end-to-end latency rather than benchmark figures, and combines speculative decoding with caching, batching, and model routing where each fits, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To reduce AI latency without changing output quality, message FISTA on WhatsApp, or read how to reduce ai costs.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Does speculative decoding change the output?

No, when implemented correctly. The verification step ensures the accepted sequence matches what the target model would have generated on its own. That exactness is the property that makes the technique safe to apply without re-running evaluations.

02Where does the speedup come from?

From verifying several proposed tokens in one forward pass instead of generating them one at a time. Generation is normally memory-bandwidth bound rather than compute bound, so checking several candidates costs little more than producing one.

03When does it not help?

When the draft model is frequently wrong, which happens on novel, creative, or highly domain-specific text. Rejected tokens waste the draft computation, and acceptance rates below a certain point make the whole arrangement slower than plain generation.

04What does it cost?

Memory for two resident models, and operational complexity in keeping the draft and target compatible. Under heavy batching the benefit also shrinks, because the hardware is already busy and the spare capacity the technique exploits is no longer spare.

05Should teams implement it themselves?

Usually not. Major serving frameworks and hosted providers implement it already, and a correct implementation requires care around the acceptance criterion to preserve output equivalence exactly. For nearly all teams it is a configuration decision — enable it, pick a draft model, measure — rather than an engineering project.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project