← all articles

Speculative decoding explained: how it works and why

Speculative decoding is a trick for making a large language model produce text faster without changing what it says. A small, cheap model guesses the next several words, and the big model checks all those guesses at once instead of writing them one at a time. When the guesses are right, you get several words for the price of one step. When they are wrong, you lose almost nothing.

I run AI tools for a living out of Singapore, and this is one of the few speed-ups I have seen that does not come with a quality trade-off attached. If you host your own models or wonder why one provider streams faster than another, it is worth understanding.

what it is

Speculative decoding is an inference technique. It does not change how a model is trained, and it does not change the model’s weights. It changes how the finished model is run to generate text.

Two models are involved:

  • the target model: the big, slow, high-quality model you actually want answers from
  • the draft model: a much smaller model, often from the same family, that is fast and usually right about easy words

The draft model proposes a short run of tokens. The target model then verifies that run and keeps the part it agrees with. The output is meant to be statistically identical to what the target model would have produced on its own. It is a speed technique, not a quality compromise.

The idea was published in 2022 by Yaniv Leviathan, Matan Kalman and Yossi Matias at Google, in the paper Fast Inference from Transformers via Speculative Decoding. A team at DeepMind independently described a very similar method, which they called speculative sampling, in Accelerating Large Language Model Decoding with Speculative Sampling. Both papers report roughly two to three times faster generation on the models they tested, with the same output distribution.

how it works

To see why this helps, you need one fact about how language models generate text. They produce one token at a time. A token is a chunk of text, usually a word or part of a word. To write a 200 token answer, the model runs 200 times in sequence, and each run needs the token from the run before it.

Here is the odd part. Each of those runs is slow mostly because the hardware has to load billions of weights from memory, not because the arithmetic is hard. On a modern GPU the compute units sit partly idle while they wait for memory. Checking five tokens in one pass costs only a little more time than generating one, because the weights are loaded once either way.

Speculative decoding exploits that idle capacity. The loop goes like this:

  1. the draft model quickly generates a short run of tokens, say four or five
  2. the target model takes the prompt plus those draft tokens and scores all of them in a single forward pass
  3. going left to right, the system compares what the target model would have chosen with what the draft proposed
  4. every draft token that passes is accepted
  5. at the first token that fails, the target model’s own choice is used instead, and the rest of the draft is thrown away
  6. the loop repeats from that point

Even in the worst case, where the very first draft token is rejected, you still get one good token from the target model, which is what you would have had anyway. In the best case you get the whole draft plus a bonus token. So the floor is normal speed and the ceiling is several tokens per step.

The acceptance rule is what keeps the output honest. In the sampling version, a draft token is accepted with a probability based on how likely the target model finds it compared with the draft model. A rejected token is replaced by a sample from a corrected distribution. The papers prove this gives the same distribution as sampling from the target model directly. The proof is in the Leviathan paper linked above.

a worked example

Say the prompt ends with “the capital of France is”. A tiny draft model will happily propose “Paris, which is located”. The target model checks those four tokens in one pass, agrees with “Paris”, “which” and “is”, but would have written “situated” instead of “located”. You keep three tokens and swap in the target’s fourth. That is four tokens from one big-model pass instead of four passes.

Now say the prompt is a tricky legal clause. The draft model is often wrong, most guesses get rejected, and you drift back to roughly normal speed. The technique helps most where text is predictable, and boilerplate, code and structured output are very predictable.

variants you will run into

Running a second model has a cost, since you need memory for it and it has to be reasonably well matched to the target. So researchers have built variants that skip the separate draft model:

  • prompt lookup or n-gram drafting: copy likely continuations straight from the prompt, which works well for summarising or editing text you already supplied. Hugging Face describes its version of this in its assisted generation post.
  • Medusa: add extra prediction heads to the target model itself so it drafts its own next few tokens, described in the Medusa paper.
  • EAGLE: train a small draft component that works on the target model’s internal features rather than on raw text.

Serving engines such as vLLM and Hugging Face transformers expose some of these as settings; check the current documentation, since option names change quickly.

why it matters

Speed is the obvious answer, but not the only one.

  • latency you can feel: for chat and coding tools, the time it takes text to stream matters more than the total throughput. Faster token generation makes an assistant feel responsive. This connects to something I wrote about in coding assistants judged on review time, not typing time, where waiting on the model is only one part of the real cost.
  • cost per answer on your own hardware: if you host a model yourself, you are paying for GPU time. Getting more tokens per big-model pass means fewer GPU seconds per response. For anyone shopping from the best open-source LLMs you can self-host in 2026, this is one of the cheaper ways to get more out of the same card.
  • no retraining and no quality loss: most speed-ups ask you to accept a smaller or quantised model. This one keeps the target model as it is. If your evaluation results were good yesterday, they should be the same today.
  • long, predictable outputs: code generation, JSON, templated reports and edits to existing documents all have stretches where the next tokens are easy to guess.

If you are self-hosting mainly because you want your data to stay on your own machines, the theprivacywire.com blog covers that side of the decision.

common misconceptions

“It makes the answers worse because a small model is involved.” No. The small model only proposes. The big model has the final say on every token, and the published method is built so the result matches what the big model would have produced alone. A badly chosen draft model slows you down.

“It always makes things faster.” Not always. If the draft model is poorly matched to the target, most guesses are rejected and you pay for the extra draft work with little gained. The speed-up depends on the acceptance rate. Under heavy load on a shared server, where the GPU is already busy serving many users at once, the spare capacity that speculative decoding relies on may not exist. I would benchmark on your own prompts before trusting a paper’s number.

“It is the same thing as using a smaller model.” No. A smaller model gives you smaller-model answers. Speculative decoding uses a small model as a helper and still delivers the big model’s answers. You need both models loaded, so it costs extra memory.

“Hosted APIs are doing this, so I need to set it up.” Maybe, maybe not. Providers do not always say which inference tricks sit behind an endpoint. For hosted APIs it is the provider’s problem, and you feel it only as speed and price; it becomes your decision when you self-host. If you do switch providers and the streaming speed changes, what changes when you move from one provider to another is a good checklist for what else to compare.

where to go from here

Once the basic idea clicks, these are the topics I would read next.

You can also browse everything else on the blog index.

The short version: speculative decoding spends a little spare hardware capacity to get several tokens per big-model step, and it is designed to leave the answer unchanged. Try it if you self-host, measure it on your own prompts, and do not expect a magic number.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-30.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →