Concept
Decode is memory-bound: a step spends its time reading the weights, and the arithmetic units mostly wait. So a step that checks several positions costs about the same as one that produces a single token. Speculative decoding exploits that.
- A small, fast draft model (or a cheap extra head on the big model) guesses the next tokens, one after another.
- The large target model runs one forward pass over all guesses at once, as if they were a short prompt, which gives its own distribution at every one of those positions.
- A verify rule walks the guesses in order, accepting each with probability , where is the target's probability of the guessed token and the draft's. At the first rejection it draws a replacement from the leftover distribution and stops. If all are accepted, the target's pass also gives one more token for free.
Each target pass therefore yields between 1 and tokens. And the rule is built so that the tokens come out with exactly the target model's distribution. Speculative decoding changes the speed, never the output distribution: the draft only decides how much of the target's work is reused. (Under greedy decoding this reduces to "keep the guesses that match the target's argmax".)
What it costs: the draft's own steps, and some wasted target work on rejected guesses. It pays when the draft agrees with the target often (acceptance high) and is cheap (cost ratio low). It helps most at small batch sizes; at large batches the target step is no longer so memory-bound, and the verified positions compete with other users'.
Variants. Medusa adds several extra decoding heads to the target itself, each predicting a token further ahead, and verifies a tree of candidates in one pass. EAGLE drafts at the level of the target's hidden features rather than tokens, with a light head. Both avoid running a separate draft model.
Interactive
Draft γ tokens, verify them in one pass
Top: Leviathan et al.'s closed forms (i.i.d. acceptance α; one draft step costs c of a target step). Bottom: 4,000 rounds of the real accept/reject rule on an 8-token toy vocabulary.
Grey: the target distribution p. Blue: what speculative decoding actually emitted. They match whatever the draft is: only the speed changes.
Maths
Let the target be and the draft . A drafted is accepted with probability , so the per-token acceptance is
On rejection the replacement comes from . The emitted token's distribution is then
exactly the target's (Leviathan et al., §3; Chen et al.). With i.i.d. acceptance and drafts, the expected tokens per target pass are
and if a draft step costs target steps, the expected speed-up is (their Theorem 3.8). For example , gives 3.36 tokens per pass, and with a 2.80× speed-up.
Code
// src/lib/inference/speculative.ts (excerpt)
export function verify(
p: Vector,
q: Vector,
x: number,
u: number,
rng: () => number,
) {
if (u < Math.min(1, p[x]! / q[x]!)) return { accepted: true, token: x };
return { accepted: false, token: sampleFromProbs(residual(p, q), rng) };
}
The tests check the emitted distribution equals to , and run 40,000 Monte Carlo rounds whose acceptance, tokens per pass and token histogram match the formulas.