Speculative Decoding Speeds Up LLMs by Shifting GPU Work from Memory-Bound to Compute-Bound
Key point
Speculative decoding accelerates LLM inference by increasing total FLOPs to utilize idle compute resources, shifting the workload from memory-bound to compute-bound at low batch sizes.
Details
Speculative decoding does not speed up LLM inference by reducing the amount of work; in fact, it increases the total number of floating point operations (FLOPs). The speedup occurs because standard single-token decoding is heavily memory bound, requiring the GPU to stream all active parameters (e.g., 27B for Qwen3.5-27B) for every step, leaving arithmetic units idle. By using a small draft model to propose tokens that the large model verifies in parallel, speculative decoding allows the system to process multiple tokens per forward pass. This effectively trades breadth (many sequences) for depth (many tokens in one sequence), utilizing "free" compute capacity at low batch sizes.
However, this efficiency is context-dependent. At high batch sizes, the GPU becomes compute bound, and speculative decoding can waste resources, especially if draft tokens are rejected. Consequently, frameworks like vLLM allow users to disable speculation at high throughputs using the --speculative-disable-by-batch-size parameter. This insight has also led to research into dynamic strategies like dSpark, which adjust speculation depth based on real-time batch size and available compute.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.