WaveFront Decoding Achieves Up to 4.81x Speedup for Looped Language Models
Key point
The training-free framework reaches 4.81x speedup on Huginn-3.5B by concurrently batching drafting and verification.
Details
WaveFront Decoding (WFD) is a training-free self-speculative decoding framework designed to reduce the high latency of looped language models, which increase effective depth by repeatedly applying weight-shared blocks.
How It Works
WFD exploits two architectural properties: weight sharing allows token states at different positions and recurrence depths to be processed in a single batch, and shallow-depth states can effectively draft predictions for new positions. The system organizes these mixed-depth states into a wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike traditional draft-then-verify schedules, WFD concurrently batches drafting and verification within the same recurrent calls, correcting rejected drafts using full-depth predictions.
Performance Benchmarks
Across six Spec-Bench task categories, WFD consistently outperforms standard draft-then-verify methods:
- Ouro-2.6B: Achieves a 2.42x speedup over autoregressive decoding.
- Huginn-3.5B: Achieves a 3.54x speedup.
- Optimized Huginn-3.5B: With cross-recurrence KV sharing to reduce traffic, the speedup increases to 4.81x.
The code is available on GitHub, and the paper is published on arXiv.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.