Domino: 5.8x Faster Inference
Key point
Domino, a technique that separates modeling from drafting to maximize the efficiency of Speculative Decoding, has been unveiled.
Details
Domino, a new methodology that separates Causal Modeling from Autoregressive Drafting to maximize the efficiency of Speculative Decoding, has been released.
In tests on Qwen models, it achieved up to a 5.8x improvement in throughput compared to existing methods.
The related paper, source code, and model weights are all currently available via arXiv, GitHub, and Hugging Face.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.