BeeLlama v0.2.0: Faster Inference with DFlash
·2026.05.23 02:34
Key point
The BeeLlama v0.2.0 update significantly boosts inference speed for Qwen and Gemma models through DFlash optimization.
Details
BeeLlama v0.2.0 has been released. The core of this update is improved inference performance through a major overhaul of the DFlash implementation.
Key update details are as follows:
- Efficient DFlash implementation for Gemma 4 31B and Vision feature support.
- Qwen 3.6 27B performance optimization: reduced DFlash overhead, improved prefill processing, drafter K/V projection caching, and safe CUDA execution support.
- DFlash GGUF architecture support and various bug fixes.
Benchmark results on an RTX 3090 environment showed dramatic speed improvements over the previous version:
- Qwen 3.6 27B: up to 163.9 tps (a 4.40x improvement over the previous version).
- Gemma 4 31B: up to 177.8 tps (a 4.93x improvement over the previous version).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.