BeeLlama v0.2.0 Released, Significantly Boosting Inference Speed
·2026.05.23 02:37
Key point
The BeeLlama v0.2.0 update improves inference speed for Qwen 3.6 and Gemma 4 models by up to 4.9x.
Details
BeeLlama v0.2.0, an inference optimization tool, has been released. This update focused on performance optimization through the DFlash implementation.
Key updates include:
- Gemma 4 31B support: Efficient DFlash implementation and added vision capabilities.
- Qwen 3.6 27B optimization: Reduced DFlash overhead, improved prefill handling, K/V projection caching, and enhanced CUDA execution stability.
- Other improvements: DFlash GGUF support, strengthened reasoning and tool-call boundaries, and optimized verifier paths.
Benchmark results on an RTX 3090 environment recorded dramatic speed improvements compared to before:
- Qwen 3.6 27B: up to 163.9 tps (a 4.40x improvement over before).
- Gemma 4 31B: up to 177.8 tps (a 4.93x improvement over before).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.