AI Briefing
KO

ExLlamaV3 Major Update

·2026.05.11 16:05

Key point

ExLlamaV3 has added Gemma 4 support, DFlash optimization, and bug fixes.

Details

ExLlamaV3 has rolled out consecutive updates from v0.0.29 to v0.0.33. Gemma 4 support, cache efficiency improvements, DFlash support/quantization, along with additional bug fixes and efficiency improvements, have followed.

DFlash showed significant speed improvements across multiple workloads.

  • Coding: 59.21 t/s → 177.67 t/s (3.00x)
  • Agentic code: 55.98 t/s → 140.61 t/s (2.51x)
  • Agentic curl: 54.03 t/s → 125.94 t/s (2.33x)
  • Translation: 58.11 t/s → 75.73 t/s (1.30x)
  • Translation (reasoning): 58.08 t/s → 119.43 t/s (2.06x)
  • Creative writing: 59.10 t/s → 89.19 t/s (1.50x)

Also, 3090, 4090, 5090, 6000 Pro measurements for Qwen3.5, Trinity-Nano, and Gemma4 family configurations around 4bit were released, and DFlash model quantization, additional bug fixes, and further optimizations continue to progress on the dev branch.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.