AI Briefing
KO

Qwen3.8 vs Qwen3.6 vs Gemma 4 on 24GB GPUs (10-minute read)

·2026.08.18 09:00

Key point

In a 24GB GPU environment, Qwen3.8-27B demonstrated the best performance in coding and reasoning tasks.

Details

For local AI workloads on 24GB GPUs (e.g., RTX 4090), Qwen3.8-27B is evaluated as the strongest default choice among the three models. Under the same Q4_K_M quantization environment, it maintains decoding speeds similar to Qwen3.6-27B while passing all 12 coding tasks in a single seed and achieving high accuracy in reasoning and document QA.

Gemma 4 31B-it delivered the best results in structured tool calling (90/90) and vision tasks (19/20), but had drawbacks including high VRAM usage and slow initial token generation. Specifically, when using a 64K context, it required switching from an F16 K/V cache to Q8_0 to operate stably.

  • Coding agents and repository editing: Qwen3.8-27B (12/12 pass@1)
  • Strict tool calling: Gemma 4 31B-it (100% success in both single/multi-step)
  • Vision-centric tasks: Gemma 4 31B-it (19/20)
  • 64K context headroom: Qwen3.8 or Qwen3.6 (4.20 GiB headroom with F16 K/V)

If existing integrations are stable, it is advantageous to stick with Qwen3.6-27B, but for new deployments, Qwen3.8-27B offers a practical upgrade.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.