AI Briefing
KO

22 tokens/sec on CPU

·2026.04.18 19:01

Key point

Q4_K_M Qwen 3.6 35B A3B ran at 22 tokens/sec on CPU.

Details

The Q4_K_M GGUF quantized version of Qwen 3.6 35B A3B was evaluated in a CPU-only environment.

  • Environment: 32 vCPU, 125GB RAM, no GPU
  • Runtime: llama-cpp-python / Unsloth GGUF
  • Benchmarks: HumanEval, HellaSwag, BFCL
  • Sample count: 1,264

The results were as follows.

  • HumanEval: 47.56%
  • HellaSwag: 74.30%
  • BFCL: 46.00%

Speed was around 22 tokens per second, with the best performance seen on commonsense reasoning.

Code generation and function calling were relatively weaker, staying in the mid-40% range, but the key point is that an active 3B MoE model can run on CPU at this kind of speed.

The evaluation was carried out by Neo AI Engineer, who found the appropriate quantized version and chat template and built a unified eval harness for the 3 benchmarks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.