AI Briefing
KO

Qwen3 4B outperforms cloud agents on code tasks

·2026.04.28 06:41

Key point

On the Mahoraga benchmark, Qwen3 4B outperformed cloud agents.

Details

Mahoraga is an open-source orchestrator that routes local and cloud AI agents using a LinUCB-based contextual bandit.

On a 16GB MacBook Pro M-series, it ran 192 tasks, comparing 4 local Ollama models and 4 cloud CLIs via forced round-robin.

Evaluation was scored using a 4-layer heuristic.

  • novelty ratio
  • structural checks
  • embedding similarity
  • length ratio

Qwen3 4B nothink mode produced the best results on code and refactoring. Throughput was 33.8 t/s, average latency was 6.1 seconds, and cloud agents' own code scores generally stayed around 0.650.

No LLM-as-judge was used in the evaluation, and no API costs were incurred. The author interpreted this to mean that, in this setup, local models are not simply a cheaper alternative but can outperform cloud models on certain code generation tasks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.