AI Briefing
KO

Qwen Dominates Compatibility

·2026.04.18 10:03

Key point

On M3 Ultra, Qwen 3.6 35B achieved a 100% tool-calling success rate across most of the 5 agent frameworks.

Details

This is a tool-calling compatibility benchmark comparing 5 agent frameworks and 7 models on an Apple M3 Ultra(256GB).

  • Frameworks tested: Hermes Agent, PydanticAI, LangChain, smolagents, OpenClaude/Anthropic SDK
  • Models tested: Qwen 3.6 35B, Qwen 3.5 35B, Qwopus 27B, Qwen 3.5 27B, Llama 3.3 70B, DeepSeek-R1 32B, Gemma 4 26B
  • Evaluation items: single/multiple tool selection, multi-turn, streaming, stress testing, multiple tool injection, no-leak check

The key result is the dominance of the Qwen family. Qwen 3.6 35B(4bit) scored 100% on Hermes, PydanticAI, smolagents, and OpenClaude, and 93% on LangChain, with a speed of 100 tok/s.

By contrast, non-Qwen models showed large variance across frameworks. Gemma 4 26B scored 67% on PydanticAI and 80% on OpenClaude; DeepSeek-R1 32B dropped to 55% on Hermes, 50% on PydanticAI, and 40% on OpenClaude; and Llama 3.3 70B also only managed 45% on Hermes and 67% on LangChain.

In the speed table, Qwen3.5-4B(168 tok/s), GPT-OSS 20B(127 tok/s), and Qwen3.5-9B(108 tok/s) were presented as fast options, and Qwen 3.6 35B delivered 100 tok/s using about 20GB RAM while maintaining 100% tool-calling. The conclusion is that for agentic local LLMs on Apple Silicon, the Qwen family is the most stable choice.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.