Qwen3.6-27B SimpleQA 95.7%
·2026.05.02 20:21
Key point
Qwen3.6-27B achieved 95.7% on SimpleQA in a single RTX 3090 local environment.
Details
This is a fully local web-search agent combining a single RTX 3090 24GB, Ollama (qwen3.6:27b), and a langgraph_agent strategy based on LangChain create_agent().
- It used tool-calling, parallel subtopic decomposition, and up to 50 iterations.
- Grading was done via qwen3.6:27b self-graded method.
- Results were SimpleQA 95.7% (287/300) and xbench-DeepSearch 77.0% (77/100).
- The comparison table also presented Qwen3.5-9B 91.2% (182/200) and gpt-oss-20B 85.4% (295/346).
- These are agent + search scores, not closed-book performance.
- Caveats noted include possible SimpleQA contamination, LLM-judge noise, and no repeated re-testing; BrowseComp / GAIA figures are not yet available.
- It also added that xbench-DeepSearch is a Chinese-language benchmark, which may favor the Qwen family.
- It was compared as being at a similar level to Perplexity Deep Research (93.9%) and Tavily (93.3%).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.