Why Weibo's tiny VibeThinker-3B is stirring up benchmark controversy again
Key point
China's Weibo has released a 3B-scale VibeThinker model that has become controversial after posting benchmark scores that surpass ultra-large models.
Details
A research team at China's social media company Sina Weibo has released a language model called VibeThinker-3B, which has just 3 billion (3B) parameters. The model is shocking the AI industry by showing reasoning performance on par with or exceeding existing ultra-large AI models from Google DeepMind, OpenAI, Anthropic, and others.
Key achievements are as follows:
- Math ability: It scored 94.3 on AIME 2026, matching the 671B-scale DeepSeek V3.2 and surpassing Google's Gemini 3 Pro (91.7).
- Coding ability: It scored 80.2 (Pass@1) on LiveCodeBench v6, and achieved a high 96.1% accuracy rate on the latest LeetCode problems.
- Reasoning optimization: When a test-time scaling technique called Claim-Level Reliability Assessment is applied, the AIME score rises to 97.1.
These results raise a fundamental question: is scaling up model size the only path to intelligence? However, there is also skeptical opinion suggesting that the model's performance may have been abnormally inflated, and that AI benchmarks are being 'gamed' and failing to properly measure model performance.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.