Open-source 30B model outperforms GPT-5.5
Key point
An open-source 30B model surpassed GPT-5.5 on the VibeWorlding benchmark, setting a new standard for 3D world generation.
Details
The VibeWorlding benchmark, released by Yansong Ning's research team, evaluates whether multimodal agents can build 3D open worlds end-to-end from natural language prompts.
State-of-the-art frontier models such as GPT-5.5 and Qwen3.8-Max recorded success rates (Pass@1) of less than 60% on this task, revealing that precise 3D world editing remains a major bottleneck.
In contrast, VibeWorlder-30B-A3B, an open-weight 30B model, achieved the highest Pass@1 among the evaluated models through reinforcement learning based on verifiable rewards.
These results suggest that an RL (reinforcement learning) approach leveraging sandbox tools and verifiers can significantly enhance 3D content generation performance, rather than relying solely on parameter scale.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.