AI Briefing
KO

ByteDance and Tsinghua University Release DAPO, an Open-Source RL System Achieving 50 Points on AIME 2024

·2026.09.21 09:00

Key point

ByteDance and Tsinghua University have released DAPO, a fully open-source RL system that achieved 50 points on AIME 2024.

1 / 3

Details

ByteDance Seed and Tsinghua University AIR have released DAPO, a fully open-source system for large-scale LLM reinforcement learning. The system includes algorithms, code infrastructure, and datasets, achieving 50 points on AIME 2024 with 50% of the training steps compared to the previous SoTA, DeepSeek-R1-Zero-Qwen-32B. DAPO is based on the Qwen2.5-32B model; the initial version (excluding token-level PG loss and dynamic sampling) scored 44 points, while the full version achieved 50 points. The system is built on the verl framework, with experiments conducted on the Volcano Engine platform.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.