Towards Effective Process Supervision for Mathematical Reasoning
Key point
Qwen has released a new PRM and benchmark called ProcessBench that identifies step-by-step errors in mathematical reasoning.
Details
LLMs can make calculation mistakes or logical errors during mathematical reasoning, and even when the final answer is correct, the intermediate process is often wrong, causing reliability issues. To address this, Process Reward Models (PRMs), which identify and mitigate errors in intermediate steps, play a key role.
Qwen has released ProcessBench, a new benchmark for measuring step-by-step errors in mathematical reasoning. This benchmark includes 3,400 competition- and olympiad-level math problems, consisting of step-by-step solutions with error locations annotated directly by experts.
In addition, Qwen2.5-Math-PRM-7B and 72B models were released together.
- Best-of-N (BoN) evaluation: The 7B model outperformed existing PRMs across 7 tasks including GSM8K and MATH, and exceeded majority voting (maj@8) performance on all tasks.
- ProcessBench evaluation: The 7B model outperformed all open-source models, and recorded higher performance than the paid model GPT-4o-0806.
Meanwhile, Qwen2.5-Math-RM-72B, an outcome-based reward model, also showed considerable ability to identify step-by-step errors, demonstrating the potential of reward models beyond rule-based mechanisms.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.