AI Briefing
KO

Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

·2026.08.18 09:00

Key point

A large-scale analysis of reasoning performance and cross-lingual transfer effects in non-English and multilingual environments using GRPO.

Details

Reinforcement Learning with Verifiable Rewards (RLVR), particularly GRPO (Group Relative Policy Optimization), is a key technique for enhancing the reasoning capabilities of language models, but previous research has been heavily biased toward English.

A large-scale empirical analysis of GRPO in non-English and multilingual settings, utilizing various base models, training languages, and reasoning language rewards, confirmed the following key findings:

  • Native Language Reasoning Performance: When reasoning is trained in a specific language, the performance gap compared to training in English is very small.
  • Cross-Lingual Transfer: Training in one language demonstrated strong transfer effects, improving performance across multiple other languages.
  • Risk of Language-Specific Performance Degradation: There are cases where training in a specific language severely degrades out-of-domain capabilities in other languages, with significant variation depending on the model and language.

While extending RLVR to languages other than English can yield broad performance improvements, it must be accompanied by comprehensive evaluations to detect language-specific performance regressions.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.