GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings
Apple Research Blog
Read full postResearchers conducted a large-scale study on Group Relative Policy Optimization (GRPO) applied to non-English and multilingual language models. They found that training models to reason in native languages yields performance close to English and observed significant crosslingual transfer effects, though results vary by model and language. The study highlights the need for broad evaluation to detect language-specific regressions in reinforcement learning with verifiable rewards beyond English.




