GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

Apple Research Blog
Read full post
Researchers conducted a large-scale study on Group Relative Policy Optimization (GRPO) applied to non-English and multilingual language models. They found that training models to reason in native languages yields performance close to English and observed significant crosslingual transfer effects, though results vary by model and language. The study highlights the need for broad evaluation to detect language-specific regressions in reinforcement learning with verifiable rewards beyond English.

More in LLM & Text Generation

DeepSeek V4.1 Flash now available on AI Gateway

Covered by 2 sources

Harvey raises $550M more to develop AI tools for legal teams

SiliconANGLE