A collaborative study by Apple and the Hasso Plattner Institute, based on over 200 sets of GRPO reinforcement learning reasoning training experiments, has revealed that the performance gap between high-resource languages trained in their native language for reasoning and those trained in English is generally within 2%. Specifically, the gap for Chinese is a mere 1.1%. This finding challenges the conventional wisdom that reasoning must be conducted in English, demonstrating that reasoning models for non-English high-resource languages, such as Chinese, can be effectively trained independently. The experiments also highlighted that cross-language transfer is generally effective, with training in low-resource languages sometimes enhancing generalization capabilities. Additionally, multilingual mixed training offers high cost-effectiveness. However, it was noted that certain model-language combinations may lead to a significant decline in performance for specific tasks, underscoring the need for language-, task-, and model-specific risk assessments.
