Samsung's large model team, in collaboration with the University of Oxford and Peking University, has jointly proposed a new trust region-based On-Policy Distillation method—TrOPD. This method aims to address the issues of supervision signal distortion and training instability that arise when small end-side models inherit capabilities from large models due to significant capability gaps between teacher and student models. The research reveals that the key to successful OPD training lies in the teacher's ability to provide reliable supervision for each token generated by the student, rather than the choice of divergence formula. TrOPD identifies the teacher's reliable supervision regions, employs conventional policy learning within the trust region, adopts a gentler supervision approach in outlier regions, and introduces off-policy guidance to actively steer student-generated content closer to the trust region. Experimental results demonstrate that TrOPD outperforms existing OPD improvement methods across multiple benchmark tests, including mathematics, coding, instruction following, and STEM, with its advantages extending beyond the training domain. The greater the capability gap between teacher and student models, the more pronounced TrOPD's advantages become, offering a new pathway for low-cost deployment of cutting-edge AI to hundreds of millions of devices.
