It is widely assumed that dense token-level supervision is indispensable in the post-training of large models. To challenge this notion, the research team carried out experiments with OPD (Online Policy Distillation) as the starting point. The results of these experiments were quite revealing: retaining an extremely sparse gradient signal from just one token in a single reasoning trajectory is not only sufficient to prevent training collapse but also leads to a significant improvement in the reasoning capabilities of large models.
The team then conducted nine sets of tests involving various teacher-student model configurations. They found that choosing 1 to 2 tokens with the highest output divergence between the teacher and student models in the trajectory as supervision signals produced even better outcomes compared to full-signal training. The student models that emerged from this process were able to outperform their teacher models. Moreover, the performance enhancements in the student models were not simply a result of mimicking the teacher models.
This phenomenon has been substantiated across different scenarios, including code reasoning tasks, Llama series models, and the RLVR post-training paradigm. Mechanistic investigations have shown that the improvement in reasoning performance does not have a straightforward, monotonic relationship with the density of supervision signals. In fact, updating only 10% of the network parameters with supervision signals that occur once in every ten thousand instances can substantially elevate reasoning capabilities. The varying effects are also influenced by the relative strengths of the representational capabilities of the teacher and student models.
Drawing a parallel to human learning processes, the research team posits that large models, having already gone through stages like pre-training, already possess fundamental reasoning abilities. A small number of crucial supervision signals can act as guides, prompting the models to make adjustments and significantly enhance their reasoning performance.
