Primitive Rhythm has unveiled NeoHorse-1, an open-source model available in two configurations: 4B and 9B. This model leverages the comprehensive trajectory data generated by the previously released OpenSquilla routing system, encompassing demand forecasting, routing decisions, execution processes, and environmental outcomes, to serve as instructional materials. NeoHorse-1 adopts the reasoning, action, and feedback loop of the Agent as its fundamental training unit. It accomplishes a single-round closed-loop training process through a series of meticulously designed steps, including trajectory filtering, curriculum learning tailored to routing capability requirements, and routing-guided On-Policy Distillation. This innovative training methodology substantially bolsters the model's proficiency in tasks involving Harness Agent operations and tool interactions. It is widely recognized as a replicable and effective single-round closed-loop validation of Recursive Self-Improvement (RSI). The project receives robust technical backing from Infinigence AI, with algorithm research spearheaded by a collaborative team from Peking University and Tsinghua University. The model developed, in conjunction with the Harness flywheel mechanism, is anticipated to fuel the Agent's continuous growth and advancement during execution.
