The AIBuildAI team has unveiled a groundbreaking recursive self-improving agent specifically designed for the autonomous post-training of large language models—aptly named the PostTrain Agent. Moreover, they have generously open-sourced all its code and the supporting knowledge base. Users are simply required to input the base model, desired target capabilities, and computational budget. The system then takes over, seamlessly completing the entire post-training process without any human intervention, ultimately delivering a fully trained model.
Harnessing the power of a comprehensive knowledge system, the PostTrain Agent is capable of making informed decisions grounded in reliable evidence. Through the application of meta-search technology, it autonomously designs a search execution graph tailored precisely to the task at hand. This innovative approach effectively overcomes the limitations inherent in traditional linear and tree searches, which often struggle to adapt to the intricate and complex nature of post-training workflows.
On the autonomous post-training benchmark, PostTrainBench, the PostTrain Agent has achieved an impressive score of 46.6 points, securing the top position and outperforming all state-of-the-art models and agent systems. Its performance is remarkably close to the 51.1-point mark attained by human experts. When utilizing the same underlying model, Claude Opus 5, the PostTrain Agent scored a notable 11.6 points higher than directly employing Claude Code to accomplish tasks. This significant improvement underscores the substantial benefits derived from the knowledge system and meta-search capabilities, even surpassing the impact of upgrading the underlying model to the next generation.
This pioneering work signifies a pivotal evolution in the post-training landscape for large models. It marks the transition from relying on human experts to operate tools to AI autonomously completing the entire research and development closed loop. With immense potential for future expansion to encompass the entire AI R&D process, this innovation is poised to drive the formation of a robust infrastructure that enables AI to independently develop AI.
