A collaborative effort from researchers at Huawei, The Chinese University of Hong Kong, and The Hong Kong University of Science and Technology has led to the development of LegoFlow, a cutting-edge framework for code data engineering. This innovative framework automates the entire code data pipeline, transforming each process into standardized plugin skills that coding agents can directly utilize. The research team unveiled the LegoFlow-SWE dataset, a meticulously curated collection of 5,000 validation tasks selected from a pool of 12 million candidate PRs. This dataset spans across 8 programming languages and includes 2,780 high-quality trajectories. Remarkably, by training the Qwen3.5-35B-A3B-Base model with just 1,000 of these trajectories, the team achieved exceptional results in relevant benchmark evaluations.
The LegoFlow framework is designed with modularity in mind, with each module dedicated to specific tasks such as task construction, trajectory generation, fine-tuning training, and benchmark evaluation. Additionally, it features a dashboard that provides real-time monitoring of the entire process's progress. The research also sheds light on several key insights, including the fact that task quality outweighs quantity, the necessity of implementing measures to prevent model evaluation cheating, and the critical importance of the inference depth of SFT trajectories over their coverage.
Following two rounds of recursive iteration with autonomous agent adjustment strategies, the model's performance on the SWE-bench Verified saw significant enhancements. Presently, the project is expanding its horizons to encompass terminal tasks, complex long-term software engineering tasks, and recursive self-improvement. All relevant code, data, and documentation have been made openly available to the public.
