The research team has successfully developed a sophisticated multimodal agent, VideoGen-Agent, leveraging the capabilities of the Qwen3-VL-8B-Instruct model. This innovative agent is designed to autonomously invoke a variety of tools—such as retrieval, generation, and verification—to accomplish complex video generation tasks with precision and efficiency. Initially, the agent acquires proficiency in utilizing these tools through a process of supervised fine-tuning. Subsequently, it refines its decision-making abilities by employing multi-task reinforcement learning techniques. In rigorous testing conducted on the self-developed VABench benchmark, the agent achieved a remarkable score of 75.6 under the Toolset 1 configuration. This score represents a significant improvement of 19.1 points over the base video generator. Furthermore, after enhancing the generation tools, the agent's score soared to 86.1, all without the need for retraining its strategic approach. Human evaluations underscored the agent's prowess, with 84.3% of respondents expressing a preference for the videos it generated. Moreover, the agent exhibits a remarkable ability to autonomously combine tools, enabling it to tackle unfamiliar composite generation requirements with ease. It also demonstrates scalability, allowing for the integration of new toolchains to adapt to diverse scenarios, such as robot operation video generation. Its decision-making model, which prioritizes the completion of missing information prior to generation, effectively elevates the quality of video generation under complex and demanding conditions.
