On October 27, 2025, the Meituan LongCat team made an official announcement, introducing and open-sourcing the LongCat-Video video generation model. Built upon the Diffusion Transformer architecture, this model cleverly differentiates tasks based on the 'number of conditional frames'. It boasts versatile capabilities, including text-to-video conversion, image-to-video transformation, and video continuation, thereby creating a seamless and comprehensive task closed-loop system.
The model is capable of generating high-definition videos at 720p resolution and 30 frames per second. It can produce lengthy videos, with a maximum duration of up to 5 minutes, while maintaining impeccable cross-frame temporal consistency and ensuring that the physical motions depicted are highly plausible.
The base model, which incorporates 13.6 billion parameters, sets a new benchmark in open-source SOTA (State-of-the-Art) performance for both text-to-video and image-to-video tasks. Moreover, it showcases a remarkable 10.1-fold improvement in inference efficiency. Presently, the model is available for open-source access on popular platforms like GitHub and Hugging Face.
