Qianli Intelligent Driving has forged partnerships with a number of prestigious universities to jointly unveil and open-source the World Action Model WA-JEPA. This groundbreaking initiative paves a fresh avenue for autonomous driving models to comprehend scenes, forecast future events, and generate driving actions. Rather than striving to produce lifelike future images, the model zeroes in on predicting information that is intimately tied to driving planning within a high-dimensional latent space. WA-JEPA brings three significant enhancements to the original V-JEPA, boasting a parameter scale of around 481 million. Its initial training phase solely necessitates continuous driving videos, eliminating the need for labor-intensive manual annotations. In evaluations, WA-JEPA showcased outstanding performance and remarkable zero-shot transfer capabilities. Presently, the research paper, code, model weights, and evaluation procedures have all been released to the public.
