XPENG Group has recently introduced the TuringViT, an efficient visual encoder tailored for the VLM/VLA (Vision Language Model/Vision Language Assistant) age. This innovative encoder represents a paradigm shift in the architecture, data handling, and training methodologies of visual encoders. It is designed to seamlessly integrate with XPENG's three primary business domains: intelligent driving, intelligent cockpit systems, and the IRON humanoid robot. Moreover, TuringViT offers the industry a replicable and cost-effective technical approach for training state-of-the-art (SOTA) visual Transformers.
TuringViT marks significant advancements across three critical facets: architecture, data utilization, and training efficiency, thereby establishing a comprehensive technical framework. It comes in two variants, TuringViT-18L and TuringViT-24L, both showcasing remarkable efficiency benefits, especially at high resolutions.
The VISTA-Curation multimodal data governance pipeline employed by TuringViT achieves outstanding performance across multiple benchmark tasks, utilizing a mere 0.85 billion image-text pairs—roughly 10% of the training data scale used by SigLIP2-L. Furthermore, TuringViT adopts a four-stage progressive native dynamic resolution training approach. This strategy is instrumental in enhancing intelligent driving systems, refining cockpit interaction recognition accuracy, and serving as the 'visual retina' for humanoid robots, such as being the core visual encoder for the second-generation VLA model.
