On July 21, XPENG Group proudly announced the launch of its cutting-edge TuringViT efficient visual encoder, tailored specifically for the Visual Language Model (VLM)/Visual Language Assistant (VLA) era. This innovative encoder represents a significant leap forward, as it reimagines the architecture, data paradigm, and training methodology of traditional visual encoders.
TuringViT features a design primarily centered around linear attention mechanisms, offering two distinct versions: the 18L and the 24L. Both variants have undergone rigorous empirical testing, showcasing their superiority in high-resolution inference tasks. This translates into a substantial boost in deployment efficiency on the end-user side, ensuring smoother and more responsive visual processing.
Furthermore, this state-of-the-art encoder is set to provide comprehensive support across XPENG's three core business domains: intelligent driving systems, intelligent cockpit experiences, and the groundbreaking IRON humanoid robot. By integrating TuringViT into these key areas, XPENG aims to redefine the boundaries of visual perception and interaction, ushering in a new era of intelligent technology.
